Loading an agent skill does not keep its requirements active in long contexts — a code-audit task passed 8/10 runs in an ~11k-character clean context but only 3/10 at 299k characters (trend-level, p=0.0698), while a detailed external checklist passed 10/10 versus 5/10 for generic self-checking (p=0.0325); a second task passed all conditions, so no universal context-length threshold is supported.
When and How Context Rot Appears in Coding Agents: A White-Box Study of Agent Skills in Code AuditingWhat it means For long agent sessions, externalize compliance into explicit checklists the agent must tick against rather than trusting a loaded skill or self-review to stay active; small run counts, so treat effect sizes as indicative.
arXiv preprint · JUL 20