A brief confession of developer hubris to begin.
When my engineering team first began deploying autonomous agent pipelines across our repositories, we operated under the comfortable illusion that governing a large language model was simply an exercise in drafting authoritative English prose. We approached it like overzealous schoolmasters and assembled what we fondly regarded as the definitive system prompt: 3,500 words of majestic, unyielding instruction.
“You are an elite principal architect. You must ALWAYS write modular, clean, enterprise-grade TypeScript. You must NEVER use any. You must NEVER modify the database schema without asking. You must NEVER create new files unless strictly necessary. Do not hallucinate. Do not write bugs. Be concise.”
We committed this to our workspace rules, sat back, and received the model’s solemn digital assurance: “Understood. I will strictly adhere to all enterprise standards.”
Twenty turns later, the agent had imported three deprecated packages, converted a performant SQL stored procedure into an N+1 query loop, removed half a test suite, and generated five new files titled temp_fix_final_v2.js in the project root.
We spent several expensive months treating system prompts like Christmas wish lists before sitting down to conduct a proper forensic autopsy on what was actually happening under the hood.
Along the way, our worst rule-authoring habits cost us hundreds of thousands of wasted tokens, blown API allocations, and several late evenings rescuing git trees. Here are the three most expensive failure modes we diagnosed, the transformer attention physics behind them, and how to write invariants that actually hold.
Trap 1: The “Pink Elephant” Failure Mode (Bare Negatives as Attention Magnets)
Tell a colleague in casual conversation, “Whatever you do, do not think of a pink elephant,” and their mind will dutifully summon a fluorescent pachyderm.
In transformer self-attention, the outcome is noticeably worse. Softmax attention distributions do not possess a native grammatical NOT operator. When you write:
[MISTAKE]“Do NOT create any new files on disk.”
The attention heads calculate cross-attention across the semantic tokens. The tokens receiving the highest entropy weight in that phrase are create, new, and files. The tiny token NOT exerts diffuse, negligible cross-attention across a 30,000-token context window.
By deploying a bare negative prohibition, you have effectively turned the forbidden action into a semantic attention magnet. Under pressure, the model attends heavily to create new files and executes the exact behaviour you requested it to avoid.
-
What it cost us: Hours spent pruning phantom scratch files and untangling infinite retry loops where an agent, forbidden from touching a file, panicked and devised inventive workarounds in unrelated directories.
-
The Engineering Fix (The Paired Invariant Standard):
Never leave a negative constraint in isolation. Every negative boundary must be explicitly coupled to an affirmative mandate:[PAIRED INVARIANT]
Positive Mandate: Restrict all modifications strictly to existing lines withinsrc/controller.js.
Negative Constraint: Absolute prohibition on creating new files or modifying parent directories.
When the model is given a defined railway track to follow, the negative boundary functions as a guardrail rather than an irresistible semantic prompt.
Trap 2: The “Kitchen Sink” Monolith and the 30k Attention Dip
Our second mistake was the Monolithic Kitchen Sink. We maintained a single 4,200-token markdown document attempting to regulate everything at once:
- Database connection pooling
- CSS flexbox naming conventions
- Licensing headers
- Git commit message regexes
- Unit test mocking standards
In Turn 1, the model appeared impeccably compliant. By Turn 28, it had succumbed to the well-documented “Lost in the Middle” attention dip.
As multi-turn context accumulates (tool outputs, compiler diffs, stack traces), the attention distribution flattens. The middle 80% of your prompt enters an attention twilight. The model has not deleted the rule; it is experiencing attention interference. It simply loses track of whether it was instructed to use vanilla DOM APIs or permitted to pull in an external utility library.
- What it cost us: A substantial Token Tax. Injecting 4,200 fixed tokens of rules across a 60-turn session burns 252,000 input tokens purely on reciting regulations before the model reads a single byte of repository code.
- The Engineering Fix (Scope Anchoring & ANIR Compression):
- Strict Context Scoping: A front-end layout task should never carry database transaction rules in its prompt. Constraints should be partitioned into domain-specific workspace families or surfaced dynamically via tool schemas.
- High-SNR Vectorisation: We began compiling verbose natural language into high-contrast Agent-Native Intermediate Representation (ANIR):
- Natural Language (42 tokens): “Whenever you write database queries in this project, you must never write N+1 select loops in the application layer, but instead always use set-based batch operations via Table-Valued Parameters.”
- Compiled ANIR (9 tokens – 78% reduction):
[DB: SET_BASED(TVP|XML_SHRED) !N1_LOOP]
Trap 3: Compaction Evaporation (The Amnesia of Context Summaries)
This was the most insidious failure mode we encountered, requiring weeks of trace archaeology to isolate.
When an autonomous coding session extends and approaches the context threshold, the agent environment triggers automated compaction (the <CONTEXT_SUMMARY> block).
The vulnerability lies in the fact that summarisation models exhibit an inherent positive narrative bias.
When an LLM summarises a 40-turn trajectory, it faithfully records affirmative actions:
- “Created controller.ts”
- “Executed npm test”
- “Refactored auth handler”
The element it reliably discards: your negative operational constraints.
The summariser treats “Remember not to touch the billing schema” as conversational preamble rather than an active system invariant. Once the context is compacted, your negative boundaries evaporate. At Turn 42, the newly compacted agent reverts to base pre-training priors and promptly modifies the protected table.
- What it cost us: Production rollback drills and corrupted local test databases.
- The Engineering Fix (Root Ingestion & Protocol Gates):
Operational rules cannot be left to survive inside the volatile in-chat conversational stream. Critical boundaries must either reside at the immutable root instruction layer (re-anchored after every compaction cycle) or, preferably, be enforced through physical protocol gates in an MCP server (such as an execution guard) that reject unauthorised disk writes until explicitly unlocked by the developer.
Three Core Invariants for Rule Design
Having paid for these lessons in compute and developer hours, three core invariants now govern our rule design:
- The Paired Invariant Rule: Never deploy a negative constraint in isolation. If you instruct an agent what not to touch, explicitly point it to the affirmative track it is permitted to edit instead.
- The 120-Word Atomicity Limit: The moment an individual rule requires four nested sub-clauses, semicolons, and three “and also remember…” qualifications, it has ceased to be an operational constraint and become a sprawling, unindexed novella. Decompose it into atomic, single-domain assertions.
- Deterministic Gates Over Prompt Pleading: If an agent violating an instruction would break production or corrupt a database, never rely on prompt prose to prevent it. Enforce it through physical tool gating (such as an MCP barrier or pre-commit hook). A prompt is probabilistic; a rejected filesystem call is absolute.
Observations & War Stories
We arrived at these patterns through the unglamorous process of watching autonomous agents devise inventive ways to misinterpret plain English.
- What is the most stubbornly creative rule evasion you have caught an agent executing?