Most agent failures are not model failures. They are context failures.
A modern model can reason, call tools, and hold a natural conversation. But it still cannot act on information it does not have, and it gets worse when too much irrelevant information competes for attention. That is why production agents are less about stuffing more into the prompt and more about delivering the right context at the right time.
Table of contents
Open Table of contents
- The core idea
- Why giant context windows are not the answer
- Progressive disclosure is the practical fix
- Conditions are what make it work
- Why this matters more at production scale
- Flows still matter, but they should not run the whole system
- A practical way to think about agent design
- The real takeaway
The core idea
Context engineering is the discipline of deciding:
- what the agent should know now,
- what it should not know yet,
- and what should be unlocked only after a condition is met.
That sounds simple, but it changes how you build the system.
Older support systems worked like phone trees. Then AI systems moved to decision trees and SOP-driven flows. Those approaches are fine when the task space is small. They become brittle when the number of journeys, policies, and edge cases grows.
Production agents need a looser but more disciplined model:
- goals instead of rigid scripts,
- guardrails instead of hardcoded every-step logic,
- and progressive disclosure instead of one giant context blob.
Why giant context windows are not the answer
More context is not automatically better context.
As the context window fills up, the model has to decide what deserves attention. Every irrelevant paragraph, policy, or country-specific rule competes with the information that actually matters for the current turn. The result is predictable:
- worse recall,
- weaker decisions,
- more hallucination,
- and wasted tokens.
This is why agent quality often drops long before the model itself hits a hard capability ceiling.
Progressive disclosure is the practical fix
The most useful pattern in this whole discussion is progressive disclosure.
Start with the minimum viable context:
- general policy,
- brand voice,
- basic tools,
- and the current task.
Then reveal more only when the conversation earns it.
If a user asks about an international shipment, the agent does not need rules for every country upfront. Once it learns the shipment is going to Germany, Germany-specific guidance becomes relevant. Before that, it is noise.
This is the difference between giving the model a library and handing it the one page it needs.

A good production agent starts small, then unlocks tools, policies, and workflows only when the conversation makes them relevant.
Conditions are what make it work
Progressive disclosure only works if the system knows when to unlock more context.
Those unlocks usually come from two kinds of conditions.
State-based conditions
These come from system state:
- the user is authenticated,
- a tool returns a specific result,
- an account or subscription is loaded,
- a payment dispute is detected.
Observation-based conditions
These come from what the model observes in the conversation:
- the user mentions cancellation,
- the user asks about a refund,
- the user references a specific product,
- the user signals confusion, urgency, or intent.
Once the condition is met, the system reveals the next layer of context.
Why this matters more at production scale
A small agent can often survive with sloppy context management. A large one cannot.
An agent with five supported journeys may work fine with loose prompts and a couple of tools. An agent with fifty journeys, multiple systems, and segment-specific policies needs discipline. Without it, the model gets overloaded and the user experience degrades fast.
This is why context engineering is not prompt polish. It is systems design.
At scale, it improves:
- accuracy,
- cost efficiency,
- naturalness,
- operational control,
- and rollout safety.
It also makes the system more durable. If your architecture already separates context into blocks and unlock conditions, stronger future models can take advantage of that structure immediately. If the system is hardcoded into brittle flows, model upgrades help less than people expect.
Flows still matter, but they should not run the whole system
This is not an argument against workflows.
Some flows are necessary:
- regulated intake,
- compliance-heavy verification,
- irreversible operations,
- or high-risk handoffs.
But in strong agent systems, a workflow becomes one more contextual tool, not the master architecture for every conversation.
That distinction matters. It lets the agent stay flexible in the common case while still switching into a strict path where the business really needs one.
A practical way to think about agent design
If you are building production agents, a useful mental model is:
- Keep the base context small.
- Break additional knowledge into separate blocks.
- Define the condition that unlocks each block.
- Reveal only what is needed for the next decision.
- Treat rigid workflows as special-purpose modules, not the default operating mode.
That is a much better scaling path than trying to predict every branch in advance.
The real takeaway
Model quality still matters. But model quality alone does not create a good agent.
The real leverage comes from context architecture: deciding what the agent sees, when it sees it, and what must stay out of view until it becomes relevant.
That is what turns general model capability into production-grade behavior.
Based on Neil Rahilly’s 2026 post on context engineering, this rewrite focuses on the architectural idea that matters most in practice: good agents are built by controlling context, not by overfeeding it.