CONTEXT ENGINEERING: WHAT IT IS AND WHY IT MATTERS.
This article explains what context engineering is, how it differs from prompt engineering, why context is the real bottleneck, and how to start governing it inside your organization.
IN THIS ARTICLE
By Juan Antonio Casado · Published July 2026 · Contexta
For two years, the conversation about artificial intelligence revolved around the model: which one was more powerful, which one reasoned better, which one had just shipped. That conversation hasn't gone away, but it's no longer the only one, because models keep converging toward each other and improving in lockstep. What separates a useful output from a generic one is no longer which model you use, but what knowledge you feed it and how you feed it. That discipline now has a name, context engineering, and it's what decides whether an AI agent produces something you can put in front of a client or an elaborate waste of time.
IN ONE SENTENCE
The model is a commodity, context is the asset. What you can't copy is your organization's knowledge, well structured, verified, and current.
What context engineering is
Context engineering is the practice of designing, structuring, and maintaining all the knowledge a language model consumes to produce a response. Not just the question you ask it, but the entire information environment: the documents it reads, the data it retrieves, the memory of what happened before, the rules it follows, and the tools it connects to.
The term took off through 2025, once it became clear that the missing piece in AI systems wasn't reasoning capacity but context quality. Andrej Karpathy helped popularize it, Anthropic published its guide Effective Context Engineering for AI Agents, and by mid-year the first formal taxonomy of the discipline showed up in the academic literature. In a short span it went from niche shorthand to the standard way of naming the real work behind an agent that actually works.
How it differs from prompt engineering
Prompt engineering deals with the instruction, how to phrase the question so the model answers better. That was useful when interacting with AI meant a text box and a reply, but an agent doing real work doesn't live in a text box. It reads files, queries a CRM, remembers decisions from earlier sessions, chains tasks together. The prompt is barely a fraction of what it consumes.
A simple image captures the difference. Prompt engineering fine-tunes the question walking in the door; context engineering designs the entire house, what's in each room, what's in plain sight and what's tucked away, what's current and what's expired, who's allowed to touch what. You can ask the perfect question and still get garbage if the house is a mess. And the reverse holds too: with the house in order, a mediocre question produces good results, because the model has what it needs right in front of it.
Prompt engineering fine-tunes the question walking in the door. Context engineering designs the entire house.
Why context is the bottleneck
There's a simple thesis behind all of this. The model is a commodity, context is the asset. Models improve every quarter and converge, anyone can use the same one their competitor is using. What can't be copied is your organization's knowledge, well structured, verified, and current. Two companies running the same model produce radically different results depending on the quality of their context.
This isn't a hunch anymore, it's starting to be measured. In a Vercel experiment, the same agent, without changing the model or the prompts, went from 53% to 100% effectiveness just by changing what knowledge it consumed and how that knowledge was organized. In another documented case, a chemist with no programming background produced more than a hundred thousand lines of working code relying solely on an agent and well-structured context. Context did the work, not the model.
There's a contradiction here with what instinct usually tells us. When an agent produces a bad result, the natural reaction is to edit the output or tweak the prompt. The right move is different: ask what the context was missing. The quality of what AI produces is a symptom of the quality of the knowledge feeding it. Fix the symptom instead of the source, and the error comes back tomorrow wearing a different shape.
The hidden risk
There's one failure mode that makes bad context especially dangerous: ghost documentation. Documents that look alive but are dead inside. Impeccable form, weak substance, stale data, unverified sources, claims that stopped being true a while ago.
A person with judgment spots a document like that and treats it with caution. An agent doesn't, it treats it as valid, cites it with confidence, summarizes it eloquently, and propagates the error at machine speed. The result is an illusion of productivity, lots of output and little impact, a team that feels like it's moving forward because the agent delivers smoothly, on foundations nobody has verified. Governing context is, first and foremost, an immune system against this kind of document.
The three levels of context
It helps to separate three levels, because each is governed differently and organizations tend to pay attention only to the first.
The first level is immediate context: what you give the model within the interaction itself, the instruction and the data attached to that specific request. This is the territory of classic prompt engineering. It matters, but it's the most superficial layer and the one that creates the least differentiation, because it gets rebuilt in every conversation.
The second is session context, memory. What the agent has learned during the work, the preferences it has picked up, the decisions made earlier that shape what comes next. Without this level, the agent starts from zero every time and re-asks questions that already had answers. With it, it accumulates competence. But uncurated memory turns into noise, so this level requires rules about what gets kept, what gets discarded, and how much space it takes up.
The third, the one organizations work on least and the one that produces the most differentiation, is organizational context: the knowledge of the entire company, its data, its policies, its way of doing things, its judgment. This is where context engineering stops being an individual technique and becomes a governance problem. It's not enough for one employee to know how to give the AI good instructions, the knowledge the whole organization depends on has to be organized, verified, and accessible. This level doesn't get solved with prompts, it gets solved by governing knowledge.
This isn't new, and that's the advantage
The part the tech industry tends to overlook is that these problems have been solved for decades. Organizing knowledge so the right information reaches the right consumer at the right moment is, literally, the definition of library science. There's no need to invent new taxonomies, what's needed is applying the ones already validated over a century.
Three generations of information science thinkers arrived at the same conclusion AI teams are rediscovering today. Ranganathan formulated the laws of library science in 1931 and championed the open shelf, letting the reader browse the sections down to the book instead of asking an intermediary for it. It's exactly the pattern Anthropic landed on when building Claude Code after three iterations, letting the agent navigate knowledge in layers instead of injecting all of it at once. Wurman showed in 1989 that there are only five ways to organize information, and no sixth one. Morville established in 2005 that if information can't be found, for all practical purposes it doesn't exist, and that metadata isn't bureaucracy, it's invisible infrastructure.
The only thing that's changed is the consumer. It used to be a human researcher searching a catalog, someone who tolerated a slightly outdated document because they applied judgment. Now it's an agent that consumes knowledge programmatically, treats it as truth, and propagates the error at scale. The discipline that has governed scientific knowledge for more than twenty years, with its taxonomies, its thesauri, its expiration controls, hasn't become obsolete because of AI. It has become more necessary than ever. Context engineering is that craft's update for the age of agents, not a recent discovery by consultancies that met RAG last year.
What a context engineering framework is
Once the problem moves from immediate context to organizational context, loose best practices aren't enough. What's needed is a decision framework that answers concrete questions. What knowledge enters the system and what stays out. How you tell verified knowledge apart from content generated without review. When a document expires and who decides whether it's still current. What always gets loaded and what only loads when the agent needs it. Who's accountable for each piece staying true.
A context engineering framework organizes these decisions into layers. Governance sets the knowledge quality rules before anything gets connected to any agent. Architecture translates those rules into a structure an agent can consume effectively, because knowledge that's well governed but poorly organized still produces bad results. Connection handles how that knowledge reaches the agent at the right moment, without overloading it. And operation keeps the system alive, because a context system nobody curates degrades on its own: knowledge expires, sources change, and real-world use reveals gaps the design didn't anticipate.
At AlterBiblio we've built this framework under the name CONTEXTA, drawing on two decades of governing scientific knowledge and on recent research into agents. That's not the subject of this article, but if the third level, the organizational one, is where you recognize your problem, that's where the real work begins.
What happens when context fails and when it's governed
The difference shows up best in specifics.
When context isn't governed, an agent pulls up a three-year-old protocol still filed as if it were current, because nobody gave it an expiration date or an owner. It finds three versions of the same figure in different places and cites whichever one shows up first, which turns out to be the wrong one. It retrieves a lengthy report full of polished narrative and numbers that no longer hold. It produces a convincing deliverable, well written and well structured, built on information nobody has verified. The team finds out late, once the error is already circulating.
When context is governed, the opposite happens. Live knowledge is kept separate from the archive, so the agent defaults to consuming only what's verified and current. Every critical figure has a single authoritative source, and anything derived from it links back rather than duplicating it. What gets stored is data anchored to its source, and narrative gets generated on the fly, so there are no mummified reports propagating dead numbers. Every piece has an owner and a date, so it's clear what carries weight and what doesn't. The agent still has access to the historical record if it needs it, but doesn't mistake it for what's current. The result isn't just more reliable, it's auditable, you can trace where every claim came from.
Compression counts too. More context isn't better context. Vercel cut theirs from forty kilobytes down to eight, an eighty percent reduction, without losing effectiveness, because context is a finite resource with diminishing returns, and flooding it with noise makes an agent perform worse than having none at all. Governing isn't accumulating, it's distilling.
Where order begins
Bringing order to your context doesn't require a platform or a big budget to get started. It requires judgment, and it rests on a handful of principles that put everything else in order.
- The first is hygiene. No document should enter live knowledge without three minimum data points: status, date, and owner. Without an owner, a date, or a status, a document exists but can't be maintained, and it ends up as a ghost.
- The second is separating what's live from what's archived. Not all knowledge deserves to be within the agent's reach by default. Working knowledge, verified and current, gets consumed by default, and everything else sits in a zone that's queried on demand but never injects itself. That boundary is what stops the old from passing itself off as current.
- The third is expiration. Live knowledge is born with a shelf life, and when it arrives, someone decides whether it's still valid or lets it degrade. This is what fixes the quietest failure mode of all: a three-year-old protocol looking as current as yesterday's simply because nobody has looked at it again.
- The fourth is storing data, not reports. The asset isn't the twenty-page document, it's the verifiable data point with its source. AI generates excellent narrative on the fly, what it can't generate is the underlying data.
- And the fifth is one authoritative source per topic. The moment there are two versions of the same knowledge, the agent doesn't know which one is good and may cite either.
Stating these principles is easy. Sustaining them in a real organization, with hundreds of documents, systems that don't talk to each other, and several agents writing at once, is something else entirely. That's where a one-line principle turns into a job of taxonomies, migration, and ongoing maintenance, and where having done it before makes the difference. But the starting point doesn't change, and it doesn't depend on which model you use. While the industry argues over which one is best, the real difference is still made by whoever has their knowledge in order.
Context engineering isn't a tech buzzword, it's the recognition that AI's differentiating asset isn't the model but the knowledge it consumes. And that knowledge doesn't govern itself. The good news is that the discipline for doing it has existed for a century, it just needed updating for a new kind of consumer.
If you want to go deeper into structuring your organization's knowledge so your agents produce reliable results, that's the work we do with CONTEXTA, our knowledge governance framework for AI.
Want to govern your organization's context so your agents produce reliable results? Meet CONTEXTA, our knowledge governance framework for AI.
Meet CONTEXTA →KEEP READING
MORE INSIGHTS
INSIGHT · EDITORIAL
WHAT PEER REVIEW IS
The mechanism that holds up the scholarly publishing system.
READ POST