Prompt engineering did not die. It got demoted to a component of something larger. Context engineering is the practice of deciding what an agent knows at the moment it acts: which business definitions, which history, which tool outputs, and which past results land in front of the model before it answers. The distinction matters commercially because the two disciplines have different owners. Wording a prompt is a task for whoever is using the tool. Deciding what is true about your funnel, your ICP, and your channel economics, and keeping that current where an agent can read it, is a marketing leadership job that nobody is currently assigned.

This expands on a post I published on LinkedIn about how prompt engineering actually evolved.

Key Takeaways

  • Context engineering is not a rebrand of prompt engineering. Anthropic defines it as curating and maintaining the optimal set of tokens during inference, which includes everything that reaches the model outside the prompt itself.
  • The value moved in five steps: ask and answer, saved skills, goals instead of steps, context pulled on every call, and now the curation of that context as the actual work.
  • More context is not better context. Chroma's Context Rot report found that across 18 models, performance grows increasingly unreliable as input length grows, even on trivial tasks.
  • Position matters as much as inclusion. The Lost in the Middle paper found a U-shaped curve, where models use information best at the start or end of a context and worst in the middle.
  • For a marketing org, the context that decides output quality is four things: funnel definitions, ICP, channel economics, and the record of what has already been tried and failed.
  • Ownership splits cleanly. Deciding what is true stays with the CMO. Encoding it and keeping it current is a marketing ops job. Neither is a prompt problem.
  • If that context is not in front of the agent, no prompt rescues the output. If it is, the wording barely matters.

How Prompting Actually Evolved: Five Stages

"Prompt engineering is dead" trends every few months, and it is wrong in a specific way. The skill did not disappear. It kept getting absorbed into the layer beneath it, and each absorption moved value further from the words you type.

Ask and answer. One question in, one answer out. The entire quality of the output depended on how you phrased the request, which is why prompt phrasing briefly looked like a career.

Saved skills. The prompts that worked got kept and reused. Phrasing stopped being an act of authorship and became a library. The moment a good prompt is written down and shared, writing it is no longer the scarce skill.

Goals instead of steps. Rather than specifying a procedure, you specify an outcome and the agent works out its own steps. This is the point where prompt wording stops being the main lever, because you are no longer describing the work.

Context pulled on every call. The agent stopped waiting to be told things. It began retrieving from memory, tools, and past results at the moment of each decision. What it fetched mattered more than what you typed.

Curation of that context. Once retrieval is automatic, the question becomes what is available to be retrieved and what gets selected. That is context engineering, and it is where the work sits now.

Each step moved value away from the words you type and toward what the agent already knows about you. The term for this settled during 2025, and Anthropic's engineering write-up from September 2025 gives the cleanest working definition: prompt engineering is "methods for writing and organizing LLM instructions for optimal outcomes," while context engineering is "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference." The second one contains the first.

Why More Context Is Not Better Context

The obvious reaction to all of this is to give the agent everything. Larger context windows arrived, so pour the whole data warehouse in and let the model sort it out. This is the part most teams get wrong, and it is measurable rather than a matter of taste.

Chroma's Context Rot report, published in July 2025 by Kelly Hong, Anton Troynikov, and Jeff Huber, evaluated 18 models across the Claude, GPT, Gemini, and Qwen families. The finding was consistent: models "do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." That held even on tasks as simple as repeating words back, where accuracy degraded as the input grew from 25 to 10,000 words. Capability at the 100th token does not predict capability at the 10,000th.

Position compounds the problem. The Lost in the Middle paper by Liu and colleagues, published in TACL in 2024, tested how models use information depending on where it sits in the context. They found performance "is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models." A U-shaped curve, not a flat one.

Put those together and the practical rule is uncomfortable for anyone hoping to solve this by dumping data: an agent handed your entire marketing stack performs worse than an agent handed the correct slice of it. Curation is not a cost-saving measure applied to context. Curation is the mechanism by which context works at all. This is also why agent memory is not the same thing as search over your documents: retrieving ten plausible documents and retrieving the two authoritative ones are different operations with different outcomes.

What Your Agent Actually Sees About Your Business

Here is the question I would ask any marketing leader running agents today. Not "what prompt do I write," but "what does my agent actually see about my business on every single call?"

In practice, for a marketing organization, four things decide the answer.

Funnel logic. What counts as a lead, an MQL, an opportunity, and at which point. If two agents hold two different definitions of a qualified lead, they will produce two defensible and incompatible reports, and someone will spend a week reconciling them. Definitions are context, and undocumented definitions are missing context.

ICP. Who you are actually selling to this quarter, not who you were selling to when the deck was last updated. An agent that believes last quarter's ICP will quietly and competently ship this quarter's campaigns to the wrong people. It will not error. It will just be wrong at scale, on schedule.

Channel economics. What a conversion costs where, what payback looks like by channel, and which constraints are real. Without this an agent optimizes toward whatever metric is nearest, which is usually volume, and volume is the easiest thing to buy and the least useful thing to have.

The record of what already failed. This is the one that is almost never wired in, and it is the one that pays the most. Every marketing team has a list of things that were tried and did not work. If that list lives in a deck, a Slack thread, or someone's memory, your agent will propose the same failed test again, and it will propose it confidently, because nothing in its context says otherwise.

None of those four are model problems. They are all questions about whether your organization has written down what it knows in a place an agent can reach. Most have not, which is a large part of why AI adoption in marketing keeps producing activity without profit impact.

Talk to an Improvado expert about giving your agents governed marketing context they can actually read.

Context Engineering Is an Operating Discipline, Not a Prompt Library

The reason this gets stuck is that it does not have an owner, and it looks technical enough that marketing leaders hand it to engineering, where it promptly dies. Engineering can build the retrieval. Engineering cannot decide what your ICP is.

The split that works is straightforward.

Deciding what is true stays with the CMO. Which definitions are authoritative, which segments are in play this quarter, which economics are the ones we manage to. This is a judgment call about the business and it does not delegate. Handing it to an agent is not automation, it is abdication.

Encoding it and keeping it current is a marketing ops job. Getting those decisions written into a place agents read, versioned when they change, and retired when they stop being true. This is unglamorous and it is the entire difference between a working agent and a confident one.

Somebody signs their name under what ran overnight. Responsibility does not compress just because the execution did. If an agent spends budget while everyone sleeps, a named human owns that outcome in the morning. The number of agents in the org can go up. The number of accountable humans cannot go to zero.

You likely already have the people for this. What is usually missing is the explicit assignment and a durable place to put the answers. When that place does not exist, each agent keeps its own private copy of the strategy, which is how you end up paying the re-briefing cost over and over and calling it a model limitation.

How to Audit What Your Agents See

This is testable in an afternoon, and the results are usually unflattering in a useful way.

Ask the definition question cold. Open every agent or assistant your team uses and ask each one what counts as a qualified lead. Do not prompt it toward the right answer. If you get three different answers, you do not have a prompting problem, you have three systems operating on three versions of your business.

Ask what it thinks your ICP is. Then check when that was last true. The gap between those two dates is your context debt, measured in quarters.

Propose something you already know failed. Ask for a recommendation in an area where you ran a test that did not work. If the agent enthusiastically proposes the failed approach, your institutional learning is not in its context, which means every agent you add will rediscover the same dead ends at full speed.

Trace one number back. Take a single figure from an agent-generated report and ask where it came from. If nobody can trace it to a source and a definition, the output is not analysis, it is a plausible sentence.

Check who can change the answer. If updating the definition of an MQL requires an engineering ticket, the context layer will go stale, because the people who know when it changed are not the people who can edit it.

The fix is rarely a better prompt or a bigger model. It is putting funnel definitions, ICP, channel economics, and prior results into a governed layer that agents read on every call and that marketing can maintain without a deploy. That is what we build at Improvado: unified data across your channels, your taxonomy and metric definitions applied consistently, and an agentic layer that reads governed data your team owns rather than guessing from whatever happened to be pasted into a prompt.

Prompt writing was the interface. Context is the product.

Talk to an Improvado expert about auditing what your marketing agents currently see.

Frequently Asked Questions

What is context engineering?

It is the practice of curating what an agent knows at the moment it acts. Anthropic's September 2025 engineering write-up defines it as the set of strategies for curating and maintaining the optimal set of tokens during LLM inference, which covers everything reaching the model beyond the prompt: retrieved documents, tool outputs, conversation history, business definitions, and prior results. In a marketing organization the practical version is narrower and more concrete: making sure funnel definitions, ICP, channel economics, and the record of what has already been tried are available and current wherever your agents read from.

Is context engineering just prompt engineering with a new name?

No, and the difference is about scope and ownership rather than fashion. Prompt engineering is writing and organizing the instructions you send. Context engineering governs everything else the model sees, which in a production agent is the overwhelming majority of what lands in the window. Prompt engineering is a task for whoever is operating the tool. Context engineering is an operating discipline that requires someone to decide what is authoritative about the business and someone else to keep that encoded and current.

Is prompt engineering dead?

It is not dead, it is subordinate. Wording still matters at the margins, and it matters most when the context is thin, which is the trap: teams with weak context feel large gains from prompt tinkering and conclude that prompting is the lever. Once the agent has the right context, phrasing changes produce small effects, because the answer is being determined by what the model knows rather than how it was asked. If a rewrite of your prompt substantially changes the answer to a business question, treat that as a signal that the context layer is not doing its job.

Should I just give the agent more context since windows are large now?

No. Chroma's Context Rot report found that across 18 models spanning the Claude, GPT, Gemini, and Qwen families, performance grows increasingly unreliable as input length grows, including on trivial tasks. The Lost in the Middle paper found a U-shaped curve, where models use information best at the beginning or end of a context and worst in the middle. Both point the same way: a large window is capacity, not comprehension. An agent given the correct slice of your data outperforms one given all of it.

Who should own context engineering in a marketing org?

Split it. Deciding what is true, which definitions are authoritative and which segments are in play, stays with the CMO or the marketing leader, because it is a judgment call about the business rather than a technical task. Encoding those decisions, versioning them when they change, and retiring them when they stop being true belongs to marketing ops. Engineering builds the retrieval but should not be deciding what your ICP is. Separately, a named human signs off on what agents did overnight, because responsibility does not compress just because execution did.

How do I tell whether my agents have the right context?

Ask each agent your team uses what counts as a qualified lead, without steering it. Different answers across tools means different versions of your business are running in parallel. Then ask what it believes your ICP is and check when that was last accurate. Then ask it to recommend something in an area where you already ran a test that failed; if it proposes the failed approach confidently, your institutional learning is not in its context. Finally, take one number from an agent-generated report and try to trace it to a source and a definition. If you cannot, the output is a plausible sentence rather than analysis.

Where do we start if none of this is written down?

Start with the definitions that cause the most rework, which for most teams is the funnel: what a lead is, what a qualified lead is, and where the handoff sits. Write those once, put them somewhere agents read on every call rather than in a deck, and give marketing ops the ability to change them without a deploy. Add ICP next, then channel economics, then the record of prior tests. The order matters less than the property that makes it work: one place, current, editable by the people who know when it changed, and read automatically rather than pasted in by whoever remembered.