Denis Baciu canonical archive · est. 2026

Writing

Should the agent even look that up?

source: linkedin — original ↗
published: 2026-10-10 · status: canonical · expanded from the original post
Should the agent even look that up?

Most agents treat memory like a reflex. Every turn, they pull whatever looks relevant, stuff it into context, and hope the model uses it. I understand why. Retrieval feels safer than omission. The developer imagines the agent failing because it did not look something up, and that feels worse than failing because the context was too cluttered. But the result is often noise, not signal. When you add memory to an agent, you create a new default: always go look. That default is easy to implement and hard to question.

Most retrieval calls are wasted tokens, not intelligence.

The issue is not just the token count, though that matters at scale. Every chunk you inject competes for attention inside the model. A clean prompt with the query alone can sometimes produce a better answer than a prompt padded with a dozen retrieved passages, most of which are only loosely connected. Retrieval helps, but only when the question actually requires it. The rest of the time, it dilutes the signal and gives the model more ways to go wrong.

Why reflex retrieval is expensive

Retrieval is not free. It adds latency, consumes the context window, and increases the chance that the model will follow a tangent from a retrieved chunk that looks relevant but isn't. In a RAG pipeline, the default is usually to retrieve on every turn, because it is easier to justify adding context than to risk missing something. But that default quietly erodes accuracy. The context window is a scarce resource; spending it on irrelevant chunks leaves less room for reasoning.

The cost shows up in subtle ways. A model trying to reconcile eight retrieved passages spends effort deciding what to ignore. That effort could have gone into reasoning about the actual question. When retrieval is called without a specific reason, it is overhead, not assistance. Over time, this overhead compounds across every turn, every user, every session. Some teams compensate by increasing the context window or reranking, but those only soften the symptom.

Illustration from the original article

A decision model for memory

The alternative is a decision gate. Instead of retrieving by default, run a lightweight model that asks: does this turn need external memory? If yes, retrieve. If no, answer directly. The gate can be a small classifier, a set of rules, or even a heuristic; the point is that it runs before retrieval and has a single job: decide whether external memory is likely to change the answer. Building that gate requires you to know what retrieval is for, which is itself a useful exercise.

In one case, this decision model reduced injected context by 85% and cut five of eight memory calls with no accuracy loss. It is like checking whether a question needs the filing cabinet before opening it. You do not open the drawer for every question; you open it only when the answer is actually inside. Those results came from a model that learned when to look, not from a smaller model asked to do more with less.

A decision gate between query and memory, deciding when retrieval is necessary.
Fig 1. A decision gate between query and memory, deciding when retrieval is necessary.
Illustration from the original article

Instrument what changes the answer

For teams running RAG in production, the real win is not fewer tokens. Fewer tokens are nice, but they are a side effect. The real win is knowing which retrievals change the answer. If a retrieval step can be skipped without changing the output, it was never adding value. That is the signal to optimize.

Instrumenting this is straightforward. Log each retrieval call, what it returned, and whether the final answer used any of it. Over time you see patterns. Some queries always need memory. Others never do. The decision gate learns from those patterns and stops spending effort where it does not matter. Once you have that visibility, the decision gate becomes a natural next step.