The distinction that actually matters.
Retrieval-augmented generation answers questions from your content. An agent takes actions in your systems. Everything else people argue about downstream of that — frameworks, orchestration, memory — follows from which of those two you are building. The reason the distinction gets blurred is that agents usually contain retrieval: an agent that cannot look anything up is working blind. So the real question is not RAG or agents, it is whether answering is sufficient or whether something must also be done.
When retrieval is enough.
Retrieval is the right answer when a human acts on the output. Supporting a support team rather than replacing it, helping someone find the clause in the contract, summarising what is known about an account before a call. In these cases the system's failure mode is bounded: it gives a poor answer, a person notices, and the cost is a few wasted minutes. That bound is what makes retrieval cheap to deploy relative to anything agentic — you need good retrieval, citations and an evaluation set, but you do not need approval gates, audit trails, rollback paths or compensating actions.
What actually makes retrieval work.
Most disappointing RAG systems fail at retrieval rather than generation. If the right passage never reaches the model, no amount of prompt engineering recovers it. The work concentrates in chunking strategy, embedding choice, reranking and evaluation — unglamorous decisions that determine the ceiling on everything built above them. There is also a failure that is not technical at all: retrieval quality is bounded by content quality, and these projects routinely surface that the source material is contradictory, out of date or scattered across four systems. That discovery is uncomfortable and genuinely valuable, because it would have undermined any system built on top.
When you actually need an agent.
You need an agent when the work is not finished until something changes in a system — a ticket updated, a record created, a message sent, a job scheduled. The tell is simple: if a person has to read the output and then go and do something mechanical with it, and that mechanical step is repetitive and well understood, an agent can take it. If the step requires judgement about consequences, it probably should not, or at least not without a human confirming each action.
The cost of moving up.
Going from retrieval to action is a larger jump than it looks, and almost none of the extra cost is model work. It is authentication against systems never designed for a non-human caller. It is failure handling, because a half-completed multi-step action leaves your systems in a state nobody designed. It is observability, since you cannot debug an agent whose decisions you cannot replay. And it is approval gating in the execution layer rather than the prompt, because a system-prompt instruction is a preference and will hold until the one context where it does not. Budget for those four and the estimate stops looking optimistic.
The hybrid most teams should build.
In practice the useful system is usually retrieval plus a small, enumerated set of actions, with a human confirming anything consequential. The agent looks things up, proposes what to do, and executes only from a fixed list. That shape keeps the safety property that makes retrieval attractive — a bounded failure mode — while removing the mechanical work that made people want an agent in the first place. It is also the version that earns the right to more autonomy later, because it accumulates a track record you can point at.
How to decide.
Ask what happens after the system produces its output. If a person reads it and thinks, build retrieval. If a person reads it and then performs the same mechanical steps every time, build a bounded agent. If nobody reads it at all and the system must be trusted to act alone, you are in a different cost bracket entirely, and it is worth being certain the value justifies it. Most organisations should start at the first and move up only when the limitation is genuinely the bottleneck — not because the second sounds more ambitious.
