
RAG in Production: What Actually Breaks
Retrieval-augmented generation (RAG) is the architecture most enterprises reach for when they want to put LLMs to work on their own data. It's a sound instinct: rather than relying on a model's parametric memory, you ground its answers in your documents at inference time.
In a demo, RAG feels like magic. You point it at a knowledge base, ask a question, and get a fluent, accurate answer. Then you ship it, and things start breaking in ways the demo never showed.
Here's a field guide to what actually breaks—and what to do about it.
1. Retrieval quality is the whole game
The model can only answer from what it retrieves. If the right chunk isn't in the top-k results, the answer is wrong, no matter how capable the LLM. Most teams spend 90% of their effort on the generation side and 10% on retrieval. In production, that ratio should be inverted.
The failure modes are subtle:
- Chunking boundaries cut critical context in half. A sentence that spans two chunks loses its meaning.
- Embedding models drift in quality across domains. A general-purpose embedding model may perform poorly on medical or legal text.
- Top-k is too small or too large. Too few results miss the answer; too many dilute it with noise.
Invest in your chunking strategy. Evaluate embedding models on your actual data, not a benchmark. And measure retrieval recall before you ever look at generation quality.
2. The model hallucinates confidently from partial context
This is the failure mode that erodes user trust fastest. The model retrieves a relevant-but-incomplete chunk and fills in the gaps with plausible-sounding fabrication. The answer looks right, and it isn't.
Mitigations that actually work in production:
- Always return the source. Every claim should link to the chunk it came from. Users—and auditors—need to verify.
- Add an explicit "I don't know" path. If retrieval confidence is low or the context doesn't support the question, the system should say so rather than guessing.
- Use structured prompts. Constrain the model to answer only from the provided context, and penalize external knowledge. This isn't perfect, but it raises the floor.
3. Your knowledge base is a moving target
Documents change. Policies get updated. Old versions linger in the index. A RAG system that was accurate in October can be confidently wrong in December because it's retrieving a stale chunk.
Production RAG needs a data pipeline, not a one-time ingestion. That means:
- Versioned ingestion with timestamps on every chunk.
- A strategy for handling document updates—do you re-embed the whole document, or patch the changed chunks?
- Monitoring on retrieval freshness, not just retrieval relevance.
4. Latency compounds in ways you don't expect
In a demo, a 4-second response feels thoughtful. In production, with a user waiting, it feels broken. And RAG has more latency-sensitive stages than most teams account for: embedding the query, vector search, context assembly, and generation all happen sequentially.
Each stage is individually fast. Together, under load, they compound. Profile the full path early. Cache embeddings for repeated queries. Consider streaming the generation so the user sees the first token in under a second. And set a latency budget before you launch, not after.
5. You have no idea what users are actually asking
The questions users ask in production are not the questions you tested with. They're messier, more ambiguous, and often outside the scope of your knowledge base entirely. A RAG system with no visibility into real queries is a system you can't improve.
Log every query, every retrieval, every response. Build a review loop. The single highest-ROI activity after launch is sitting down once a week, reading a sample of real queries, and asking: what did we get wrong, and why?
The takeaway
RAG is a production-grade architecture. But the gap between a working demo and a reliable system is almost entirely in the parts that aren't glamorous: retrieval quality, data pipelines, latency budgets, and observability. The teams that win with RAG are the ones who treat it as an engineering problem, not a model problem.
Spend your time on the boring parts. That's where production lives.