Your AI assistant has read all your documents. Ask it the same question twice and you get two different answers. Ask it about your refund policy and it quotes the one you replaced in March. This is not the model's fault. Retrieval-augmented generation (RAG), the technique behind most company chatbots and assistants, misses information for four predictable reasons: it ranks text by similarity rather than truth, it does not know what is current, it loses facts in the middle of long contexts, and small changes in wording surface different passages. All four are fixable, and the fix is structure before retrieval.
How RAG works, in plain English
RAG is a simple idea. Your documents are cut into passages. Each passage is turned into a set of numbers that places it on a map of meaning, so similar passages sit close together. When someone asks a question, the system finds the passages closest to it on that map, pastes them into the prompt, and asks the model to answer from them.
It is one of the layers I describe in you're not talking to an LLM, you're talking to a system. It works remarkably well in a demo, where the documents are few, clean and current. It works much less well when you point it at seven years of shared drives, wikis, email exports and PDFs, which is exactly what most companies do next.
The model at the end of the pipeline is rarely the weak link. It answers from what it is handed. The problem is what it is handed. Three of the four failure modes below happen before the model sees a word. The fourth is about how much we hand it. All four are fixed in the knowledge and the context, not the model.
Reason one: similar is not the same as relevant
Retrieval ranks passages by how close their meaning is to the question. That sounds right until you notice what else is close in meaning.
Last year's pricing sheet and this year's are nearly identical in wording. So are the draft contract and the signed one, the old onboarding process and the new one, the policy and the email arguing about the policy. To a similarity search they are all excellent matches. Nothing in the data tells it that one of them is the truth and the rest are history.
So the model receives a confident mix of current and obsolete material and does its best. Sometimes its best is right. Often it blends the two, which is worse than being plainly wrong, because it sounds authoritative and is hard to spot.
Reason two: nothing tells it what is current
This is one of the most common failure patterns reported in production RAG, and it has a distinctive signature: answers that were true for the company as it was a year or so ago. One engineer's post-mortem, published in July 2026, described exactly this, answers "right for the state of the company twelve to eighteen months ago but wrong for the current state", and found that three weeks of manual triage separating current sources from superseded ones cut wrong answers by around 60% (engineer post-mortem, dev.to, 9 July 2026).
That result is telling. Nobody changed the model. Nobody changed the prompt. They changed what the system knew about which documents were still true.
Most document stores simply do not record this. A file is a file. Unless currency is captured, as a status, a supersedes link or an effective date, and used when retrieving, stale knowledge competes with current knowledge on equal terms.
Reason three: attention has limits
The instinctive response to missed facts is to retrieve more passages, or to use a model with a bigger context window and put everything in. It rarely helps.
Research on long contexts has shown, repeatedly, that models recall information at the start and end of their input better than information in the middle (Liu et al., "Lost in the Middle", 2023). The effect has shrunk in newer models but not disappeared, and recall still degrades as context grows, even on simple tasks (Chroma, "Context Rot", July 2025). Anthropic's own guidance on building agents recommends the opposite of stuffing (Anthropic, "Effective context engineering for AI agents", 2025): the smallest set of high-signal context that does the job.
People work the same way. Hand a colleague two hundred pages and ask one question, and they will miss the answer on page ninety-three. More pages do not make them better informed. They make the relevant page harder to find. More context is more dilution, and more cost on every single call.
Reason four: the right passage only sometimes makes the cut
Retrieval is sensitive to wording. Part of the variation you see is by design, because most assistants deliberately vary how they phrase an answer. But when the facts themselves change between runs, the cause is usually retrieval. Phrase a question slightly differently, and a different handful of passages makes the cut. When the right passage sits just around the cut-off, it appears in some answers and not others. That is why the same question can get two answers.
Agentic systems try to compensate by searching again, rephrasing, and searching again. That can help, but it multiplies token cost, and retrying a search over unstructured text does not reliably converge on the truth when nothing marks what is current. It searches the same haystack more times.
This is also why these problems stay hidden. Teams watch their systems far more than they measure them. LangChain's State of Agent Engineering survey found that 89% of teams monitor their agents, but only 37% evaluate them on live production traffic (LangChain, November to December 2025). A dashboard showing thousands of happy conversations tells you nothing about how many answers were wrong.
The fix is structure before retrieval
Every failure above is a knowledge problem, so the fix lives in the knowledge, not the model.
- Resolve the entities. Customers, suppliers, projects, products and documents become consistent records, so "Smith Mfg" and "Smith Manufacturing Ltd" are one company. I explain how in what is an entity registry.
- Record what is current. Status, effective dates and "supersedes" links, used as filters when retrieving.
- Map the relationships. Which contract belongs to which project and which supplier, so an agent can follow a path to the answer instead of hoping similarity finds it. For relationship-heavy questions this is where a knowledge graph beats plain vector search, as I compare in GraphRAG vs vector RAG.
- Keep context small and relevant. Which is also why grounded agents are cheaper to run.
When we built the system that resolved 67% of customer service cases without a human at a European insurance brokerage, retrieval was engineered as a first-class problem, on a data layer built before any models. That order is not a detail. It is the whole method, and it is the core of the argument in your AI does not need a bigger model, it needs to know your business.
How to test your own system this week
You can find out how much of this applies to you in a few days, without new tools.
- Collect 50 real questions your people actually ask, with answers someone trusted has confirmed.
- Ask each one three times, the way a colleague would phrase it.
- Score three things: was the answer correct, did the three answers agree, and did any answer rely on a superseded document?
If correctness is high and consistent, you have a good system and a useful baseline. If not, you now know which of the four reasons you are dealing with, and that tells you where to spend. Either way you will know more than many teams do, because many teams never run this test on the questions that matter. It is also the first thing I run in a Knowledge Audit.
If your assistant keeps answering from what your company knew last year, I would be glad to look at it with you. Let's talk.
Related: Your AI does not need a bigger model. It needs to know your business · GraphRAG vs vector RAG. Which one does your business need? · You're not talking to an LLM. You're talking to a system