A company builds a chatbot on its own help articles, and the bot still answers some questions wrongly. The natural reaction is to blame the language model and try a bigger one. In most cases that changes little, because the model was not the problem. The search step in front of it was.
How these chatbots work
A chatbot grounded in company knowledge works in two stages. First it searches your documents for passages related to the question. Then it gives those passages to the model and asks it to answer using them. This approach is called retrieval-augmented generation, or RAG.
If the search returns the wrong passages, the model writes a fluent answer from the wrong material. The answer reads well and is incorrect. From the outside this looks like the model inventing things. Usually it faithfully used what it was given.
Find out where the error happens
Take twenty questions the bot got wrong and look at the passages that were retrieved for each. Two outcomes are possible.
- The correct passage was not in the retrieved set. This is a retrieval problem.
- The correct passage was there and the answer was still wrong. This is a generation problem.
In most systems the first is far more common. Knowing which one you have saves weeks of adjusting the wrong part.
Common retrieval problems
Chunks cut in the wrong place. Documents are split into pieces before indexing. If a split falls in the middle of a procedure, the steps and the condition they apply to end up in different pieces, and neither is useful alone. Split along the document's own structure: headings, sections, list items.
Chunks with no context. A passage that says "this applies only to the annual plan" is useless without knowing which product page it came from. Attach the document title and section heading to every chunk.
Search that only matches meaning. Semantic search is good at related ideas and weak at exact terms such as product codes, error numbers and names. Combining it with ordinary keyword search covers both.
Too many or too few passages. Retrieve too few and the answer is missing. Retrieve too many and the relevant passage is buried. A second ranking step that reorders the candidates and keeps the best few usually helps.
Content problems that look like bot problems
Some wrong answers are accurate reports of bad content. Two articles that contradict each other. An old pricing page nobody removed. A policy that changed and was updated in one place only. The bot will repeat whichever it finds. No retrieval technique fixes this. Someone has to own the content and retire what is out of date.
When the model is the problem
If the right passage was retrieved and the answer is still wrong, tighten the instructions: answer only from the provided passages, quote the source, and say clearly when the passages do not contain the answer. A bot that is allowed to say "I do not know, here is how to reach a person" is more useful than one that always produces something.
Summary
Check retrieval before changing the model. Fix how documents are split, add context to each chunk, combine keyword and semantic search, and clean up contradictory content. This diagnosis is part of how we build AI chatbots. If yours is giving answers you cannot trust, we can take a look.