AI, usefully · 4 min read ·

Your AI found a policy. Was it the right one?

You ask an assistant whether a customer can return a sale item. It answers clearly, quotes a policy and includes a link.

Then you open the link. The policy is from last year.

The assistant found something relevant. It just wasn’t the rule that applied.

This is an easy mistake to overlook because we tend to focus on the answer: whether it sounds reasonable, whether it has citations, whether it follows our instructions. But an assistant answering from documents has another job to get right first. It has to find the appropriate evidence.

That process is called retrieval. It becomes especially useful when the information lives across dozens of documents, rather than in a short passage you can paste into a chat.

What happens between the question and the answer?

Imagine a small retailer has a folder containing return policies, staff guidance and old announcements.

A common AI setup searches that collection, selects relevant passages and supplies them to a language model alongside the question. The model uses those passages to compose an answer. This pattern is often called retrieval-augmented generation, or RAG: search provides additional material for the response. Microsoft’s overview explains the approach.

It does not require retraining the model on every policy update. Instead, the searchable collection can be updated.

That is useful, but it leaves two separate questions to investigate when an answer fails: did the system retrieve the right material, and did the model interpret that material correctly?

Changing the wording of the final prompt may help with interpretation. It will not restore a missing policy that never reached the model.

“Relevant” has several meanings

A search for “Can I return something bought on sale?” might find a passage containing those exact words.

It might also find a policy headed “Refund eligibility for discounted merchandise.” A search based on meaning can help connect those different phrases.

But relevance also depends on details that have little to do with wording. Which country was the purchase made in? Was it online or in a shop? When was it purchased? Is the document approved guidance or a draft?

A passage can closely match the question while applying to the wrong situation.

For a reliable setup, I would want each policy to carry a few labels: region, sales channel, effective date, version and whether it is current. Search can then narrow the eligible documents before choosing passages. Microsoft documents this use of filters alongside search, including filters based on document metadata.

Those labels are ordinary information management. They matter just as much when the reader happens to be an AI system.

Keep the exception beside the rule

Long documents are often divided into smaller passages, called chunks, so a search can return the relevant section.

Suppose one chunk says:

Returns are accepted within 30 days.

The next paragraph says:

Final-sale items are excluded.

If the system retrieves only the first passage, the model may give a perfectly reasonable answer based on incomplete evidence.

Document splitting needs to preserve meaning: headings, exceptions, table labels and references that help a passage stand on its own. Anthropic’s work on contextual retrieval examines this problem and describes adding explanatory context to passages before indexing them.

For the retailer, the practical question is simple: does the material returned for a sale-item question include both the general return rule and its exclusions?

Ask to see the material before the conclusion

You can try this with a document assistant that supports searching uploaded files:

Find the passages needed to answer this question. Before giving the answer, show the document title, effective date, relevant rule and any exception.

Flag conflicting versions or missing context. If the documents do not establish which policy applies, identify what needs clarification.

Question: Can this customer return a discounted item purchased online last week?

Inspect what it retrieves. Did it find the current online policy? Does “discounted” mean the same thing as “final sale” in these documents? Is the applicable region known?

Then ask it to answer from the checked passages.

This prompt makes retrieval easier to review. It does not create version controls or access restrictions by itself. Those need to be supported by the application and the document collection.

A small test worth doing

Make three short fictional documents: an old return policy, a current policy and an exception for final-sale items. Give them clear titles and dates.

Ask your document assistant three questions: a straightforward return, a final-sale return, and a purchase where the region is unspecified.

Check whether it finds the right passages and asks for clarification when necessary. Include a question the documents cannot answer. A system that always finds something may still need a better way to recognize insufficient evidence.

Keep two notes about each failure: what was retrieved and what was concluded. That distinction tells you where to start fixing it.

A cited answer deserves a closer look than “there is a link.” Open the source, check whether it applies, and look for the exception the answer might have left behind.