AI, usefully · 4 min read ·

AI can have the answer in front of it—and still miss it

One of the more interesting findings in AI research came from changing something surprisingly ordinary: the order of the reading material.

In a study published in 2024, researchers gave language models questions and documents containing the answers. They then changed where the relevant information appeared.

Performance was often strongest when the answer was near the beginning or end. When it sat in the middle, performance could drop significantly—even for models designed to handle long inputs.

The paper is called Lost in the Middle. Its finding makes an uncomfortable distinction: information can be available to a model without being used reliably. These were results for the models and tasks tested, not a permanent rule for every AI system. But they give us a useful question to ask of any assistant: does it still find the evidence when we stop making that evidence easy to find?

That matters when we upload a long report, add another attachment or continue a conversation for days. It is tempting to assume that supplying more material makes the answer better informed. We also need to check whether the important parts survive the extra material.

A context window is the amount of information a model can work with in a request. Capacity is useful, but it does not measure how consistently the model uses every part of that information.

Think about the difference between two questions you might ask of a lengthy report:

“Find the date this programme started.”

“Explain whether the programme met its goals, accounting for the revised targets and exceptions.”

The first requires locating a fact. The second requires finding several pieces, recognising their relationships and applying the right qualifications. Success on the first question tells us little about the second.

This is why a demonstration that retrieves one hidden sentence should leave you curious about what happens when the answer requires several passages.

There is a practical consequence for agent design here, too.

An agent accumulates material as it works: search results, tool responses, failed attempts, instructions and intermediate conclusions. Deciding what to carry into its next step is part of the engineering.

Anthropic’s context-engineering guidance discusses this problem and techniques such as compacting conversation history and keeping structured notes. The aim is to preserve useful information as the work grows.

But compression creates a question of its own: what did the summary leave out?

A qualification that looks like a minor detail can change a recommendation. If we shorten the material, we need a way to return to the original evidence. I would want a working note to retain the relevant source references, unresolved questions and exceptions alongside the current conclusion.

For everyday use, you can test this without building an agent.

Choose a public report you have actually read. Write three questions whose answers you can verify: one about a straightforward fact, one about an exception and one that requires combining two sections.

Record the answers and their supporting passages before asking AI.

Then prepare three versions of the same material. Keep the wording and section labels intact, but move the relevant sections towards the beginning, middle and end. Use a fresh conversation for each version and keep the question unchanged.

Try this instruction:

Answer using only the supplied material. For each conclusion, quote the supporting passage and identify its section. Include any qualification that changes the answer. If the evidence is incomplete or contradictory, say what remains unresolved.

Check more than whether the final answer sounds right. Did it find the correct passage? Did it preserve the exception? Does the quotation actually support the conclusion?

Repeat the exercise a few times. A single response can vary for reasons beyond document order. This is a small diagnostic exercise, not a benchmark of the model’s overall ability.

If the answer changes when the sections move, you have found a weakness worth investigating. For that task, you might first ask the assistant to locate the relevant passages, check those passages yourself, and then request an explanation using that smaller evidence set. Keep the surrounding qualifications; removing context carelessly can create a different error.

The insight I find useful is that “I uploaded it” is only the beginning of the check. We should be able to see which parts of the material shaped the answer—and whether the answer holds up when the evidence is less conveniently placed.