A supplier can appear in the first AI answer and become unsuitable three follow-up questions later. The buyer may add a budget limit, remove a preferred feature, or reveal an integration requirement. The assistant must update its recommendation while preserving requirements that still apply. A single-prompt visibility test cannot show whether that happens. Multi-turn GEO evaluation follows the purchasing decision through a conversation and checks each answer against the buyer's current constraints. This article explains the research behind that concern, a practical way to test it, and how to write pages that help an assistant distinguish an eligible offer from an attractive but unsuitable one.
A purchasing conversation changes the task
Real buyers often begin with an incomplete question. They are trying to learn which details matter.
A facilities manager might ask for a monitoring system, then explain that the facility uses an older gateway, then add that installation cannot interrupt production. Later, procurement asks for a support arrangement. Each turn changes what a defensible answer needs to include.
There are two obligations. The assistant needs to incorporate new information and retain earlier requirements that have not been withdrawn. A recommendation can fail either obligation.
For a brand, this creates a different visibility problem from simple discovery. Being named early is useful only if the description and suggested offer remain appropriate as the conversation becomes more specific. A broad capability claim can survive into later answers even after the buyer has introduced a condition the product does not meet.
The unit of evaluation should therefore be the decision sequence, with the state of the requirements recorded at each turn.
What the research establishes
Laban and colleagues compared fully specified single-turn tasks with simulated conversations in which information arrived over multiple turns. Their paper, *LLMs Get Lost In Multi-Turn Conversation*, reports an average performance drop of 39% across six generation tasks for the tested models. They identify premature assumptions and reliance on earlier attempted solutions as important failure patterns. These were simulated benchmark tasks, not an observation that 39% of commercial recommendations fail. Original paper.
More recent research explores other explanations and remedies. The February 2026 preprint *Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation* proposes separating intent interpretation from task execution. The May 2026 preprint *Found in Conversation* reports improvements from training on selected multi-turn trajectories in its tested model families. These papers suggest that interaction design and training matter; they do not establish that current consumer products have implemented the proposed solutions. Intent Mismatch, Found in Conversation.
The reasonable conclusion for GEO is narrower than “AI cannot handle follow-ups.” Multi-turn reliability deserves its own evaluation. Performance on a fresh, complete prompt does not establish performance when the same facts arrive gradually.
Build a requirement state for each turn
A list of every sentence the buyer has ever written is not enough. Some statements are requirements; others are preferences, tentative ideas, or obsolete information.
Use four categories:
| Category | Example | How it affects the next recommendation |
|---|---|---|
| Active hard requirement | Must integrate with Gateway R | Offers lacking confirmed compatibility cannot be presented as qualified |
| Active preference | Prefer remote installation | Can influence the choice without automatically excluding alternatives |
| Withdrawn condition | On-site service is no longer required | Should stop constraining the answer |
| Unresolved information | Gateway revision is not yet known | May require a question or a conditional answer |
This is an audit model proposed here. It is not a description of the hidden state used by a particular assistant.
Record how each condition became active or inactive. When the buyer says, “Our budget can increase if the integration is included,” the old ceiling has not simply disappeared. It has changed into a conditional allowance. An evaluator who keeps scoring against the original ceiling will mislabel a correct response as a failure.
It also matters who is speaking. A sales representative's promise does not automatically override a buyer's requirement or a technical limitation in official documentation. Preserve the origin of material statements.
Design a conversation that can reveal a mistake
A useful test contains a condition that actually distinguishes the available offers. If every candidate satisfies every requirement, constraint retention will be difficult to assess.
Choose one decision and prepare a reference file with the relevant, publicly supported facts. Then write a sequence in which the buyer gradually supplies the required information. Keep the language natural and avoid steering the assistant toward a named brand.
For a fictional procurement task, the sequence might be:
- “Which vendors should we consider for a monitoring pilot?”
- “The pilot has to work with Gateway R, revision 2.”
- “Production cannot stop during installation. Which options remain?”
- “Remote support is acceptable. We no longer require an on-site technician.”
- “Compare the remaining options and explain what needs confirmation before purchase.”
The third turn adds a hard constraint. The fourth withdraws one service preference. The fifth checks whether the assistant can summarize the current decision rather than replay the first answer.
Do not add arbitrary distractions merely to force a failure. A long fictional conversation about holidays, office politics, and unrelated purchases may test a different problem from the buyer journey you intend to study.
Compare the conversation with a complete prompt
Create a second condition containing the final, fully specified purchasing brief in a new conversation. It should contain the same active requirements as the end of the staged sequence.
If the complete prompt correctly excludes a candidate but the staged conversation retains it, the evidence points toward a conversation-handling issue. If both answers make the same unsupported compatibility claim, the problem may instead involve source availability, retrieval, or interpretation.
That comparison does not isolate the model's internal cause. The two conditions may issue different searches and retrieve different sources. Preserve citations and visible search behavior where available, and avoid attributing every difference to memory failure.
A third, optional condition introduces a factual correction: “The official revision table says this integration is unavailable for revision 2.” The next answer should reconsider the recommendation. It should also verify the correction when verification is possible, rather than accepting any user-supplied assertion as technical truth.
The design tests three related abilities: retaining active requirements, withdrawing obsolete ones, and updating an answer after credible new evidence.
A worked audit of an unsuitable recommendation
Imagine that three fictional offers have the following published properties:
| Offer | Gateway R revision 2 | Installation during operation | Support arrangement |
|---|---|---|---|
| Atlas Standard | Unsupported | Available | Remote |
| Beacon Industrial | Supported | Requires planned downtime | Remote or on-site |
| Cedar Pilot | Supported | Available under stated conditions | Remote |
These values are invented for the example and are not assessments of real suppliers.
After the second conversation turn, Atlas Standard should not be described as meeting the integration requirement. After the third, Beacon Industrial cannot be presented as satisfying the no-downtime requirement without resolving the contradiction. Cedar Pilot may remain a candidate, but its installation conditions still need review.
Suppose the assistant's final answer lists all three as “good fits” and cites the first product overview it found. A brand count would record three mentions. A decision audit would record two contradicted eligibility claims and one conditional candidate.
Suppose instead that the assistant says no option can be confirmed until the installation conditions are checked. That may be a useful result. The evaluator should not penalize every request for clarification as weak brand visibility.
The audit concerns the quality of the purchasing path, including when an answer stops short of a final choice.
Score failures by their consequence
At each turn, review the recommendation against the active requirement state. Record the specific condition, the offer, the source, and the error.
Several measures can help:
- Requirement retention: how many relevant active requirements are explicitly respected in the decision.
- Eligibility violations: recommendations that conflict with an active hard requirement.
- Withdrawal errors: obsolete conditions that continue to exclude an otherwise eligible offer.
- Correction recovery: whether the answer revises the affected claim after new evidence.
- Unsupported completion: a final selection presented despite unresolved decisive information.
An omitted preference and an incompatible integration should not carry the same consequence. Decide severity categories before reviewing the outputs. The categories should reflect the purchasing task, not whichever weighting makes the brand look better.
Keep the denominator visible. “Two eligibility violations in ten reviewed conversations” describes a sample. It is not a claim about all buyers, and repeated turns within one conversation should not be treated as independent people.
An answer can be commercially disappointing and technically correct. If a product does not meet the buyer's hard requirement, exclusion is a success for decision quality.
Give each offer a clear applicability record
The most useful website change may be a concise record of prerequisites and exclusions for each offer.
A product page should connect the offer to a supported version, deployment condition, market, and evidence. An integration page should say whether support is confirmed, conditional, unavailable, or not yet evaluated. Those states should not be blended into “works with leading systems.”
Write exceptions so they travel with the claim. “Supports Gateway R” is unsafe when the supporting documentation limits that statement to revision 3. “Supports Gateway R revision 3; revision 2 is not supported” gives a retrieved passage the distinction it needs.
Treat non-compatibility as useful product knowledge. Removing exclusions to make a page more persuasive may increase unsuitable recommendations and create more work for sales or support.
Cross-link an overview to the compatibility record and the applicable installation guide. The buyer should be able to verify the same conditions that the assistant uses in its reasoning.
How to use the findings in a GEO program
Sort observed failures by the action they imply. A missing specification calls for public documentation. A misleading overview calls for rewriting. A source that clearly states the limitation but is repeatedly omitted calls for a retrieval and source-selection investigation. A correct complete-prompt answer paired with a faulty staged answer calls for continued conversation testing.
The proposed deliverable is a set of reviewed journeys with requirement states, source evidence, and remediation decisions. It complements the market-specific answer review described on Xindar's GEO audit page. The engagement should state whether follow-up conversations are included; a standard single-prompt baseline cannot be assumed to cover them.
After fixing a page, rerun the relevant sequence and the complete-prompt condition. A better first answer is insufficient if the final answer still recommends the wrong package.
For a small team, begin with three purchasing sequences: an added hard requirement, a withdrawn preference, and a credible correction. That set can reveal a content gap that a much larger list of isolated discovery prompts would miss.
Frequently asked questions
Does the 39% research result measure GEO failure?
- It describes the tested generation tasks and model configurations in the paper. It motivates a separate commercial evaluation but cannot supply its result.
Should we force the assistant to summarize requirements after every turn?
That can be a useful intervention test, but it changes the conversational condition. Keep ordinary and summary-assisted sequences separate so you can see what the intervention changes.
Can our content make an assistant remember every condition?
No page can guarantee that. Clear prerequisites and exclusions provide better evidence for a correct decision; conversation handling still depends on the product and circumstances.
Is the final shortlist the only output worth reviewing?
Earlier turns matter because they can introduce unsupported assumptions that later answers preserve. Capture the path and identify where a mistaken claim first appeared.
Research notes and sources
Sources were reviewed on October 9, 2026. The conversation, offer table, scoring categories, and workflow are proposed evaluation methods using synthetic examples. The 2026 papers cited below are preprints; their reported improvements are limited to their tested settings.
- Laban et al.: LLMs Get Lost In Multi-Turn Conversation — simulated single-turn and multi-turn comparison.
- Liu et al.: Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation — an alternative explanation and mediator architecture.
- Chen et al.: Found in Conversation — training-based improvements in selected model families.
