← All posts
Engine changes

The answers changed in June and nobody reviewed them. Also: the incident that wasn't.

Two retrieval changes shipped, one enterprise and one consumer. Neither is health-specific, both move which sources get named in health answers, and no manufacturer saw either coming. Plus: why we found no verified June incident, and why that is not reassurance.

VizLoop Measurement notes · 2026-09-02

Every material AI engine change passes through someone's review process before it ships: the vendor's. What never happens is review by the people whose regulated products the changed answers describe. June 2026 offered two clean examples of how answers about drugs and devices shift as a side effect of product decisions made for entirely other reasons, and one instructive absence: a full month in which no press-covered incident of a wrong AI answer about a named drug or device could be verified at all.

Deep citations: provenance gets granular

Microsoft's June update to 365 Copilot introduced "deep citations," which link an answer to the specific portion of a referenced file rather than to the file as a whole, initially in Word and PowerPoint, with meetings, web, and PDF sources to follow. Microsoft's framing is trust: let the user trace a claim to its exact source.

This is a first-party announcement and the feature is not health-specific, but it matters to anyone whose internal documents feed enterprise AI. Granular provenance is genuinely better than file-level provenance; it is also a new surface for a familiar failure. A citation that points to one paragraph of a medical information letter or a formulary document confers precision on whatever the model extracted from that paragraph, including an extraction that dropped a qualifier. The AJHP study covered in our companion piece found that the dominant failure mode in drug-information answers is not fabrication but incompleteness. A deep citation under an incomplete answer does not complete the answer. It makes the incompleteness look verified.

The pattern to expect, on no timeline anyone outside Redmond controls: citation display that started in enterprise documents extends toward web and PDF sources, and the convention spreads, because provenance features are competitive table stakes now. Which sources get linked, at what granularity, is becoming a design choice made engine by engine, quarter by quarter.

Perplexity goes deeper, which means different sources win

The second change is consumer-side. On June 19, Perplexity reportedly brought a more powerful Deep Research mode into its agentic assistant, fanning out across many more searches and reading more material before producing a sourced report. One sourcing note we owe the reader: this item currently traces to a changelog aggregator rather than a first-party Perplexity announcement, so treat the details as provisional.

The reason it earns a place in a health roundup despite that caveat is structural. Perplexity is a citation-first engine; naming sources is the product. When such an engine changes how deep it retrieves, it changes which pages get read, and therefore which pages get named, and therefore whose account of a drug reaches the person asking. A patient's question about a therapy might have surfaced the label and a health-system page in May, and in July surface a deeper pull that includes an advocacy site, a decade-old review article, or a payer document. No one at the manufacturer approved the May mix, and no one was notified of the July one. Retrieval depth is a dial that reassigns authority over your product's public description, and it turned this month.

A June 17 Pharmacy Times commentary supplied the right vocabulary for this: as agentic AI enters clinical workflows, the binding constraint is not model capability but whether systems can reach authoritative, real-time medication data. Grounding infrastructure, not model size, determines whether the drug facts in an answer are current and complete. Both of June's changes are grounding changes. That is exactly why they matter and exactly why they arrived without ceremony.

The incident that wasn't

We looked, specifically and repeatedly, for a June 2026 press-covered incident of an AI engine giving a wrong or omitted answer about a specific named prescription drug or medical device. We found none. The most-cited real case in this genre, The Guardian's investigation of Google AI Overviews, dates to January 2, 2026 and concerns dietary advice for pancreatic cancer patients and liver-test reference ranges: serious, clinical, and not product-specific.

It would be convenient for a monitoring company to imply the absence means we simply haven't caught the engines yet. The honest reading is more specific and more useful. Peer-reviewed measurement says incomplete drug-information answers are the majority outcome, so the raw material for incidents exists in volume. What is missing is documentation: a named product, a captured answer, a date, and someone who compared the answer to the label and published the gap. Journalists sample this surface occasionally. Regulators do not monitor it at all; no June FDA promotional letter cites AI-generated content, and the agency's AI guidance for drug development remained draft through the month. Manufacturers mostly learn what the engines say about their products when someone shows them.

An unwatched surface does not generate incident reports. It generates surprises, later, with worse timing.

What a review process would even mean here

No promotional review process can approve an AI engine's answers, and nothing should claim otherwise: the engines are not yours to submit, and the verbs available to a manufacturer here are not approve and correct. They are monitor, document, and escalate.

What June's changes clarify is what such monitoring has to look like to be worth anything. It has to be dated, because answers drift when retrieval changes, and a finding without a timestamp cannot be compared to anything. It has to capture answers verbatim, because a paraphrase of an omission is not evidence. It has to be scored against the current label, because the label is the one document that defines what a complete answer contains. And it has to re-measure on a schedule, because both of this month's changes, a citation format and a retrieval depth, altered the answer surface without touching any health policy, and next month's equivalents will too.

The month's summary is short. The engines got better at naming sources and better at reading more of them. Nobody, anywhere, checked what that did to answers about any particular drug. The gap between those two sentences is the entire case for treating AI answers as a monitored surface rather than an occasional curiosity, and it is where we work.

This is one of three deep dives on the month. The regulatory chronology is in Three dates moved in June; the full June record is in the June 2026 roundup.

provenance retrieval monitoring AI engines
Read the latest posts

Email subscriptions are paused. All posts remain available on the blog.

Related posts