How we score an AI answer against a versioned FDA label
The unit of measurement is the omission. What that requires: a frozen prompt set, a versioned label, and a scoring rubric that never grades prose quality.
When a patient asks an AI engine about a prescription medicine, the answer is checked by no one. It is not submitted for review, not versioned, and not retained. The engine may answer the same question differently an hour later. To say anything defensible about these answers you need a measurement procedure: not an opinion about tone, but a count of what the answer contains against what the label requires.
This post describes the procedure behind every Safety Coverage Report: how an answer is captured, scored, and recorded, and why the score is always attached to a specific label version and a specific timestamp.
The denominator is the label
Scoring starts from the FDA label in force on the day of capture, decomposed into required safety elements across four sections: boxed warning, contraindications, warnings and precautions, and adverse reactions. Each section carries a published weight of 3.0, 2.0, 1.5 and 1.0 respectively, so a missing boxed warning costs three times what a missing adverse reaction costs.
For Tarquent, the fictional product used in every VizLoop example, the current label decomposition yields 14 elements. That number is the denominator of every coverage figure on this page, and it changes only when the label does.
Each element is stored with the label section it came from and the label version it belongs to. Labels are versioned by fetch date: when a re-fetch returns a changed label, the decomposition is re-run and the element set is re-versioned, while prior measurements keep their original denominator. A finding never silently changes meaning because the label moved underneath it.
Capture is verbatim or it is nothing
Answers are captured on a schedule, from a clean session, and the full response is retained. Scoring operates on the verbatim text. Excerpts shown in reports are marked as verbatim and truncated only with an explicit count:
Note what the excerpt does not contain. In this snapshot the element coded C2 is a contraindication the label requires, and it does not appear anywhere in the captured answer. That absence is the finding.
Three states, and no credit for prose
Each required element is scored across every answer in the snapshot rather than against one of them. An element is found when every answer surfaced it, partial when some answers surfaced it and others did not, and omitted when none did. An engine's score is the minimum across its answers, because each person sees one answer, not the average of four.
The rubric never grades fluency, empathy, or formatting. A well-written answer that omits a contraindication scores exactly like a clumsy one.
The example engine surfaced 8 of 14 required safety elements for Tarquent in this snapshot. The sentence before this one is the entire editorial position of a coverage report. Everything else is the evidence.
Every finding gets a name and a date
A scored snapshot is written once and never edited: an identifier, a UTC timestamp, the engine, the prompt set version, the label version, and the methodology version. Omitted elements are routed to a named owner. When the next snapshot runs it is a new record, so coverage over time is a series of dated points rather than a mutable score.
The procedure is deliberately boring. It has to be: the answers change, the labels change, and the engines change without notice. The only stable thing in the loop is the measurement.
Email subscriptions are paused. All posts remain available on the blog.