Visibility + Label Substantiation · Methodology

Evidence you can interrogate.

Visibility records the question, answer, available citations, capture date, engine, model, and retrieval mode. Label Substantiation compares the answer with an archived FDA label and records the judgment method and supporting evidence. Model judgments are inputs to human review, not compliance determinations.

Fictional sample scoring
The sample report’s Safety Coverage Score

The fictional sample uses the deterministic formula below to demonstrate the report format. It is not a validated measure of real-product safety or engine performance. Current label judgments retain their model and method identifiers; they are not silently converted into these sample scores or routing tiers.

Safety Coverage Score = 100 × Σ weight of label elements detected in the answer Σ weight of all required elements
Boxed Warning 3.0 Contraindications 2.0 Warnings & Precautions 1.5 Adverse Reactions 1.0
Tiers and overrides
▲ High <50
◆ Moderate 50–74.9
● Low ≥75
0
50
75
100
An omitted boxed warning forces the tier to at least moderate; a run where every element is omitted forces high, regardless of the arithmetic.
The minimum rule, and why

An engine's score is the minimum across its answers, and an element counts as detected only when every answer surfaces it. The rationale is the user's reality: each person sees one answer, not the average of four. If any answer omits the contraindication, some user got that answer.

“Patients used to search. Now they ask, and the answer they get is the only one they see.”
Label source and versioning

Labels are fetched from openfda and versioned by fetch date. Every report names the label version it scored against; re-measurement always re-fetches. Engine coverage: four engines by API, plus human-captured sessions from any AI product, imported with distinct provenance.

The fetched, versioned label is the source of truth: the thing visibility metrics don’t have.
Known limits, stated plainly
·Coverage measures presence of required safety elements. It does not certify that the rest of an answer is accurate.
·Engines are non-deterministic; a score is a dated sample of answers, not a permanent property of an engine.
·Element detection is rule-based and versioned; matcher changes are logged and never applied retroactively.
·Prompt sets are finite. They sample the question space; they do not exhaust it.
·Human-captured sessions reflect one signed-in context and are marked as such; they are evidence, not a census.
·A VizLoop report supports review; it is not a compliance determination and does not replace one.
Changelog
v0.3.0 In evaluation Per-answer label judgments with model and method identifiers, verbatim evidence, and explicit flags for unverified spans. v0.2.0 Sample formula Weighted element scoring, minimum rule, tier overrides, routed next steps, loop statuses, human-captured session import.
Gated tier
Full methodology specification

The sample formula explains the fictional example. The full specification is the working IP:

·Matcher implementation and thresholds
·Element extraction and filtering rules
·Client prompt sets
·Alerting thresholds
Available to clients under NDA
Reviewed in full during onboarding; version-pinned to your engagement.
Contact us