L2 = per-sentence verification of the patient summary. L1 = per-entity grounding of the
extraction (panel below). They measure different things — a 100% L2 pass rate next to a 17%
L1 ABSTAIN rate is normal, not contradictory.
—
L3 faithful rate
—
L3 questionable
—
L3 fabricated
—
L3 judge coverage
—
Held (auto / manual)
L3 = post-generation faithfulness judge. Checks whether each claim in the summary traces back to the source
document. Fabricated = at least one claim contradicts or has no basis in the source.
—
L1 consistency errors
—
L1 consistency warnings
L1 clinical consistency = pure-logic checks on extracted data (staging contradictions, ICD/description mismatches, therapy date inversions, implausible lab values). No LLM — runs in <1 ms per doc.
Top Omissions
Loading…
Cluster triage
Label this cluster:
Reviewer labels
Faithfulness Judge
Extraction Grounding (L1 verifier — Q7)
Loading…
Top near-misses below — entities the verifier almost let through.
Sorted by fuzzy score descending. Anything ≥ the PARAPHRASE threshold passed;
ABSTAIN entries got suppressed. Use this to decide whether the threshold is
rejecting real content.