Benchmarking Truthfulness Metrics in Health-Oriented Large Language Models via Automated Fact Verification Pipelines
Main Article Content
Abstract
The proliferation of large language models (LLMs) in health-related applications has intensified the need for rigorous evaluation frameworks capable of quantifying factual accuracy and mitigating the propagation of medical misinformation. Despite substantial advances in general-purpose truthfulness benchmarking, systematic assessments tailored to the clinical and biomedical domain remain critically underdeveloped, leaving practitioners without reliable instruments for deployment-readiness evaluation.
This work introduces a comprehensive benchmarking methodology that evaluates truthfulness metrics across a curated ensemble of health-oriented LLMs using automated fact verification pipelines. We construct a domain-specific evaluation corpus comprising clinically grounded claims spanning pharmacology, epidemiology, diagnostic reasoning, and evidence-based treatment guidelines. Each claim is processed through a multi-stage verification architecture that integrates retrieval-augmented evidence sourcing, natural language inference scoring, and calibrated confidence estimation to produce composite truthfulness scores.
We systematically compare five distinct truthfulness metrics---including entailment-based fidelity, semantic consistency under paraphrase, source-grounded precision, hallucination rate, and calibration error---across seven prominent LLMs spanning both general-purpose and biomedically fine-tuned architectures. Experimental results demonstrate that biomedically specialized models achieve statistically significant improvements in source-grounded precision ($\Delta = 0.14$, $p < 0.01$) yet exhibit elevated calibration error relative to general-purpose counterparts, revealing a previously undercharacterized accuracy--confidence trade-off in the health domain.
Our findings establish that no single metric suffices for comprehensive truthfulness assessment, motivating a composite evaluation paradigm. We release our benchmark dataset, verification pipeline, and evaluation toolkit to facilitate reproducible research and responsible deployment of LLMs in high-stakes healthcare environments.