Benchmarking Truthfulness Metrics in Health-Oriented Large Language Models via Automated Fact Verification Pipelines

Main Article Content

Kiran Das
Rahul Singh

Abstract

The proliferation of large language models (LLMs) in health-related applications has intensified the need for rigorous evaluation frameworks capable of quantifying factual accuracy and mitigating the propagation of medical misinformation. Despite substantial advances in general-purpose truthfulness benchmarking, systematic assessments tailored to the clinical and biomedical domain remain critically underdeveloped, leaving practitioners without reliable instruments for deployment-readiness evaluation.


 


This work introduces a comprehensive benchmarking methodology that evaluates truthfulness metrics across a curated ensemble of health-oriented LLMs using automated fact verification pipelines. We construct a domain-specific evaluation corpus comprising clinically grounded claims spanning pharmacology, epidemiology, diagnostic reasoning, and evidence-based treatment guidelines. Each claim is processed through a multi-stage verification architecture that integrates retrieval-augmented evidence sourcing, natural language inference scoring, and calibrated confidence estimation to produce composite truthfulness scores.


 


We systematically compare five distinct truthfulness metrics---including entailment-based fidelity, semantic consistency under paraphrase, source-grounded precision, hallucination rate, and calibration error---across seven prominent LLMs spanning both general-purpose and biomedically fine-tuned architectures. Experimental results demonstrate that biomedically specialized models achieve statistically significant improvements in source-grounded precision ($\Delta = 0.14$, $p < 0.01$) yet exhibit elevated calibration error relative to general-purpose counterparts, revealing a previously undercharacterized accuracy--confidence trade-off in the health domain.


 


Our findings establish that no single metric suffices for comprehensive truthfulness assessment, motivating a composite evaluation paradigm. We release our benchmark dataset, verification pipeline, and evaluation toolkit to facilitate reproducible research and responsible deployment of LLMs in high-stakes healthcare environments.

Article Details

Section

Articles

How to Cite

Benchmarking Truthfulness Metrics in Health-Oriented Large Language Models via Automated Fact Verification Pipelines. (2026). International Journal of Computational Health & Machine Learning, 4(2). https://ijchml.com/index.php/ijchml/article/view/257

References

Similar Articles

You may also start an advanced similarity search for this article.