Hallucination, Monofacts, and Miscalibration: An Empirical Investigation

Significance We show that hallucination in large language models can be controlled through deliberate manipulation of training data frequency distributions. By sampling training data from heavy-tailed distributions that naturally reduce rare facts, and by strategically repeating small subsets of training examples during fine-tuning, we reduce hallucination rates by up to 40% without sacrificing accuracy. Our findings challenge the widespread practice of deduplicating training data and reveal that the distribution of fact frequencies could fundamentally affect model reliability. This work establishes training data composition as a primary lever for hallucination control, offering practitioners a simple, interpretable alternative to complex post hoc intervention methods that operate on model internals.

Paper

Similar papers

© 2026 NYSGPT2525 LLC