This paper presents a novel framework in which multiple advanced large language models—GPT-4-0125-preview, Meta-LlaMA-3-70B-Instruct, Claude-3-Opus, and Gemini-1.5-Flash—jointly generate and answer complex, doctoral-level probability problems without relying on predefined ground truth. Instead of measuring accuracy against fixed answers, we interpret inter-model agreement as an empirical indicator of response reliability and question quality. Statistical analyses—including chisquare tests, Fleiss’ κ, and bootstrap confidence intervals—are used to quantify alignment, variability, and stability across models. Results show that Claude and Gemini produce questions with higher coherence and narrower confidence intervals, leading to stronger consensus among answering models. In contrast, LLaMA exhibits broader uncertainty ranges and lower agreement, reflecting greater inconsistency in its formulations. These findings demonstrate that collaborative reasoning among heterogeneous LLMs can enhance both answer dependability and question design evaluation, offering a scalable, data-driven approach to truth-free validation in multi-model AI reasoning systems.
Paper
References (58)
Scroll for more · 38 remaining