The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives

Instruction-tuned large language models (IT-LLMs) exhibit strong zero-shot reasoning, yet their ability to execute simple, self-contained instructions remains underexplored, despite being foundational to more complex instruction-following. We evaluate 20 IT-LLMs on modified MMLU and MMLU-Pro benchmarks by systematically varying option-label formats (alphabetic, numeric, Roman) while preserving semantic equivalence across four experimental paradigms. (1) With explicit instructions, semantically invariant label changes induce large performance shifts (e.g., -30.45% for Roman versus numeric), revealing strong instruction-format bias. (2) Removing instructions further degrades performance (up to -10.84%) and amplifies label sensitivity, highlighting the importance of explicit guidance. (3) When option content is removed, models fail to consistently meet random-choice baselines except with numeric labels, indicating weak adherence to atomic directives under semantic underspecification. (4) Three-shot exemplars yield no significant gains in robustness or output fidelity, and generation analyses highlight persistent label violations, particularly for non-numeric formats. Although larger models achieve higher overall accuracy, instruction adherence remains inconsistent across scales, demonstrating that improved task performance does not entail reliable instruction-following. These findings expose limitations of current instruction-tuning paradigms and motivate evaluation methods and training strategies that explicitly target atomic instruction-following.

Paper

Similar papers

© 2026 NYSGPT2525 LLC