Task Framing Outweighs Model Scale in Contract Clause Extraction: Pre-Registered Experiments with Frontier LLMs and Small Fine-Tuned Specialists on CUAD

We benchmark four frontier language models (Claude Opus 4.8, Claude Opus 4.6, GPT-5.2, GPT-4o) and a series of small fine-tuned specialists on CUAD contract clause extraction, scored with the official AUPR metric through one identical pipeline, with every raw prediction file published. Three findings emerge. First, the frontier gap is category-shaped rather than uniform: a fine-tuned 14B specialist that trails the best frontier model by 17 points overall still beats GPT-4o in 25 of 40 clause categories and beats all four frontier models simultaneously in six. Second, task framing dominates model scale. Reframing the specialist from answer generation to extractive span prediction, with the base model, data, and budget held fixed, moves test AUPR from 0.291 to 0.389, while growing the extractive backbone 28x (0.5B to 14B) moves dev AUPR only from 0.284 to 0.303. Frontier scaling shows the same pattern: GPT-4o and GPT-5.2, a full model generation apart, score within 0.002 of each other. Third, we pre-registered a public prediction that the extractive 14B would clear 0.50 test AUPR; it missed, and an analysis of the miss localizes the remaining gap to the deliberately lean training recipe rather than architecture or scale. All code, predictions, and the trained adapter are public at https://github.com/ashish24142/nightwing

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC