Vertical Scaling of Domain-Specific Small Language Models in High-Stakes Financial Compliance
This paper evaluates the efficacy of synthetic regulatory instruction datasets on the performance of small language models (SLMs) in structured statutory reasoning and compliance verification tasks. We fine-tune a 3.8B parameter model derived from the Phi-3-mini architecture using Parameter-Efficient Fine-Tuning (PEFT) via QLoRA on a programmatically generated, multi-turn corpus spanning complex U.S. financial regulations (e.g., SEC Rule 10b-5, Regulation D, Regulation SHO, and Basel III). The fine-tuned artifact, Nova-FinLex-Phi3, is open-sourced and accessible via HuggingFace.1 To establish strict operational performance bounds, we benchmark this model against two large-scale foundational architectures: NVIDIA Nemotron-3-Super (120B) and Owl-Alpha. Our investigation isolates whether highly targeted, low-variance synthetic prompt engineering can induce specialized regulatory compliance execution without manual, expert-annotated training data. The empirical results show that while domain fine-tuning significantly improves statutory identification precision (+22.5% over the unaligned base architecture), it introduces systemic behavioral vulnerabilities under extended generation limits. Specifically, we document a recurrent behavioral phenomenon marked by structural delimitation slippage, phrase truncation, and autoregressive token looping. These findings illuminate a fundamental optimization trade-off between task-specific semantic alignment and generalpurpose instruction adherence in constrained parameter regimes
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex