This paper examines the extent to which QLoRA fine-tuning can improve the performance of the large language model Mistral-7B-Instruct-v0.3 on Serbian high school mathematics competition tasks.Based on a dataset of tasks in Serbian, a fine-tuned model, Math-SRB-Mistral-7B, was developed and compared with the base model.The responses were evaluated using Claude 3.7 Sonnet as a judge, according to multiple criteria, including final answer accuracy, logical coherence, explanation quality, and an aggregate score.The results suggest that the applied fine-tuning did not lead to improved performance; instead, the fine-tuned model achieved slightly lower scores across all evaluated dimensions.This finding suggests that parameter-efficient adaptation of general-purpose LLMs on small and challenging mathematical datasets does not necessarily result in better generalization to new tasks.At the same time, the results highlight the importance of multi-criteria evaluation in the analysis of mathematical reasoning generated by LLMs.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex