Averroes-Q: Training a 32-Billion Parameter Bilingual Arabic-English LLM on a Single Apple M2 Ultra Workstation
We present Averroes-Q, a 32-billion parameter bilingual Arabic-English large language model fine-tuned entirely on a single Apple M2 Ultra workstation with 192 GB of unified memory. Built upon the Qwen2.5-32B-Instruct architecture, Averroes-Q is the first publicly released Arabic-capable LLM trained exclusively on consumer-grade Apple Silicon hardware. Using parameter-efficient fine-tuning via Low-Rank Adaptation (LoRA) and Apple's MLX framework, we demonstrate that a complete bilingual instruction-tuning pipeline—including training, adapter fusion, quantization, and deployment—can be executed on a single machine without any cloud GPU infrastructure. The model is trained on the Averroes Corpus, a curated bilingual dataset of 5.7 million instruction-response pairs, with a filtered high-quality v2 subset of 1.2 million examples. Over four training iterations spanning adapter ranks 8 and 16 with up to 5,000 steps, we achieve a best validation loss of 1.440 and a training loss of 1.629. Peak memory consumption reaches 67.4 GB, well within the M2 Ultra's 192 GB capacity. Averroes-Q serves as the primary bilingual backbone for multiple production systems, including the SAIF framework, the Rushd assistant, and the Hayula Bot. We release all model variants, training configurations, and evaluation benchmarks to the research community via HuggingFace, where the model has received over 156 downloads. Our results demonstrate that competitive bilingual LLMs can be developed on consumer hardware, significantly lowering the barrier to entry for low-resource language AI research.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex