Small Open-Weight Language Models versus Frontier Models for High-Stakes Clinical Triage in Low-Resource Settings: Two Case Studies and a Multi-Model Research Plan
By mid-2026, small open-weight language models (SLMs) with 1–10 billion parameters are good enough to deploy in production for clinical decision support. Builders pick them for two reasons: they run offline on commodity hardware, and they sidestep the data-residency and per-call cost penalties of frontier-model APIs. The clinical-capability trade-off, however, is poorly characterized at the system level. Public benchmarks evaluate models in isolation, not the pipelines in which they actually carry production traffic. We characterize this trade-off using two deployed case studies built by the first author in rural Tamil Nadu — Sentinel Health (Gemma 4 e4b on a clinic laptop, offline emergency triage) and Path to Care (Gemma 4 31B-it on a single AMD MI300X, LoRA-tuned dermatology classification) — and a planned multi-model evaluation. Exploratory pilots: three frontier providers (Claude Opus 4.7, GPT-5, Gemini 2.5 Flash) returned RED-tier triage at 9/9 sensitivity including on an image-only suspected snake-bite case where Gemma 4 e4b returned YELLOW with "no acute condition identified." Six local SLMs ranged from 0/5 to 9/9 sensitivity on the same cases. The atypical-MI text case was misclassified to YELLOW by 3 of 6 SLMs. The two case studies converged on the same architectural primitive — a deterministic safety scaffolding pattern — without coordination. We pre-register a 12-model × 250-case evaluation (SentinelEval-250) with seven hypotheses, a defined statistical methodology, sample-size power analysis, and a 30-week timeline. We propose an architectural taxonomy of five compensations — deterministic safety nets, KB-grounded JSON-Schema output, two-pass vision-as-sensor pipelines, selective frontier escalation, and domain LoRA adaptation — that we hypothesize allow SLMs to carry production clinical traffic. The framing claim is that the right unit of evaluation is the pipeline, not the model. Code, prompts, JSON schemas, safety-net rule sets, pilot results, and the SentinelEval-250 case construction at https://github.com/SankarSubbayya/sentinel-health (Apache-2.0). Path to Care sister case study at https://github.com/SankarSubbayya/amd_hackathon (Apache-2.0).
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex