A Method for the Architecture of a Medical Vertical Large Language Model Based on Deepseek R1

In recent years, despite foundation models like DeepSeek-R1 and ChatGPT demonstrating significant capabilities in general tasks, professional knowledge barriers, computational resource requirements, and deployment environment limitations have severely hindered their application in actual medical scenarios. Addressing these challenges, this paper proposes an efficient lightweight medical vertical large language model architecture method, systematically solving the lightweight problem of medical large models from three dimensions: knowledge acquisition, model compression, and computational optimization. At the knowledge acquisition level, a knowledge transfer pipeline is designed from the fine-tuned DeepSeek-R1-Distill-70B teacher model to the DeepSeek-R1-Distill-7B student model, and Low-Rank Adaptation (LoRA) technology is adopted to precisely adjust key attention layers. At the model compression level, compression techniques including 4-bit weight quantization are implemented while preserving the core representation ability for medical reasoning. At the computational optimization level, inference optimization techniques such as Flash Attention acceleration and continuous batching are integrated, and a professional prompt template system is constructed to adapt to different types of medical problems. Experimental results on medical question-answering datasets show that the method proposed in this paper maintains professional accuracy while reducing memory consumption by 64.7\% and inference latency by 12.4\%, providing an effective solution for the application of medical large models in resource-constrained environments such as edge computing devices.

Paper

References (27)

05Specialized vs. general large language models in medicinenpj Digital Medicine
06Benchmarking medical large language models: Challenges and opportunitiesJournal of Biomedical Informatics
07AI-assisted clinical decision support systems: A reviewJournal of Clinical Medicine
08Medical question answering systems: A comprehensive surveyACM Computing Surveys
09Progressive knowledge transfer for medical concept learningProceedings of AMIA Annual Symposium
10CUDA graphs for accelerated deep learning inferenceProceedings of Machine Learning and Systems
11Med-PaLM 2: Towards expert-level medical question answering with large language modelsNature
12DistillBERT-Med: A distilled biomedical language representation modelProceedings of ACL BioNLP Workshop

Scroll for more · 15 remaining

Similar papers

© 2026 NYSGPT2525 LLC