Intelligent conversational systems have provided improved access to initial healthcare data by using artificial intelligence. The present paper describes a multimodal AI-based medical chatbot that facilitates the interaction by text, voice, medical image, and document with pretrained models. The system is deployed based on a Flask backend and Google Gemini models to understand medical queries and the generated response via Groq-based large language inference and prominently the models. Voice response is supported by using Vosk speech-to-text and pyttsx3 text-to-speech. The chatbot keeps track of the history of conversation through the context management of its sessions thus allowing consistent follow-up replies. Gemini vision and Hugging face models are used to analyze medical images and documents. The system is fully based on real-time inference, which offers low-latency timely, caring, and thoughtful medical support.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex