Open Source AI in Healthcare: A Practical Guide to Models, Use Cases and Where to Start
Open weight models let health organizations run AI inside their own security boundary. A guide to MedGemma, Meditron, Llama and DeepSeek, real use cases, and how to start safely.
Published 2026-10-01 · By Shahzad Asghar
Open source AI in healthcare lets a health organization run capable models inside its own security boundary, so patient data does not have to leave its environment. Strong open weight options already exist, including MedGemma, Meditron, Llama 3.1 and DeepSeek R1. The main advantage is not lower cost. It is control over data, deployment and governance.
Artificial intelligence in healthcare is no longer limited to proprietary cloud platforms. Open weight large language models now give hospitals, ministries of health, public health agencies, researchers and humanitarian organizations a real choice: run the model on their own infrastructure, adapt it to their own clinical guidance, and keep control of patient information.
This guide is written for clinicians, health information system specialists, public health experts and health leaders who want a clear answer to three questions. What is available? What can it actually do? And where should an organization start?
Key takeaways
- Open weight AI lets a health organization run the model on its own servers, so patient data does not have to leave its environment.
- Strong options already exist, including MedGemma, Meditron, Llama 3.1, DeepSeek R1 and medical models from John Snow Labs.
- The best first use cases are low risk: documentation, summarization, coding support and guideline search, always with human review.
- Benchmark scores are not clinical approval. Validation, governance and accountability are still required for every use case.
- The main advantage of open weight AI in healthcare is not cost. It is control.
What is open source AI in healthcare?
Open source AI in healthcare refers to AI models whose weights, meaning the trained parameters of the model, are published so that organizations can download them, run them on their own infrastructure, and adapt them to medical tasks. Most of these models are more accurately described as open weight rather than fully open source, because the training data and code are not always released.
For a health system, the difference that matters is practical. With a proprietary AI service, you send data to the provider and receive an answer. With an open weight model, the model comes to your data. It can run in your data center, your private cloud, or on a single server in a hospital with limited connectivity.
Why does open weight AI matter for health systems?
Healthcare is different from most sectors. Patient information is highly sensitive, clinical decisions require clear accountability, and many health systems are legally or ethically unable to send protected data to external AI services. Open weight models address several of these constraints directly.
Data privacy and residency. Patient records can stay inside the organization's own security boundary, which simplifies compliance with HIPAA, GDPR and national data protection laws, as well as humanitarian data protection standards.
Customization to local practice. A model can be connected to approved clinical guidelines, national treatment protocols, formularies and local terminology, or fine tuned on de-identified local data.
Transparency and auditability. Because the model runs under your control, you can log every interaction, version the model, test it against local protocols, and reproduce results during an audit or incident review.
Cost predictability. Instead of paying for each request, the organization pays for infrastructure. At scale, and for steady workloads, this can be more predictable than usage-based pricing.
Reduced vendor lock-in. If a better model appears, the organization can switch without rebuilding every integration around a single provider.
Low-connectivity and field settings. Smaller models can run offline, which matters for district hospitals, field clinics and humanitarian operations where internet access is unreliable.
Local language support. Models can be adapted to the languages patients and health workers actually use, including languages that commercial products serve poorly.
Open weight or proprietary: an honest comparison
Proprietary services are usually faster to start with, require no hardware, and the leading commercial models still perform at the top of many benchmarks. They suit organizations with strong contractual protections, approved cloud environments and limited technical staff.
Open weight models require more internal capability: infrastructure, engineering and ongoing maintenance. In return they offer control over data, deployment and governance. Many health organizations will end up with a mix, using open weight models for sensitive patient data and commercial services for work that carries no patient information, such as drafting public communications or analyzing published literature.
Which open weight AI models matter for healthcare?
Five are worth knowing in detail.
MedGemma from Google is a purpose-built medical model for text and images, and the most realistic starting point for most health organizations. Meditron from EPFL is built on clinical guidelines and fits global health and humanitarian settings. Llama 3.1 405B from Meta is a general model with strong diagnostic reasoning results in research, but it needs large infrastructure. DeepSeek R1 is a reasoning model suited to complex analytical work, with smaller distilled versions that are more practical. John Snow Labs publishes smaller healthcare-focused models designed for clinical text processing inside hospital infrastructure.
MedGemma: a medical model built for text and images
Google developed MedGemma specifically for medical text and image applications, by training a medically optimized image encoder and then training 4 billion and 27 billion parameter versions of Gemma 3 on medical data.
The MedGemma 27B text model achieved 87.7 percent on the MedQA benchmark, up from 74.9 percent for the general model it was built from, and within roughly three points of DeepSeek R1 at approximately one tenth of the inference cost.
Potential healthcare uses include clinical document review, medical question answering, radiology and imaging support, longitudinal patient record analysis, and medical education.
On deployment, the 27B model can run on a single high-end data center GPU, and the smaller multimodal model runs on far more modest hardware. Google itself states that MedGemma requires validation and adaptation before use in clinical workflows.
Meditron: built on clinical guidelines for global health
Meditron was developed by researchers at EPFL with medical and humanitarian partners. The original 7 billion and 70 billion parameter models were adapted from Llama 2 using a medical corpus that included PubMed material, medical papers and tens of thousands of clinical practice guidelines.
Its design makes it relevant to global health and humanitarian settings. A ministry of health, United Nations agency, non-governmental organization or field hospital could adapt a model of this kind around approved guidance, national treatment protocols, emergency health procedures or local reference material.
This does not mean such a model should independently provide medical advice. Its developers explicitly warn that additional alignment, testing and real-world evaluation are required before medical deployment.
Llama 3.1 405B: general purpose, with strong diagnostic research results
Meta designed Llama 3.1 405B as a general purpose model, yet medical research has shown notable results. A study led by Harvard Medical School with clinicians at Beth Israel Deaconess Medical Center and Brigham and Women's Hospital compared Llama 3.1 405B with GPT-4 across 92 diagnostically challenging clinical cases from the weekly case records of The New England Journal of Medicine. The open model performed on par with GPT-4 as assessed by physicians. The results were published in JAMA Health Forum.
This makes Llama attractive for organizations that want one model platform supporting both clinical and administrative workflows.
Potential healthcare uses include diagnostic reasoning research, clinical case analysis, medical knowledge retrieval, document summarization, and retrieval systems connected to approved clinical guidance.
On deployment, the 405B model needs a multi-GPU server. The smaller 8 billion and 70 billion parameter versions are far easier to host and are often enough for summarization and retrieval.
DeepSeek R1: a reasoning model for complex analysis
DeepSeek R1 is a 671 billion parameter reasoning model, with 37 billion parameters active during inference. Its development placed strong emphasis on reinforcement learning and step-by-step reasoning.
In healthcare, its strongest role is likely complex analytical work rather than direct diagnosis. Potential uses include summarizing long patient histories, comparing clinical documentation across multiple encounters, preparing draft case summaries, and explaining technical medical information in simpler language for patients or non-specialist staff.
Running the full model locally requires substantial GPU capacity. The smaller distilled versions are a more realistic option for hospitals with limited computing resources.
John Snow Labs medical models: practical models for clinical text
John Snow Labs provides medical language models designed around healthcare-specific workloads. Some are relatively small, including models in the 7 billion parameter class, which makes local deployment far more practical than models with hundreds of billions of parameters.
Potential uses include clinical text processing, information extraction from clinical notes, medical question answering, summarization, medical coding support, and integration with existing health information systems. Licensing terms vary across these products, so procurement and legal teams should review them early.
Also worth watching
The open weight landscape moves quickly. General purpose open models have reported strong results on health-related evaluations, and new medical fine tunes appear regularly on model hubs. A good governance process should make it straightforward to evaluate new models against your own test cases rather than relying on headline scores.
What are the practical use cases by role?
For doctors and clinicians: drafting discharge summaries, referral letters and clinical notes for clinician review; summarizing long patient histories before a consultation; searching clinical guidelines in plain language; preparing patient education material in simpler language or local languages; supporting case-based learning.
For health information system and digital health teams: extracting structured data from free-text clinical notes; supporting diagnostic and procedure coding; improving data quality, for example detecting duplicates or inconsistencies in records; building retrieval systems over approved clinical knowledge; adding natural language search to electronic health record and health management information systems.
For public health experts: analyzing surveillance reports and outbreak narratives at scale; summarizing research evidence for policy briefs; classifying community feedback, hotline logs and rumour tracking data; producing multilingual risk communication drafts.
For hospital management: automating routine correspondence, scheduling and reporting; summarizing incident reports and quality reviews; supporting procurement, policy and training documents.
For ministries of health and humanitarian organizations: adapting models around approved national and international guidance for health workers in the field; offline decision support for low-connectivity facilities; faster reporting during emergencies, with data kept under organizational control. The wider picture of how these systems are governed is covered in my work on AI governance in the United Nations, and the clinical applications in more depth under AI in health.
Which use cases should come first?
Not every use case carries the same risk. A sensible approach is to move up a ladder of risk only as experience and evidence grow.
Low risk: administrative drafting, summarizing published literature, internal knowledge search, training content.
Medium risk: summarizing patient records, documentation support and coding suggestions, where a qualified person reviews every output.
High risk: anything that influences diagnosis or treatment. These uses require formal clinical validation, and in many jurisdictions they fall under medical device regulation.
Most organizations should begin at the low risk level, build governance and confidence, and only then consider clinical applications.
What does it take: infrastructure, people, data and governance
Infrastructure. Small models of roughly 4 to 8 billion parameters can run on a single GPU server or a capable workstation. Mid-sized models of roughly 27 to 70 billion parameters typically need one or more data center GPUs. Very large models are usually practical only for national-level or well-funded institutions. Quantized versions reduce hardware requirements at some cost to quality.
People. At a minimum: a clinical lead who owns the use case, an engineer or data scientist who can deploy and evaluate models, an information security officer, and a data protection lead. Small teams can achieve a great deal if roles are clear.
Data. You need clean, approved reference content such as guidelines, protocols and formularies, and de-identified test cases that reflect your real patients, languages and workflows.
Governance. Every production use case needs a named owner, a documented purpose, access controls, audit logging, human review rules, model version control, monitoring for errors and bias, and an incident process.
Validation. Test the model against local cases before deployment and continue testing afterwards. A model that performs well on a United States examination benchmark may perform differently on your patient population, language and disease burden.
Where should a health organization start?
- Pick one problem, not a technology. Choose a specific, measurable pain point, such as the time clinicians spend writing discharge summaries.
- Classify the risk. Confirm whether the use case is administrative, supervised clinical support, or clinical decision-making.
- Form a small multidisciplinary team. Include clinical, information technology, security, data protection and, where relevant, patient representation.
- Shortlist two or three models. For many organizations this will be a MedGemma variant, a smaller Llama model, or a domain model such as Meditron.
- Build a local test set. Use de-identified real examples and define what a good output looks like before testing.
- Pilot in a controlled environment. Run the model inside your own infrastructure, with human review of every output and full logging.
- Measure and decide. Assess accuracy, time saved, user satisfaction and safety events, then decide to scale, adjust or stop.
A focused pilot of this kind can often be completed in around three months and produces evidence that leadership can actually use.
What questions should leadership ask before starting?
- What specific problem are we solving, and how will we measure success?
- Will patient data be used, and where exactly will it be processed and stored?
- Who is clinically accountable for each output?
- What is the licence of the model, and does it allow our intended use?
- What hardware and skills do we need, and do we have them?
- How will we validate the model on our own patients and languages?
- How will we detect and report errors, bias or harm?
- Does this use case fall under medical device regulation in our jurisdiction?
- What happens if we need to switch models later?
Regulation and ethics: what health leaders should know
Open weights do not remove regulatory obligations. Key reference points include the World Health Organization guidance on the ethics and governance of large multi-modal models in health, the European Union AI Act, which treats AI used in medical devices as high risk, the United States Food and Drug Administration framework for AI-enabled software as a medical device, and national data protection laws such as HIPAA and GDPR. Humanitarian organizations should also apply humanitarian data protection standards, because health data about displaced or vulnerable populations carries additional risk.
Limitations to keep in mind
Large language models can produce confident but incorrect answers. They can reflect biases in their training data. Benchmark results such as MedQA measure examination performance, not safe patient care. Open weight models also shift responsibility to the organization running them: security patching, monitoring and maintenance become your job. None of these limitations is a reason to avoid AI, but each is a reason to deploy it carefully.
Frequently asked questions
What is the best open source AI model for healthcare? There is no single best model. MedGemma is a strong general starting point for medical text and images, Meditron suits guideline-based global health settings, and smaller Llama or John Snow Labs models are practical for clinical text processing. The right choice depends on your use case, hardware and validation results.
Is open source AI safe to use with patient data? It can be safer from a data protection perspective because patient data stays within your infrastructure. However, safety also depends on access controls, audit logging, validation and human review. Running a model locally does not make it clinically safe by itself.
Can open source AI diagnose patients? Research shows some open models perform well on difficult diagnostic cases, but no general open weight model is approved to diagnose patients independently. Any use that influences diagnosis or treatment requires clinical validation and, in many jurisdictions, regulatory approval.
What hardware do hospitals need to run open weight AI? Small models can run on a single GPU server. Mid-sized models such as MedGemma 27B usually need one high-end data center GPU. Very large models need multi-GPU servers. Many organizations start small and scale once a use case proves its value.
What is the difference between open source and open weight AI? Open weight models publish their trained parameters so you can run and adapt them. Fully open source models also release training data and code. Most healthcare-relevant models today are open weight.
How should a health organization start with AI? Start with one well-defined, low risk problem, form a small multidisciplinary team, test two or three models on de-identified local cases, run a controlled pilot with human review, and measure results before scaling.
The direction is clear
Health organizations no longer have to choose between sending patient data to proprietary services and avoiding generative AI altogether. Open source AI in healthcare offers a third path, one that gives organizations far greater control over infrastructure, data, integration and governance.
For health systems responsible for sensitive patient information, that control may become one of the most important characteristics of the next generation of healthcare AI. The organizations that benefit most will be those that start small, govern carefully and build evidence from their own context.
---
*Shahzad Asghar is Head of Data and Digital Solutions at UN-ESCWA, working on AI governance, data and digital transformation, with experience across humanitarian health, information management and data systems in the Arab region, East Africa and South Asia. The views expressed here are his own and do not represent those of the United Nations. More detail is available on his professional profile.*
Written by Shahzad Asghar — Head of Data and Digital Solutions at UN-ESCWA, with 20+ years building AI and data systems across UNHCR, UNICEF, and UNOCHA. His team built UNHCR’s first global IVR appointment system, serving 700,000+ refugees. He created the Last-Mile AI Framework. Read more about this UN AI expert