Open-Weight vs Open-Source LLMs, Fine-Tuning and RAG: A Plain-English Guide
Weights are the model, open weight is about who holds it, fine-tuning changes it, and RAG changes what it reads. A plain-English guide to all four, with examples.
Published 2026-10-05 · By Shahzad Asghar
Weights are the model, open weight is about who holds the model, fine-tuning changes the model, and RAG changes what the model reads when it answers. Those four sentences settle most of the confusion in the room.
I did my master's in AI back in 2003. AI then meant expert systems, search algorithms and small neural networks. Nobody talked about models with billions of parameters that you could download.
So I will be honest. When terms like open weight, fine-tuning, LoRA and RAG started coming up in my work, I struggled to understand them properly. Most explanations I found were either too technical or read like sales pitches for one approach. I even went back and forth with AI chatbots and still did not get a straight answer to a simple question: which tool fits which problem?
So I worked it out myself, and wrote this guide the way I wish someone had explained it to me.
If you sit in any meeting about AI today, four terms come up again and again: weights, open weight, fine-tuning and RAG. They are often used loosely, and decisions get made on half-understood definitions.
This guide explains each one in plain English, with an analogy and a concrete example. It is written for managers, policy staff and technical people who want the basics right before choosing an approach. It does not argue for one option. The goal is to help you see which tool fits which problem.
What are model weights?
A large language model (LLM) is, in the end, a very large set of numbers. These numbers are called weights (or parameters). A small model has a few billion of them; the largest have hundreds of billions or more.
During training, the model reads huge amounts of text and keeps adjusting these numbers until it gets good at predicting the next word. When training ends, the numbers are frozen. Everything the model can do, such as understanding a question, writing in Arabic or English, summarising or reasoning, is stored in those numbers.
Analogy: think of a sound mixing desk with billions of sliders. Training is the long process of setting every slider to the right position. The finished settings are the weights. Copy the settings and you have copied the sound.
Two practical points follow from this:
- Whoever holds the weights holds the model. If you have the weights file and the hardware to run it, you can run the model yourself.
- Weights are not a database. The model does not store documents it read like files in a folder. It stores patterns as numbers, which is why it can recall facts only approximately and sometimes gets them wrong.
Open-weight vs open-source vs closed LLMs
The difference between these three is simply how much of the model you are given.
| Closed model | Open-weight model | Fully open-source model | |
|---|---|---|---|
| Examples | GPT, Claude, Gemini | Llama, Mistral, Qwen | OLMo |
| Weights you can download | No | Yes | Yes |
| Training data published | No | No | Yes |
| Training code published | No | Sometimes, partly | Yes |
| Where it runs | Vendor's servers | Your servers or any cloud | Your servers or any cloud |
| Licence | Vendor terms of service | Varies, often with conditions | Usually permissive |
Analogy: a closed model is a restaurant. You order, the meal arrives, and the kitchen stays closed. An open-weight model is a cooked meal you can take home, reheat and add your own spices to, but the recipe stays secret. A fully open-source model gives you the meal, the full recipe and the list of where every ingredient came from.
Open weight is not the same as open source
This is the most common confusion. Most well-known "open" models are open weight only. You can run them and adapt them, but you cannot see what data they were trained on, and you cannot rebuild them from scratch. The Open Source Initiative's Open Source AI Definition asks for much more than downloadable weights: it requires enough detail about the training data for a skilled person to build a substantially equivalent system.
The licence also matters. Some open-weight licences restrict commercial use, very large user bases or certain purposes. Read the licence before building anything on top of a model.
Why organisations care about open weight
The main reason is control over data. With an open-weight model, you can host it inside your own infrastructure, so sensitive information never goes to an outside vendor. For public institutions, humanitarian organisations and health systems handling personal data, this can decide whether AI is usable at all. The sector-specific version of that argument, including which open models are credible for clinical work, is in open source AI in healthcare.
The trade-off is that you now carry the work the vendor used to do: GPUs or cloud compute, security patching, monitoring, and keeping up with newer model versions.
What is fine-tuning, and what does it actually change?
Fine-tuning means training an existing model further on your own examples so that it behaves differently. It changes the weights. Nothing else in the model changes: not its structure, and no document store is added.
Analogy: a newly hired interpreter already speaks Arabic and English well. You then give them months of practice on one regional dialect and one type of interview. They do not carry a reference book afterwards. The new skill is now part of how they work.
How fine-tuning is done, step by step
- Prepare examples. Collect pairs of input and correct output, usually a few hundred to a few thousand. Most of the effort goes here, because inconsistent examples produce an inconsistent model.
- Choose a base model. Usually an open-weight model that is already reasonably good at the task and small enough to run affordably.
- Train. The model sees an input and predicts an output. Its prediction is compared with the correct answer and the error is measured. Each weight is then nudged slightly in the direction that would have reduced the error. This repeats across all the examples, typically for a few full passes.
- Test. Run the model on examples it never saw during training. Check that it improved on the task and did not get worse at things it used to do well.
- Deploy. Host the fine-tuned model on your own infrastructure, or on the vendor's platform if you fine-tuned a closed model through a vendor service.
A training example for classifying citizen feedback might look like this:
Input: "The health centre had no insulin for two weeks." Output: Category: Health, medicine stock-out
After enough examples like this, the model sorts new messages into your categories consistently, without being told the categories each time.
Full fine-tuning vs LoRA
There are two common ways to change the weights:
- Full fine-tuning adjusts all of the model's weights a little. It is powerful but needs significant GPU capacity, and the result is a complete new copy of the model.
- LoRA (Low-Rank Adaptation) freezes the original weights and trains a small add-on set of new weights, often well under 1% of the model's size. At run time, the add-on is loaded on top of the base model. It is much cheaper, a modest model can be tuned on a single good GPU, and you can keep several add-ons for different tasks on one base model. If something goes wrong, you remove the add-on and the original model is untouched. The method was introduced in LoRA: Low-Rank Adaptation of Large Language Models by Hu and colleagues in 2021.
In both cases, what changes is the model's probabilities. After fine-tuning, for a given input, the most likely output is the one you trained it to produce.
Can you fine-tune a closed model?
Sometimes. Some vendors offer fine-tuning of their closed models as a service. You send your training data to the vendor, and the tuned model stays on their platform. With an open-weight model, both the data and the tuned model stay with you.
What is RAG (retrieval-augmented generation)?
RAG leaves the model completely unchanged. Instead, when someone asks a question, the system first finds the relevant passages in your documents and hands them to the model together with the question. The model reads them and answers.
Analogy: fine-tuning is sending a staff member on a long course. RAG is putting the right folder on their desk just before they answer the phone.
How a RAG answer is produced
- Your documents (policies, manuals, contracts, reports) are split into short passages and indexed for search.
- A user asks a question, for example: "How many days of annual leave can I carry over?"
- The system searches the index and pulls the few passages most likely to contain the answer.
- Those passages, the user's question and an instruction (such as "answer only from the passages and cite them") are sent to the model together.
- The model writes the answer and points to the source passage.
Many enterprise AI assistants, including document chatbots and the assistants built into office software, work this way. When they answer from your files, they are retrieving those files at that moment, not drawing on a model trained on them.
Does the model forget its training when you use RAG?
No, and it does not need to. The work is divided:
- The weights provide the ability: understanding the question, reading the passages, reasoning and writing a clear answer.
- The retrieved passages provide the facts: what your policy actually says today.
The instruction tells the model to take its facts from the passages, and good models mostly follow it. It is not guaranteed. If the passages do not contain the answer, or the search pulls the wrong ones, the model may fill the gap from general knowledge, for example by describing a typical leave policy instead of yours. That is why well-built RAG systems show their sources and are told to say "I could not find this" when the documents are silent.
The practical lesson: a RAG system is only as good as its search and its documents. Outdated or contradictory documents produce outdated or contradictory answers, however strong the model is.
Fine-tuning vs RAG: when to use which
One question settles most cases: is the problem about what the model knows, or about how the model behaves?
- Knowledge problems (facts, rules, figures that change or must be cited) point to RAG.
- Behaviour problems (a dialect, a fixed format, a narrow task done the same way every time) point to fine-tuning.
| RAG | Fine-tuning | |
|---|---|---|
| What it changes | What the model reads at question time | The model's weights |
| Updating content | Replace the document | Retrain the model |
| Citing sources | Yes, it can point to the passage | No |
| Respecting user permissions | Yes, search can filter by user | No, everyone gets the same model |
| Main effort | Clean documents and good search | Good training examples and ML skills |
| Main risk | Wrong or missing passages retrieved | Model learns facts approximately and they go stale |
Before this becomes a technical decision it is usually an institutional one, covering correction speed, data deletion, auditability and residency. Those questions are set out in RAG or fine-tuning: the governance questions first.
Examples where RAG fits
- HR policy assistant. Staff rules get amended, people need the exact clause, and some content is restricted to certain roles.
- Contract and lease questions. Every contract is different and new ones arrive all the time. Nobody retrains a model each time a contract is signed.
- Donor or partner intelligence. Priorities shift every year, and access often needs to be limited to certain teams.
- Guidance for teachers or health workers. The answer must match the current official guideline, with a source.
Examples where fine-tuning fits
- Regional dialect transcription and translation. No document can teach a model a dialect. It has to learn from many verified examples.
- Reading handwritten historical records into a fixed structure. Human verifiers' corrections become training examples that steadily improve accuracy.
- Classifying complaints or feedback into fixed categories. High volume, must be consistent, and a small tuned model can run cheaply on your own server.
- Standardising report formats. If the problem is that reports come out in different structures, tuning can teach one fixed format.
Examples that need both
- A protection or case-management advisory tool. Fine-tune for dialect understanding, then use RAG to supply the current rules and procedures, which change and must be cited.
- A medical report assistant. Fine-tune for the report structure, then use RAG for the current clinical guidelines.
Examples that need neither
Drafting emails, summarising meeting notes, producing first drafts of lesson plans, or reviewing web pages against a clear rule. A capable model with a well-written prompt usually handles these. Always try good prompting first; it costs the least and often turns out to be enough.
Where hosting fits in
Open weight vs closed is a separate decision about where data goes and who runs the infrastructure. RAG works with either. Fine-tuning works with either, but only an open-weight model keeps both the training data and the tuned model fully in your hands.
Common mistakes to avoid
- Fine-tuning a model on your policies so it "knows" them. It sounds logical but works badly. The model absorbs the rules approximately, can mix up clauses, cannot say which clause it used, and needs retraining every time a rule changes. Anyone using the model can also draw out what it learned, whatever their clearance. Facts belong in RAG.
- Treating open weight as open source. You still cannot see the training data, and the licence may limit how you use the model.
- Assuming self-hosting is automatically cheaper. There are no per-question fees, but GPUs, engineers and maintenance cost real money. Compare total cost, not just licence cost.
- Blaming the model for bad RAG answers. Most RAG failures come from poor search or outdated documents. Fix the content and retrieval before switching models.
- Skipping evaluation. Whichever approach you pick, test it on a set of real questions with known correct answers, before and after every change.
Frequently asked questions
What is the difference between an open-weight and an open-source LLM?
An open-weight LLM lets you download and run the trained model, but its training data and full training process stay private. A fully open-source LLM also publishes the training data and code, so anyone can inspect or rebuild it.
Does fine-tuning change the model's weights?
Yes. Full fine-tuning adjusts all of the weights slightly. LoRA keeps the original weights frozen and adds a small set of new weights on top. Either way, the model's behaviour changes because its numbers change.
Is RAG better than fine-tuning?
Neither is better in general. RAG is the better fit for facts that change or must be cited. Fine-tuning is the better fit for teaching a skill, a dialect or a fixed format. Many systems use both.
Does RAG stop an AI model from hallucinating?
It reduces hallucination but does not eliminate it. If the right passage is not retrieved, the model may still guess. Source citations and a clear "not found" instruction help users catch this.
Can I use RAG with ChatGPT, Claude or Gemini?
Yes. RAG works with closed and open-weight models alike, because it only changes what the model reads, not the model itself.
Do I need my own GPUs to use an open-weight model?
You need somewhere to run it: your own GPUs, or a cloud or hosting provider that serves open-weight models. Smaller models can run on a single server; the largest need serious hardware.
The bottom line
- Weights are the model: billions of numbers that hold everything it can do.
- Open weight means you can hold those numbers yourself; open source means you also get the data and the recipe.
- Fine-tuning changes the numbers, so the model behaves differently.
- RAG leaves the numbers alone and changes what the model reads when it answers.
Before choosing, ask whether your problem is knowledge or behaviour, how often the content changes, who is allowed to see what, and where the data is allowed to go. Those four answers usually point clearly to the right approach.
I write about AI governance and digital transformation in international organisations. If this guide was useful, the AI governance in the United Nations guide covers how these decisions get approved and audited inside a public institution, and the rest of the writing is in the blog archive.
Written by Shahzad Asghar — Head of Data and Digital Solutions at UN-ESCWA, with 20+ years building AI and data systems across UNHCR, UNICEF, and UNOCHA. His team built UNHCR’s first global IVR appointment system, serving 700,000+ refugees. He created the Last-Mile AI Framework. Read more about this UN AI expert