Why AI Pilots Fail in Government and UN Agencies
AI pilots fail in government and UN agencies when they prove a model but not an operating model. A responsible pilot needs a named owner, lawful and documented data, language testing, decision rights, procurement control, a stopping condition, and a path from prototype to production before the demo becomes an institutional dependency.
Published 2026-08-17 · By Shahzad Asghar
Most AI pilots do not fail because the model is weak. They fail because nobody designed the path from demo to institution.
I have seen AI pilots impress a room and then disappear before they touched an operational workflow. The demo worked. The slides were convincing. The prototype answered questions, classified documents, summarized complaints, or generated dashboards. Then the project slowed down, moved into a committee, lost its sponsor, failed a data review, or remained permanently under evaluation.
That pattern is common in government and United Nations environments. It is also misunderstood. The problem is rarely that public institutions are too slow or that AI vendors are too ambitious. The problem is that a pilot is often treated as a technical experiment when it is actually an institutional test.
An AI pilot in a public institution has to prove more than model performance. It has to prove ownership, lawful data use, workflow fit, language coverage, procurement control, risk management, and a route to production.
If those conditions are missing, the pilot is not really a pilot. It is a demo.
Key takeaways
- AI pilots fail when they prove a model but not an operating model.
- Public institutions need decision rights before tools.
- Data readiness is usually weaker than model capability.
- A pilot without an owner, budget line, risk tier, and stopping condition is not a pilot. It is a demo.
- The most important question is not whether the model works. It is whether the institution can responsibly use the output.
1. The pilot solves a presentation problem, not an operational problem
Many AI pilots begin with a question like: Can we use AI for this?
That is the wrong starting point.
The better question is: Which decision, service, or workflow is failing today, and what would improve if AI worked?
A chatbot that answers policy questions may look impressive in a meeting. But if staff do not know when to trust it, where the source documents come from, who updates the knowledge base, or how errors are reported, it will not change the work.
A document classifier may produce attractive accuracy numbers. But if case officers still need to review every file from the beginning, and the classification does not change routing, triage, prioritization, or workload, the model has not improved the operation.
Public institutions do not need more AI demonstrations. They need systems that change a defined workflow without increasing risk. That is the operating discipline behind public sector AI governance and AI governance in the United Nations.
A good pilot starts with the operational pain:
- People are waiting too long for a service.
- Staff are manually reading thousands of records.
- Decision makers cannot see trends early enough.
- Feedback from communities is arriving faster than teams can classify it.
- Data errors are weakening every downstream process.
If the pilot cannot name the workflow it will change, it is not ready.
2. Nobody owns the decision after the vendor leaves
A vendor can build a prototype. A consultant can run a workshop. A lab can produce a proof of concept.
But only the institution can own the decision.
This is where many AI pilots fail. During the pilot, responsibility is distributed across innovation, IT, data protection, procurement, programme teams, and senior leadership. Everyone is interested. Nobody is accountable.
That structure works during a demo because the consequences are low. It fails in production because an AI system eventually affects a real workflow, a real staff member, a real beneficiary, or a real public decision.
Before an AI pilot begins, three roles should be named:
- Business owner: the person accountable for the operational outcome.
- Technical owner: the person accountable for the system, data flow, integration, monitoring, and security.
- Risk owner: the person with authority to approve, pause, or stop the use case.
If those names do not exist, the pilot will drift.
Ownership cannot be assigned after launch. By then, the system already has users, expectations, and political weight. The hard decisions must be made before the prototype becomes popular.
3. The data is not ready, documented, or lawful for the intended use
AI inherits the condition of the data underneath it.
This is one of the most common reasons pilots fail in government and UN agencies. The institution has data, but the data is fragmented, duplicated, incomplete, undocumented, or collected for a purpose that does not support the proposed AI use.
The model may be capable. The dataset may not be.
Common data problems include:
- No authoritative source.
- Multiple systems holding conflicting records.
- Missing fields in exactly the place the model needs signal.
- Poor metadata and unclear definitions.
- Sensitive personal data mixed with operational data.
- Retention rules that were never enforced.
- Data-sharing agreements that do not cover AI processing.
- No clear legal basis for the intended inference.
In humanitarian and public-sector settings, this is not just a technical issue. It is a protection and trust issue.
If the people described in the data cannot meaningfully consent, appeal, or choose another provider, then data responsibility becomes part of the system design. Data minimization, access control, retention, audit logs, and purpose limitation are not compliance decorations. They are safeguards.
This is why AI-ready data should be assessed before the model is selected. A pilot that skips data readiness may still produce outputs. It just cannot responsibly move into production.
4. The model works in English but fails in local languages or dialects
Language is one of the easiest risks to underestimate.
Many AI pilots are tested in English because the vendor interface, project documents, and evaluation examples are in English. Then the system is expected to serve Arabic speakers, refugees, field staff, government counterparts, or communities using local dialects and mixed-language communication.
That is where performance changes.
A model that performs well in English may struggle with:
- Arabic dialects.
- Transliteration.
- Code-switching between languages.
- Informal speech.
- Voice notes.
- Low-resource languages.
- Names and locations with multiple spellings.
- Humanitarian terminology.
- Local administrative terms.
For UN agencies and governments, language performance is not a user-experience detail. It is an equity issue. If the model performs better for English-speaking staff than for affected communities, the pilot may deepen the gap it was meant to close.
Language testing should happen before approval, not after deployment.
The question is not: Does the model support Arabic?
The question is: Has it been tested on the Arabic, dialects, names, documents, and speech patterns our users actually produce?
This is especially important in the Arab region, where AI governance has to account for Arabic-language performance, not only translated policy language.
5. Governance arrives after the pilot, not before it
Many institutions treat governance as the phase that comes after experimentation.
That is backwards.
Governance should begin before the pilot because the pilot itself creates facts on the ground: data is moved, vendors gain access, staff expectations form, outputs influence judgment, and leadership begins to see the system as real.
For public institutions, AI governance does not need to start with a large framework. It needs to start with decision rights.
Before a pilot begins, the institution should know:
- Who may approve the pilot?
- Who may approve production use?
- What risk tier does the use case fall into?
- What data may be used?
- What data may not be used?
- Who reviews the outputs?
- What happens when the system is uncertain?
- What must be logged?
- What is the stopping condition?
- Who can switch the system off?
A principles document cannot answer these questions by itself. Governance becomes real only when it produces operational decisions.
If governance is added after the pilot succeeds, it will be seen as a blocker. If governance is built into the pilot from day one, it becomes part of delivery. The AI governance playbook turns these questions into an approval gate, model register, and risk-tiering process.
6. The procurement contract does not give the institution enough control
Public institutions often buy AI systems they cannot fully inspect.
That is not always avoidable. Governments and UN agencies will use commercial platforms, cloud services, and vendor-built tools. But when the system is procured rather than built, the contract becomes part of governance.
Many pilots fail because procurement focused on capability and price, but not control.
The institution needs answers to practical questions:
- Where does inference happen?
- Under which jurisdiction is data processed?
- Is institutional data used to train or improve the vendor model?
- Can the system be evaluated on local data before purchase?
- Can logs be exported?
- Can decisions be audited?
- Can the institution change vendors without losing records?
- What happens if the vendor changes the model?
- What happens if the vendor changes pricing?
- Can the institution suspend use immediately?
A pilot may look successful while hiding future dependency.
The risk is not only vendor lock-in. It is operational lock-in: staff redesign work around a tool the institution cannot govern, explain, or exit.
In public institutions, a good AI contract should preserve the right to understand, evaluate, pause, and leave.
7. There is no path from prototype to production
A prototype can be run by a small team. A production system cannot.
Production needs support, security, training, monitoring, documentation, budget, integration, and ownership. If those are not planned early, the pilot will stop at the exact moment it becomes useful.
The path to production should be visible from the beginning.
That means answering:
- Where will the system run?
- Who will maintain it?
- Who pays for it after the pilot?
- Which workflow will absorb it?
- What system will it integrate with?
- Who trains users?
- Who handles incidents?
- Who updates prompts, models, or knowledge bases?
- What metrics define success after 30, 60, and 90 days?
- What happens if the pilot works?
The last question is often ignored.
Many institutions know what happens if a pilot fails. It is closed. But they do not know what happens if it succeeds. There is no budget line, no team, no procurement route, no security approval, and no business owner ready to absorb it.
A pilot without a production path creates frustration. It proves value and then strands it.
Demo questions vs production questions
| Demo question | Production question |
|---|---|
| Does the model work? | Who is accountable when it is wrong? |
| Is the output impressive? | Can the workflow absorb it? |
| Can the vendor show a benchmark? | Has it been tested on our data and languages? |
| Can users try it? | Can affected people appeal or reach a human? |
| Can it launch? | Can we switch it off? |
| Is the interface polished? | Is the data flow lawful, secure, and logged? |
| Does leadership like it? | Who owns it after the pilot budget ends? |
| Can it automate the task? | Should this task be automated at all? |
Before approving an AI pilot, ask these 10 questions
- What decision, service, or workflow will change if the pilot succeeds?
- Who owns the operational outcome?
- What data will be used, and where is the authoritative source?
- What is the legal basis for using that data in this way?
- Which languages, dialects, and user groups must the system support?
- What happens when the model is uncertain or wrong?
- Who reviews the output before it affects a person or public decision?
- What is the stopping condition?
- What does success mean after 90 days?
- What budget, role, and process carry the pilot into production?
If the institution cannot answer these questions, the pilot should not be approved as an operational pilot. It may still be useful as research, learning, or market exploration. But it should be named honestly.
What makes an AI pilot worth running?
A good AI pilot in government or the UN system has five characteristics.
First, it is tied to a real operational problem. It does not begin with a tool. It begins with a workflow that needs to improve.
Second, it has a named owner. Someone is accountable for the outcome, not just the experiment.
Third, it uses data that is ready enough, lawful enough, and documented enough to support the task.
Fourth, it is governed before it is celebrated. The approval route, risk tier, review process, and stopping condition are known.
Fifth, it has a path to production. If the pilot works, the institution knows what happens next.
That does not guarantee success. But it prevents the most common failure: building something impressive that the institution cannot responsibly use.
For institutions starting this work, the practical next step is not another demo. It is a short intake process, risk tier, data readiness check, and production decision. The AI use case discovery playbook is the fastest way to do that.
FAQ
Why do AI pilots fail in government?
AI pilots in government usually fail because they prove a technical capability without proving institutional readiness. The model may work, but the data is not ready, the workflow is unclear, the owner is missing, the legal basis is weak, or the procurement route does not allow production use.
Why do AI pilots fail in UN agencies?
AI pilots in UN agencies often fail when they do not account for multilingual users, sensitive data, protection risks, auditability, mandate constraints, and field conditions. A system that works in a headquarters demonstration may not work in a refugee operation, regional commission, or country office without strong governance and operational design.
What is the difference between an AI pilot and an AI demo?
An AI demo shows that a tool can do something. An AI pilot tests whether an institution can responsibly use that tool in a real workflow. A pilot needs an owner, data controls, success metrics, user testing, risk review, and a path to production.
How should public institutions evaluate AI pilots?
Public institutions should evaluate AI pilots by asking whether the system improves a defined workflow, uses lawful and documented data, performs across required languages, protects affected people, can be audited, and has a responsible owner. Model performance is only one part of the evaluation.
What governance is needed before an AI pilot?
Before an AI pilot begins, the institution should define decision rights, risk tier, data permissions, review requirements, logging, escalation, and stopping conditions. Governance should decide who can approve the pilot, who can approve production, and who can switch the system off.
How do you move an AI pilot into production?
To move an AI pilot into production, the institution needs a business owner, technical owner, budget, security approval, user training, monitoring plan, support model, procurement route, and success metrics. The production path should be designed before the pilot starts, not after it succeeds.
Final verdict
The model is rarely the hard part.
The hard part is building the institutional machinery around it: ownership, data, controls, language coverage, procurement, monitoring, and the authority to stop.
That is why AI pilots fail in government and UN agencies. They are treated as model experiments when they are really tests of institutional readiness.
A successful AI pilot does not only answer: Can this technology work?
It answers the harder question:
Can this institution use it responsibly, repeatedly, and at scale?
For more on the delivery method behind this view, read the Last-Mile AI Framework, the UN AI expert profile, and the practical AI playbooks.
Written by Shahzad Asghar — Head of Data and Digital Solutions at UN-ESCWA, with 20+ years building AI and data systems across UNHCR, UNICEF, and UNOCHA. His team built UNHCR’s first global IVR appointment system, serving 700,000+ refugees. He created the Last-Mile AI Framework. Read more about this UN AI expert