Why Government AI Projects Fail

Last updated: 2026-08-15

Government AI projects rarely fail because the model was inadequate. They fail because the institution could not supply data of sufficient quality, because no single person owned the outcome, because the pilot was never designed to survive the operating environment, or because governance arrived after an incident rather than before launch.

The recurring causes of failed public-sector AI projects, and the portfolio discipline that prevents them: knowing what you run before deciding what to build.

Why Government AI Projects Fail

Public-sector AI failure is rarely dramatic. A pilot demonstrates well, a budget is committed, and eighteen months later the system is either unused or quietly running without anyone able to explain how it decides. The post-mortem, when there is one, usually blames the technology. The technology is almost never the cause.

The four recurring causes

The data could not support the task. This is the most common and the most predictable. The records exist, but they are duplicated across systems, incomplete in exactly the fields the model needs, undocumented so nobody can say how they were collected, or held under a legal basis that does not cover the intended use. None of these is repaired by a better model; the model inherits them and expresses them at scale. The remedy is to assess readiness for the specific use case before approval rather than after disappointment, which is what the AI readiness assessment exists to do.

Nobody owned the outcome. A steering committee is not an owner. When a system's business owner is "the department", there is no individual who will stop the project when the evidence turns, and no individual who carries the consequence when it goes wrong. Projects with distributed accountability do not get cancelled; they get quietly extended.

The pilot never met the operating environment. Pilots run under conditions the production system will never see: clean data prepared for the pilot, the most engaged users, good connectivity, one language, and a team watching. Nothing about that predicts behaviour in a field office at month nine. The discipline for closing that gap is set out in the Last-Mile AI framework.

Governance arrived after an incident. Where approval, tiering, and a stopping condition are added retrospectively, the institution spends its scarce governance capacity on defending a system rather than on deciding whether to keep it.

The precondition nobody budgets for

An institution that cannot list its systems cannot govern them, and cannot sensibly decide what to build next. Most public bodies do not have that list. Applications accumulate through projects, grants, and departmental initiative, each rational in isolation, and the estate becomes something no single person can describe.

A digital discovery inventory is the work of establishing what actually exists: every application and database, who owns it, what it holds, whether it is still used, what depends on it, and a lifecycle verdict, which is a decision to keep, consolidate, replace, or retire each item.

A discovery of this kind across UNESCWA inventoried 181 databases and 145 portals and assigned each a lifecycle verdict. The output secured director-level endorsement to proceed to remediation, which is the point: an inventory is not documentation, it is the evidence base that makes a consolidation decision defensible.

How to rationalise a portfolio

Rationalisation fails when it is run as a cost exercise, because every system has a constituency and cost arguments invite negotiation. It works when it is run as a decision exercise.

Inventory first, judgement second. Establish what exists before anyone argues about what should go. Mixing the two produces a list shaped by who was in the room.

Assign a lifecycle verdict to every item, with no null values. Keep, consolidate, replace, retire. An item with no verdict is an item that survives by default, which is how estates grow.

Attribute an owner to every item, by name. Systems without owners are the ones that turn out to be load-bearing at the worst moment.

Sequence by dependency, not by size. The satisfying win is retiring the largest system. The safe sequence retires the ones nothing else depends on.

Publish the verdicts and let them be challenged before acting. A verdict that survives challenge is executable. A verdict imposed quietly is reversed the first time someone senior notices.

What this means before starting an AI project

Three questions, answered honestly, remove most failures before they consume a budget. Does the data for this specific use case exist at a quality that supports the task, verified by inspecting a sample rather than by asking the system owner? Is there one named individual who will carry the outcome and can stop the work? And has the pilot been designed to run under the conditions the production system will face, including the connectivity, languages, and staff turnover of the place it will actually operate?

A use case that cannot answer these is not ready, and the four-axis scoring sheet is the mechanism for saying so before commitment rather than after. The wider governance context is on public sector AI governance.

Written by Shahzad Asghar, Head of Data and Digital Solutions at UNESCWA. See all articles, the playbooks, and the templates.