Skip to main content
← Resources

The Failure Modes Public Bodies Are Being Instructed to Build

Designed-In Blind Spots in UK Government AI and Data Guidance

Anthony Lawton, Fit to Care (Front Foot MI Ltd), United Kingdom

August 2026, revised 8 September 2026. Working paper v1.2. Not peer reviewed. Correspondence: anthony.lawton@ffmi.co.uk

Questions or want to discuss a finding? anthony.lawton@ffmi.co.uk

Abstract

This paper asks a narrower question than most reviews of UK government AI guidance: not what the guidance gets wrong, but what a public-sector organisation still cannot catch when it follows the guidance exactly as written. Three recurring patterns are identified across a structured review of the cross-government AI Playbook and assurance family, data quality and transparency standards, departmental guidance for schools, the NHS and local government, and ICO guidance, together with six further findings drawn from the independent audit record of the National Audit Office and the Public Accounts Committee. Those six are evidenced by that audit record rather than by the author's own reading. The pattern recurs in four of the five document families reviewed, and the exception is reported in full. The pattern is a substitution: the existence of a review, a record or an assurance step is treated as proof that it functions, and nothing downstream checks whether it actually did. Human review is repeatedly required by competence, never by capacity. Assurance is repeatedly delivered by self-attestation, with no independent threshold at which self-report stops being sufficient. Transparency records go stale because nothing requires anyone to refresh them. And the government's headline figure of £45 billion a year from digital adoption is, on the government's own account, an estimate of maximum technical potential carrying no stated timeframe for realisation, while the reporting that does exist cannot show what share of it digital and AI have delivered, which is what the National Audit Office has now asked HM Treasury to put right. A disconfirming sweep is reported alongside the findings: one document tested clean against every hypothesis on its own text, and its weakness sits at adoption rather than design, consistent with what the Ministry of Housing, Communities and Local Government reported in July 2026 about the wider guidance landscape councils navigate. The paper closes by naming what independent audit alone was able to establish that textual analysis could not, and by stating plainly where inference ends and citation begins throughout.

1. What this paper is and is not

This paper does not argue that published UK government AI guidance is badly written or badly intentioned. Most of it is competent. The Local Government Association's procurement guidance in particular is genuinely well designed, and it is treated as such below. The question tested is narrower and less forgiving than "is this guidance good": when a public-sector organisation follows it exactly as written, what failure mode does the guidance's own prescribed architecture leave structurally unguarded? Not what the text gets wrong. What a fully compliant reader still cannot catch.

The distinction matters because most critique of government AI policy treats non-compliance as the risk to manage. This paper's finding is that compliance itself, on the guidance as currently written, does not close the gap that matters. A trust, council, school or department that does everything the guidance asks of it can still deploy a system with no functioning check on whether it works as claimed, and no independent means of finding out until something has already gone wrong.

2. Method

Four parallel research sweeps were run: the cross-government Playbook and assurance family; data quality and the Algorithmic Transparency Recording Standard (ATRS); departmental guidance for schools, the NHS and local government; and ICO guidance together with the National Audit Office and Public Accounts Committee audit reports. Each sweep confirmed current document versions by live retrieval rather than from recall, and each was required to report where a hypothesis was tested and killed rather than force a weak finding. Three such disconfirmations are reported in section 5.

The sweeps were conducted with AI assistance, under a protocol requiring live retrieval of the currently published version of every document. The quotations reproduced in this paper were re-checked against their primary sources on 8 September 2026 using a separate retrieval pass with declared controls, and the corrections arising are incorporated in this version. That check was AI-assisted and is not a substitute for peer review; errors that remain are the author's. Guidance documents were read at the versions listed in the References.

National Audit Office and Public Accounts Committee material was treated as a priority stream, separate from the author's own reading of primary guidance. An independent audit body naming a failure mode is a different order of evidence from a reviewer's own inference, and the two are kept distinct throughout this paper. Findings anchored in NAO or PAC conclusions appear in section 4. Where section 4 sets out the government's own account of the forty-five billion pound figure, it does so from two named government publications, and that is flagged in the text. Findings drawn from the author's own reading of guidance text appear in section 3, grouped into three recurring patterns, with the document and the quotation given so a reader can check the claim against the source. Section 3 does not carry paragraph-level references for every quotation. Each quotation is reproduced in full with its document named, and the versions used are listed in the References, so every one can be located by search in the source. Paragraph-level references will be added to this page as they are completed. Section 3 should be read as three patterns evidenced by the quotations shown, not as an enumerated finding count. Sections 3 and 4 keep the two categories apart throughout.

One qualifying note on currency. The two older NAO and PAC reports (HC 612, 15 March 2024; HC 356, 26 March 2025) are twenty-nine and seventeen months old respectively as at 8 September 2026, in a field where ownership of the adjacent cross-government guidance has moved twice in the same period, from the Cabinet Office to the Department for Science, Innovation and Technology, and on again within that department. The findings drawn from those two reports should be read as dated evidence of a pattern, not as a live measurement of today's state. The third audit source (HC 267, 15 July 2026) is current. The ICO guidance cited in section 3 is a consultation draft: the consultation ran from 31 March to 29 May 2026 and closed, and the guidance is not in force.

3. The substitution that recurs across four of the five families reviewed

Read across cross-government, schools, NHS and data protection guidance, the same substitution appears in a consistent shape: the existence of a check is treated as proof the check functions. A named reviewer is treated as proof review happens. A published record is treated as proof the record is current. A sent alert is treated as proof it was acted on. No document in those four families asks the harder question: does the review clear its queue, does the record still describe the live system, was the alert seen. The fifth family, local government, is the exception, and it is set out in section 5.

Human review, prescribed by competence, never by capacity. The AI Playbook for UK Government tells organisations that they "should fully test the product before deployment, and have robust assurance and regular checks of the live tool in place" and "have systems in place that allow users to report issues and prompt a human review" (Principle 4). Nothing in the text sets a caseload ceiling, a maximum time to review, or an escalation trigger for overload. The same shape recurs in the Department for Education's product safety standards, its separate monitoring standard for safeguarding alerts, and, more narrowly, in the Information Commissioner's Office's draft guidance on automated decision-making, which states that "using ad hoc spot-checking isn't sufficient because some automated decisions won't receive a check and therefore don't have meaningful human involvement" and expects the reviewer to be "suitably trained and qualified to understand the system's logic, outputs, limitations, and risks", but never asks what happens when that trained reviewer faces a decision volume the available time cannot cover. A compliant organisation can name one qualified person, log every decision as reviewed, and pass every written test while the review itself takes thirty seconds or thirty minutes with no way for anyone outside the organisation to tell which.

There is a structural reason this gap cannot be closed by writing a better standard. A duty that applies to every decision, without limit, guarantees the exact situation the ICO's own guidance tries to rule out: the reviewer eventually faces more decisions than the available time can review. The ICO's draft guidance says the human involved "should apply these non-exhaustive criteria every time they make a decision about a person" (draft ADM guidance, consultation closed 29 May 2026); the duty scales with how many decisions the system makes, while the reviewer's time does not scale with it at all. When that gap opens, something gets skipped, by somebody, on criteria nobody wrote down. The question worth putting to any board: when your reviewers cannot clear the queue, who decided what got skipped, and is that written down anywhere?

Assurance delivered by self-attestation, with no stated point at which it stops being enough. DSIT's Introduction to AI Assurance sets out, in the worked table in section 5.1, one route for an organisation to show it has "understood the potential risks of AI systems it is buying", and that route is measured against "(Self) assessment against proprietary framework or responsible AI toolkit". The same column names a UKAS accredited conformity assessment body as the provider, so the worked example is not pure self-attestation. What is left open is the benchmark, which the table allows to be an organisation's own framework, with no threshold above which that stops being acceptable. The AI Management Essentials tool produces a rating that is, by DSIT's own description, "calculated on self-assessment answers", while the same guidance states plainly that "AIME does not provide formal certification", and also says that "in the future, there may be opportunities to explore embedding AIME into public sector procurement frameworks for AI products and services". The NHS's Digital Technology Assessment Criteria is completed by the supplier, and DTAC does not itself verify the supplier's claims against the product. When the form was updated on 24 February 2026, NHS England recorded that "the previous version of the DTAC form required the clinical safety officer named in the Clinical safety section to undertake training provided by NHS Digital" and that "this requirement no longer stands". A general requirement remains, that the officer "must be a clinician, have a current registration with a professional body and be trained in clinical risk management", but the tie to one named course has gone and DTAC has no mechanism for checking either. Under the Department for Education's standards, the requirements are framed overwhelmingly as what the department "expects" of suppliers, including that "a clear risk assessment is conducted for every product to assure safety for educational use", with no verification route attached to any of them. A full-text search of the thirteen-section document on 8 September 2026 returned no instance of accreditation, certification, audit or independent review.

Transparency and quality records that nothing requires anyone to refresh. The Algorithmic Transparency Recording Standard's guidance states that organisations "should update the ATRS template" when "substantive details change", with "substantive" judged by the same team that owns the tool, and its mandatory scope and exemptions policy names only one circumstance in which a record must be updated, retirement. The guidance for public sector bodies gives three further examples, a pilot moving to production, new training datasets, and a change to the surrounding operational process, but none of them is a cadence, and none requires anyone to look again at a record where nothing has been reported as changing. The consequence is visible in a live public record. The transparency record published for QCovid, a clinical risk-scoring tool, under the names the record itself carries, the Department for Health and Social Care and NHS Digital, gives the phase in its header metadata as "Retired" while field 2.6 of the same record gives the status as "Production/run and maintain". The record still shows a worked example "generated on 25 February 2022" and still states "as of April 2022, there are no further planned updates to the tool". Nothing in the framework required any of that to be refreshed, and nothing has been. Record accessed 8 September 2026. The Government Data Quality Framework asks organisations to "benchmark and regularly assess levels of data quality over time", and its companion guidance addresses how often and then declines to answer, saying the frequency "will be specific to your organisation and data". Neither document names an independent checker.

4. What independent audit alone confirms

Six findings in this review are anchored in the published conclusions of the National Audit Office and the Public Accounts Committee rather than in the author's reading of guidance, and they carry more weight for exactly that reason: they were reached and evidenced by an external, statutory audit body through its own process rather than through this review's. Where the sixth finding sets out the government's own derivation of the forty-five billion pound figure, it draws on two published government documents, identified as such in the text, and the audit finding itself remains the National Audit Office's.

No department owns AI adoption accountability. The National Audit Office's March 2024 report states plainly that the government's draft AI adoption strategy "does not set out which of these departments has overall ownership and accountability for its delivery", and that AI-adoption governance sits "largely separate from the cross-government governance structure established to oversee wider AI policy delivery led by DSIT" (HC 612, summary paras 9 to 10).

Assurance runs on self-report at scale. Eighty-seven organisations responded to the NAO's survey, of which thirty-two had deployed AI. Almost half of those thirty-two kept no register of live AI use cases at all. Across all eighty-seven respondents, thirty per cent reported having risk and quality assurance processes that explicitly covered AI risk, and a further forty-six per cent said only that they had plans to put one in place (HC 612, para 3.28). The thirty per cent figure is itself a self-report, unverified by the survey.

Take-up of the transparency register has been low. Only eight of the thirty-two organisations with deployed AI reported being "always or usually compliant" with ATRS (HC 612, para 3.20). DSIT announced in February 2024 that it intended to make the standard mandatory for all government departments "during 2024". Eleven months later, with that deadline passed, the Public Accounts Committee found that "at January 2025, only 33 records had been published" (HC 356, para 12).

A remediation plan lost its funding without anything catching it. Of the seventy-two highest-risk legacy digital systems prioritised in the 2022 to 2025 digital and data roadmap, twenty-one "still lack remediation funding" (HC 356, conclusion 1), and money allocated for legacy remediation "had too often been reallocated elsewhere" (HC 356, para 8). The risk register correctly named the hazard. Nothing caught the point at which a system recorded as funded stopped being funded.

The Committee asked government the same question this paper asks of guidance. Cabinet Office told the Committee that a departmental reorganisation had "pretty comprehensively addressed" the NAO's concerns about accountability and complexity. The Committee, in the same paragraph, did not accept that as settled: "it is early days and we will be looking for more evidence that these changes will address the NAO's concerns around complexity and accountability fully" (HC 356, para 25). The Committee is not asking whether the reorganisation happened. It is asking what changed as a result, which is the harder and more useful question.

Nobody reports the digital and AI share of the headline savings figure, and the auditor's own fix is to require exactly that. The National Audit Office records that "the government expects substantial efficiencies from digital transformation and AI, amounting to £45 billion each year" (HC 267, para 1.8). The Department for Science, Innovation and Technology produced that figure in January 2025 as an estimate that "over £45 billion per year of unrealised savings and productivity benefits, 4-7% of public sector spend, could be achieved through full potential digitisation of public sector services". Asked for its working, the government supplied it to the Science, Innovation and Technology Committee in April 2025: a bottom-up estimate of "the maximum technical productivity gains to the UK public sector annually", built on "a standard assumption of 90% maximum digital uptake", whose benefits "do not have a specific timeframe for realisation, but they are anticipated to be realised over the long term".

So the figure is an estimate of a ceiling, published with no date for realisation. Departments do already report efficiency delivery to HM Treasury under the Government Efficiency Framework, and from 2026-27 they will also report efficiency savings in their annual reports and accounts. What that reporting does not isolate is the digital and AI share, which is the part the £45 billion is made of. The NAO does not test how the £45 billion was derived. What it finds is that "the published efficiency plans do not provide details of how departments derived their expected workforce efficiencies" (HC 267, para 1.6), and its Recommendation 13 asks HM Treasury to update the Government Efficiency Framework so that its guidance covers "how departments will be expected to calculate estimated workforce efficiencies and cost savings from digital adoption, and report realised efficiencies to HM Treasury". An estimate of maximum technical potential, published without a date for realisation and reported through a line that records efficiencies in aggregate but not the digital and AI share, is the same substitution this paper describes in guidance: the figure exists, and the reporting that exists cannot yet show what part of it has been realised.

5. Where the hypothesis did not hold

A structured review that only reports confirming evidence is not a structured review. Three tests were run against this paper's own working hypothesis and did not hold.

The Green Book route that the Playbook points large AI business cases towards already mandates benefits realisation and post-implementation review. Citing "no post-hoc evidence standard" as a gap here would misdescribe a document family that already carries the counter-mechanism this paper is looking for elsewhere.

The concept of a risk register as a compliance artefact, which recurs as a failure mode in other government digital programmes, could not be tested against the Playbook at all, because the Playbook never uses the term. The hypothesis was dropped rather than forced onto text that does not support it.

"Responsibly buying AI", published by the Local Government Association with LOTI and developed with the Information Commissioner's Office and the Equality and Human Rights Commission, is the only document reviewed that tested clean against every hypothesis on its own text. It tells councils that "you also need to review", among other things, "whether the expected benefits are being achieved", sets defined review triggers for equality and data protection impact assessments, and asks "Is human review meaningful?... Are human reviewers adequately trained, and do they have the authority to override decisions?". Where it restates duties that are already legally binding, including the Public Sector Equality Duty and the requirement to consult the ICO where risks cannot be sufficiently reduced, those duties bind whatever the guide says. Its own prompts carry no force of their own, and its weakness sits one level down, at adoption rather than design: recurring review is assigned to a "contract owner" role with no resourcing named anywhere in the guidance. Separately, the Ministry of Housing, Communities and Local Government ran its own engagement sessions with sixty councils, reported in July 2026, and reported, in the department's own words, what those councils told it about the wider guidance landscape they navigate, not this guide specifically, which MHCLG does not sponsor (the LGA publishes it, with LOTI, ICO and EHRC): "there is a lot of guidance, which can be confusing, fragmented and hard to navigate", it is "too high-level and not practical enough to translate into day-to-day decisions", and there is "[a] lack of skills and training across councils to procure, use and govern AI tools". This is a government department reporting, in its own words, what councils said about the guidance landscape they navigate, not this paper inferring it, and not a statement about the LGA guide itself. It is also the most encouraging finding here: where a document is well designed, the problem visibly moves from the page to the resourcing behind it, which is a fixable problem of a different and more tractable kind.

6. Where inference ends and citation begins

Every finding above that quotes guidance text directly is a citation: the source says what is claimed, in the words given. A smaller number of claims are inference rather than citation. Rather than assert that they are marked, this section lists the principal ones. The list is not exhaustive, and any sentence in this paper that asserts an absence, a count across documents, or a consequence is the author's reading rather than a quotation.

The inferential claims in this paper are these. That a compliant organisation could staff one reviewer against a caseload the available time cannot cover, log every decision as reviewed and still pass every written test. That the absence of a caseload ceiling, an escalation trigger, a risk threshold for self-assessment, an accreditation route, a stated cadence or an independent checker means the guidance leaves that failure mode unguarded, rather than that it is guarded elsewhere. That retirement is the only update trigger the ATRS framework operationalises. That the QCovid record's two conflicting status fields matter, rather than being a presentational artefact. That a ceiling estimate with no timeframe, reported without a digital and AI breakdown, cannot show whether the saving is being realised. And that the substitution can be counted across document families at all, rather than being a judgment about a small and unrepresentative corpus. Each of those is this paper's reading of what the published text permits or omits, tested by full-text search of the versions listed in the References, and each is offered so a reader can disagree with the reading while checking the quotations independently.

No public body is named or implied to have acted in bad faith, and no claim is made that any organisation has exploited any gap described here. Where a specific published document or record is discussed by name, including NHS England's DTAC form, the transparency record published for QCovid, the Department for Education's product safety standards and DSIT's AI Management Essentials guidance, the criticism is of the published text and of nothing else. Corrections are welcome and will be made on the same page.

7. Limitations

This is a textual review of published guidance, not an audit of live deployments, and it should not be read as one. The three patterns in section 3 rest on the author's own reading of primary documents. Every quotation is given so a reader can check it independently, and every quotation has been re-verified against its primary source by independent agents, but the reading built on those quotations has not been through academic peer review. Section 3 does not carry paragraph-level references for every quotation, though every quotation is reproduced in full and can be located by search in the document versions listed in the References. Two of the three audit-sourced reports are dated, twenty-nine and seventeen months old, in a policy area whose ownership has moved twice since. Whether the capacity gaps identified in sections 3 and 4 have caused a specific, documented harm is not something this review can establish; the National Audit Office and Public Accounts Committee have not yet published a report at that level of granularity, and until one exists, this paper's claim is that the architecture is unguarded, not that a named harm has occurred through the gap it describes.

8. Discussion

The pattern found here suggests that the practical test for any public body evaluating its own AI governance is not "does a review step exist" but "can the review step actually clear its queue", and not "has assurance been completed" but "what happens in the world if the assurance was wrong". Every gap identified above shares the same shape: a check that can only confirm, never refute, and a record that describes a past state treated as though it described the present one.

This has a direct and unglamorous implication for procurement and board oversight: a supplier's self-attestation, a named reviewer's job title, and a published transparency record are each necessary and each, on their own, worth close to nothing as evidence that a system is safe today. The question worth asking in every case is what, independent of the organisation deploying the system, would notice if the record and the reality diverged. In most of the guidance reviewed here, the honest answer is currently nobody, and the National Audit Office's own recommended fix for the savings figure, a reporting line that would show the digital and AI share of efficiencies actually realised, is evidence that this problem is recognised inside government's own scrutiny function and not only by outside reviewers.

Competing interests. Anthony Lawton is a management accountant and a director of Front Foot MI Ltd, trading as Fit to Care, which provides assurance and analysis services to NHS, local government and multi-academy trust clients. This paper reviews published guidance only. It describes no client engagement and uses no client data. No organisation whose guidance is reviewed here has been a client of Front Foot MI Ltd. Its clients are NHS provider organisations, councils and multi-academy trusts, none of which authored any document reviewed here.

References

National Audit Office (2024). Use of artificial intelligence in government. HC 612, Session 2023-24, 15 March 2024.

Committee of Public Accounts (2025). Use of AI in Government. HC 356, Session 2024-25, 26 March 2025.

National Audit Office (2026). Government workforce planning: lessons learned. HC 267, Session 2026-27, 15 July 2026.

Department for Science, Innovation and Technology (2025). State of digital government review. CP 1251, 21 January 2025.

Department for Science, Innovation and Technology (2025). Letter and methodology note from Feryal Clark MP to the Science, Innovation and Technology Committee on the derivation of the £45 billion figure. 10 April 2025.

HM Treasury (2025). The Government Efficiency Framework. Updated 24 November 2025.

GDS/DSIT (2025). AI Playbook for the UK Government. 10 February 2025.

DSIT (2026). Guidance for using the AI Management Essentials tool. Updated 6 February 2026.

DSIT (2024). Introduction to AI Assurance. 12 February 2024.

Cabinet Office/GDS (2020, updated 2026). The Government Data Quality Framework. 3 December 2020.

GDS and DSIT. Algorithmic Transparency Recording Standard: guidance for public sector bodies (published 5 January 2023) and Mandatory Scope and Exemptions Policy (published 17 December 2024). Both read at the versions live on 8 September 2026.

Department for Education (2025, updated 2026). Generative AI: product safety standards. 22 January 2025, updated 19 January 2026.

NHS England (2026). Digital Technology Assessment Criteria (DTAC), updated form published 24 February 2026, mandatory from 6 April 2026.

Local Government Association, with LOTI, ICO and EHRC (2025). Responsibly buying AI. 16 April 2025.

Information Commissioner's Office (2026). Automated decision-making, including profiling. Draft guidance; consultation ran 31 March to 29 May 2026 and has closed. Not in force.

Department of Health and Social Care and NHS Digital. QCovid algorithm, Algorithmic Transparency Recording Standard record. Accessed 8 September 2026.

Ministry of Housing, Communities and Local Government (2026). From engagement to delivery: supporting responsible AI adoption in councils. MHCLG Digital blog, 13 July 2026.

The quotations in this paper are the evidence, and each can be checked against the document versions listed above. Paragraph-level references will be added to this page as they are completed.

Questions or want to discuss a finding?

anthony.lawton@ffmi.co.uk