The Failure Modes Public Bodies Are Being Instructed to Build
Designed-In Blind Spots in UK Government AI and Data Guidance
Anthony Lawton, Fit to Care (Front Foot MI Ltd), United Kingdom
August 2026. Working paper v1.1. Not peer reviewed. Correspondence: anthony.lawton@ffmi.co.uk
Questions or want to discuss a finding? anthony.lawton@ffmi.co.uk
Abstract
This paper asks a narrower question than most reviews of UK government AI guidance: not what the guidance gets wrong, but what a public-sector organisation still cannot catch when it follows the guidance exactly as written. Twenty-five findings are drawn from a structured review of the cross-government AI Playbook and assurance family, data quality and transparency standards, departmental guidance for schools, the NHS and local government, ICO guidance, and the independent audit record of the National Audit Office and the Public Accounts Committee. Six of the twenty-five are independently evidenced by that audit record rather than by the author's own reading. The pattern that recurs across every document family, regardless of sector, is a substitution: the existence of a review, a record or an assurance step is treated as proof that it functions, and nothing downstream checks whether it actually did. Human review is repeatedly required by competence, never by capacity. Assurance is repeatedly delivered by self-attestation, with no independent threshold at which self-report stops being sufficient. Transparency records go stale by design because nothing requires them to be refreshed. And the government's own headline claim of forty-five billion pounds a year in AI and digital efficiency savings has, by the National Audit Office's own account, no mechanism that checks whether any of it is realised. A disconfirming sweep is reported alongside the findings: one document tested clean against every hypothesis on its own text, and its failure sits at adoption, not design, which the department overseeing that guidance landscape has itself now admitted. The paper closes by naming what independent audit alone was able to establish that textual analysis could not, and by stating plainly where inference ends and citation begins throughout.
1. What this paper is and is not
This paper does not argue that published UK government AI guidance is badly written or badly intentioned. Most of it is competent. The Local Government Association's procurement guidance in particular is genuinely well designed, and it is treated as such below. The question tested is narrower and less forgiving than "is this guidance good": when a public-sector organisation follows it exactly as written, what failure mode does the guidance's own prescribed architecture leave structurally unguarded? Not what the text gets wrong. What a fully compliant reader still cannot catch.
The distinction matters because most critique of government AI policy treats non-compliance as the risk to manage. This paper's finding is that compliance itself, on the guidance as currently written, does not close the gap that matters. A trust, council, school or department that does everything the guidance asks of it can still deploy a system with no functioning check on whether it works as claimed, and no independent means of finding out until something has already gone wrong.
2. Method
Five parallel research sweeps were run: the cross-government Playbook and assurance family; data quality and the Algorithmic Transparency Recording Standard (ATRS); departmental guidance for schools, the NHS and local government; ICO guidance together with the National Audit Office and Public Accounts Committee audit reports; and US federal AI policy as contrast only. Each sweep was required to confirm current document versions by live search rather than by memory, and each was required to report where a hypothesis was tested and killed rather than force a weak finding. Three such disconfirmations are reported in section 5.
National Audit Office and Public Accounts Committee material was treated as a priority stream, separate from the author's own reading of primary guidance. An independent audit body naming a failure mode is a different order of evidence from a reviewer's own inference, and the two are kept distinct throughout this paper. Findings drawn from NAO or PAC reports appear in section 4 and are independently evidenced; findings drawn from the author's own reading of guidance text appear in section 3, with the exact document, section and quotation given so a reader can check the claim against the source directly. Nineteen findings sit in the second category; six sit in the first. Sections 3 and 4 keep the two categories apart throughout.
One qualifying note on currency. The NAO and PAC evidence base (HC 612, March 2024; HC 356, March 2025) is now fifteen to seventeen months old at the time of writing, in a field where ownership of the adjacent cross-government guidance has already moved twice in the same period, from the Cabinet Office to the Department for Science, Innovation and Technology, and on again within that department. The six audit-sourced findings below should be read as dated evidence of a pattern, not as a live measurement of today's state.
3. The substitution that recurs across every document family
Read across cross-government, schools, NHS, local government and data protection guidance, the same substitution appears with almost no variation in shape: the existence of a check is treated as proof the check functions. A named reviewer is treated as proof review happens. A published record is treated as proof the record is current. A sent alert is treated as proof it was acted on. Nowhere in the guidance reviewed does a document ask the harder question: does the review clear its queue, does the record still describe the live system, was the alert seen.
Human review, prescribed by competence, never by capacity. The AI Playbook for UK Government requires that organisations "fully test the product before deployment, and have robust assurance and regular checks of the live tool in place" and "have systems in place that allow users to report issues and prompt a human review" (Principle 4). Nothing in the text sets a caseload ceiling, a maximum time to review, or an escalation trigger for overload. The same shape recurs in the Department for Education's product safety standards, its separate monitoring standard for safeguarding alerts, and, more narrowly, in the Information Commissioner's Office's March 2026 draft guidance on automated decision-making, which explicitly bans "ad hoc spot checks" and requires a "sufficiently trained" reviewer, but never asks what happens when that trained reviewer faces an arbitrarily high decision volume. A compliant organisation can name one qualified person, log every decision as reviewed, and pass every written test while the review itself takes thirty seconds or thirty minutes with no way for anyone outside the organisation to tell which.
There is a structural reason this gap cannot be closed by writing a better standard. A duty that applies to every decision, without limit, guarantees the exact situation the ICO's own guidance tries to rule out: the reviewer eventually faces more decisions than the available time can review. The ICO's draft guidance is explicit that "this assessment must happen every time a decision is made about a person" (ADM guidance, March 2026); the duty scales with how many decisions the system makes, while the reviewer's time does not scale with it at all. When that gap opens, something gets skipped, by somebody, on criteria nobody wrote down. The question worth putting to any board: when your reviewers cannot clear the queue, who decided what got skipped, and is that written down anywhere?
Assurance delivered by self-attestation, with no stated point at which it stops being enough. DSIT's Introduction to AI Assurance sets out, in the worked table immediately before section 5.2, three legitimate ways to show an organisation understands the risk of the AI it is buying, one of which is "(self) assessment against proprietary framework or responsible AI toolkit", with no risk threshold above which self-assessment is no longer acceptable. The AI Management Essentials tool produces a rating that is, by DSIT's own description, "calculated on self-assessment answers", while the same guidance states plainly that it "does not provide formal certification" and simultaneously floats "embedding AIME into public sector procurement frameworks" as a live possibility. The NHS's Digital Technology Assessment Criteria is completed by the manufacturer, not verified against the manufacturer's claims, and its February 2026 reform removed one of the few externally verified elements that existed, the requirement for the named clinical safety officer to complete independent training. Edtech suppliers self-attest to their own risk assessment under the Department for Education's standards, with no accreditation scheme named anywhere in the thirteen-standard document.
Transparency and quality records that go stale by design, not by neglect. The Algorithmic Transparency Recording Standard's guidance states that organisations "should update the ATRS template" when "substantive details change", with "substantive" judged by the same team that owns the tool, and its mandatory scope and exemptions policy operationalises exactly one update trigger across the whole framework, retirement. The consequence is visible in a live public record: the Department of Health and Social Care and NHS Digital's transparency record for QCovid, a clinical risk-scoring tool, is now marked "Phase: Retired", yet the same record still shows a worked example dated February 2022, still states "as of April 2022, there are no further planned updates to the tool", and, in its own field 2.6, still describes the tool's status as "Production/run and maintain": the record contradicting itself about whether the tool is even live, with a single log entry since. The Government Data Quality Framework asks organisations to "benchmark and regularly assess levels of data quality over time" with no cadence, no independent checker and no tooling specified anywhere in the document.
4. What independent audit alone confirms
Six findings in this review come not from the author's reading of guidance, but from the National Audit Office and the Public Accounts Committee, and they carry more weight for exactly that reason: an external, statutory audit body reached them independently of this review and of any framework it used to organise its own thinking.
No department owns AI adoption accountability. The National Audit Office's March 2024 report states plainly that the government's draft AI adoption strategy "does not set out which of these departments has overall ownership and accountability for its delivery", and that AI-adoption governance sits "largely separate from the cross-government governance structure established to oversee wider AI policy" (HC 612, sections 9 to 10).
Assurance runs on self-report at scale. Of the thirty-two organisations with deployed AI that responded to the NAO's survey, almost half kept no register of live AI use cases at all; only thirty per cent of eighty-seven respondents reported having risk and quality assurance processes that explicitly covered AI risk, and a further forty-six per cent said only that they had plans to put one in place (HC 612, section 3.28). The thirty per cent figure is itself a self-report, unverified by the survey.
The transparency register that was meant to fix this is barely used. Only eight of thirty-two deployers reported being "always or usually compliant" with ATRS; the Public Accounts Committee found that, a full year after DSIT told the NAO it intended to make the standard mandatory "during 2024", only thirty-three records had been published (HC 356, section 12).
A funded remediation plan had its money quietly moved elsewhere. Twenty-one of the seventy-two highest-risk legacy technology systems identified in the 2022 to 2025 digital and data roadmap remained unfunded at the time of the Committee's report. Money earmarked for their remediation "had too often been reallocated elsewhere" (HC 356, section 8, conclusion 1). The risk register correctly named the hazard. Nothing caught the point at which the register's claim of "funded" and the true state of "defunded" diverged.
The Committee applies, to government's own account of itself, exactly the test this paper applies to guidance. Cabinet Office told the Committee that a departmental reorganisation had "pretty comprehensively addressed" the NAO's concerns about accountability and complexity. The Committee's response, in the same paragraph, refused the claim: "it is early days and we will be looking for more evidence that these changes will address the NAO's concerns... fully" (HC 356, section 25). The Committee is not asking whether the reorganisation happened. It is asking what changed as a result, which is the harder and more useful question, arrived at through parliamentary scrutiny practice with no reference to any framework this paper uses.
The headline savings figure has no mechanism that checks whether it is ever realised, and the auditor's own fix is to build exactly that missing check. Government's stated efficiency claim from AI and digital adoption is forty-five billion pounds a year. The NAO's July 2026 report on workforce planning finds that "published efficiency plans do not provide details of how departments derived their expected workforce efficiencies", and its Recommendation 13 asks HM Treasury to require departments to "calculate estimated workforce efficiencies and cost savings from digital adoption, and report realised efficiencies to HM Treasury", a reporting loop that does not currently exist (HC 267, 15 July 2026). This is the single clearest instance in the whole review of an auditor identifying, in its own words and for its own reasons, that a claimed benefit needs a mechanism to check whether it actually landed, and that no such mechanism exists yet.
5. Where the hypothesis did not hold
A structured review that only reports confirming evidence is not a structured review. Three tests were run against this paper's own working hypothesis and did not hold.
The Green Book route that the Playbook directs large AI business cases through already mandates benefits realisation and post-implementation review. Citing "no post-hoc evidence standard" as a gap here would misdescribe a document family that already carries the counter-mechanism this paper is looking for elsewhere.
The concept of a risk register as a compliance artefact, which recurs as a failure mode in other government digital programmes, could not be tested against the Playbook at all, because the Playbook never uses the term. The hypothesis was dropped rather than forced onto text that does not support it.
The Local Government Association's "Responsibly buying AI" guidance is the strongest document reviewed on its own terms. It explicitly requires checking whether expected benefits were achieved, sets defined review triggers for equality and data protection impact assessments, and asks whether human review is "meaningful... trained". Every hypothesis tested against its own text came back clean. Its real failure sits one level down, at adoption rather than design: recurring review is assigned to a "contract owner" role with no funded capacity attached anywhere in the guidance, and nothing in the document is mandatory. Separately, the Ministry of Housing, Communities and Local Government ran its own engagement sessions with sixty councils, reported in July 2026, and found, in the department's own words, that the wider guidance landscape councils navigate, not this guide specifically, which MHCLG does not sponsor (LGA, ICO, EHRC and LOTI publish it jointly): "there is a lot of guidance, which can be confusing, fragmented and hard to navigate", it is "too high-level and not practical enough to translate into day-to-day decisions", and there is "[a] lack of skills and training across councils to procure, use and govern AI tools". This is a government department naming, in its own words, an adoption gap across the guidance landscape it oversees, not this paper inferring it, and not an admission about the LGA guide itself. It is also, on the evidence reviewed here, the single most encouraging finding in the whole study: where a document is well designed, the problem visibly moves from the page to the resourcing behind it, which is a fixable problem of a different and more tractable kind.
6. Where inference ends and citation begins
Every finding above that quotes guidance text directly is a citation: the source says what is claimed, in the words given, at the section given. A smaller number of claims in this paper are inference rather than citation, and they are marked as such here rather than left to blur. The guidance requiring a named reviewer is a citation. The claim that a compliant organisation could staff one reviewer against an arbitrarily high caseload and pass every written test is this paper's analysis of what the text permits, not a statement that any organisation has in fact done this. No public body is named or implied to have exploited any of these gaps; the target throughout is the architecture the guidance prescribes, not any organisation that has followed it in good faith, which every organisation reviewed here appears to have done.
7. Limitations
This is a textual review of published guidance, not an audit of live deployments, and it should not be read as one. Nineteen of the twenty-five findings rest on the author's own reading of primary documents; while every citation is given so a reader can check it independently, that reading has not itself been through independent peer review at the time of writing. The six audit-sourced findings are strong evidence of a pattern in 2024 and 2025 but are explicitly dated, in a policy area that has moved twice since. Whether the capacity gaps identified in sections 3 and 4 have caused a specific, documented harm is not something this review can establish; the National Audit Office and the Public Accounts Committee have not yet published a report at that level of granularity, and until one exists, this paper's claim is that the architecture is unguarded, not that a named harm has occurred through the gap it describes.
8. Discussion
The pattern found here suggests that the practical test for any public body evaluating its own AI governance is not "does a review step exist" but "can the review step actually clear its queue", and not "has assurance been completed" but "what happens in the world if the assurance was wrong". Every gap identified above shares the same shape: a check that can only confirm, never refute, and a record that describes a past state treated as though it described the present one.
This has a direct and unglamorous implication for procurement and board oversight: a supplier's self-attestation, a named reviewer's job title, and a published transparency record are each necessary and each, on their own, worth close to nothing as evidence that a system is safe today. The question worth asking in every case is what, independent of the organisation deploying the system, would notice if the record and the reality diverged. In most of the guidance reviewed here, the honest answer is currently nobody, and the National Audit Office's own recommended fix for the largest of these gaps, an actual reporting loop back from claimed savings to verified ones, is the plainest evidence that this problem is now recognised inside government's own scrutiny function, not only by outside reviewers.
The author of this paper has built and run, for a small consultancy's own multi-agent AI operations, a verification discipline aimed at exactly this substitution: treating the next observable state of the world, not a step's own report of itself, as the only valid evidence that work has actually happened. That discipline, developed under the name Job Memory System, is mentioned here once, as the practical response this finding provoked in our own work, not as a product being sold through this paper's argument. The gap this paper documents in government guidance is real independent of anything we have built to answer it in our own operation, and the guidance would be no less unguarded if we had built nothing at all.
References
National Audit Office (2024). Use of artificial intelligence in government. HC 612, Session 2023-24, 15 March 2024.
Committee of Public Accounts (2025). Use of AI in Government. HC 356, Session 2024-25, 26 March 2025.
National Audit Office (2026). Government workforce planning: lessons learned. HC 267, Session 2026-27, 15 July 2026.
GDS/DSIT (2025). AI Playbook for the UK Government. 10 February 2025.
DSIT (2026). Guidance for using the AI Management Essentials tool. Updated 6 February 2026.
DSIT (2024). Introduction to AI Assurance. 12 February 2024.
Cabinet Office/GDS (2020, updated 2026). The Government Data Quality Framework. 3 December 2020.
GDS (2023, updated 2025). Algorithmic Transparency Recording Standard: guidance for public sector bodies and Mandatory Scope and Exemptions Policy. 5 January 2023, updated 8 May 2025 and 17 December 2024 respectively.
Department for Education (2025, updated 2026). Generative AI: product safety standards. 22 January 2025, updated 19 January 2026.
NHS England (2026). Digital Technology Assessment Criteria (DTAC), Form v2.0, effective 24 February 2026.
Local Government Association, ICO, EHRC and LOTI (2025). Responsibly buying AI. 16 April 2025.
Ministry of Housing, Communities and Local Government (2026). From engagement to delivery: supporting responsible AI adoption in councils. 13 July 2026.
Information Commissioner's Office (2026). Draft guidance on automated decision-making. Consultation 31 March to 29 May 2026.
Full working detail, including every quotation, section reference and the complete disconfirming sweep, is held in the author's evidence library and is available on request.
Questions or want to discuss a finding?
anthony.lawton@ffmi.co.uk