Skip to main content
Fit to CareCare into Action
← Resources

Did they answer? A public record of Prime Minister's Questions

Open codes, session scoreboard, and monthly party balance from January 2016

1. What this is

Every Wednesday the Prime Minister answers questions in the House of Commons. This page records, for every question since January 2016, whether the question was answered: in full, in part, or not at all, with the cases where the two readers could not agree shown separately. It records the same thing for every question asked, by every MP, under the same rules.

Fit to Care, a public sector consultancy, publishes it as a public service. No party, person or client commissioned or paid for it. It does not say what anyone meant or intended. It counts and it quotes. Every quoted exchange in any article links to its Hansard speech. This page shows the totals; the view that shows every question with its own link follows in a later release.

Coverage: 329 sessions, 9,480 questions and answers, 6 January 2016 to 9 September 2026. 0 sessions are held back and listed in section 6 with the reason.

2. The scoreboard

Colours are codes for answers, never for people. Each Prime Minister is shown as the share of their answers under each code, per session, with the range the data supports.

Prime MinisterFullQ-FPartialQ-PPartial or nonepartial_or_nonNoneQ-NJustified challengeQ-JConstraintQ-CInterruptedQ-IUnclearQ-UNot judgedNC
Cameron48.3 (44.7 to 52)21.9 (18.8 to 25)5.3 (3.6 to 6.9)15.8 (13 to 18.8)0.7 (0 to 1.6)0 (0 to 0)1.2 (0.3 to 2.3)0 (0 to 0)6.7 (4.8 to 8.6)
May36.3 (34.6 to 38.1)23.5 (21.9 to 25.1)7.8 (6.8 to 8.8)24.7 (23.1 to 26.3)1.1 (0.8 to 1.5)0.3 (0 to 0.6)0.9 (0.5 to 1.3)0 (0 to 0)5.4 (4.6 to 6.3)
Johnson36.2 (34.3 to 38.2)23.1 (21.4 to 24.8)6.5 (5.4 to 7.6)27.9 (25.6 to 30.3)1.1 (0.7 to 1.5)0.3 (0.1 to 0.5)0.2 (0 to 0.4)0.1 (0 to 0.3)4.6 (3.8 to 5.5)
Trusstoo few sessions to compare30.3 (26.7 to 33.3)23.6 (16.7 to 27.6)4.4 (0 to 10)34.8 (33.3 to 36.7)1.1 (0 to 3.4)0 (0 to 0)0 (0 to 0)0 (0 to 0)5.6 (3.4 to 6.7)
Sunak32.3 (29.9 to 34.8)18.3 (16.5 to 20.3)7.4 (6 to 8.8)35.6 (33.2 to 38)0.9 (0.4 to 1.5)0.5 (0.1 to 1)0.3 (0.1 to 0.6)0.1 (0 to 0.2)4.5 (3.6 to 5.4)
Starmer33.3 (31 to 35.5)22.7 (20.2 to 25.3)8.3 (7 to 9.7)30.3 (27.7 to 33)0.9 (0.5 to 1.4)0.3 (0.1 to 0.5)0.2 (0.1 to 0.5)0 (0 to 0)4 (3 to 4.9)
Burnhamtoo few sessions to compare39.7 (33.3 to 46.2)32.1 (30.8 to 33.3)1.9 (0 to 3.8)18.1 (15.4 to 20.8)0 (0 to 0)0 (0 to 0)0 (0 to 0)0 (0 to 0)8.2 (3.8 to 12.5)
All PMs35.8 (34.9 to 36.9)22.4 (21.4 to 23.3)7.2 (6.7 to 7.8)27.8 (26.9 to 28.8)1 (0.8 to 1.3)0.3 (0.2 to 0.4)0.5 (0.3 to 0.6)0 (0 to 0.1)4.9 (4.4 to 5.3)

The columns are the answer codes of the rulebook the record was scored with, in that rulebook's own published words (section 4), as written by the export. Each row adds to 100 across all nine columns. Two columns are easy to miss: "Partial or no answer" is where the two readers split between partial and none, and "Could not be judged" is where the answer could not be assessed, including every answer that one reader read as full and the other as no answer. Neither is ever resolved by hand. Shares are averages per session. The range in brackets is a 95 per cent bootstrap interval over sessions. Prime Ministers with fewer than five sessions are shown but not compared.

By year:

YearSessionsFull answerNo answer
20163243.4 (40.6 to 46.2)18.3 (16 to 20.5)
20172737.9 (34.4 to 41.4)22.6 (20 to 25.4)
20183437.3 (34.7 to 40.4)25 (22.2 to 27.7)
20192732.9 (30 to 36.1)28.7 (25.2 to 32.5)
20203737 (34.3 to 39.8)23.4 (20.5 to 26.5)
20213237.3 (34 to 40.9)29 (25.5 to 32.6)
20223132 (28.9 to 35.2)34.2 (30.5 to 37.8)
20232833.8 (30.7 to 37)34.7 (31.7 to 37.8)
20242732.5 (29.5 to 35.5)33.7 (30.7 to 36.8)
20253334.2 (31.2 to 37.3)29.2 (25.3 to 33.1)
20262132.8 (29 to 36.2)29.8 (26.2 to 33.2)

Selection rule for anything Fit to Care writes about from this board, in order: materiality, novelty, evidence readiness, timing. Heat is not a criterion. Fit to Care may decline to write about a finding; it may not remove it from the board.

3. Party balance

Published monthly whatever it shows. For each party, the table counts the questions asked by that party's MPs and shows the share that received a full answer.

Party of the MP askingAnswers judgedFull answers (%)
Alliance1100 (100 to 100)too few to compare
Conservative1533.3 (28.6 to 37.5)
Labour1855.6 (37.5 to 70)
Liberal Democrat742.9 (33.3 to 50)too few to compare
Plaid Cymru20 (0 to 0)too few to compare
Scottish National Party333.3 (0 to 50)too few to compare

Last updated September 2026. Next update 1 October 2026.

4. How a question is judged

Six codes for the answer and five for any factual claim, adapted from published academic work on question avoidance (Bull and Waddle, among others). The codes and their definitions are open. The engine that applies them at scale, the way disagreements are adjudicated, and the ranking behind any story are Fit to Care's own work and are not published.

Answer codes

  • Q-F. Full answer. Supplies the requested information or clearly states its absence where that is the answer. Factual accuracy is assessed separately.
  • Q-P. Partial answer. Addresses only part of the information request. Identify the missing part.
  • Q-N. Non-answer. Substitutes another issue, or gives no requested information despite a usable opportunity.
  • Q-J. Justified challenge. Challenges a materially false or ambiguous premise and explains the problem. Does not automatically settle any separable valid question.
  • Q-C. Recognised constraint. Declines to give the requested information and names, as its reason, a constraint from the closed list in the rules below that applies to what was asked.

    The closed list of constraints is part of the scoring method and is not published.

  • Q-I. Opportunity interrupted. Cannot fairly assess completion because the speaker was cut off or the available clip is incomplete.
  • Q-U. Unclear. Ambiguous question, inaudible content or uncertain reference prevents assessment.

Factual codes

  • F-S. Supported. Relevant reliable evidence supports the proposition within its stated scope and time.
  • F-R. Refuted. Relevant reliable evidence contradicts the proposition within the same scope and time.
  • F-C. Conflicting evidence. Material credible evidence supports competing conclusions and the conflict is unresolved.
  • F-U. Insufficient evidence. Available evidence cannot establish support or refutation.
  • F-N. Not a verifiable factual claim. Value judgement, preference, prediction or rhetorical utterance; extract embedded factual claims separately.

Derived codes

  • NC. Could not be judged. The two readers could not agree a code, or the answer could not be assessed.
  • partial_or_non. Partial or no answer. The two readers disagreed between a partial answer and no answer; counted as not a full answer and not resolved by hand.
  • no_clear_ask. Could not be judged at the question gate: no identifiable ask in the question.
  • points_elsewhere. Could not be judged at the question gate: the answer points to a document, report or earlier statement the reader was not given.

Every session is read twice, by two independent AI readers that never see each other's work, each instructed to err on the strict side. Where they disagree on whether an answer was full or partial, the record shows partial. Where they disagree on whether an answer was partial or absent, the answer is shown as partial or no answer. Where one reads a full answer and the other no answer, the answer is shown as could not be judged, and the number of such answers is published in section 6. Nobody at Fit to Care adjudicates a disputed answer by hand. The agreement rate between the two readers across the record: 82.2 per cent.

Humans check the machine, not the other way round. The two readers agree with each other on 82 per cent of answers; agreement beyond chance (Cohen's kappa, where 0 is chance and 1 is perfect) is 0.74, which the usual scale calls substantial. Anthony Lawton, the project owner, marked 50 answers blind on 29 September 2026: his first choices matched the readers on 29 of 49 judged answers (59 per cent) and 38 of 49 (78 per cent) after a second look; where he differed, he was more generous than the readers 13 times and stricter 7 times. A second marker who works with Fit to Care has marked 25 of the 50 so far, matching on 15 (60 per cent) and leaning stricter. The readers are more consistent with each other than either human is with them, and the two humans lean in opposite directions, which is why the published numbers come from the readers under fixed rules and not from anyone's marking. One marker with no link to Fit to Care is still sought. (The export copies this paragraph from the ruling that records the check; the agreement figures come from the analysis scripts of 2 October 2026, kept with the record.)

One answer, two readings. 8 February 2023. Feryal Clark (Labour) asked when further aid for the Turkey and Syria earthquake would be announced, and what discussions the Prime Minister was having with international counterparts. The Prime Minister described his call with President Erdogan, the search and rescue teams on the ground, and the Foreign Secretary's contact with the United Nations.

Reader A: full answer. "Fully addresses the questions by detailing his own and the Foreign Secretary's discussions with international counterparts and explaining the current and ongoing nature of the UK's aid commitments."

Reader B: partial answer. "Addresses the second part of the question by outlining discussions. However, he does not state when an announcement on further aid commitments can be expected."

Published code: partial, because when the readers split between full and partial the record shows partial. Both human markers chose full answer.

Two rulebooks: the ten-year record stays fixed under the rules it was scored with, v0.5. New sessions are scored under the current rules, v0.7. Both are listed in section 7 with what changed and why, so a reader can see the system changing and can judge it.

5. Right of reply and corrections

Anyone named on this page, or their office, may reply. Before Fit to Care publishes any article drawn from this record, the people named in it are sent the relevant findings at least 48 hours in advance and their reply is published with the article.

To ask for a correction, write to anthony.lawton@ffmi.co.uk. A correction that shows the record is wrong (for example, a question answered later in the same session) is made within 48 hours. The original wording stays visible, struck through, beside the correction. Every correction is listed below with its date, whatever it shows.

Corrections log

No corrections have been requested or made since this page went live on 2026-10-01.

6. Sessions held back, and the widest splits

No session is currently held back.

212 of 9,480 answers in the record (2.2 per cent) were read as a full answer by one reader and as no answer by the other. They are shown as could not be judged. The rule: Where one reader records a full answer and the other no answer, the answer is not classified and the count is published; from rubric v0.7 new sessions go to a third reader under the three-step rule.

Until 2 October 2026 a session with such a split was held back from the board pending a review that never changed a code. From rulebook v0.7 (section 7) a new session with such a split goes to a third reader, and no session is held back for this reason. A session with incomplete Hansard text is still held back and listed here with the reason.

7. The rules this page keeps

RuleText
1Scope: elected national politicians in Hansard text, and public sector documents. Never a named public sector officer, NHS or council employee, or private individual.
2Facts, never intent.
3Same rubric for everyone, every session scored in full.
4The whole scoreboard is public. Nothing is withheld.
5Two independent passes before anything leaves.
6Every claim carries a link.
7Right of reply and a public corrections log, on this record's own page.
8Party balance published monthly.
9Numbers and quotes, no adjectives about people.
10No claim of impact without a measurement.
11Licence and attribution on every output.
12Open codes, protected engine.
13The kill switch.

(The thirteen credibility rules: the rule column only, verbatim from Credibility Rules v0.2, with the project name replaced by "this record". The "in practice" and "example" columns are not published: they carry typed figures and quoted words that the page's own gates forbid.)

v0.5. What changed: Appendix 3 refusals and commitments, Q-C recognised constraint, tactics on Q-P and Q-N, and asked-again flag; readers A and B prompts v0_5. Why: Option A pilot and ten-year corpus scoring; fixed for the published record.

v0.7. What changed: Full against no answer splits use a third reader and the three-step rule instead of session holds; step 1 NC reasons no clear ask and answer points elsewhere. Why: October 2026 ruling: no human adjudication on full against no answer splits.

The three-step rule of rulebook v0.7, applied by a third reader shown the question and answer and neither earlier reading, when one reader records a full answer and the other no answer:

  1. Question gate. List every ask in the question. No identifiable ask: could not be judged, reason "no clear ask". An answer that only points to a document, report or earlier statement the reader has not been given: could not be judged, reason "answer points elsewhere".
  2. Ask by ask. For each ask: answered, addressed without the information, or not addressed (including answering a different question, or speaking about the questioner rather than the question).
  3. Result. All asks answered: full answer. Some answered or addressed: partial answer. None addressed: no answer.

No step records why the person answered as they did. The ten-year record stays as scored under rulebook v0.5; the rule applies to sessions scored from October 2026.

8. Licence and sources

Contains Parliamentary information licensed under the Open Parliament Licence v3.0. Hansard text via TheyWorkForYou (mySociety). Every quoted exchange on this page and in any article links to its Hansard speech id and date.

This record measures whether a question was answered. It does not measure whether what was said was true. Fact checking, where Fit to Care does it, is published inside an article with its sources, never as a list on this page.