Will AI make the state smaller - or simply better?

Written by
Rhys Spence

In Plain Sight - 06

In Plain Sight is a weekly blog from Rhys that we will publish every Thursday at 8am UK / 9am CET. Each edition will be focused on a thorny topic within investing, startups, learning and work and policy.

In Plain Sight is partially a personal attempt to think independently and to avoid leaning too much on AI for answers to our questions. Will we use AI to edit and polish the text? Yes. But more importantly, will we come up with all of the ideas and analysis? Yes.

Here follows Edition 06 - 'Will AI make the state smaller - or simply better?'.

------

Barely a week goes by without tech giants or multi-national corporations (MNCs) announcing layoffs, restructuring, or a hiring freeze it links, directly or indirectly, to AI. The pattern is now familiar enough that regulators have started responding: the EU AI Act's transparency provisions and a growing body of EU and UK employment law around algorithmic decision-making are pushing employers toward much greater disclosure when AI shapes a redundancy or restructuring decision. As is well-documented, entry-level roles have taken the hardest hit. Ops, customer service, people, strategy, finance, legal - the functions that used to absorb graduates while they learned the business - are exactly the functions AI tools are augmenting first.

This is mostly discussed as a private sector story, but presumably it's a topic the public sector is grappling with too, publicly or privately.

Every function being reshaped inside big tech and MNCs also exists, at enormous scale, inside government. Every department has an ops team, a finance team, a legal team, a people function, a policy and strategy layer. If AI can compress a 12-person customer service team into four at a fintech, there's no structural reason it can't do something similar inside the Department for Work and Pensions or the Home Office (there are caveats that vary by department of course, depending on the relationship between the state and its users - the general public).

Which raises a question government hasn't had to answer publicly yet: how will the state capitalise on these efficiencies, or will it choose to be less efficient in order to preserve jobs? It's not going to be hypothetical for much longer: many governments are already running versions of this experiment. We briefly explore a couple of cases below, starting with DOGE and then leading to the UK case.

DOGE is the starker version, which cut 271,000 federal jobs in under ten months - a 9% reduction, the sharpest peacetime workforce cut in US history. By the usual logic, that should lead to savings. It didn't: spending in the first eleven months of 2025 came in roughly $248bn higher than a year earlier, because federal salaries are only about 8% of total spend and most of the budget is transfer payments that headcount doesn't shape. Brookings has argued DOGE's own cost-cutting mission crowded out the slower work of building the AI systems that would have let those roles be substituted rather than simply vacated - it cut the people without doing the work to understand where AI could do the job instead, and a chunk of those cuts are now being reversed.

Compare that to the UK's version of the same bet. Andy Burnham hasn't commented in detail on this subject yet but Keir Starmer and Peter Kyle framed their civil service AI push around a potential £45bn in savings, with an early proof point already running - an AI assistant built with Citizens Advice that halved response times on complex enquiries. Kyle's language is notably sequenced: a smaller civil service is "almost certain," but as a consequence of the tooling working, not the headline goal.

So the real argument probably isn't left versus right - crudely, you might expect the left to prioritise saving jobs and keeping a larger state and the right to present it as an opportunity to reduce taxes and shrink the state. DOGE was the more ideologically efficiency-driven project, run by a Republican administration that promised to shrink the state, and it's the one that failed to produce savings and is now being partially reversed. The UK's Labour government, nominally the more jobs-protective side, was the one actually being disciplined about proving the tool works before reducing headcount. This arguably represents evidence that the real variable is institutional patience for the slower, less headline-friendly path, rather than the government's ideological stance.

Part of why AI's use in key government flows now feels inevitable is growing awareness of AI 'clogging' the state. The Economist gave this its own name in early August: "agentic flooding." Citizens are using AI to file appeals, claims and complaints at a volume and quality the state was never built to absorb. UK employment tribunal backlogs are up 55% in a year, "in large part due to AI-fuelled claims," and demand for emergency injunctions has surged roughly 100-fold. Producing a well-argued appeal used to cost time or money, which rationed how many people bothered pursuing one. AI removes that friction and so the volume it unlocks is arriving faster than departments built for a slower caseload can process. This is worst where citizen submissions are long-form i.e. precisely the bit that generative AI is utterly transforming.

In the UK and elsewhere in Europe, there are concerns that governments are fuelling this 'clogging' whilst taking limited, ineffective steps to counteract the clog. For example, from January 2027 in the UK, the same employee protections that were previously afforded to employees with more than two years' service will be afforded to them after six months. This will lead to greater opportunities to appeal for employees who are dismissed after six to twenty-four months. This isn't a comment on the intention of this policy - it's more to highlight that clogging a flailing system further without relieving pressure elsewhere deserves the same scrutiny as the sequencing of AI tooling and staffing decisions within government departments.

The Economist framed this as a tragedy of the commons: everyone getting access to sharper tools collectively breaks the system all of them rely on. We'd add that it's also a preview of what happens when only one side of an interaction has AI. Right now it's citizens who've adopted faster than the departments processing their claims. But if this imbalance can be corrected or even reversed (a state whose own adoption of AI in its casework outpaces the public's), then states can drive enormous efficiency savings and reduce persistent issues like waiting times for cases to be adjudicated in our courts.

Closing that gap, though, means AI doing more of the state's own deciding - which is exactly why we ought to be keen to avoid AI acting as adjudicator or arbiter in civic cases. Evidence as to why this is the case is already rearing its head at Sainsbury's, the UK supermarket chain. In early 2026, a customer was wrongly pulled aside at the chain's Elephant and Castle branch after Facewatch's facial recognition system flagged him; in August, it happened again to another shopper at the East Dulwich branch - mid-checkout, full basket scanned - after staff were told the system had linked him to a prior theft.

What's notable both times isn't that the algorithm was wrong (it was) - Facewatch claims 99.98% accuracy, and both parties insist it flagged correctly, with staff simply acting on the alert badly. Assume that's true: the real point of failure is the human safeguard on top of the model, the "trained manager" meant to check the match before anyone gets escorted out, who instead deferred to the alert under time pressure, twice within months. That's what should worry us in government, where the stakes are a benefits sanction rather than a supermarket ejection, and a caseworker faces the same incentive to defer to an AI fraud flag. Holding AI adjudication to a higher standard than we'd apply to a human is a direct response to that oversight layer failing the same way, predictably, more than once.

Set this adjudication risk to one side, though and there's a genuinely long list of places inside government where AI is a clear net positive.

High-volume, transactional processing is the obvious start: tax returns, passport renewals, benefit applications - repetitive and rules-based at enormous scale, where AI can triage, pre-populate and flag exceptions rather than making a human read every case from scratch. Frontline citizen service is close behind, per the Citizens Advice example mentioned above.

Backlog and case triage in courts and tribunals is another - as described, not AI deciding cases, but summarising and surfacing the ones needing urgent attention, which is also the state's best defence against the agentic-flooding problem. Fraud and compliance detection, pattern-spotting across tax, benefits and procurement data at a scale no investigator could manually cover, is a strong fit precisely because it produces leads for a human to check, not verdicts. And policy modelling and risk analysis lets AI stress-test a proposed change against demographic, fiscal and behavioural data while keeping the judgment call where it belongs.

The common thread is that AI is doing the preparation, rather than making the decision - we think this represents a useful test for any department wondering where to start.

The risks are real. Civil servants arguably have less mobility than their private-sector counterparts and a public sector redundancy doesn't come with the same retraining ecosystem a tech worker gets. Uneven adoption between departments, or between the state and the public it serves, risks creating a two-tier system where service quality depends on who got the tooling first. And DOGE is a useful warning that cutting headcount and calling it AI efficiency isn't the same as actually deploying AI to do the work - the gap between the two is where the wasted savings and the reversed cuts live.

We'd love to talk to founders building the tooling that supports governments to work efficiently on both fronts: proving AI is really substituting for tasks rather than just for people and making sure the human still meaningfully checks the machine rather than rubber-stamping the guidance it provides.

I’m building what’s next

Share your deck and a few lines about what you’re building.

Submit a pitch
I have a question

For partnerships, media and general enquiries, we’d love to hear from you.

general enquiries