AI Engineer Europe 2026
Rewiring the State
Read the talk
Rewiring the State
Inside Number 10’s experiment with technical fellowships, embedded engineers and reusable AI tools—and the institutional changes needed to make their work scale.
From a talk by Eoin Mulgrew
What can a small team do about public-service backlogs?
How can a small technical team at the center of government help a public-service system struggling with waiting lists, court backlogs and slow planning decisions? Eoin Mulgrew leads cross-government transformation and the fellowship program in the Number 10 Data Science team, 10DS. Established during the pandemic, partly in response to it, the team’s core responsibility is to bring the best available evidence to consequential national decisions. It is now expanding its AI engineering capability, both inside Number 10 and across strategically important parts of the state.
The scale of the problem needs careful denominators. Mulgrew opens with an NHS waiting-list figure of 7.25 million; the contemporaneous NHS England release counts waiting-list entries, representing approximately 6.13 million unique patients in England in January 2026. He also cites roughly 350,000 backlogged court cases, without specifying the jurisdiction or court types. His description of only one in five planning decisions arriving on time similarly needs a narrower scope: England’s July–September 2025 statistics report 19% of major applications decided within the statutory 13 weeks, rising to 90% when agreed extensions are included. That is not a rate for all planning applications.
Behind these delays sits a public-sector productivity problem, worsened since the pandemic. Mulgrew cites a Tony Blair Institute estimate of £40 billion in potential annual productivity gains from AI in government. This is an estimate of opportunity, not savings already delivered. His broader framing is useful: government behaves more like a complex industry than a single organization. His shorthand of 400,000 people understates the civil service measured in the March 2025 statistics, which record 549,660 people and exclude NHS employees. The organizational challenge remains one of many different services, professions and operating environments.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The constraints are organizational as well as technical
Government has struggled to build and retain high-performing technical teams. Pay makes recruitment and retention difficult; hierarchy and bureaucracy can make delivery slow. Some barriers are perceived, while others are real. Some also exist for good reasons: public bodies are accountable to the public and Parliament, and their safeguards cannot simply be treated as unnecessary friction. Together, however, these conditions can repel precisely the impatient, capable engineers a transformation program wants to attract.
Changing the whole system exceeds the remit of a small central team. Mulgrew assigns that wider reform agenda to the Chief AI Officer, Kal Beer, and the Department for Science and Technology. For 10DS, the immediate opening is narrower: political leaders want new technology to improve services, so what can a small team do with that backing?
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An exceptional team inside the system
The answer is an insurgency model: a small unit at the center with different operating conditions. Its design combines several permissions that normally sit apart.
- Political mandate: Number 10 backing to enter departments and work on consequential problems.
- Viable compensation: market rates within reason, without promising compensation comparable to Meta.
- Autonomy: freedom to pursue opportunities discovered while working inside a department.
- Technical selection: a recruitment process designed around technical ability rather than the standard civil-service process.
- Outside recruitment: fellows come exclusively from outside government, bringing experience the system would otherwise struggle to acquire.
Mulgrew reports an applicant success rate of about 0.7–0.8%. The longer-term benefit is not confined to a fellow’s initial placement: some stay in government and establish teams of their own.
Political backing creates room to work; it does not make implementation easy. A ministerial mandate cannot, by itself, dissolve data silos or make departments cooperate. If authority alone were sufficient, this model would already be commonplace.
The recruitment offer has nevertheless attracted people from AI labs, big technology companies and research institutes, alongside Y Combinator founders and serial entrepreneurs. The combination matters: economically viable pay, consequential decisions and an environment where technical people can do effective work. Mulgrew describes the desired recruits as “missionaries, not mercenaries.” Public purpose must sustain them when the work becomes difficult; compensation alone will not.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Start with the workflow, then choose the deployment
Legacy workflows offer many relatively simple opportunities for AI. For those, 10DS usually builds directly. Mulgrew describes Number 10’s first forward-deployed engineers working alongside policy advisers, lawyers, communications staff and pollsters. Their procedure starts with the people doing the work:
- Observe the existing workflow and its pain points.
- Co-design a solution with the team that will use it.
- Take the idea through implementation and put the capability into users’ hands.
Mulgrew describes some simple use cases taking days, with this embedded route typically moving from idea to implementation in a couple of weeks.
Large public-service backlogs require a different commitment. They are not just a collection of quick automation tasks, so 10DS deploys people into partner teams or departments, sometimes for prolonged periods.
| Work | Deployment approach |
|---|---|
| Bounded internal workflow | Build directly with its users |
| Complex departmental problem | Embed engineers through a sustained partnership |
The distinction is the depth and duration of the engagement, not whether engineers talk to users: both approaches depend on working inside the operational setting.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Model policy choices and make legal analysis reusable
Much of Number 10’s internal workflow automation is sensitive, but the policy simulation example shows the basic interaction. A policy team can explore a proposed change before making a decision. The demonstrated case concerns Universal Credit and its effects on household finances. The tool supports human analysis rather than necessarily replacing it; Mulgrew reports that more decisions can receive modeling support, and receive it faster. The useful capability is access to analysis during policy development, while choices are still open.
The next example changes the economics of repeated analysis. Mulgrew reports that the Cabinet Office was preparing to spend £1.5 million on external lawyers to analyze the entire UK statute book. He describes the volume of legal text as a stack four African elephants high. Instead, one engineer embedded with the in-house lawyers for a couple of weeks to build a tool they could use themselves.
The benefit was not only avoiding a proposed commission. The external analysis was expected to take longer than the pace at which new laws and regulations appeared. Its output would therefore age during production and eventually need to be repeated.
| Proposed commission | Internal tool |
|---|---|
| Analysis delivered after an external engagement | Analysis rerun by the legal team |
| New legislation makes the output stale | New legislation can prompt another run |
| Repeated work requires another engagement | Capability remains with the users |
The durable output is the ability to repeat the analysis. Mulgrew also raises the possibility of open sourcing the tool and sharing it across government; those are potential next steps, not completed releases.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Challenge delivery reports and expose progress
Number 10 receives reports on major projects and manifesto commitments across government. A delivery red-teaming tool gives its delivery teams something like a project management office in their pockets. Mulgrew says the team built it a couple of weeks earlier and that it is already used daily. Beyond interrogating an individual report, it questions the reporting organization’s judgment:
- Optimism bias: does this team or department habitually present an overly positive picture?
- Risk ratings: does it disproportionately classify risks as amber?
- Mitigations: have its proposed mitigations usually worked?
This adds a second judgment about the source of a report, rather than treating each report as an isolated account.
In-house engineering also supports public transparency. Mulgrew reports publishing two public-facing delivery dashboards in two months, describing them as a departure from previous practice. The named example is the AI Opportunities Action Plan dashboard, tracking commitments from Matt Clifford’s plan, including compute rollout and AI adoption. Its official landing page confirms publication on 29 January 2026; it does not establish the broader claim that government had never previously published a public delivery dashboard.
A further example is deliberately unnamed: a new public service awaiting a ministerial launch. Mulgrew says its idea originated two months earlier, with launch expected two and a half weeks after the talk and millions of prospective users. He contrasts that timetable with a conventional government discovery phase that might last a year or more. At this point in the recording, launch and public adoption are still expected outcomes.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Build capability through partner teams
The partnership model extends beyond Number 10 through three examples: the AI Safety Institute, the Incubator for AI and Justice AI. Mulgrew describes AISI as the world’s first government body for evaluating frontier models and says early fellows helped establish its cybersecurity workstream. He credits fellow Dr. Harry Coppock, placed there from the beginning, with leading work on Inspect.
Inspect supports evaluation of models and agents, including prompts, tools and multi-turn interactions. Mulgrew introduces it through the need to test what agents actually do when given autonomy and tools in an isolated environment. The technical distinction is that Inspect is an evaluation framework with sandboxing support, not an automatic isolation boundary around every tool call. Its sandboxing documentation distinguishes code running in the evaluation process from work explicitly routed through a sandbox interface.
The Incubator for AI, now within DSIT, illustrates a second form of institutional growth. Mulgrew describes it as a spin-out of the fellowship program, with fellows making up much of its original technical team. Its role is to incubate AI solutions for use across the public sector; the relationship with 10DS continues as those solutions scale.
Extract is one such collaboration. Built with DeepMind on Gemini, it digitizes planning materials, including handwriting and hand-drawn maps. It was unveiled at London Tech Week in June 2025, and Mulgrew describes rollout to every local authority in England as underway. Digitization addresses a bottleneck in the planning process; automatic AI decisions on planning applications are a future aspiration he raises, not the tool’s stated current capability.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Evaluate tutoring before broad adoption
AI tutoring offers the possibility of making strong educational support available regardless of a child’s socioeconomic background. But access alone is not enough: classroom use requires safeguards and evaluations suited to children. The team is evaluating frontier models against benchmarks for both safe interaction and educational behavior.
The illustrated criterion is cognitive load placed on the student. The slide pairs evaluation criteria with a student–AI tutor conversation, making the interaction itself the object of assessment. The question is not merely whether the model can produce an answer, but how demanding its tutoring behavior is for the learner.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Embedded engineers on the frontline
Justice AI takes the embedded model into the Ministry of Justice. Its founder, Dan James, is a former fellow, though Mulgrew explicitly declines to call the team a fellowship spin-out. Instead of working beside policy advisers and communications staff, its forward-deployed engineers work alongside parole officers and prison wardens. The examples include using AI to interrupt drug flows into prisons, reduce labor-intensive manual processes and improve prison safety and security. The operational details are not disclosed.
A fellow named Will makes that placement model tangible. A few months earlier he had been in California, having left Harvard and founded a company accepted into Y Combinator. Mulgrew describes a photograph of him outside HMP Wandsworth in the rain, in his second week on the job, holding the prison’s actual keys. The recruitment proposition is unusually concrete: experienced industry practitioners can enter government and receive access to consequential operational work.
Mulgrew calls the program an early experiment. He points to savings, faster public-service delivery, frontline reform and new internal AI capabilities as evidence that small, highly capable teams can accomplish useful work. The prepared talk ends with a hiring invitation through an on-screen QR code and conversation afterward. The questions that follow test what would be needed to make the experiment reliable and scalable.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
What if the policy tool tells users what they want to hear?
The first question returns to policy simulation. A user who already believes a policy is excellent may steer a conversational model toward confirming that belief. Mulgrew illustrates the failure with the idea of cutting income tax to zero and receiving enthusiastic agreement. Sycophancy would turn an analytical aid into a source of apparently authoritative reassurance.
The response has two parts. Before release, the team red-teams models specifically for this behavior. Afterward, it coaches users on the risks of the tools they are receiving. Those users may be lawyers, sociologists or professors with substantial domain expertise but little experience of AI failure modes. Mulgrew says the team has encountered relatively little sycophancy so far and attributes that partly to testing before deployment; he does not dismiss the risk.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make the exception normal, then work across services
The next question asks how the model scales across central and local government, including different political contexts. Fellows staying in government and founding teams help, but Mulgrew judges that organic growth too slow for the urgency of the problems. Scaling requires strategic changes to how the rest of government operates.
The bargain with ministers was to let a small team operate under different rules and use its results as a proof point. The intended destination is for those conditions to become business as usual. Mulgrew describes the present arrangement as a way around the system; institutional reform is needed so effective technical delivery no longer depends on an exception.
The other route to scale is horizontal work: capabilities applicable to repeated processes across services rather than one targeted use case at a time. Mulgrew sets this as an ambition for the next 12–24 months. His underlying distinction is between the visible policy workforce around Westminster and the much larger operational workforce. He ranges across call-center staff, prison wardens and nurses to illustrate frontline public service, although the latter sit outside civil-service statistics. The concrete opportunities are police transcription and the large call centers in DWP and HMRC. A reusable improvement to a common process could reach much further than a tool serving one team.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The remaining gaps: motivation and shared experience
An edtech practitioner challenges the tutoring example from another direction: even an excellent tutor may fail if a child does not want to learn. What would school deployment look like, and how would the government address motivation? Mulgrew clarifies that the current plan centers on benchmarks and guardrails for schools choosing products, rather than building a government tutor to compete with vendors. Student uptake has received little work so far.
Mulgrew says the initial tutor testing involved about 70 teachers role-playing their pupils. That supports a different kind of assessment from observing children’s motivation or sustained use directly. He invites the questioner to share practical experience afterward, leaving adoption as an open area for collaboration rather than a solved part of the program.
The final question comes from a Norwegian attendee asking about international exchange. Mulgrew reports some collaboration, but none with Norway so far. He names Tech Force and parts of the US Digital Service as broadly similar initiatives, and says the team talks frequently with Singapore. There is room to do more, and he welcomes contacts in the Norwegian government: the experiment can learn from other countries as well as from the departments where its engineers are embedded.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The June 2025 announcement explains Extract's use of Gemini and its planned rollout to English councils.
AISI's introduction to its open-source framework for evaluating language models and tool use.
Official access page for tracking progress against the government's AI Action Plan commitments.
Further reading
- AI action plan for justiceArticle
The Ministry of Justice's approach to AI adoption, specialist recruitment, workforce training and partnerships.
Updates since the talk
Current instructions and examples for running evaluation workloads in isolated environments.
Read the complete timestamped transcript
- 0:00
[upbeat music] Good to go?
- 0:17
Can people hear me okay?
- 0:18
Yeah.
- 0:19
Awesome. Cool. Thank you. I'm slightly embarrassed by the grandiose title, uh, now that I've had to leave it up for a few seconds. Um, but anyway, um, who here works in government, can I ask?
- 0:31
Show of hands. Very good. That's what I was hoping for. You might be very well acquainted with some of the stuff that I'm gonna grumble about. Who can't think of anything worse? [laughing]
- 0:41
All right. Good. What gets measured gets improved, so I'll do another one at the end. Actually, I won't in case more hands go up. Um, hi, everyone. I'm Eoin Mulgrew.
- 0:51
Um, I work in the data science team just down the road in [REDACTED:location_address]. Um, I run our cross-government transformation work, uh, including our fellowship program, which is predominantly what I'm going to talk about today.
- 1:03
Um, yeah, I wanted to come down here, tell you what we're doing in the hope that some of you might decide you want to be a part of it, which would be quite cool.
- 1:11
Um, quick bit about us. So Number 10 Data Science team, 10DS, we were set up, um, sort of during the pandemic, partly in a response to the pandemic. Um, our core business is making sure that the most important decisions in the country are informed by the best possible evidence.
- 1:29
Um, however, we are in the process of quite radically scaling up our own AI engineering and development capability, not just with the intention of driving AI adoption within Number 10 itself, but also across strategically important parts of the state.
- 1:45
Um, and the way that we're going about doing that is quite novel in itself. Before we get into that, just a little bit of context to set the scene.
- 1:53
I know some of you are flying in from the West Coast, et cetera. Uh, some of you might not follow the news. Um, believe it or not, there are some challenges when it comes to public s- public service delivery in the UK.
- 2:06
Um, I say that half in jest, but it's pretty serious. You know, at the moment, there are seven and a quarter million people that are on NHS waiting lists.
- 2:16
There are, I think, about three hundred and fifty thousand court cases that are stuck in a backlog. Um, only one in five planning application decisions in this country are currently decided on time.
- 2:28
Um, sitting behind all of this, um, is a public sector productivity crisis, um, that was bad and has only been exacerbated since the pandemic.
- 2:42
There are different figures for the sort of extent of this crisis. Um, I've gone for a Tony Blair Institute figure, um, that says there's a sort of forty billion prize annual productivity gains from AI in government.
- 2:56
Um, but it's clear to anybody that works in the system and most, most of society, um, that if any industry is ripe for disruption over the next few years, government is one, and I call it an industry rather than an organization.
- 3:10
It's a big, complex industry of four hundred thousand people. Um, and I think we should look at it through that lens. Unfortunately, however, government has traditionally
- 3:23
not been great at building and nurturing high-performing technical teams. Um, a lot of these issues are not specific to the UK. A lot of our American friends will be familiar with them.
- 3:33
Um, but just, uh, some of the commonly cited ones, um, pay is a, you know, an obvious one. It makes it quite hard for us to compete for the best talent out there and then retain it.
- 3:43
Um, but also some barriers that are both real and perceived. Um, so government is a very hierarchical organization in many parts. There's a lot of bureaucracy. Um, as a result, it can often move incredibly slowly.
- 3:58
Um, not just for those reasons, but also the fact that, you know, there are regulations and safeguards in place that are very sensible because we're ultimately accountable to the public and to Parliament.
- 4:08
Um, but all of this can result in a system that is not always that appetizing for high-performing technical people to join, especially the sort of people that we want, people that are impatient to leave their mark on the world.
- 4:23
So what do we do that? Or what do we do about that? Um,
- 4:28
changing a lot of these things, uh, it's quite like a systemic challenge. It's like turning an oil tanker, which is a bit of a tired cliché, but it's a good one.
- 4:36
Um, we're a little team at the center. Um, you know, turning that oil tanker is beyond the remit of, of any one team, let alone a scrappy little startup like ours.
- 4:46
However, um, I think Kal Beer, our Chief AI Officer, is going to be closing out the conference this evening. That's very much his job, and he's doing great work at the Department of Science and Technology to do just that, so I recommend everybody goes along to it.
- 5:00
But there is a lot of political will at the moment, um, to get stuff done and to make sure that this time around we are seizing new technology to actually make a dent in some of those problems I just mentioned.
- 5:12
So the question put to us was, "Well, what can you do about it?"
- 5:18
And this was sort of the answer. Um,
- 5:23
as I said, we're, like, quite a small team at the center. Um, in terms of what we can do, um, we said, "Okay, well, let us take the shackles off.
- 5:33
Let us basically set up a small insurgent unit, unit at the very center that is not, um, sort of burdened by some of the constraints that I just described to you."
- 5:45
What do I mean by an insurgency model? Um, so we're setting up a new team. Um, it operates with a mandate from Number 10. We operate with an unusually high level of political backing to go into departments and get stuff done.
- 6:00
We're able to- Pay market rates within reason. We're not paying, like, meta money necessarily. Um, but the thing is, a lot of people will happily take pay cuts if we make it economically viable to come in and work on some of these challenges because they're interesting, right?
- 6:16
Uh, we operate with an unusually high level of auto-autonomy as well. We're able to be fairly opportunistic about the challenges we take on, where we go into a department and see opportunities to have impact.
- 6:27
Also, this one's pretty r-crucial. Um, the civil service recruit-- standard recruitment process is optimized for a lot of things, but not necessarily recruiting exceptional technical talent. Uh, we've been allowed to recruit our own way.
- 6:41
We've got a f-fairly grueling selection process that's laser-targeted on technical skills. Uh, we've got a success rate of about naught point seven, naught point eight percent. Um, and in-- most interestingly of all, and this is what differentiates us, uh, we recruit exclusively outsiders.
- 6:58
Um, one of the best ways that I can have impact is by getting some people of the likes of this conference into government, um, because what has happened in the past is they tend not to leave, and some of them end up setting up their own teams.
- 7:10
I'll get into that later. Just when we're on this, though, I don't want to make this sound, like, overly simplistic. Um, a lot of people from the outside, particularly, particularly the tech industry, think that this bit alone is the only important thing that you need, that if you have a big enough stick for ministers, you can go
- 7:28
in, you can, you know, break down data silos, you can do what you want. In practice, it's a lot harder than that, otherwise everybody would be doing it.
- 7:39
And it's really early days. Like, we're, we're only setting out in this journey, but it turns out there's a huge amount of appetite for it. So we have been taking people from the labs, we've been taking people from big tech, from top research institutes.
- 7:52
We've been taking YC founders, serial entrepreneurs, um, people who probably did not think they would be working in the civil service this time last year. Um, but when you think about it, the decisions that go across a minister's desk are, like, some of the most important things you could possibly work on.
- 8:10
So if you make it economically viable, and you promise people that you're gonna put them in an environment where they can do their best work, it makes it really interesting.
- 8:19
Um, it's also worth pointing out as well, we do want to recruit missionaries, not mercenaries. Um, so the pay matters, but it's not alone because a paycheck is not going to get you out of bed in the morning when stuff gets hard, and doing the stuff that we do, um, does tend to be difficult.
- 8:37
In terms of how we operate then, um, again, a bit different from normal government teams. Um, some of the--
- 8:45
There's, like, an abundance of low-hanging fruit around the system. As you can imagine, it's a legacy organization. Um, there's lots of simple AI use cases that you can do in a few days to save money, to improve service delivery, all that good stuff.
- 8:58
When it comes to that, we largely do it ourselves. Um, that's the easiest and most satisfying part. Um, as we speak, we've actually got the first forward-deployed engineers in the history of [REDACTED:location_address] embedding themselves with policy and operational teams, teams of policy advisors, teams of lawyers, teams of comms people, pollsters, and everything in between.
- 9:20
They're observing their workflows, their pain points, co-designing solutions with them to help them do their job more efficiently and effectively, and generally taking things from idea to implementation, uh, in a couple of weeks and getting new capability into the hands of users quickly.
- 9:36
Um, and then some of the other problems we talked about. I mentioned, you know, some of the huge backlogs in the system. That is not low-hanging fruit. That's really complicated stuff.
- 9:45
And normally, when it comes to those, we take more of a partnership model where we will deploy some of our people into another team or another department, sometimes for prolonged periods.
- 9:56
And I'm gonna give you examples of both of these in a minute.
- 9:59
I'm gonna start with some of the low-hanging fruit. It's worth pointing out, actually, um, a lot of the stuff that we do in Number 10, I can't really show you.
- 10:08
Um, I know that sounds like an easy get-out-of-jail card, but trust me. Um, so-some of this stuff is a little bit sensitive. Um, we're doing a lot of workflow autom-automation, augmenting existing teams, as you can imagine.
- 10:22
And here are a few other examples of stuff we've done just in the past few weeks. Um, so policy simulation, that's turned out to be really interesting. Um, so here we can, uh, allow policy teams in the building to test out the impact of different policy decisions before they're made.
- 10:39
Um, I think in this one here, we're looking at different decisions around universal credit and how they might impact... Oh, I've paused it.
- 10:48
How they [chuckles] how they might impact household finances, amongst other things, but this can be applied to a broad range of stuff. Um, not replacing human analysis necessarily. We're not putting ourselves out of a job.
- 11:01
Um, but what it has meant is that far more decisions in the building are being informed by high-quality modeling and at a far faster rate than otherwise would have been the case.
- 11:15
And this is another one just from the past couple of weeks. Um, so the Cabinet Office was about to spend one and a half million pounds on getting an outside firm of lawyers to come in and do analysis of the entire UK statute book.
- 11:28
Uh, granted, the statute book is the height of four African elephants, um, of legalese, um, but still pretty obvious AI use case. So we were gonna spend one and a half million.
- 11:38
Instead, one of our engineers embedded with that team o-of in-house lawyers for a couple of weeks. Um, the benefits of this is not just money saved. Obviously, one and a half million is n-is not nothing, um, but also speed.
- 11:53
Um, so the issue with the analysis that we were going to pay for is that it could, it, it was gonna be done slower, uh, than the pace at which new laws and regulations are made, uh, which means you're gonna have to do it again after a certain period of time.
- 12:06
So now we've got this tool, uh, that that team can use, and they can do it whenever they want at the drop of a hat, and we can also, uh, potentially open source it and share it with other teams in government.
- 12:18
Um, and then this is another one. Um, so in Number 10, we're responsible for the delivery of every major project and manifesto commitment in government. That means a lot of reports come in on how various things are doing.
- 12:31
Um, this is a little sort of delivery red-teaming tool, uh, that the team s-spun up a couple of weeks ago that is now being used every day. Um, it's essentially a PMO that we've put in the pockets of delivery teams in Number 10, not just so that they can interrogate the delivery ports that-- reports that are coming
- 12:51
across their table, but also give a second judgment on the teams that are reporting them. Um, so it will flag up to decision-makers in Number 10, you know, does this team, does this department normally have a bit of optimism bias?
- 13:05
Uh, do they tend to disproportionately rate their risks as amber? And are their mitigations usually effective or not?
- 13:15
Also, aside from, like, AI adoption, having this capability in-house is really good. I think transparency is one thing that this go- that this country can do a bit better at.
- 13:25
Up until a couple of months ago, um, the government had never published a public-facing dashboard so that you lot can actually see how we're doing, uh, when it comes to delivery.
- 13:36
Um, but now we've published two in as many months. I think some of you might be familiar with the one on the left. This is the AI Opportunities Action Plan that Matt Clifford drafted about a year ago.
- 13:47
This is how the UK is doing when it comes to rolling out compute and generally setting up the UK to be a leader in AI adoption. Um, and now you can go online and see how, how we're actually doing.
- 13:58
Um, yeah. Um, also, another thing that I can't show, uh, but in two and a half weeks' time, we-- one of our ministers is going to launch a new public service that millions of people in the country are going to use.
- 14:09
I can't go into more detail and steal their thunder. Um, but, um, in their words, it's hard to believe this didn't already exist. I assumed something like it already existed.
- 14:20
Um, that's something that we thought of two months ago and is now going to be live and used by the public. It is not an understatement to say that normally in government, that pro- a project like that might be in discovery for a year or more.
- 14:32
Um, so yeah. Let's get into the meatier stuff then. So, um, that's nice low-hanging fruit within the building. Um, now I wanna talk about some of the work that we're helping, um, other teams with across different parts of the ecosystem.
- 14:48
Uh, for the purposes of this, I'm just gonna focus on, um, three of our partners: um, the AI Safety Institute, the Incubator for AI, and Justice AI.
- 15:00
ASE, I think pretty much everybody in this room will be familiar with it. Massive win for the UK. It's a great thing that we set up. Um, we're a leading government body for evaluating frontier models, and it was also the world's first, and we were really proud from day one to support it by putting a couple of
- 15:16
our fellows in there to help them set, set up their cybersecurity workstream, amongst other things. I'll not dwell too much on this, uh, but one of our early fellows was Dr.
- 15:25
Harry Coppock. I don't think Har-Harry is here, but we put him into ASE from day one. Um, and he led on their Inspect tool, amongst other things. Um, so this is a safe, isolated environment for testing what AI agents actually do when you give them autonomy and tools.
- 15:43
And the Incubator for AI, um, which now sits in DSA. Um, the-- Who here is familiar with the Incubator? A few people. Um, so the Incubator is essentially a spin-out of our program.
- 15:57
Um, it's a team that exists in the Department for Science and Technology that does what it says on the tin. It incubates new AI solutions for usage across the public sector.
- 16:06
Um, its original f- uh, founding team, most of the technical team were our fellows, and what's really cool now is not just seeing the work that they produced while they were there, but also the fact that we're able to collaborate with them when it comes to scaling up some of that work.
- 16:21
Um, here's one recent example. Um, so this tool is Extract. Um, so a bunch of our people have worked on Extract. It's a collaboration with DeepMind. It's built on Gemini, and it essentially digitizes large swathes of the planning application process, especially those bits that are currently, um, l-largely handwritten, including handwritten, uh, hand-drawn maps there as well.
- 16:45
Um, this was unveiled by the Prime Minister at London Tech Week last year, and we're currently in the process of rolling it out to every local authority in England.
- 16:53
As I said, only one in five planning applications are currently decided on time. That has a massive impact on economic growth, and economic growth is basically the biggest challenge this country faces right now.
- 17:03
So anything we can do to make a dent on that is really significant. I think aspirationally as well, this will hopefully get us to a place where more and more planning applications can be decided, um, by AI, um, automatically.
- 17:18
And then another interesting one, um, this is very current. Um, the education gap, uh, is a big problem, not just in the UK, but elsewhere. Um, many of you will have read the papers about AI tutors.
- 17:30
It's a really exciting moment, um, the, the prospect of being able to level the playing field somewhat and put world-class tutors in front of every child, regardless of their so-socioeconomic background.
- 17:42
Um, but it's something that has to be done really carefully. Um, so currently we're working, um, uh, we're working on producing safeguards and evaluating various frontier models against benchmarks, um, not just to make sure that children can interact safely with these in a classroom environment, but also measuring them against various metrics.
- 18:04
I think in this one, uh, the relevant benchmark is the cognitive load placed on the student.
- 18:11
And then last but not least, the new kids on the block, uh, Justice AI. Some of you might have been here for the Justice AI talk yesterday. Was anybody here?
- 18:21
Okay, good. For most of you, this is new. Um, Justice AI are a new team that have been set up in the MOJ, some of whom are over there.
- 18:29
Hello. Um, um, again, I wouldn't say a spinoff from the fellowship, that gives us way too much credit. Uh, but the founder of Justice AI is one of our former fellows, Dan James, who's doing brilliant work in there, and they are deploying forward-deployed engineers into prisons and into other parts of the criminal justice system.
- 18:49
Um, so kind of taking an approach that we're doing in Number 10 with, like, policy people and comms people and lawyers, but instead they're embedding with parole officers and prison wardens, and they're doing loads of really interesting work.
- 19:04
I can't go into too much detail, um, but most of it is around using AI to stop the flow of drugs into prisons, to find efficiencies where currently there's quite manual processes involving lots of people, and generally improving the, uh, security and safety within the prison system.
- 19:23
Um, and one of those FDAs is over there. It's Will. Uh, Will's one of our current fellows. Sorry, Will, I've embarrassed you. I, I just added in your photo last night because I thought this was, um, sort of a good point to end on.
- 19:36
Um, Will is... So to give you an idea, a few months ago, Will was in California getting a tan. Uh, that's him outside HMP Wandsworth on a rainy day.
- 19:49
Um, but yeah, Will dropped out of Harvard, started a company, got it into Y Combinator, made a bit of money, but wanted to come and work for us, and that's his second week on the job, and he's standing outside a prison with the keys to that actual prison about to go in.
- 20:04
And that is, that is exactly what we're trying to do through this program. You've maybe done good stuff in industry. That's brilliant. Come join us, and we'll give you the keys to the state and see what you can do.
- 20:16
So yeah, um, look, it's really early days. Um, it's sort of an experiment what we're doing. Um, but I think the proof points so far have been that actually, like, small elite teams can actually achieve quite a lot.
- 20:31
Um, we're already, already saving money. We're already shipping new public services at a unprecedented speed. Um, we're already reforming frontline public services, and we're already, um, putting new AI capabilities into the, into the hands of other teams, um, at the top of government.
- 20:50
So yeah. And surprise, surprise, this was a recruitment pitch. We are hiring. Um, so please do sign the, uh, scan the QR code, and I'll be here the rest of the day if you wanna come up and chat.
- 21:00
Thank you. [audience applauding] I, uh, I think I've got time for a couple of questions, possibly. Somebody can tell me if not.
- 21:16
Cause?
- 21:19
Um.
- 21:19
Oh. [chuckles]
- 21:25
Uh, so on, on the, on one of the earlier example you showed, uh, there was this, uh, chat t-to explain policies and, um, not explain, but try different, like, projection and see how th-they would behave.
- 21:36
Um, do you have to deal with, uh, sy-sycophant- sycophancy? Uh, with the fact that, you know, like,
- 21:44
if, like, you have a user that's not necessarily very well-versed in AI, like, just wants to hear what he wants to hear, can direct the tool towards, "Oh, look, I'm an absolutely brilliant mastermind.
- 21:55
My policy is going to be fantastic," uh, despite the policy being actually bad, but the idea-
- 21:59
Yeah, yeah.
- 22:00
-is basically.
- 22:01
Yeah. Should I cut income tax to zero percent? You're absolutely right.
- 22:05
Exactly. [laughs]
- 22:06
Um, yeah. That, that's a very real risk. Um, so it's, it's not something that we have encountered too much, but it's only because we have, um, sort of red-teamed the models for that before we've put it into the hands of users.
- 22:19
We also provide quite a bit of upskilling. Um, so, like, a lot of the teams that we work with, we're creating to-tools for them. They're possibly lawyers, they're possibly sociologists, professors, whatever.
- 22:29
So we do coach them on some of the risks that this presents. Um, but yeah, it's a good question.
- 22:36
Thank you.
- 22:41
Hello. Thank you very much for the speech. Uh, I'm Jack from Accenture. I think you beat our pitch for hiring, but, uh, we'll try as well. Um, but my question on the, the policy side is, as you progress with this FD type of model actually making an impact, how did you see...
- 22:59
Or, uh, as you said, it's early day experiment. How did you see you start to
- 23:04
scale and essentially start dealing with the center, cen- uh, central government or with local governments, the different party lines, so on and so forth? So it's kind of, you know, real kind of-
- 23:13
Yeah, yeah.
- 23:14
-governmental kind of, you know, human kind of stance. How, how do you see that's gonna play out?
- 23:19
Yeah, 100%. So in terms of how this scales. So I suppose at the end I said, "Oh, you know, we can do quite a bit," and some of the people that join us then set up their new teams.
- 23:29
That's great. Is it enough to turn the oil tanker itself? No. Possibly over time, but it would take a long time, and we need to solve these problems quicker than that.
- 23:38
Um, so we have been thinking about that. Um, I think realistically, some of this stuff requires strategic intervention, um, so we need to change the way the rest of government operates.
- 23:51
Part of the reason, part of the bargain that we basically made with ministers was, you know, let us take the shackles off. Let us set up a small team at the center that abides by different rules and use it as a proof point.
- 24:04
So this is, this is almost like a pilot. I would like to see a lot of what we're doing, uh, become the norm, become BAU. This, at the moment, is basically a hack to get around the system, so we, we need to change that first of all.
- 24:17
Um, I think also another... Like, if we're, if we're talking really about scale Uh, we, we talked about some, like, fairly targeted use cases there. Um, I think what we want to do over the next sort of 12 to 24 months is do more horizontal work, uh, looking at processes.
- 24:35
So I should explain this to people as well. Like, um, when you think of the civil service, you probably think of policy people working in some of the buildings around here that you can see out the window.
- 24:46
That is a very small sliver of the civil service. It's about 400,000 people. Most of them are call center operators. They're, uh, prison wardens, they're nurses, et cetera, et cetera.
- 24:58
Um, there are a lot of processes out there, whether that's like transcription, uh, that every poli- police person will tell you is the bane of their existence, or it will be those massive call centers in DWP, HMRC.
- 25:11
So yeah, if we want to dial up the ambition, I would like to see us going after more of those, like, horizontal use cases that can be applied en masse across the system.
- 25:20
Um, yeah, sorry, that was a bit of a long answer.
- 25:24
No, you're good. I think you got it all out, but that's good. [laughing] [laughing]
- 25:27
You can turn it off now.
- 25:30
Uh, I think this is the last one. Yeah.
- 25:32
Should we wait for the microphone? Or should we shout?
- 25:38
You can shout if you want. I'm... It's up to you.
- 25:40
That's all right.
- 25:44
Uh, yeah. So I work for an edtech company that, among other things, is making AI tutors, so I'd love to talk to you more about that. But one thing we find, I mean, the big problem w- with most kids is that, uh, you know, you can make the best AI tutor in the world, but, y- um, the
- 25:58
real problem is motivation. You know, if you sit a kid down in fr- if you sit a [REDACTED:age] down in front of a computer, they're, they're gonna do everything they can to avoid learning.
- 26:06
So my question is, like, how-- I mean, first of all, I'm just interested to know more about, like, what, what the vision is of the govern- the government for this.
- 26:13
Like, is this gonna be going into schools? Are kids gonna be sitting down in front of computers and using it? Um, and secondly, like, how do you solve that motivation problem?
- 26:20
Yeah, 100%. So, um, at the moment, I think our plan is largely not to necessarily develop products that compete with yours, but rather set-
- 26:30
Good use. [chuckles]
- 26:31
But rather set benchmarks and guardrails for how schools can then adopt whatever, whatever products they want. Um, in terms of student uptake, though, that's not something we've done too much on yet.
- 26:43
Um, the test that you just saw, uh, was our initial testing, which has been done, I think, with 70 teachers who were then role-playing their pupils. Um, so actually your experience might be quite valuable.
- 26:55
Um, so we'll-- we should chat after.
- 26:58
Ah, just one more.
- 26:59
Yeah.
- 27:00
So, uh, I'm from Norway. Uh, so it's very great to see your ambitions, and I'm sure that many countries across Europe are doing exactly the same thing. Are you doing any form from co- collaboration with other countries, experiencing and sharing ideas, et cetera?
- 27:16
Yeah, a bit. Um, Norway, no. But if you've got contacts, I'd be very happy to chat to them. Um, yeah, we, we do a bit. Um, there are a couple of teams that are, like, sorta similar to what we're doing, albeit a little bit differently.
- 27:32
Uh, there are a couple of different initiatives underway in the US government, which are not a million miles away, um, stuff like Tech- Techforce, um, parts of the US Digital Service.
- 27:42
Um, Singapore as well. We talk quite a bit with Singapore. Um, but yeah, we, we could do more. Um, so yeah, if you have any contacts in the Norwegian government, I'd be very open to it.
- 27:59
Done. [audience applauding] [upbeat music]