AI Engineer Europe 2026
How Building with AI Can Double the Throughput of Your Engineering Team
Read the talk
Doubling engineering throughput means changing more than code generation
Intercom’s 2X program combined a shared agent platform, tested internal skills and organizational change—then confronted the review and infrastructure bottlenecks that followed.
From a talk by Brian Scanlan
Before you start: Familiarity with pull requests, continuous integration and coding agents will help; no prior knowledge of Intercom’s internal tools is required.
An AI business still has to reinvent how it builds
Intercom’s opening chart poses a concrete question: how does an established SaaS company return to growth while its peers slow down? Brian Scanlan presents a U-shaped recovery in Intercom’s revenue growth against declining growth among publicly traded SaaS companies. The company had pivoted toward AI the weekend ChatGPT appeared. By this point, the privately held Irish American business had roughly 1,400 employees across six cities, with R&D led from Dublin and engineering concentrated in Europe, increasingly supplemented by forward-deployed engineers.
Scanlan follows the chart with a joke about shutting Intercom down live onstage, recalling a previous live deployment. The substantive reinvention is its customer-support agent, Fin, rather than an autocomplete field attached to an existing product. He points to New York Times coverage of Intercom’s transformation and reports more than 8,000 Fin customers, describing its average resolution rates as industry-leading.
Scanlan describes Fin revenue as approaching 100 million, without specifying a currency or reporting period in the talk. Fin’s original announcement coincided with GPT-4’s release; that announcement described internal testing and a forthcoming limited beta, so the timing should not be read as general availability that day. Intercom had already been building AI features since approximately 2018. Modern language models expanded what it could do with support questions, and Fin’s customers included Anthropic, Snowflake, Linear, Glean and LaunchDarkly.
The next step was a specialized model. Scanlan says Intercom’s own model serves all Fin English text conversations and is cheaper, faster and better at its customer-support task than frontier models such as Sonnet. He reports approximately two million resolutions per week and says direct access to the model suite is available. These are company-reported deployment and performance claims; the talk does not provide the evaluation conditions behind the model comparison.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Shipping culture meets marginal improvements
Inside engineering, the transformation was less immediate. Scanlan introduces himself as a senior principal engineer with 12 years at Intercom, working in the platform group. Its remit includes uptime, performance, security, cost management, observability, internal developer productivity and the company’s largely Ruby on Rails monoliths. Fast, iterative shipping was already central to its conception of product quality, captured in Shipping is your company’s heartbeat—an essay Scanlan remembers Honeycomb turning into stickers.
That made AI-assisted development an obvious investment. Engineers used GitHub Copilot, adopted Cursor and explored Augment. Yet by the middle of the preceding year, the results were disappointing: some tasks were marginally better or more enjoyable, but the organization had not fundamentally changed how it built software. The conviction that models and agent harnesses would transform knowledge work remained stronger than the results of simply distributing tools.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the change a measurable target
The response was 2X: double engineering throughput within a year without doubling the team. Intercom chose code changes per R&D person as its primary productivity measure, alongside developer surveys and tools such as DX. This made the ambition organizational: better individual experiences with an assistant should eventually produce a visible increase in delivery across R&D.
The measure was deliberately imperfect. Counting changes cannot capture every dimension of useful engineering work, and making a measure a target changes behavior. But those limitations were not a reason to avoid expecting a substantial throughput increase. The initiative, announced around the preceding June, became both a named project and a dedicated team.
The program also coincided with a major improvement in coding models and harnesses. Scanlan recalls a principal engineer’s astonishment around the Christmas break as capabilities changed, and explicitly credits that shift with contributing substantially to 2X’s success. The eventual outcome therefore reflects a changing technical baseline as well as Intercom’s organizational work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make adoption an organizational responsibility
Leadership made the expectation explicit. Intercom updated job descriptions so that adopting AI was part of meeting expectations for engineers, designers and product managers. The message had to recur across forums; a single announcement would not change established habits.
The supporting work combined reinforcement, practice and staffing:
- Visible examples: Automated Slack announcements surfaced skill updates, while teams celebrated useful work and shared techniques.
- Time to learn: Hackathons and AI immersion days gave people opportunities to change their working methods.
- Dedicated support: A growing, full-time 2X team helped hundreds of R&D staff adopt the tools.
For a medium or large organization, Scanlan recommends assigning strong engineers to this work full time. Mandating AI use without providing that support leaves each employee to solve the same adoption problems independently.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose a platform, then onboard the agent
Intercom standardized on Claude Code, replacing an approach in which people chose among several editors and agents. Scanlan compares the decision to concentrating on one cloud provider: spreading investment across platforms can prevent a team from accumulating deep operational knowledge and reusable infrastructure. Specific, consequential needs can justify exceptions, but continually reconsidering the preferred model should not consume the effort needed to make a platform work.
The ambition was to make Claude capable of acting like a senior Intercom engineer on technical work. Anything an engineer could do on a laptop should be accessible to the agent. That included access to real systems, governed by the company’s existing permissions, controls and audits—not unrestricted permission to alter databases. Those controls provided the basis for granting agents useful capabilities.
Access alone would leave the agent missing the knowledge that makes an experienced employee effective. Intercom had to teach it Rails conventions, architecture, React patterns, testing standards and security rules accumulated over years of development. The operating loop was straightforward:
- Use the agent for technical work.
- Observe where it takes the wrong path or misses a convention.
- Update shared guidance so the correction benefits subsequent work.
- Capture reusable knowledge in skills and use hooks to enforce required behavior.
A correction becomes more valuable when it improves the shared platform, rather than only rescuing one session.
Delivering that shared platform was itself operational work. Intercom pushed internal Claude plugins directly to employee laptops, bypassing Claude Code’s update mechanisms. Scanlan compares debugging installations across hundreds of laptops to managing Python installations: a capable agent still depends on reliable local setup and distribution.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Move beyond code production
The scope includes debugging, testing and planning as well as writing code. Engineers initially drive the agent closely, then aim to intervene less as the environment and guidance improve. The purpose remains delivering working products to customers. Scanlan believes the existing building blocks already support moving substantial amounts of the software lifecycle to an agent-first approach, even without further improvements in models or harnesses.
Written principles help hundreds of people interpret that ambition consistently. Giving an agent access to production can feel uncomfortable, but the larger shift is in the engineer’s role. Scanlan recalls working as a Unix sysadmin: visiting data centers, racking servers, running cables and configuring networks. Cloud infrastructure moved that work toward SRE and automation, with greater impact and higher pay. He sees the current transition as another move up the stack, compressed into a much faster industry-wide change.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Put durable knowledge into small, testable skills
Intercom’s technical conservatism carries over from its Rails monoliths: use a small number of tools extremely well. That favors investment in reusable knowledge over a proliferation of custom multi-agent orchestrators and opinionated workflows. Specific tools will change; a clear description of how to do work correctly at Intercom is likely to remain useful.
The preferred unit of investment is a small, durable, testable skill. Historical code changes, incidents and completed work supply examples against which to backtest it. Instead of judging a skill only by whether its instructions sound reasonable, the team can examine whether it handles known work at a high standard. Finding the right internal knowledge remains a challenge, so discoverability matters alongside correctness.
Continuous improvement includes making skills capable of updating themselves. Meanwhile, the platform should remain able to adopt capabilities as Anthropic ships them, rather than becoming trapped behind custom infrastructure. Intercom may eventually change providers; the durable investment is its accumulated working knowledge, not a commitment to rebuild every feature of the agent harness.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the agent a problem it can investigate
Explicitly asking an agent to run a particular skill is still useful, and Scanlan does it himself. But the desired interface is a problem description: let the agent infer the intent and choose the relevant capabilities. This removes the requirement that the person asking for help already know every available internal procedure.
A security incident made the distinction tangible. Someone had accidentally published Snowflake table metadata to a public GitHub repository. Scanlan opened Claude Code, asked it to join the incident’s Slack channel and investigate, and discovered that an internal skill already captured the company’s data-breach policies, assessment criteria and response procedure. He had not known the skill existed.
Claude downloaded the files, analyzed the exposure, concluded that it was innocuous and supplied next steps. Scanlan reports that the investigation took about two minutes, compared with his estimate of 20 minutes to locate the policy and do the analysis manually. The useful mechanism was the combination of accessible evidence, discoverable policy and a well-written skill; the timing is his account of this particular incident.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Progress from personal automation to a better environment
Adoption remained uneven even inside Intercom. A maturity rubric, which Scanlan compares with Steve Yegge’s discussion of engineering maturity, helps people see a path beyond occasional assistance. The progression is to use Claude Code broadly, automate recurring work, capture that work in skills, become proficient at writing skills and then improve the skills themselves.
The later stages move the focus beyond the individual session. Engineers write skills for others, improve them using evaluations and session data, and optimize the environment in which agents operate. That can mean changing software architecture, documentation or working methods to make effective behavior easier for today’s agents. The advanced stages on the slide describe contribution to a shared capability, rather than simply more frequent tool use.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Higher throughput moves the bottleneck to review
Intercom made the single-platform decision in December and began rolling it out in January. Scanlan reports that Intercom doubled PR throughput in less than a year. He presents a sharp inflection after standardization, although that sequence does not isolate the effects of platform choice from the capability improvements and organizational changes already described. A separate Claude Code PR dashboard figure is described only as being in the nineties, too imprecisely to establish an exact percentage here.
Code review then became the bottleneck. Scanlan reports a 17.6% automatic-code-approval rate in the dashboard shown during the talk. The measurement window and denominator are not specified. Getting there required more than asking Claude whether a change looked acceptable: Intercom backtested approvers against historical work, had humans label outputs, and used those results to establish confidence. It also shaped pull requests toward safe, simple changes that could be approved automatically.
Scanlan says Intercom worked with its auditors on SOC 2, ISO 27001 and HIPAA requirements. His account is that human approval is not inherently required for every change, but the organization must understand the approval process precisely and maintain appropriate audit controls. This is a description of Intercom’s approach, not a blanket compliance determination for other systems.
The approval system uses a tested suite of agents, including Codex for code review. That is an explicit exception to platform consolidation: a focused multi-model review process can justify additional tools. Scanlan expresses confidence that these controls prevent the automated process from degrading the environment or adding risk.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Observe usage and inspect what actually happens
Scanlan goes further, suggesting that well-defined agents can remove risks that human reviewers miss. Assessing agent behavior requires evidence, however, and he flags the earlier skill-invocation figures as unreliable. Intercom uses two distinct data paths:
| Data | Destination | Purpose |
|---|---|---|
| Basic skill-invocation events from hooks | Honeycomb | Understand which skills are used and where |
| Session transcripts | S3 | Mine sessions, write reports and assess skill effectiveness |
Scanlan says the internally available usage telemetry contains no private information. That statement concerns the basic telemetry; the separately collected session transcripts are a richer source of behavioral evidence.
The session archive closes the loop between written instructions and observed behavior. It can help the team see whether a skill is effective, rather than merely popular. Scanlan says the loop is already useful, while acknowledging that Intercom could extract more value from the data.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Look beyond the number of changes
Defect trends provide another view of the outcome. Scanlan says defects had been increasing until recently, but were now being closed faster than ever. Some teams deliberately pursued backlog zero and worked through large inventories of defects. Other reductions appeared to follow naturally from making the work faster. The observation combines planned cleanup with an improvement in day-to-day repair capacity.
Scanlan also reports improving code quality according to metrics used by a Stanford research group to which Intercom supplied its code. He does not identify the group or explain its methodology in the talk, so the claim supplies a directional quality signal rather than a reproducible quality benchmark.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Build the skill by doing the work
The shared plugin system had hundreds of contributors and thousands of lines of code. Base plugins handled session transcripts, session syncing and safety hooks. On top of that infrastructure, specialized skills captured particular kinds of engineering work.
One example was flaky-spec repair. Intercom had hundreds of thousands of tests, and frequent shipping had encouraged engineers to push past intermittent failures. Scanlan did not create the repair skill by sitting down and attempting to enumerate every possible cause of a flaky test. He gave the agent a goal, guided it through actual repairs and used the repeated work to develop the procedure.
The resulting skill organized accumulated knowledge into lookup tables, shortcuts and progressively disclosed detail. A compact Markdown outline illustrates that organization: keep the initial workflow easy to find, then expose the relevant repair knowledge when the investigation needs it.
markdown
# Repair a flaky spec
## Goal
Find and repair the cause of an intermittent test failure.
## Investigation
- Inspect the failing spec and the available failure evidence.
- Consult the relevant entries in the repair lookup table.
- Develop and validate a repair against the observed failure.
## Repair lookup table
Record recurring symptoms, their established causes, and
links to the detailed repair guidance.
## Improve this skill
After a validated repair, update the relevant guidance
with what the investigation established.
The important asset is the repair knowledge accumulated through use. Scanlan says the skill produces fixes that would impress him if they came from senior Rails engineers.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The consequences extend into infrastructure and product roles
More throughput also exposed infrastructure limits: Scanlan says CI became overwhelmed and had to be fixed, without detailing the remediation. Meanwhile, Claude Code spread beyond software engineering. That raised a broader organizational question: if the same agent platform supports engineering, design and product management, should the boundaries between those roles change?
Intercom was experimenting with single-person product teams. Scanlan describes shipping code that lets other people’s agents sign up for Intercom, while using internal skills to perform product-management work himself. The closing slide places these experiments alongside broader Claude Code adoption, replacing runbooks, remote agents and automatic approvals. The scope has expanded from producing code faster to changing who can carry a product idea through to delivery.
Scanlan expects these practices to become widespread soon. He closes by directing readers to his personal site and describing ways to interact with Fin through its messenger and a CLI. The substantive ending is the role experiment: once organizational knowledge and technical capabilities are available through a shared agent platform, an individual can attempt work that previously required coordinating several specialties.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Darragh Curran’s essay on continuous shipping, customer feedback, and engineering culture.
Further reading
Brian Scanlan explains concrete Claude Code workflows, including skills and hooks for pull-request creation and CI monitoring.
The March 2026 Apex announcement describes Fin’s specialized answering model and its reported deployment scale.
Fin’s original announcement explains its GPT-4 architecture, grounding in support content, limitations, and planned beta.
Updates since the talk
A follow-up defining Intercom’s per-person PR throughput measure and reporting changes in delivery, defects, and structural code quality.
Intercom’s automated-review design, pilot results, change-size restrictions, audit evidence, and human accountability.
Read the complete timestamped transcript
- 0:00
[upbeat music] Uh, hey, I'm Brian from Intercom.
- 0:17
Uh, and this has been such a great conference so far. I've learned so much alpha and inspiration from the talks and all the chats with people. Um, so Intercom is a 15-year-old privately held Irish American B2B SaaS startup that has pivoted to be an AI company the weekend that ChatGPT came out.
- 0:33
Uh, we've got about 1,400 people across Dublin, London, Berlin, SF, Chicago, Sydney. R&D is led from Dublin. Uh, engineering is almost entirely across Europe, four deployed engineers have kind of changed that up a bit.
- 0:44
Um, and this graph compares our revenue growth to the growth rate of publicly traded SaaS companies over the last few years. Um, and you can see publicly traded SaaS companies kind of on the down.
- 0:55
Intercom, this l- amazing U. Um, and like we're bucking this downward trend that SaaS companies, uh, have been suffering from recently. Uh, and now I'm gonna shut Intercom down live on stage. [laughs]
- 1:08
Uh, I, I once did a live deployment during a talk, and I thought that was impressive. Um, so like Intercom has become the poster child for companies redefining themselves in the age of AI.
- 1:18
New York Times recently did a article about SaaS companies reinventing themselves that prominently featured Intercom. Um, and being an AI company means a lot more just than s- slapping on, you know,
- 1:30
lightweight wrappers or, like, autocompletes in text field or whatever. Um, our agent, our AI agent for customer support, Fin, has over 8,000 con- customers, industry leading average resolution rates.
- 1:44
Uh, revenues like approaching 100 million. Launched the day GPT-4 came out, first product actually released on GPT-4. Uh, and we've been building AI features since about 2018 or so.
- 1:55
Um, and but, you know, the modern LLM models have unlocked huge capabilities for dealing with customer support questions, completely obvious. Um, and companies like Anthropic, Snowflake, Linear, Glean, LaunchDarkly use Fin for their customer support.
- 2:09
Um, so maybe SaaS isn't dead. Uh, and it works well for all size businesses. Um, also we recently announced that we have our own model serving 100% of Fin, uh, like English text-based conversations, um, outperforming frontier models like Sonnet.
- 2:25
Um, cheaper, faster, better, uh, and we're at like about two, 2 million resolutions a rate, uh, resolutions a week. Um, and we're also happy to sell direct access to our suite of models.
- 2:37
Uh, I'm not talking about any of this, though. Uh, so I'm a senior principal engineer at Intercom, been there for 12 years, and I'm on our platform group, and we take care of Intercom's uptime, performance, security, cost management, observability, uh, and our majestic monolith applications that we love, mostly Ruby on Rails, um, and all internal developer productivity.
- 2:56
Um, and another thing about Intercom is that we are obsessed with shipping. Uh, ship- sh- shipping fast and iteratively is the best way to build high-quality products that customers love to use.
- 3:05
Um, and so developer productivity is something we've always invested in. Shipping is the heartbeat of your company. It's great blog post that we po- did, like many, many years ago, and Honeycomb made cool stickers of it.
- 3:16
Um, and obviously for the last few years, I've been spending a lot of time, uh, on enabling use of AI in our software development life cycle. So I'm gonna kinda talk about that.
- 3:25
Um, and so unsurprisingly, we've been very excited, uh, about AI in general. You know, ch- changed the whole company to bu- to build customer support, uh, using AI agents.
- 3:35
Um, and we've been impatient about getting its adoption and changing how we build across Intercom. Uh, you know, we went down some kind of familiar routes. You know, we're all using GitHub Copilot, and, uh, then everyone started adopting Cursor, and, like, we looked at Augment and a few other things.
- 3:52
But ultimately, you know, say middle of last year, uh, we've been dissatisfied with the results. Uh, some good signs, some kind of tasks, some work, uh, made marginally better and kind of more fun.
- 4:03
Um, but you know, w- we- we're pretty aware of where the models are going and the harnesses, and we have a strong conviction that AI, uh, like f- for many years ago, is gonna change all knowledge work.
- 4:15
Uh, so last year, uh, middle of last year, we set a simple goal. Let's double the throughput of engineering in a year. Um, and you know, we measure a lot of things in Intercom.
- 4:24
We use like... We do a lot of developer surveys. Uh, we use tools like DX and, uh, but we picked code changes per R&D person as the primary way we're measuring productivity.
- 4:36
N- every measure is bad. Once you start measuring it, it's not a measure and all this. Um, but also, like, we're, like, impatient about, uh, or, like, expect the overall throughput to increase.
- 4:47
Like, if we're, uh, like actually adopting new ways of working, putting AI into all of the different places, uh, then we should expect a large throughput increase. Um, and so 2X, what we, we call this 2X.
- 5:01
It's the name of the project, and now a team and everything. Um, this is like wildly ambitious, like when we published this back last June or something like that, um, doubling productivity without doubling team size, but also kind of wildly unambitious as well if you, like, connect the dots and see where the models and coding harnesses are
- 5:17
going. So in this talk, I'm gonna talk about how we went about this, how we think about productivity, and a sneak peek at some of our internal data and skills and stuff.
- 5:25
Um, you know, this also coincided s- the work, uh, here with, like, the most noticeable shift in model capability and coding capability. Um, and so, you know, we've all seen it, and this was, like, one of our principal engineers, uh, posting just kind of like everyone else was, uh, in around the Christmas break last year, going like,
- 5:44
"Oh my God," like, "th- things have changed massively." And so, uh, that has contributed a lot to our success on 2X. Um, so this is the kind of engineering leadership-y part of the talk.
- 5:54
Uh, and so you need to be decisive and give clear executive guidance and, uh, you know-
- 6:02
Do organizational change. Uh, and, uh, we've done a lot of things. We updated job descriptions. If you're not adopting AI in Intercom, whether you're a designer, product manager, engineer, or whatever, you are not meeting expectations.
- 6:14
Binary. Um, and yeah, you have to say the same message over and over and over, 100 times, every different forum, whatever. You just gotta stay on message and constantly talk about the urgency of us doing this.
- 6:28
Um, you gotta reward us as well. Like, when people do good stuff, you gotta, like, all the Slack channels, um, uh, showing... Like, automating where people... Like automating, um, when people update skills or do this, that, and the other.
- 6:40
It's like, it gets put into these channels. We celebrate stuff. Uh, people are kind of showing each other different techniques and what's working for them and that kind of thing.
- 6:47
Uh, we've done hackathons, we've done AI immersion days, and, you know, all of these things are necessary to kind of bring people along. Like, um, also, we staffed this full time.
- 6:57
Uh, we have a team 2X, uh, that seems to be just keeps on growing and growing and growing. Um, and, you know, we're, we're not just saying, "Hey, you gotta AI everything.
- 7:07
Best of luck." Uh, we're, like, trying to bring everyone, like the hundreds of engineers, hundreds of people in R&D, along with us. So, you know, if you're in a medium or large organization, you absolutely need to have people, and, like, your best people, uh, on this full time.
- 7:22
Um, and so we chose Claude Clo- C- Claude Code as our platform. So, uh, prior to this, we were kind of omnivorous and, like, letting people choose their favorite editor and this, that, and the other.
- 7:34
And, uh, you know, there's, like, loads of people adopting co- Claude Code, loads of people using Cursor, loads of people using Augment. Um, but, uh, we, like, we're a believer in platforms in general.
- 7:45
Um, and it kind of doesn't matter what you choose. Uh, but choosing one is important. Uh, you know, to a certain extent, you need to get away from model anxiety.
- 7:53
It's like being multi-cloud. It's like you don't get the compounding benefits of a well-designed platform if you're sending all your different work across different cloud providers or whatever. Um, you're way better off being all in on one and optimizing it and proving that it works.
- 8:08
Um, and, like, unless there's, like, very specific or impactful reasons why you need to be spread across multiple agents or whatever. Um, and so our vision on this was, like, to treat Claude or to, like, work on, to get Claude to be able to act like a senior engineer on any technical task across of Intercom.
- 8:24
Um, and our vision here was, like, connect Claude to everything. So anything I do on my laptop, Claude should be able to do that. And that means everything. Like, uh, now of course, we're not reckless.
- 8:33
We're not, like, just trying to, uh, let the thing go off and delete all of our databases, but we're, like, a mature company. We've got plenty of controls and permissions and, uh, audits and everything like that.
- 8:45
Uh, that gives us a lot of confidence to be able to, like, unleash Claude in the same way that we unleash our engineers in our environments. Um, and you know, we gotta onboard it.
- 8:53
We gotta teach it all the stuff that we teach people when they join Intercom. Uh, all of our Rails conventions, like our architecture, React patterns. Like, we've built a lot of software in 15 years.
- 9:03
Um, but, like, stand- testing standards, security rules, all this, Claude h- absolutely has to know the Intercom-specific information to be able to do the job. Um, and, uh, and most importantly, start using the platform for all technical work, and it doesn't get things right first time, hits an issue, goes down the wrong path, update the guidance.
- 9:21
Like, this is a flywheel that we're all contributing to. Um, and so we've encapsulated a lot of this knowledge in context and engineering, captured the skills, guidance, hooks to force these things.
- 9:29
We spent a lot of time cajoling Claude Code to work well. Uh, we do things like, uh, push out our internal Claude plugins, um, to everyone's laptops, like bypassing the, all the Claude Code updates, uh, mechanisms.
- 9:43
Um, because just, it's, you know, it's, uh, you spend a lot of time debugging, uh, Claude Code installs on like hundreds of laptops. It's like trying to install Python or manage Python installs or something.
- 9:53
Um, and so ultimately though, like every single part of technical work, so it's not just code production. It's not like more advanced autocomplete. It's everything. Um, so debugging, testing, planning, all this kind of stuff, uh, it should just be you driving Claude, and ideally, like, driving it less and less and moving higher up the, the, the, the
- 10:11
food chain. Um, and it, you know, delivers real value, delivers the code, products, whatever, to, uh, customers. So everything's in scope. And, like, we think that even if the models and harnesses do not improve at all, uh, which is not, definitely not happening, um, if anything, like, this capability curve is, uh, is accelerating.
- 10:31
But, like, the building, we have the building blocks today to, um, to improve... Like, basically move vast amounts of work, uh, in our software development l- lifecycle to be agent-first.
- 10:40
Like, they could just pause everything, and we've just, uh, got this flywheel, and we're going through everything and looking at every single piece of work. And, like, the tools are good enough today to do this.
- 10:52
Um, so we wrote some principles to help guide us along the way. You know, when you have hun- you're trying to get hundreds of people to change how they work or understand what we're trying to achieve, you need to write things down and help them out.
- 11:03
Um, and you know, different principles should apply in different places. But, like, um, you know, we believe that all of engineering is changing. Everything that you can do, the agent must be able to do.
- 11:13
Uh, and that, that can feel weird as well, like, when you're first connecting it into production systems or whatever. Um, and, uh, yeah, like the, our job is moving up the stack as engineers, as product builders, whatever.
- 11:25
And, like, if I, like, a long time ago, I used to be a, a Unix sysadmin. Um, and, uh, you know, you'd, like, going into data centers, racking servers, uh, cabling things, configuring networks and all that.
- 11:38
Um, and then the cloud came along, and I moved up the stack. I, you know, and pe- people transitioned from being sysadmins to SREs. The work was more automation oriented, more impactful, higher paid.
- 11:50
Uh, and so I think this is like we're kind of speed running this 100 times faster on a full industry scale. Uh, but I kind of feel like I've been through this before.
- 11:57
Um, we at Intercom are technically conservative. We like using single tools and just using them extremely well. Um, so hence we end up with these Ruby on Rails monoliths and stuff.
- 12:07
Uh, and so we're kind of applying this thought process as well to like, uh, you know, what is the... Where should our focus be? Where is our attention? Do we want everyone writing their own multi-agent orchestrators or opinionated workflows?
- 12:19
And, you know, the, we want to build durable, testable, high-quality components and people to be considering like the lifetime value of what they produce. And like, you know, it, the, the tools, the specific implementations of these things will change over time, but, uh, I'm pretty sure that writing down how to do work in Intercom will be valuable
- 12:36
no matter what happens. Uh, maybe it might be easier to discover in the future. That's like, uh, a, a problem at the moment. Um, and so what this, what this really means in practice is that we spend our time focusing on small, high quality, durable, testable skills that do the job extremely well, that we can, you know,
- 12:52
use data, use backtesting. We've got like all of the work. We've got this huge body of work and changes in code and incidents and everything, and so we're using all of this to help form us and prove out that these skills are operating at extremely high quality.
- 13:06
Um, and, you know, we, uh, then we, and we also practice continuous improvement here, get these things to be self-updating, um, make sure that these things are very high quality.
- 13:15
Um, and, uh, yeah, we don't want to get stuck behind the curve, like getting stuck because we've implemented a lot of our own s- own, own things. We just want to use things as they become available, as in topic ship or whatever.
- 13:28
Uh, and maybe we mightn't stay on topic forever, but like, uh, we're very, we, we're eager to get the advantage of somebody else building and shipping great software and capabilities rather than us having to build everything ourselves.
- 13:38
So, uh, yeah, another thing we guide people to, to do is like you want to give problems agents, not tasks. You know, a lot of the time people even say in Intercom are saying like prompting agents, "Hey, run this skill to do a thing."
- 13:49
Uh, which is mostly fine and still kind of necessary. I still do it, um, a lot. But like, uh, we're more kind of having to like move ourselves to be kind of just describing the problem or de- de- des- describing the task, uh, and let, let the, uh, agent figure out what skills to invoke and what to
- 14:05
do here. Um, I've... Fun story. Recently, I was brought into a security incident. We had accidentally published some kind of
- 14:13
Snowflake table metadata to a public GitHub repository. Uh, and I just habitually, uh, opened Claude Code, told it to join a Slack channel, take a look. Um, and I didn't even know that a skill existed that, uh, actually perfectly encapsulated all of our like data breach policies and criteria and what to do, how to analyze this.
- 14:33
Claude just automatically downloaded, uh, the, the files, did full analysis, concluded it was innocuous, told me all next steps. Um, and I, like I didn't tell it to do this.
- 14:41
It just kind of figured it out. It was done in like two minutes. Um, and like that would've been a 20-minute task and kind of boring work. I'd have to go, "Oh, where's that policy?"
- 14:48
And take a look at this, that, and the other. Um, and like this just felt like a little... Like it was a small example, but it's like, again, I just like gave it the problem of like taking a look security incident, and it just figured out the intent, uh, and used a well-written internal skill that did this
- 15:03
job for me. Um, and yeah, it, I mentioned that it, even at Intercom, like AI adoption is unevenly distributed. I think we're ahead of the vast majority of companies.
- 15:13
Uh, but you still need to help people understand where they're at and grow towards being highly effective at using agents in their work. Uh, CV Aggie recently talked about like maturity rating for engineers, and like our internal one is kind of similar here.
- 15:24
You're kind of like trying to get through these different kind of levels, and like ultimately you kind of end up like mastering all the skills and like knowing the tool inside out.
- 15:32
And ultimate- like what we want people to do is like use Claude Code for everything, automate your work, then move that to a skill, then get really good at writing skills, and then writing s- write skills and improve the skills, uh, and then optimize the environment for agents.
- 15:45
That could be everything from software architecture, maybe just to documentation or other approaches or other ways of doing things, uh, that allows the agents to be even more effective and optimized for what they're great at today.
- 15:55
So here, here's where we're at. As you can see, yeah, wild inflection points, uh, after, uh, after going all in on one tool. That, that decision was made in December.
- 16:04
We started rolling it out in January. Um, and we've been just, like we have reached the doubling PR throughput in faster than one year. Um, here's more like data from our internal dashboards.
- 16:17
There's some interesting stuff in here. This is like, yeah, number of pull requests auto Claude Code. It's like in the 90-somethings. Um, you can see also, uh, we're starting to move into...
- 16:28
Like our current bottleneck is, uh, code review. And, uh, but you can see we have this like 17.6%, uh, approval rate, uh, of our automatic code approvals, and it's like a lot more, uh, in-depth than just like, "Hey, Claude, can you approve this?"
- 16:44
Um, we've gone through a lot of detailed work to figure out, again, using backtesting and previous data, uh, and then getting humans to kind of label the outputs and figure out, like get the confidence level of the automatic approvers and kind of shape the pull requests towards very safe and simple, uh, pull requests, which probably always should
- 17:04
have been that way. Um, but now like they're just approved automatically. And, you know, we've also worked with our auditors to ensure that we're fully SOC 2, ISO 227001, HIPAA compliant, all that.
- 17:13
You do not need humans in the loop to, uh, to meet these certifications. You need, you do need to know exactly what you're doing though and make sure you've got like auditing controls and everything.
- 17:23
Um, and so by moving approvals to an extremely well-organized, tested, and competent suite of agents, um, including codecs for code reviews, uh, I think multimodal code reviews are okay.
- 17:31
I just like completely went back on my platform thing. Um, uh, like, uh, we've got a high confidence that like this stuff is not de- degrading environment or adding additional risk.
- 17:41
In fact, I think it's removing risk because humans aren't actually as good as, uh, agents, like when they're well-defined. Um, here's like skill invocation. I actually think the earlier numbers were a bit wonky.
- 17:51
So like we hook up everything into Honeycomb, all the... We've got hooks all over the place, um, for basic information about like which skills are being, uh, invoked and things like that, and that's internally available.
- 18:04
There's no private information, uh, in this, and everyone can kind of use it to kind of get an idea of like what's being used and where. But we also pull in all s- sessions transcripts, uh, into S3 for data mining, writing reports, guiding, like also looking to see are skills effective, that kind of stuff.
- 18:18
Um, so we've got like a feedback loop using the session data, uh, which is, uh, we, we can get more out of it, but it's, we're, we're, we're doing some interesting stuff with it already.
- 18:27
And this isn't a goal, but like, and we're not particularly proud of, like, defects always increasing, uh, up till recently. But, like, defects are getting closed faster than ever.
- 18:35
And, like, some teams have been inspired by the move to AI to, uh, think about things like backlog zero or crunching through hundreds or thousands of, uh, defects. Um, so, like, some of this was, like, a bit deliberate and planned, but just in its...
- 18:50
At the same time, there's just, like, this natural deflation 'cause, uh, getting through this work, getting through all the defects so much faster these days. Um, and, uh, yeah, it's like we're, we're just seeing this naturally.
- 19:02
We've also been working with, like, Stanford. There's a research group there. Uh, we've, um, we give them all our code, and our, our code quality per their metrics has been increasing over the last while.
- 19:10
Um, okay, uh, I'm kind of running out of time at this point. Uh, we have, like, hundreds of contributors, thousands, like, thousands of lines of co- of code, uh, in our Claude Code plugins.
- 19:24
Um, it's very active. Uh, and, uh, yeah, c- I mean, Claude Co- Claude itself loves it. Um, here's an example skill. This is, like, not the most... Or, sorry, we got like b- base plugins, things that, like, do all the session transcripts, session syncing, um, uh, some safety hooks and things.
- 19:40
Um, and here's, like, a s- a skill I built which, like, it just s- it fixes flaky specs. We have hundreds of thousands of s- of tests and, you know, they get a bit flaky, uh, over time and, uh, we don't...
- 19:53
We ship a lot, so we just kinda bar- barge through the kinda flakes. Um, but this, this skill was not built, like, by me kinda sitting down and figuring out just like, "Oh, what are, what are all the things you need to do to fix flaky specs?"
- 20:04
I've worked in a feedback loop, gave the, um, gave the agent a goal, and, uh, through, like, guiding us to the right place, uh, and working with us to fix a lot of flaky specs, uh, it's written this pretty decent thing.
- 20:18
What are, like, these cheat codes or, like, lookup tables, and, uh, relatively well-organized, it's using pr- progressive disclosure and all that. Um, and, uh, it is like fixing stuff that if our more senior Rails engineers were doing this, I'd be like, "Wow, they're amazing."
- 20:33
Um, and yeah, like, other, a lot of other stuff going on, like our CI melted, we had to fix that. Um, Claude Code is actually widely used across Intercom outside of software.
- 20:42
It's gone completely viral. People are banging down our doors to, like, use cons- news consoles. Um, and, uh, yeah, we're, you know, we're thinking a lot about, like, the future of engineering pro- like, should we just merge all product manager design, everything?
- 20:55
Um, oh, yes, the single-person team product experiments have been pretty interesting as well. And I've even been shipping, like, code, like, stuff that people can use in their agents to sign up to Intercom.
- 21:04
Uh, this is stuff that, like, I, like, I've just been using our skills to act as a product manager, which is pretty wild. Um, so that's it. I wish you all the best of luck.
- 21:13
If you're not doing pretty much all of this today, you're gonna be doing it in the very near fu- future. Um, my contact details are at brian.scanlan.ie. You can interact with Fin the messenger, configurable CLI, um, and you can check out, uh, ideas.fin.ai for a lot more information about Intercom and our agents.
- 21:28
Thank you. [clapping] [outro music]