AI Engineer World's Fair 2026

AI Writes More PRs. Who Validates Them? — Ali-Reza Adl-Tabatabai, Sonar

Ali-Reza Adl-Tabatabai· Gitar (now part of Sonar)11:34

Read the talk

AI Writes More PRs. Who Validates Them?

Ali-Reza Adl-Tabatabai explains how Gitar automates the path from code review and CI failures to a green pull request—and why workflow orchestration, gradual trust and program analysis matter as much as the agent.

From a talk by Ali-Reza Adl-Tabatabai

At a glance

Ideas worth remembering

  • AI-generated PR volume moves work into CI and review, where failures can add hours or days of coordination and delay.

  • Agentic validation connects review and CI diagnosis to retries, repairs and a loop toward a green PR. Team-defined conditions govern automatic approval and merge.

  • The observed adoption path grants capability gradually: review first, then blocking, autofix and rule-controlled merging.

  • A validation-specific harness lets Gitar optimize token cost, precision and coverage, while orchestration and model routing remain separate responsibilities.

  • Combining agents with SonarQube program analysis is the proposed next step toward better precision and coverage; the talk presents the direction without a comparative evaluation.

Faster code generation leaves validation on the critical path

A pull request can be quick to write and slow to land. CI and code review sit between a proposed change and production, enforcing security, compliance and code-quality checks. Ali-Reza Adl-Tabatabai, a Gitar founder and former CEO, opens with this centralized validation phase. Gitar had recently become part of Sonar, and its focus was the work required to verify code before shipping it.

Validation gets harder as an engineering organization grows. Builds and tasks run longer; reviewers work asynchronously across sites and time zones; more tools need integration; infrastructure costs rise. The platform team inherits a system whose speed affects every developer’s ability to get changes into production.

A failure adds an “outer loop” measured in hours and days: the change needs another round of attention and validation before it can proceed. The cost includes developer time and the interruption of work already underway. Improving that loop is difficult because the underlying system consists of fragmented tools connected through bespoke configuration and scripts, often maintained by a constrained platform team. Automating it requires coordination across asynchronous workflows and integrations, beyond interpreting a single error message.

0:130:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:13 · section reference included

More and bigger PRs move the cost downstream

AI increases both the number and size of pull requests, each of which can contain defects. That moves development costs downstream: more code requires review, and missed defects can become production failures. Careful manual review slows delivery; rubber-stamping preserves movement while increasing the risk of incidents. Either outcome can erode the productivity and developer satisfaction that faster code generation was supposed to improve.

“Of course, it’s more AI” is Adl-Tabatabai’s answer to this bottleneck. Gitar applies an agent to validation, with a concrete target: deliver a green PR ready to merge, or carry it through an automatic merge. That target makes workflow completion central to the product. Finding an issue is only the beginning of the work.

2:563:27
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:56 · section reference included

Review findings and CI failures feed a repair loop

The first step is familiar: creating a PR triggers review, and the agent posts issues and inline comments. Gitar’s design goal is to keep that experience quiet enough to be useful, through both its interface and its selection of findings that matter. Teams can supply their own review rules, custom checks and workflow automations. They can also make unresolved findings block merging, turning a comment into an enforced quality gate.

CI analysis adds another source of feedback. The agent summarizes the root causes of failures and identifies flaky tests. A team can configure automatic retries for those tests. This distinguishes two responses to a failed check: retry a task believed to be flaky, or repair an issue identified in the code or CI workflow. The usefulness of the distinction depends on correctly recognizing the failure.

Consider the flaky-test example as a walkthrough of the configured behavior. A PR has a failing CI check. The agent analyzes the failure and, if it identifies a flaky test covered by a retry rule, retries that test. If unresolved review findings or CI failures remain, the autofix capability can address them; the team can request a specific fix or enable a loop that continues until the PR is green. A green PR can then receive automatic approval and merge when the team’s conditions permit it. The observable change is from a PR waiting on failures to one eligible to land, with diagnosis, retry, repair and approval doing different jobs along the way.

Selected presentation frame from AI Writes More PRs. Who Validates Them? — Ali-Reza Adl-Tabatabai, Sonar at 331 secondsOpen full source frame
A pull request interface shows an approval and merge step.

Where does the loop continue, and where does it stop? The diagram separates validation feedback from the final merge decision. Remaining issues send the PR back through repair and validation; reaching green makes it a candidate for rule-controlled approval and merge. Passing checks is therefore one condition in the workflow, rather than an unconditional instruction to ship.

How it fits togetherFrom validation feedback to a rule-controlled merge

The proposed change enters validation.

Review and CI supply issues to resolve. Autofix repeats until the PR is green; configured conditions govern automatic approval and merge.

4:044:34
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:04 · section reference included

Trust grows from comments to permission to merge

Gitar’s users do not necessarily enable the whole workflow at once. Adl-Tabatabai describes a progression that starts with reading reviews and judging whether the findings are accurate and important. Once those reviews earn trust, teams let the agent block PRs. They then use autofix to resolve review findings and CI failures. The larger leap is allowing the agent to approve and merge under explicit rules.

Selected presentation frame from AI Writes More PRs. Who Validates Them? — Ali-Reza Adl-Tabatabai, Sonar at 361 secondsOpen full source frame
A slide arranges review, block, fix, approve and merge along a rising trust path.

Each step changes the agent’s effect on delivery. Comments advise; blocking prevents a change from landing; autofix changes the PR; approval and merge complete the workflow. This progression lets teams build confidence before granting the next capability. The additional value comes from removing more of the human coordination between those steps.

5:456:15
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:45 · section reference included

Processing every PR also reveals recurring work

The same system that works on individual PRs can categorize activity across them. Two audiences get different views:

  • Platform teams: recurring CI failures. Failure categories distinguish flakiness, infrastructure problems and other causes. That helps teams decide whether to improve task reliability or address the infrastructure behind the checks.
  • Engineering leaders: the mix of landed PRs. Categories such as feature development, issue fixes, chores and refactors describe what kinds of changes the team is shipping.
Selected presentation frame from AI Writes More PRs. Who Validates Them? — Ali-Reza Adl-Tabatabai, Sonar at 422 secondsOpen full source frame
A slide shows an insights interface with a text panel and a donut chart.

These views answer different questions. CI categories help locate recurring obstacles in validation; PR categories describe the composition of engineering activity. Their value comes from aggregating information that would otherwise remain scattered across individual reviews and failures.

6:447:14
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:44 · section reference included

The runtime is built around validation economics

Gitar separates three responsibilities under the hood:

  • Control plane. An orchestration layer handles PR workflows.
  • Custom agent harness. The runtime handles multi-agent execution, context and memory management, tool calls and integrations.
  • LLM proxy. A separate component routes requests across models and handles failover.
Selected presentation frame from AI Writes More PRs. Who Validates Them? — Ali-Reza Adl-Tabatabai, Sonar at 482 secondsOpen full source frame
A system diagram connects web, control plane, agent runtime and LLM proxy components.

The custom harness is an explicit product decision. Owning the runtime lets Gitar optimize token cost, precision and coverage for validation. Those concerns pull in different directions: the system needs to find relevant issues without flooding reviews, while controlling the cost of the work needed to reach a result. Adl-Tabatabai connects these optimizations to offering outcomes at a fixed price per PR; the recording does not specify the price or the terms of the outcome guarantee.

7:448:14
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:44 · section reference included

Agents and program analysis contribute different capabilities

The technical ending turns to Gitar’s integration with SonarQube. The proposed combination brings the agentic validation workflow together with taint analysis, control-flow analysis, data-flow analysis and software composition analysis. These are algorithmic ways to examine code and its dependencies, alongside the agent’s ability to review, explain failures, make repairs and coordinate the PR workflow.

The integration aims to improve precision and coverage at an attractive cost. Adl-Tabatabai presents the combination as an opportunity and expects it to outperform either approach alone; he does not detail how analysis findings enter the agent or provide a comparative evaluation. The useful design direction is nevertheless clear: combine specialized code analysis with a system that can act on validation feedback through to merge.

The reported benefits span better code quality and security, faster delivery, less review work, improved developer sentiment and more useful tools for platform teams. These are qualitative product observations in the recording, rather than measured effect sizes. The causal ambition is to reduce waiting and grunt work while giving teams more confidence in the changes they ship.

The closing requirement is to keep validation moving at the rate AI can generate code. Gitar’s answer extends from the created PR through review, failure analysis, repair and controlled merge. Program analysis supplies another way to examine that code. Together, the workflow and the analyses aim to make verification scale with production of changes, rather than leave developers accumulating an ever-larger queue of code to inspect.

8:449:14
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:44 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:13

    Let's get started. Um, good afternoon, everyone. I'm Ali, and, uh, I'm one of the founders of Gitar.ai and the previously the CEO up until about a month ago, where we were acquired by Sonar, and now we are part of Sonar. And you've heard a lot in the past few days about the importance of verification in the age of AI. And so today, I'm gonna focus my talk on code and

  2. 0:42

    verifying code. So in m- in almost all software teams, there is this centralized validation phase that is basically CI plus code review. And as you scale your engineering, this plays-- this phase plays a very critical part, uh, especially in an AI world. Uh, first, it enforces all your quality gates centrally. So these are security, compliance, code quality gates. It's on the critical path of getting your

  3. 1:12

    changes into production, so the faster is always the better. But as you scale engineering, this becomes kind of a challenge. As, as you have longer running builds and tasks, you have asynchronous team review flows across sites and geographies, and things just get slower naturally as you grow. Um, this is also a phase where you end up spending a lot of developer infrastructure dollars. So there's a lot of expensive tools here that need to be integrated and

  4. 1:42

    scaled up as you, um, scale your organization. And it all runs on very expensive infrastructure. Uh, for these reasons, it's oftentimes the main focus of your platform organization team or your platform team. Um, but any kind of failure in this phase is very expensive. It introduces... Failures introduce an outer loop that are typically measured in hours and days. So

  5. 2:11

    failures are expensive. They costly. They really slow down your velocity. Uh, you lose developer time, and that's compounded by disruptions to your flow state. Uh, it's also very expensive to optimize. So you've got a lot of fragmented systems that have been integrated together with bespoke configuration and scripting, and you have a usually headcount restricted, constrained platform organization trying to keep it all up and running for you. And, uh, trying to address that all with

  6. 2:41

    AI is also kind of complex because it requires complex workflow orchestration, uh, involving various inter-- asynchronous interaction, team workflows, integrations across tools, and so on.

  7. 2:56

    In fact, AI kinda makes the problem even worse. Uh, you've got a lot of PRs now being generated, more PRs than before, uh, thanks to AI, and bigger PRs, uh, all of which can have defects. And so your costs in your development life cycle are kinda shifting to the right. You have a lot more code to review. You, you can now potential for a lot more failures in production. Uh, and so your choices are kind of tough here. Either you slow down and ask developers to review every PR very carefully,

  8. 3:27

    or you rubber-stamp PRs and you risk incidents in production. Either way, you've got

  9. 3:34

    less happy developers, lower sentiment, and much less productivity than, than you'd expect. So what's the answer here? Of course, it's more AI. It's, uh, an agentic approach to validation, which is essentially what we built at Gitar, um, that we're now also shipping, uh, as part of Sonar. And basically, what we've got here is an agent that fully automates all the validation steps, and as outcomes, it delivers

  10. 4:04

    for you green PRs that are ready to merge or merge automatically. Now, how does this work? Well, first, when you create your PR, it reviews your PR and posts issues and inline comments. No surprise there. But our goal has been to really make this experience as noise-free as possible, both in the way we handle the UX, as well as posting issues that are real and really matter. You can customize your review with your own rules, and you can add your own custom checks, and you can trigger your own

  11. 4:34

    custom workflow automations. And then you can also configure the system to block merging if there are any unresolved code review findings.

  12. 4:45

    It also analyzes CI failures and posts a summary of the root causes of those failures, and in particular, it's really good at catching flaky tests, and you can even set up rules that automatically retry flaky tests if you want. That's a very popular feature. You can also automatically fix any code review issues or any CI failures, any kind of, uh, failures that happen in CI. It can fix it for you as well as address any of the issues that it raises. You can either ask it to specifically

  13. 5:14

    fix an issue, or more interestingly, you can set it up to automatically loop until all the fish-- issues have been fixed in your PR, and you have a green PR ready to merge. And finally, you can automatically approve and merge PRs that are green. You can define rules and conditions that control when, uh, PRs get automatically approved by the agents and merged automatically.

  14. 5:45

    So we're seeing actually a very interesting trend in our users. There's a clear trajectory towards fully automated PRs. Uh, as users use the system, as they build trust in the accuracy of the AI, what we see is they work towards full automation. First, they get reviews, and they're happy with... They build trust in the accuracy of the reviews. They see that these reviews are not only accurate, but they're raising really important issues. So then they give teeth to the agent by

  15. 6:15

    allowing it to start blocking the PRs, and then they start using the autofix capabilities to automatically turn PRs green and fix all the issues that were raised by code review or any CI failures that come up. And then where the real interesting leap is beginning to happen now is users setting up rules that define when the agent automatically approves and merges PRs. So over time, as they build trust in the

  16. 6:44

    system, they unlock more levels of automation, and with the additional levels of automation and with the trust, they're getting more value out of the system. Now because the system processes every single PR using AI, it's also able to uncover new types of insights. For example, uh, one of the insights that are ve- is very useful to platform teams is a categorization of the top CI failures. Where are you seeing failures in your CI system?

  17. 7:14

    Where are your developers running into failures? Are these flakiness issues that you can addre- that you need to address by improving the reliability of the tasks? Or are they infrastructure issues that you need to address? Or are there other kinds of failures that they're running into? Very, very useful for platform teams. We also have... Another example of what else we have in terms of AI-enabled insights are categorization of PRs. What exactly is happening across your engineering team in

  18. 7:44

    terms of the PRs that are being landed? This is extremely useful for leadership teams. For example, what amount of PRs are going to feature development versus fixing issues, chores, refactors, and so on. Now, under the hood, if you look at the system, there's a number of components that constitute Gitar. There are three interesting pieces here. One is the control plane that acts as this orchestration layer for handling PR workflows. And

  19. 8:14

    then there's an agent runtime or, or harness, we built our own harness here, that handles multi-agent execution, context management, memory management, um, tool calling, all the integrations, et cetera. And having this, having our own agent harness has really allowed us to optimize specifically for the validation use case, and it's allowed us to optimize for token cost, precision, and coverage of the agent itself. So that allows us

  20. 8:44

    to essentially guarantee outcomes at a fixed PR price while giving really good quality results. We also have an LLM proxy that handles the routing across different models as well as things like fa- failover. Now that we're a part of Sonar, we've been integrating Gitar with SonarQube. There's a ton of opportunity here in combining an agentic approach to validation with program analysis. So combining

  21. 9:14

    agents with things like, um, taint analysis, control flow analysis, data flow analysis, software composition analysis. So all the traditional approaches combined with agents. We think there's a ton of opportunity here, and that together you can get much better precision, much better coverage at a really great price point. So the overall results of the value that we're seeing in, in all of this is much better quality, much better

  22. 9:43

    security, especially with AI-generated code, faster delivery times through all the automation going from PRs straight to approved and merged PRs, much higher productivity and sentiment by developers, so less time spent on actually reviewing code and worrying about what's happening in production with more confidence, as well as sentiment because now you... Much better sentiment be- now because you have a lot of automation in place and less grunt work. And finally, the to- the features that we've built for

  23. 10:13

    platform teams gives them a lot more leverage to be able to fix issues and to add customizations and see insights in this critical part of the software development life cycle.

  24. 10:26

    To wrap up, PR validation is really critical, even more so in AI. But AI turns validation into a friction point, and Gitar, what we've built, uh, allows you to automate this step, fully automate PR validation, go from created PRs all the way through merge, completely limiting this friction, and helps you keep up with the rate at which you're generating code from AI. And finally, the combination of

  25. 10:56

    using a- agents, uh, doing agentic validation plus an algorithmic program analysis, it will give you something that, that I think is gonna be much better than either alone. If you wanna talk more and you wanna see more, if you'd like to see a demo, come and visit us at booth number P7 right down, right down the hall there. Thank you.