AI Engineer World's Fair 2026
AI Hackers Are Faster Than Your Pen Test — Eli Cohen, Snyk
Read the talk
AI Hackers Are Faster Than Your Pen Test
Eli Cohen explains Snyk’s approach to continuous offensive security: test each code change, give specialized agents application context, and validate exploits before asking developers to fix them.
From a talk by Eli Cohen
At a glance
Ideas worth remembering
Application context matters because attackers can chain smaller weaknesses and exploit business rules that predefined payloads do not fully capture.
Cohen describes AI pen testing on every PR, with separate agents for reconnaissance, vulnerability hunting, exploit validation, remediation and reporting.
Static and dynamic scanner findings can guide AI testing, while established scanners continue to cover known attack patterns.
Evaluate AI pen testing through its cadence, application reasoning, context inputs and evidence of exploitability.
More code creates more work for defenders—and attackers
Coding agents can increase software output faster than security teams can clear the resulting backlog. Eli Cohen, former co-founder and CEO of runtime security company Helios and now working on AI security at Snyk, opens with that mismatch. Faster development adds code to examine, while attackers use the same class of tools to search for weaknesses.
Cohen cites Snyk data showing 218% more lines of code per developer, more than 85% of developers using coding agents, and 62% of LLM output being insecure or broken. He also cites successful AI attacks averaging 34 minutes, with the fastest taking four minutes, and vulnerabilities in 43% of MCP servers. The recording does not establish the populations, testing conditions or measurement methods behind these figures; they frame his urgency rather than predict the risk of a particular application.
The attack mechanism matters beyond the headline speed. Several low-severity vulnerabilities can form a critical attack when chained together. Application context helps an attacker work out how one weakness makes another useful: what the application permits, which data is reachable, and how its operations fit together. Cohen identifies this same contextual reasoning as a weak spot in generated code. The defender’s problem therefore includes relationships between flaws, not just the number of findings in a scanner.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Source patterns, running behavior and business logic need different tests
The traditional tools examine different parts of the problem:
- Static application security testing (SAST): scans source code for patterns such as SQL injection and cross-site scripting. It is relatively cheap to run on every change.
- Dynamic application security testing (DAST): sends predefined payloads to a running application’s API endpoints and examines its behavior. Runtime configuration and authorization are important reasons to test dynamically.
- Penetration testing: uses a security researcher’s understanding of the application’s purpose and business rules to find ways to break them.
Broken Object Level Authorization, or BOLA, supplies the concrete example: user A can access user B’s data. The meaningful change is from an expected restriction—A should not receive B’s data—to an observed disclosure. Probing the running application makes that consequence visible. A suspicious source pattern alone does not demonstrate that the deployed application actually returns another user’s data. Cohen makes the stronger claim that static analysis cannot find these vulnerabilities; the example specifically establishes why runtime testing matters for confirming the access failure.
A predefined payload collection can cover many known attack patterns, but business logic asks a different question: does the application allow an action that violates its intended rules? Understanding the user A/user B distinction requires knowing what access is legitimate. Human pen testers bring that contextual understanding to their attempts to break the application.
That depth costs time and money. Cohen describes companies commissioning human pen tests once, twice or perhaps three times a year. The work then continues through a PDF report, developer interpretation, repairs and retesting. His comparison to “a movie from the ’90s” targets this slow handoff. A valuable test loses freshness as new code arrives between assessments, and developers still have to determine which findings deserve action.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Move contextual testing into every code change
Snyk’s Evo Continuous Offensive Security offering brings together AI pen testing, agent red teaming and dynamic testing. Agent red teaming extends the target beyond ordinary application endpoints: it simulates multistep attacks involving prompt injection or data exfiltration. Dynamic testing continues to examine runtime weaknesses such as authorization failures. These capabilities address different ways a production system can be attacked.
The proposed cadence is every PR and every code delta. LLMs supply the contextual reasoning that previously depended on a human assessment, while developers receive findings and guidance on what to change. This is Cohen’s description of the product’s intended operation; the recording does not quantify per-PR runtime, cost or coverage, so continuous execution should not be read as a guarantee that every change receives an exhaustive assessment.
The second requirement is exploit validation. A growing backlog becomes less useful if it fills with plausible but unconfirmed bugs. The system aims to prove that a finding can be exploited before sending it onward. In the authorization example, that distinction means establishing whether user A can actually obtain user B’s data, rather than merely flagging a place where access control looks suspicious.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Separate discovery, exploit validation and repair
The architecture uses a team of agents. An LLM orchestrator follows a dynamic plan and operates the other agents. Reconnaissance collects network endpoints, APIs and fingerprints, then turns those observations into an application architecture for vulnerability testing. That intermediate representation gives the hunters a target they can reason about.
The remaining agents divide the work by purpose:
- Vulnerability hunters: specialized agents test known vulnerability classes, while others search for ways to break the application using its business context.
- Exploit-validation judge: receives candidate vulnerabilities and assesses whether they are real and exploitable, with the aim of reducing false positives and developer cognitive load.
- Remediation agent: makes findings actionable by explaining what developers can do about them.
- Reporting agent: communicates the resulting findings.
What moves between these agents? The diagram follows the progression from an application map to a suspected vulnerability, then to an exploit assessment and repair guidance. The important separation is between discovering something suspicious and deciding that it warrants a developer’s attention.
Following the BOLA example through this architecture makes the roles concrete. Recon maps the relevant APIs; a hunter investigates the user A/user B access rule; the judge assesses whether the suspected cross-user access is exploitable. Remediation and reporting then turn the finding into work a developer can act on. This is an explanatory application of the described architecture, rather than a demonstrated exploit in the recording. The judge’s testing procedure and error rate are not specified, so a separate validation stage expresses the intended filter without establishing its reliability.
Follows a dynamic plan and operates sub-agents.
The orchestrator directs specialized agents; candidate findings pass through exploit validation before remediation and reporting.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Existing scanners tell the agents where to look
“Context is everything” becomes a concrete integration decision: feed static and dynamic scanner data into AI pen testing. The agent architecture is built around LLMs, but it still uses other engines’ findings to guide the search. Those inputs help tailor testing to the application instead of asking an agent to discover everything from scratch.
The tools also retain their own strengths. Cohen favors DAST’s extensive predefined payloads for cross-site scripting, and contextual pen testing for business-logic failures. His recommendation is to combine them. Fixed payloads provide repeatable tests for known patterns; contextual reasoning investigates how the application’s rules can be broken. Neither method’s usefulness requires discarding the other.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Four questions that make the promise testable
The ending turns the architecture into questions for a vendor—or a team building AI pen testing internally. Continuous code changes and continuous attacks make assessment cadence important, but cadence alone does not establish useful testing. The questions also examine reasoning, inputs and the quality of the result.
- Continuous execution: does testing run continuously, or only as an occasional point assessment?
- Application reasoning: does the system reason about the application’s context and business logic?
- Context enrichment: what additional data can it ingest to guide its work? Cohen expects richer context to improve results and lower operating cost, though the size of that benefit is not established here.
- Exploit evidence: does it establish that a finding is exploitable, or simply flag a possible bug?
For a developer receiving the next finding, the practical standard is actionable evidence: what can an attacker do, why does the application permit it, and what should change? That connects exploit validation to remediation and to Cohen’s closing goal of helping developers and security teams focus on the issues that matter most.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Great. Hi, everyone. We're gonna talk today about continuous offensive security, okay? I'm gonna speak quite soon what does it mean. Uh, but before that, a little bit about me. I'm very excited being here today. I'm Eli. Uh, used to be co-founder and CEO of Helios, a runtime security company. Got acquired by Snyk. And in the last two and a half years, I had the privilege in working in Snyk, primarily on product management for AI products, but more recently founding field CTO, meaning I'm working with developers
- 0:42
and with security folks to better understand and to help them shape their strategy around AI security. So how do you take your AI agents and making sure they are, um, secured? How do you work with your agents in production? Um, Begder, can you set up the timer for me, please? The timer, so I'll know the time. Great. Okay, now we have the timer also. Great. Um, so just so I'll get to know a bit
- 1:12
better, please raise your hand. Who is here more from the development side of the table? Developers. Okay, great. From the security side of the table. Yeah, I already know you. Great. And other people like sales, marketing, other people running agents in production that are not from development or security. Great. Okay. So we have everyone here. So I'm gonna tailor this, uh, session towards everyone. So obviously, this is an AI engineering conference, right? So we
- 1:42
already know that the amount of code generated today is incredibly bigger than it used to be before, okay? According to our data, it's two hundred and eighteen percent more lines of code per developer than it used to be before, okay? And more than eighty-five percent of the developers are actually using coding agents. Now, probably running here today, it's more than one hundred percent, right? 'Cause I guess that each one of you probably have five different session of Claude
- 2:11
Code running simultaneously, otherwise you wouldn't be here today. And the funny thing about those agents running in production and generating code for us is that it's far less secured than we are used to, okay? So sixty-two percent of the output generated by LLMs is not secure or it's broken, okay? So just imagine this data combined together. It's meaning that we are generating more code than we ever did, and this insecure code is
- 2:41
getting into production. So that's a huge challenge for us as a community. And what it actually means is that our backlog of security issues is becoming longer and bigger and much more challenging to remediate it, okay? So we are actually creating many more issues at a pace that we simply couldn't remediate or mitigate. And not only that, AI is helping the attackers also, so they are able to
- 3:11
chain low vulnerabilities together and to form critical vulnerabilities from those, um, low, low severity vulnerabilities. And what actually is going on is that attackers are leveraging those LLMs, and they're attacking us especially where it's very difficult using, um, application context. This is actually where our LLMs are struggled the most 'cause they are creating many more vulnerabilities around those areas.
- 3:41
So obviously, we are leveraging AI, but our attackers are over-- also leveraging frontier models, okay? And the numbers are simply incredible, okay? We are familiar with an operating model that actually was able to identify more than six hundred vulnerabilities, uh, to breach six hundred firewalls in more than five-- fifty-five countries, okay? This is simply unbelievable. And the time for an AI attack to be effective is
- 4:11
constantly getting faster and faster, okay? So the average time for an AI attack to be successful is twenty-four min-- thirty-four minutes, okay? While the fastest time is four minutes, okay? It means that since the moment I began talking, there is probably an AI attack out there that was able to successfully breach a production system. And forty-three percent of all the MCP servers actually has vulnerabilities, okay?
- 4:41
So combining all of this data together, what we realize is that we are producing much more code than ever. Each one of us has a few sessions simultaneously of Claude Code or of Codex. Those-- This code is insecure. It's getting into production. And the people we are facing, the attackers, are actually leveraging the same tools, and they are acting faster than ever in trying to actually leverage those vulnerabilities. And our backlog is simply getting
- 5:11
piled up, and we are not able to get it. And this is exactly what I wanna talk about today, what we should be doing to help, to help improve our security posture, and this is-- ties heavily to how we can leverage offensive security, not just for the attackers, but for us as defenders.
- 5:33
So why traditional security can't really keep up. So there is a mixed audience today, both people from the development side, also people from the security side. So I wanna cover first a few basic concepts, okay? To make sure everyone is familiar with. So what used to be really easy about operating was static code scanning, okay? SAST. So imagine that you are pushing new code to production. There is a static scanner that is scanning that and is looking for SQL injection, cross-site scripting. So it used to be
- 6:03
relatively cheap and relatively easy to run this across every code change, and this is great. Now, what it doesn't catch, it doesn't catch things that are going into runtime, okay? So if you have any- Authorization issues. If you have any configuration issues, for that you will need dynamic application security testing, okay? You'll need something to test the application dynamically. One example is BOLA, uh, Broken Object Level, uh, Authorization, which basically it
- 6:33
means that user A can access the data of user B. Okay, those kind of vulnerabilities you can never find with static analysis. For that, you need to probe the application in runtime. So what actually a DAST scanner is doing, it just takes a predefined set of tons of payloads, and it just send into the application to the different API endpoints, and by that it managed to find vulnerabilities. So this was also something that could find new things,
- 7:03
but it wasn't good with the biggest challenge we have with the code that LLM are producing, and this is the business logic. So for that, we've been u- what we've been using as an industry, it's the pen testing. Okay, so what's the idea about pen testing? It's a human pen tester that really understands the concept and the business logic of the application, and is trying to break it and to find all the vulnerabilities. Now, it is the golden standard
- 7:33
in application security testing, but it has some challenges, right? First of all, it costs a lot of money, right? 'Cause you actually needed to have a security researcher, a human. Not a human in the loop, a human that is actually doing the testing and trying to break the application, and that costs a lot of money. Now, the second problem is because it costs a lot of money and because the process was so slow, companies can only do it once a year, two times a year, maybe
- 8:03
three times a year, okay? So combining the fact that nowadays our AI attackers are trying to continuously attack us, like, doing this pen testing two times a year is simply not good enough, okay? So what you gonna do with the rest of the three hundred and fifty days you have a year that no one tested your application? Not... Now, not only that, the results, you got the result, you got this crazy and glorified report in PDF that you had to send to your developers. They had to
- 8:33
take it back to understand what's real and what's not, and then to start fixing that, and then you had to retest, okay? It sounds like a movie from the '90s, right? It doesn't make sense to do this anymore. But this was the golden standard. This is what we were doing. Now, the thing is, this is the best method out there to actually test your application for its business logic. Okay, so what changes today? Today, we have a new technology, and this is exactly what I wanna talk. How we can
- 9:03
leverage this new area of AI tools and LLM and frontier models to actually do pen testing on a continuous manner, in a scalable manner.
- 9:17
And this is why we're introducing Evo Continuous Offensive Security, okay? So this is obviously very tied up to what we're doing within Snyk, but what I'm gonna tell you today is general, okay? You can take the principles of what I'm telling you, and you can actually verify that with any vendor out there in the expo. Okay? So I wanna, I want you to think about the following things, okay? So what you need to do today, you need to prepare your production system from the unknowns unknowns,
- 9:47
okay? And for that, you need a new paradigm, okay? So first of all, obviously you need the AI pen testing capabilities, which I will double click on soon again. But you also need to start taking the concept of red teaming, okay? So agent red teaming, the idea is to take multi-steps attacks, okay? So to try to simulate prompt injection or data exfiltration and goal agenting, and obviously to combine it with dynamic testing that is good at identifying authorization issues. So we are taking all of those
- 10:17
capabilities together. We are building one suite out of that for offensive security as part of our Evo offering, and the idea is to help you actually mitigate or help you, uh, handle the fact that the attackers are leveraging the same tools that we're using, and they are doing that in a continuous manner, twenty four/seven, finding vulnerabilities faster than before, and your backlog simply cannot... can't stand up with that. So what we are doing differently
- 10:47
within AI pen testing, okay? So we-- I told you before the great promise of pen testing was always the fact that the human in the end really understand the business context. Okay, so now we are taking this to the LLM, and if there is something that LLM are very good at, is really trying to understand the business context. So behind the scene, we are leveraging the LLMs, and we are doing that in a continuous manner, right? So no longer you're gonna do things like do pen testing twice a year or
- 11:17
three times a year. We're gonna do it on every code change, okay? Every PR that you send, every delta, we're gonna run this pen test for you, and we're gonna tell your developers what do they need to fix and how do they need to do it differently. Obviously, we leverage our understanding of the application context, but we also help reduce the false positive. For us as a industry, we have a severe problem of false positive. So we are actually proving you that the vulnerabilities that is-- that we are finding are
- 11:47
actually exploitable, and we don't just flag them as bugs. We actually pr- uh, prove that they are exploitable, and we take context for many different machines and for many different environments, and we combine them together to help guide our LLMs. Okay, so this is what we're doing differently, and I think this is really what's makes this solution so unique and so well positioned to the fact that each one of us have a few coding agents running in
- 12:17
production and pushing more code than ever. Now, how do we do it? How does it work behind the scenes? Okay? So what I want you to understand is we are not taking the static scanners or the DAST scanners and just combining with LLM. Okay? We build it from the ground up using LLM. So we have a f- a, a series of a few agents working together. The first one is the LLM orchestrator. Okay? Imagine it as the brain. This agent is working
- 12:46
according to a dynamic plan that it follows, and it knows how to operate all the different sub-agents. The first agent, or actually the second agent, is the recon. Okay? You can imagine it as your elite intelligence unit. The idea is to collect all the data it can about the different network endpoints, the different APIs, the different fingerprints, collect all of this data together, formalize an architecture out of that, and hand it to the next
- 13:16
agent. And the next agent is the vulnerability testing. It's actually, again, a series of a few specialized agents. Some of them are targeting the well-known classes of vulnerabilities. Some of them are actually hunting for vulnerabilities in the wild, and they are trying to break your application to understand the business context and to understand where can attacker can get from. And then they hand it into the judge, as I like to call it, the exploit validation. So actually,
- 13:46
every vulnerability that the previous agent find are now being handed over to this judge, and the goal of the judge is to tell us whether this is a real vulnerability or l- or not, to tell us whether this can be ex- exploitable or not. Exactly because we want to reduce the cognitive load. Exactly because we want to reduce the amount of false positive we have in the system. Coming up next is the remediation agent because everything we do, we wanna make actionable. Okay? We don't want
- 14:16
to report about a vulnerability that eventually the developers, they don't know what they can do with that. And coming after that is the reporting agent because everything needs to be reported.
- 14:31
So what makes us different? And those are the point I urge you to check with every vendor out there trying to provide you with a solution or whether you wanna implement a solution in-house for AI pen testing. This is a huge industry right now.
- 14:47
So I think the most important thing you need to remember is, and you already know that, especially the developers in the room, context is everything. Your LLM are as good as the context that you feed them with. So what we're trying to do, we are not just taking our A- AI pen testing solution as a standalone. We actually feed it with data from all the previous scanners, okay? So the, the statics code scanner, the dynamic code scanner, all of the different agent--
- 15:17
engines are being combined together, and they are being fed into AI-- our AI pen testing solution because that way we know how to tailor and we know how to guide the LLM where to look for. And as you know also from Mythos or from Fable, when those models know where to look, the results are significantly better. So that's one thing we're doing that is, you know, very unique and very differentiated. Um, and also what I want you to
- 15:47
remember is that one solution is not enough. Okay? The great thing about it is when you combine those things together. And based on the data that we see both from the labs and from real data from customers, combining the data together is really what's important, and you can't really, you can't really use just one of them. Okay? There are certain classes of vulnerabilities. This DAST is the best scanner for that, okay? 'Cause you have a set of really
- 16:17
exhaustive pre, pre, um, a, a set of really exhaustive, uh, payloads that you can send to your application. So it's very good, for example, with cross-site scripting. Okay? So DAST for cross-site sc- cross-site scripting is the best solution. But if you want, again, like business logic, then you should combine it with, uh, pen testing because it will be the best solution. So I think what I really want you to take from this, if there is one thing I want you to take from here
- 16:46
is to realize that the world is changing. You already know that. That's what made you come here. But, you know, we used to test our application once every few months because the change and the velocity of the change was so slow, and we'll-- we had people at the end that were reasoning about the application and about the change. But that world is-- does not longer exist, right? Right now we are pushing code to production twenty-four seven, and we have
- 17:16
LLMs that are trying to reason about the application context, and they are trying to break it continuously. Some of them can do it in like as fast as four minutes, okay? That's almost five time of the lecture of the session that I just did. So when you're coming to look for new solutions for this problem, you have to ask yourself four questions. Okay? First of all, do we run this solution in a continuous way, or is it just a point solution that we run every once in a while? Okay? That's a
- 17:46
major difference. Second, do they really reason about the application context? Okay? 'Cause given the nature of those LLM, application context is really where the new vulnerabilities are coming from. And third, what more context can they digest? Okay, how can they enrich the data? Because the richer the data is, the better, and by the way, cheaper the LLM operation is gonna be. And the fourth part
- 18:16
is do they really help me prove that there is a bug and to make sure that this is exploitable, or they, they just flag it as a bug? Okay, so our goal is not only to help you find the bugs faster but to actually do it better and cheaper and to help you as a developer or as the security people here focus on what matters most. Um, so thank you very much. And, uh, I used to do it in every lecture. You know, my, my,
- 18:47
my dream was to be a rock star, but I had to be a cyber entrepreneur, so instead of that, I'm gonna take pictures. So hands up, everyone.
- 18:57
Okay, great. Thank you so much.
- 19:01
And I'm here for questions if you want. Go for the booth. It's very close here. Snyk.io. Meet Evo, our new AI security offering. Thank you so much.