I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI
Read the talk
Voice-Agent Reliability: From Convincing Conversations to Verified Actions
Sumanyu Sharma explains why natural speech can conceal failed actions, how centralized changes spread mistakes, and how testing, monitoring, and adversarial evaluation fit into a continuous reliability loop.
From a talk by Sumanyu Sharma
At a glance
Ideas worth remembering
Judge action-taking agents by the actual workflow outcome. A convincing booking confirmation can coexist with a missing appointment, just as a policy answer can omit required eligibility or verification.
Prioritize failures using frequency and severity together. Shared prompts and architectures can spread one defect widely, making systematic, high-impact failures the most urgent targets.
Start with manual listening, scale repeatable checks, and analyze patterns across calls. More scoring coverage does not automatically reveal problems outside the predefined rubrics.
Verify fixes through repeated failing cases, variations in wording and voice, and regression checks. Use real-world A/B tests for hypotheses about human responses that simulations cannot adequately assess.
Reliability testing for earnest callers and adversarial testing address different risks. Sharma recommends monitoring both sides of the conversation and ongoing red teaming where failures are costly; his reported error and attack-success rates lack enough methodological detail to generalize universally.
From crime audio to agents that take actions
Sumanyu Sharma, founder and CEO of Hamming, opens with his earlier work at Citizen in New York. His team listened to thousands of hours of police radio and sent millions of alerts to users in San Francisco, New York, LA, Chicago, Baltimore, and other cities. The reports ranged from disturbing incidents to unusual thefts and a person stranded after someone took a ladder. He also describes recent alerts in the app. This experience establishes his reference point: listening to audio at scale, interpreting what happened, and distributing information that matters to people.
Voice agents make him nervous precisely because they are moving into production. When he began working on their reliability in early 2024, voice was only starting to work well. He now sees substantial improvements in the underlying infrastructure and orchestration layer. Teams can build an experience he describes as perhaps 60% good in a short period, but the remaining long tail still takes work. Faster construction of a convincing experience does not establish that it handles the less common situations reliably.
He describes teams experimenting with hybrid architectures that combine voice-to-voice capabilities and cascading stacks. Their goal is to improve reliability while keeping latency low; the recording does not specify how individual components are divided between those approaches. Meanwhile, connections to calendars, CRM, HR, and reservation systems let agents take actions. This changes the reliability question: a conversation can now alter a business workflow, so speaking well is only one part of completing the task.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A confident voice can conceal an incomplete task
Sharma calls reliability the main obstacle to voice-agent deployment at scale. His first example concerns a customer seeking trade-in information who receives confusing answers. In the embedded conversation, the customer describes uncertainty about whether they were speaking to AI or a person, while someone on the business side needs to call back and repair the situation. Sharma emphasizes the mismatch between natural, confident delivery and incorrect information. The business inherits both the original problem and the work of restoring trust.
His personal example makes task completion concrete. He believed he had booked an appointment with a physician, arrived, and discovered that he was absent from the schedule. The front desk turned him away, costing him 2 hours. The failure was a gap between the apparent outcome of the interaction and the actual appointment record. He then changes the circumstances without changing the failure mode: suppose the patient were a parent or grandparent, or the appointment were for a procedure rather than a regular checkup. The same missing booking could carry a much higher cost.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Scale, severity, and the reach of a single change
Sharma contrasts what he describes as declining crime with rapidly growing production use of voice agents. He estimates at least a trillion calls per year and predicts that conversational voice agents will handle a majority within five years. His scale calculation is conditional: at a trillion interactions, a 1% error rate yields 10 billion incidents annually. The forecast is not established by that arithmetic, but the calculation explains why even a seemingly small error rate becomes consequential at large volume.
He reports that Hamming monitors 10,000 agents and sees an error rate closer to 10%. Examples include finding a policy while skipping eligibility or verification, applying an unauthorized discount, providing incorrect information, and claiming an appointment was booked when it was not. These failures concern business logic and action completion as well as conversational accuracy. The recording does not give the sampling method, measurement window, or precise error denominator, so the reported rate should be understood as an observation from Hamming's monitored population rather than a universal rate.
An error count alone does not express harm. Sharma places repetition and failures to understand a caller toward the annoyance end of a severity spectrum. At the other end, he offers a hypothetical drive-through order involving a vegan burger and peanut allergies. Mishandling the caller's requirements can create a safety concern, particularly when the workflow is deployed widely. The example teaches teams to distinguish the frequency of a failure from the consequences of that failure in its specific setting.
His second comparison concerns how failures spread. A robbery or vehicle theft generally affects a finite set of people involved in a local event. Voice-agent deployments share centralized prompts and architectures. A single change can therefore alter behavior for millions of downstream users. That shared control creates a large blast radius: the same mistake can recur across many calls, making visibility into incidents essential.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Build a loop that discovers problems as well as scores them
Sharma introduces a framework borrowed from friends who worked on growth at Facebook. First identify problems in the conversational experience, then prioritize them by frequency and severity. Understand the fix, execute it, check whether it worked and whether it caused regressions elsewhere, and continue monitoring production. The sequence matters because changing the agent is only one step. A useful change needs evidence of improvement and continued observation after deployment.
He separates known problems from emerging behavior and asks how many conversations the team analyzes. Known issues include latency, interruptions, and automatic speech recognition problems that can be tracked over time. Emerging patterns may become apparent only across many conversations. Coverage and discovery answer different questions: evaluating more calls expands observation, while looking beyond existing checks determines whether the team can recognize a problem it had not already named.
Manual listening is his recommended starting point. Specific conversations give a team depth, context, and intuition about the experience. But listening by hand does not scale. Teams commonly turn those observations into a spreadsheet with five or 10 rubrics covering greetings, closing, validation, and core logic, then invest in evaluation tools to compute metrics and scores across more calls. The progression converts detailed human observations into repeatable checks.
Predefined scoring still tends to check consistency against known problems. Sharma argues that it does not, by itself, discover novel insights across conversations. Hamming spends substantial effort on cross-conversation analysis, and he says some of the strongest teams it works with do the same. The unit of discovery becomes a pattern across calls rather than a score on one call. He does not describe a specific algorithm for finding those patterns, so the substantive recommendation is to add this form of analysis alongside individual-call evaluation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Prioritize recurrence and consequence together
Sharma organizes priorities around whether a problem is isolated or systematic and whether its impact is low or high. An isolated, low-impact issue deserves less attention, but systematic annoyances such as repetition remain worth solving at scale and during a comparison between competing systems. A high-impact isolated failure should not become chronic. Systematic, high-impact failures are the P0 targets. His example is a financial-services agent intended to freeze credit cards: repeatedly failing to perform that action defeats a critical purpose of the workflow.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Test the failure, vary the interaction, and measure real users
Attempting a fix is, in Sharma's view, the simplest and lowest-effort part of the debugging loop. The harder question is whether the changed system actually works. His basic test takes a real failed call, such as the appointment conversation, and replays it five, 10, 20, or 50 times to estimate how often the agent passes that case. Repetition checks whether success is consistent rather than relying on one favorable run.
A stronger test preserves the intent while varying wording, conversational patterns, accents, and style, then adds another intent to the interaction. This expands coverage beyond the exact conversation that exposed the bug. The desired result is confidence that the change improves the system across related situations and produces a net benefit. Passing a repeated original case is useful evidence, but it leaves open whether the fix survives ordinary variation in how callers ask for help.
Some hypotheses are difficult to test in a synthetic setting before deployment. Sharma uses outbound calls as an example: he considers the first five seconds especially important, with vocal quality and the exact opening words affecting the interaction. He argues that real-world A/B testing is necessary to assess those choices because simulations alone cannot supply the relevant results. Synthetic tests exercise controlled cases; live experiments measure how actual recipients respond.
He closes this part by returning to the complete loop: identify, prioritize and size impact, understand and execute the fix, check that it works without breaking something else, and keep monitoring. Verification therefore includes both the targeted improvement and regressions elsewhere. Deployment continues the learning process rather than ending it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
When callers deliberately exploit the agent
Sharma then changes the assumption about the caller. Making an agent useful is already difficult when people earnestly want their problems solved. Deliberately adversarial callers introduce another class of risk: persuading an agent or a human to reveal protected health information or personally identifiable information. His broader warning about increasingly capable automated callers is a threat scenario; the recording does not establish the external capability example he invokes as a demonstrated event.
The pressure to make agents more capable expands this risk. Teams give agents more data and tools and deploy them quickly; Sharma argues that each increase in capability enlarges the surface available to an attacker. Natural, human-sounding voice can also be weaponized against people. His concern therefore covers both agents with access to sensitive systems and humans who may trust a persuasive caller.
Hamming released a red-teaming product in April to investigate how many agents it could break adversarially. Sharma reports that the team can break approximately one in five agents, with testing across financial services, healthcare, and consumer settings. He describes bypassing verification and obtaining data the testers should not have received. These are reported failures of access boundaries, although he supplies neither attack transcripts nor a precise definition of a broken agent. The rate indicates the concern within those tests without establishing an industry-wide prevalence.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Combine testing, monitoring, and continuous red teaming
His defensive program starts with deep testing before deployment. Tests can operate text-to-text or voice-to-voice; he acknowledges advantages and disadvantages to both without detailing them here. This first layer checks whether the agent can serve ordinary people who simply want a problem solved. The recording supports using both testing modes as options, but does not establish that either one alone provides complete coverage.
The second layer is production monitoring through per-call scoring, manual evaluation and listening, and cross-call analysis. Sharma expands the object of monitoring to include the users as well as the agent. Teams need to observe what the agent says and does, while also recognizing callers who try to induce prohibited behavior. That distinction helps separate mistakes during legitimate use from deliberate attempts to exploit the system.
His additional recommendation is 24/7 red teaming, especially where bad interactions can be costly. This makes adversarial testing an ongoing activity alongside ordinary monitoring. He does not specify the operating cost, test volume, or a threshold for sufficient coverage, so continuous testing is presented as a response to high stakes rather than a quantified guarantee of safety.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The promise of better calls carries personal exposure
Sharma ends by preserving the positive case for voice agents. They could make interactions feel more human than clunky IVR trees, chatbots, or waiting on hold. His warning concerns what happens as deployment grows and bad actors exploit vulnerabilities: he expects incident counts to rise and more people to experience the consequences personally. This is his closing prediction, not a measured future outcome. It explains why he regards voice-agent risk as something callers will encounter directly.
The talk closes with an invitation to work in this area and to discuss whether a deployed agent's architecture and evaluations are configured correctly. Sharma mentions Hamming's substantial token use and offers to speak with attendees afterward. The practical emphasis remains on examining both the system that produces behavior and the evaluations used to judge it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
My name is Suman Yu and I'm the founder
- 0:15
and CEO of Hamming. And before working
- 0:19
on voice agent reliability and safety, I
- 0:22
worked at a company called Citizen
- 0:25
out of New York. Anybody here use
- 0:26
Citizen app? Awesome. Thank you. Uh, and
- 0:31
at Citizen,
- 0:33
we listened to crime, thousands of hours
- 0:36
of police radio station data and sent
- 0:40
millions of alerts to users in San
- 0:42
Francisco,
- 0:44
New York, LA, uh, Chicago, Baltimore,
- 0:48
and so on.
- 0:50
Some obviously gory and pretty sad. Uh,
- 0:54
but others more funny like a person
- 0:56
stealing bags of ice cream from Safeway
- 1:00
or report of a man hanging off the side
- 1:02
of the house after a woman stole his
- 1:04
ladder.
- 1:07
If I actually take a look at the citizen
- 1:09
app right now for those who are
- 1:10
customers or users, I can see that there
- 1:13
is a man yelling at person. There's
- 1:16
indecent exposure. This is real. This is
- 1:18
real time. This is, you know,
- 1:21
couple hours ago. These are real-time
- 1:23
alerts that we're sending.
- 1:27
Now, voice agents scare me more because
- 1:29
they're finally graduating from demos
- 1:31
and PC's to production. We should be
- 1:33
super excited, but I'm nervous. I'm
- 1:35
personally nervous. Uh, they're talking
- 1:37
to users at a scale that would make Gary
- 1:39
Tan and Polygram proud.
- 1:43
When I got started in voice agent
- 1:45
reliability in early 2024, voice was
- 1:48
just starting to work. It was not quite
- 1:50
good yet, but it was just starting to
- 1:52
work. You would have to pay me a lot of
- 1:54
money for me to stop using, you know,
- 1:55
Aqua voice, Super Whisper, uh, Whisper
- 1:58
Flow, and so on. These products are just
- 2:00
getting super, super good. And a big
- 2:02
reason is because the underlying
- 2:03
infrastructure is getting better, and
- 2:05
the orchestration layer is getting
- 2:06
meaningfully better. It's getting much
- 2:09
faster to build products and voice
- 2:10
experiences that maybe are 60% good in a
- 2:14
pretty short period of time, but the
- 2:16
long tail is still Hey, Gorov. the long
- 2:18
tail is still uh wise away.
- 2:21
I think speech speech models are getting
- 2:23
better. Um teams are experiment
- 2:25
experimenting with hybrid architectures
- 2:28
of combining more voicetovoice
- 2:32
modalities and also cascading stacks to
- 2:34
make the experience reliable but still
- 2:36
pretty low latency.
- 2:39
Things are obviously getting better.
- 2:41
Agents are being connected to calendars,
- 2:42
CRM, HRs, reservation systems, and so
- 2:45
on. Voice agents can now take actions.
- 2:50
However, reliability is still the number
- 2:52
one problem holding back most voice
- 2:54
agent deployments at scale. This is
- 2:55
still the number one problem. This is an
- 2:58
example I found on Twitter pretty
- 2:59
randomly, you know, two weeks ago and a
- 3:02
person is trying to get information for
- 3:04
a tradein and gets absolutely confused
- 3:06
with the information that they're
- 3:08
receiving. Alex now has to correct for
- 3:10
this loss of trust but trying to, you
- 3:13
know, call the person and see what see
- 3:15
what happened and fix the situation. Let
- 3:17
me see if audio works here.
- 3:20
>> Screwed up with another customer. We're
- 3:22
getting it fixed, but I got to call him
- 3:24
and see if I can work it out.
- 3:25
>> I'm like, dude, half the time I'm like,
- 3:27
I don't know if I'm talking to AI. I
- 3:28
don't know if I'm talking to a person.
- 3:29
It was just confusing, but we got there.
- 3:32
>> It probably is AI and human. And
- 3:37
>> so I think voices sound very confident.
- 3:39
They sound very natural, but the
- 3:40
information provided is often, you know,
- 3:42
not correct. That's the biggest problem
- 3:44
here.
- 3:46
This example is more personal. I had
- 3:48
booked an appointment with a physician a
- 3:49
couple weeks ago or I thought I did. I
- 3:52
showed up to the appointment and turns
- 3:54
out I was not actually on the schedule.
- 3:57
So the front desk, you me turned me
- 3:58
away. I wasted 2 hours. For me, this was
- 4:01
a waste of time. But what if this was
- 4:03
actually your parent?
- 4:06
What if this was your grandparent?
- 4:08
What if this appointment was for a
- 4:10
procedure instead of a regular checkup?
- 4:13
The costs for these different
- 4:15
permutations of the same failure mode
- 4:16
can actually be super super high.
- 4:20
Now, let's compare crime to voice
- 4:22
agents. Um, I think observation number
- 4:24
one is crime is actually decreasing over
- 4:26
time. This is a good thing and I hope it
- 4:30
crosses the x- axis at some point you
- 4:32
know in the future.
- 4:35
Voice on the other hand is generally
- 4:37
taking off right we're seeing a pretty
- 4:38
fast takeoff of voice agents being
- 4:40
deployed in production. There's at least
- 4:42
a trillion calls that are done every
- 4:44
single year and majority of these will
- 4:46
be done by conversational voice agents
- 4:48
over the next you know five years. If
- 4:50
you assume a 1% error rate that is still
- 4:53
10 billion incidents per year. That's a
- 4:55
lot.
- 4:58
In practice, we currently monitor 10,000
- 5:00
agents and the error rate is closer to
- 5:03
10% in practice. These range from agents
- 5:06
saying they found the right policy when
- 5:08
they actually skipped the eligibility or
- 5:10
verification steps or applying discounts
- 5:12
when they were not really supposed to,
- 5:14
misharing what the person said,
- 5:16
providing incorrect information, or
- 5:18
claiming they booked an appointment when
- 5:19
they actually did not, just like it
- 5:21
happened for me.
- 5:23
Now, not every single call has an
- 5:25
equally, you know, bad cost. Uh, some
- 5:29
range, you know, in the crime land, some
- 5:30
range from trash fires, which are kind
- 5:33
of funny, annoying, not really hurting
- 5:35
somebody. For a voice equivalent, that
- 5:38
would be annoyances like repetition, um,
- 5:40
or just sort of not quite understanding
- 5:42
what the user is saying. all the way to
- 5:44
safety risks like mass shootings or in
- 5:47
the voice agent equivalent, it would be
- 5:49
um a drive-thru that's deploying um
- 5:52
voice agents at scale like a Taco Bell
- 5:54
or McDonald's and a person orders a
- 5:57
vegan burger with peanut allergies.
- 5:59
If one of those two situations are not
- 6:02
handled correctly, that is definitely a
- 6:04
safety concern at scale.
- 6:07
The other big difference between crime
- 6:10
and and voice agent deployments is is
- 6:13
crime generally tends to be pretty hyper
- 6:16
local,
- 6:18
tends to be very decentralized, right?
- 6:19
Things like robbery or motor vehicle
- 6:22
theft or lararseny. They're impacting a
- 6:25
finite set of individuals that are
- 6:27
involved in that um situation.
- 6:32
On the other hand, voice agents are much
- 6:34
more centralized. a single prompt change
- 6:37
or an architecture change can have
- 6:38
pretty massive implications downstream
- 6:41
for all of the millions of you know
- 6:43
users that are um in the crossfire. So
- 6:46
the blast radius is is quite quite
- 6:48
massive. So the natural question is how
- 6:51
do you make these incidents much more
- 6:52
visible and obvious? That's the kind of
- 6:54
obvious question here.
- 6:57
I'll borrow a framework from a couple of
- 6:58
my friends who were OG growth folks at
- 7:01
Facebook. So step one is to identify
- 7:04
okay what are all the challenges and
- 7:05
problems that um exist in your
- 7:08
conversation experience. Step two is to
- 7:10
prioritize an impact size. There's a
- 7:12
frequency and severity analysis that's
- 7:14
pretty important. Step three is to
- 7:16
understand okay how do we actually fix
- 7:18
this? Step four execute. Step five okay
- 7:21
did my change actually work and did it
- 7:23
cause any regressions somewhere else.
- 7:25
And lastly we continue to monitor in
- 7:27
production.
- 7:30
On the y- axis, I think it's important
- 7:32
to highlight there are known problems
- 7:34
that already exist. Things like turnover
- 7:36
latency, interruptions, um maybe some
- 7:39
ASR problems you're, you know, aware of.
- 7:41
And these are known problems that exist
- 7:44
that the team should track over time. On
- 7:47
the other axis is actually emerging
- 7:49
behavior or patterns that are only
- 7:51
obvious across lots of conversations. Um
- 7:55
on the x- axis, you have coverage just
- 7:57
like insurance. Are you analyzing few
- 8:00
conversations? Are you analyzing many,
- 8:02
many conversations? Most teams will
- 8:04
typically start by listening to calls
- 8:06
manually.
- 8:08
And I think that's the best place to
- 8:09
start. I don't think you should skip
- 8:11
that step. There's a lot of depth and
- 8:13
insights to get by actually listening to
- 8:15
specific conversations and building that
- 8:17
texture that that comes from that
- 8:19
intuition. However, it's obviously not
- 8:21
scalable. So most teams end up having a
- 8:24
spreadsheet of I don't know five or 10
- 8:27
different rubrics around greetings,
- 8:29
closing, validation,
- 8:31
um, core logic and so on. To scale that
- 8:35
up even further, you then end up
- 8:36
investing in some eval product, right?
- 8:39
You might run some element as a judge
- 8:41
and compute classic metrics and also
- 8:44
more more deterministic and stoastic
- 8:47
scoring logic. Um, but there you're
- 8:49
still stuck with checking for
- 8:51
consistency of known problems, but
- 8:53
you're not really discovering novel
- 8:54
insights that are actually happening
- 8:56
across conversations. We're spending a
- 8:58
ton of time on performing cross
- 9:01
conversation analysis, not a pattern on
- 9:03
a single call, but across conversations.
- 9:05
And some of the best teams that we work
- 9:07
with are are doing the same.
- 9:10
Now, to prioritize an impact size, I
- 9:12
think there's problems that are one-off
- 9:14
that are low impact. I mean, who cares?
- 9:17
uh even low impact and systematic
- 9:19
problems in the crime world that would
- 9:21
be a trash fire in a voice aation world
- 9:23
it could be some repetitions the team is
- 9:25
experiencing they're still annoying at
- 9:27
scale and if you are doing a bake off
- 9:29
it's still worth solving for them I
- 9:31
would not ignore these class of problems
- 9:33
oneoff and high impact well hope it
- 9:35
doesn't chronic and I think systematic
- 9:38
and high impact are obviously the P 0
- 9:40
you know target areas um for the team to
- 9:42
solve an example of that
- 9:45
would be in a fins serve capacity
- 9:47
There's a voice agent that um helps
- 9:49
users freeze their credit cards. And if
- 9:51
it doesn't do that, well, that's a
- 9:53
massive fail.
- 9:56
All right. So, understand and execute.
- 9:58
I'm pretty sure everyone's doing this.
- 10:00
Please fix my agent. Uh I think fixing
- 10:02
or rather attempting to make a fix is
- 10:05
the simplest and the lowest effort
- 10:08
component of this debugging pipeline and
- 10:11
loop. Um the next step is all right, I
- 10:14
made a change to my system. How do I
- 10:16
actually know this thing works um for
- 10:19
real? A great way that's naive is to
- 10:23
take a real call, for example, in my
- 10:25
case, I booked an appointment and it
- 10:27
didn't get scheduled and replay that
- 10:29
exact conversation and run that maybe 5,
- 10:32
10, 20, 50 times and see, okay, what is
- 10:34
my probability of passing this type of
- 10:37
issue? A better way is to keep the same
- 10:40
intent but change the wordings, change
- 10:44
the patterns, change the accents, change
- 10:46
the style, add one more intent to the
- 10:48
mix. And that gives teams much more, you
- 10:51
know, better coverage to feel confident
- 10:53
that yes, I actually made a change and
- 10:56
my changes are net positive instead of
- 10:58
net negative.
- 11:01
There are certain fixes and I guess
- 11:05
hypothesis that are very difficult to
- 11:07
test in a pre-eployment synthetic
- 11:09
setting. And so AB testing ends up
- 11:11
being, you know, pretty pretty critical
- 11:13
for those circumstances. For example, if
- 11:15
you have an outbound agent, the first 5
- 11:17
seconds of a conversation tends to be
- 11:19
the most important. And so the vocal
- 11:22
quality um and the specific words you
- 11:25
end up using, they matter the most. And
- 11:27
so AB testing that is the only way in in
- 11:29
kind of real life setting to to get
- 11:32
results. You can't really do it through
- 11:34
simulations alone.
- 11:37
And so there we have the loop. Identify,
- 11:39
prioritize, impact size, understand the
- 11:42
fix, execute, check, make sure it didn't
- 11:45
break anything, and then continue
- 11:47
monitoring.
- 11:51
So I think making voice agents useful is
- 11:54
already hard as it is. even when dealing
- 11:58
with earnest users on the other line,
- 12:01
right? These are people who who just
- 12:02
want their problem solved. They're not
- 12:03
trying to mess with you. These are like
- 12:04
legit normal people.
- 12:08
Now, what happens when mythos learns how
- 12:10
to dial?
- 12:13
So, if it can extract trade secrets and
- 12:16
uh you know, from the NSA, it can
- 12:18
certainly, you know, seduce you into
- 12:22
revealing PHI and PII data as well.
- 12:26
And I think both voice agents and humans
- 12:29
will [clears throat] be targeted here.
- 12:31
Voice agents because there's a pressure
- 12:33
to make these more capable. Give them
- 12:37
access to more data. Give them access to
- 12:39
more tools.
- 12:41
Deploy them quickly.
- 12:45
The more the capability, the bigger the
- 12:46
surface area. This is this is pretty
- 12:48
pretty common sense. And the more the
- 12:50
voice agents become natural and human
- 12:54
sounding, the more humans will be
- 12:56
tricked along the way as well for those
- 12:58
who are weaponizing.
- 13:01
Uh we ship a we shipped a red tipping
- 13:03
product um back in April just to test
- 13:05
out this hypothesis for how many agents
- 13:07
can we actually break from a adversarial
- 13:10
capacity and we can probably break one
- 13:12
in five agents at this point. We've
- 13:14
tested this across financial services,
- 13:16
healthcare, um consumer and so on. We've
- 13:19
bypassed verification. Uh we've
- 13:22
definitely had agents, you know, we've
- 13:24
been able to promject uh several agents
- 13:27
and and gotten data we should not have.
- 13:29
So this is not theoretical. This is
- 13:31
actually a real a real concern.
- 13:37
I think the only real defense against
- 13:39
the dark arts is
- 13:41
step one to invest deeply in
- 13:43
pre-eployment testing. This could be
- 13:45
textto text. This could be voice to
- 13:48
voice. There's pros and cons to both.
- 13:50
Happy to chat offline if folks are
- 13:51
interested. And this is just making sure
- 13:54
you're not self-owning, you know, when
- 13:56
you're talking to real people who just
- 13:57
want to get their problem solved. Step
- 13:59
two is to have a great monitoring system
- 14:02
of all kinds. And I've highlighted, you
- 14:04
know, different flavors of monitoring
- 14:06
per call scoring, manual kind of evals,
- 14:10
you know, listening to conversations and
- 14:11
cross call analysis. And this is helpful
- 14:14
both for monitoring what the agent is
- 14:16
saying and behaving and how it's
- 14:17
actually doing, but also the users. Are
- 14:19
the users being adversarial? Are they
- 14:21
being annoying? Are are they trying to
- 14:22
trick the agent into doing things it's
- 14:24
not supposed to be doing?
- 14:26
And I think our new recommendation now
- 14:28
is to run 24/7 red teaming um for your
- 14:31
agents, especially if you believe the
- 14:34
cost of bad interactions can be can be
- 14:36
rather large.
- 14:38
So, I think voice agents um have this
- 14:40
awesome potential of of making the world
- 14:43
feel much more human compared to
- 14:48
interacting with clunky IVR trees or
- 14:50
chat bots or worse um being stuck on a
- 14:53
on a hold.
- 14:55
And when we think about crime, we often
- 14:57
think of crime happening to somebody
- 14:59
else. You know, crime does not happen to
- 15:01
you typically with voice agents,
- 15:04
especially bad actors. As these agents
- 15:07
are deployed and as bad actors start to
- 15:09
exploit a lot of the vulnerabilities,
- 15:11
the number of incidents is about to kind
- 15:13
of go way way up. And so the reason I
- 15:16
fear voice agents more than crime is
- 15:18
that one of these incidents is going to
- 15:20
impact you. It already did for me.
- 15:24
Awesome. So it's time for me to shill.
- 15:25
Well, we burn a lot of tokens. If you
- 15:28
are interested in working in this space,
- 15:31
please come and talk to us. And if you
- 15:33
are deploying voice agents and want to
- 15:35
validate whether your architectured or
- 15:39
your eval are set up correctly, please
- 15:40
come and talk to us. We'll be outside.
- 15:42
And here's here's my number. Here's my
- 15:44
WhatsApp. Thanks everyone.
- 16:02
>> [music]