I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI

Read the talk

Voice-Agent Reliability: From Convincing Conversations to Verified Actions

Sumanyu Sharma explains why natural speech can conceal failed actions, how centralized changes spread mistakes, and how testing, monitoring, and adversarial evaluation fit into a continuous reliability loop.

From a talk by Sumanyu Sharma

At a glance

Ideas worth remembering

  • Judge action-taking agents by the actual workflow outcome. A convincing booking confirmation can coexist with a missing appointment, just as a policy answer can omit required eligibility or verification.

  • Prioritize failures using frequency and severity together. Shared prompts and architectures can spread one defect widely, making systematic, high-impact failures the most urgent targets.

  • Start with manual listening, scale repeatable checks, and analyze patterns across calls. More scoring coverage does not automatically reveal problems outside the predefined rubrics.

  • Verify fixes through repeated failing cases, variations in wording and voice, and regression checks. Use real-world A/B tests for hypotheses about human responses that simulations cannot adequately assess.

  • Reliability testing for earnest callers and adversarial testing address different risks. Sharma recommends monitoring both sides of the conversation and ongoing red teaming where failures are costly; his reported error and attack-success rates lack enough methodological detail to generalize universally.

From crime audio to agents that take actions

Selected presentation frame from I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI at 70 seconds
From crime audio to agents that take actions

Sumanyu Sharma, founder and CEO of Hamming, opens with his earlier work at Citizen in New York. His team listened to thousands of hours of police radio and sent millions of alerts to users in San Francisco, New York, LA, Chicago, Baltimore, and other cities. The reports ranged from disturbing incidents to unusual thefts and a person stranded after someone took a ladder. He also describes recent alerts in the app. This experience establishes his reference point: listening to audio at scale, interpreting what happened, and distributing information that matters to people.

Voice agents make him nervous precisely because they are moving into production. When he began working on their reliability in early 2024, voice was only starting to work well. He now sees substantial improvements in the underlying infrastructure and orchestration layer. Teams can build an experience he describes as perhaps 60% good in a short period, but the remaining long tail still takes work. Faster construction of a convincing experience does not establish that it handles the less common situations reliably.

He describes teams experimenting with hybrid architectures that combine voice-to-voice capabilities and cascading stacks. Their goal is to improve reliability while keeping latency low; the recording does not specify how individual components are divided between those approaches. Meanwhile, connections to calendars, CRM, HR, and reservation systems let agents take actions. This changes the reliability question: a conversation can now alter a business workflow, so speaking well is only one part of completing the task.

0:120:15
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:01 · section reference included

A confident voice can conceal an incomplete task

Sharma calls reliability the main obstacle to voice-agent deployment at scale. His first example concerns a customer seeking trade-in information who receives confusing answers. In the embedded conversation, the customer describes uncertainty about whether they were speaking to AI or a person, while someone on the business side needs to call back and repair the situation. Sharma emphasizes the mismatch between natural, confident delivery and incorrect information. The business inherits both the original problem and the work of restoring trust.

His personal example makes task completion concrete. He believed he had booked an appointment with a physician, arrived, and discovered that he was absent from the schedule. The front desk turned him away, costing him 2 hours. The failure was a gap between the apparent outcome of the interaction and the actual appointment record. He then changes the circumstances without changing the failure mode: suppose the patient were a parent or grandparent, or the appointment were for a procedure rather than a regular checkup. The same missing booking could carry a much higher cost.

2:502:52
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:50 · section reference included

Scale, severity, and the reach of a single change

Selected presentation frame from I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI at 357 seconds
Scale, severity, and the reach of a single change

Sharma contrasts what he describes as declining crime with rapidly growing production use of voice agents. He estimates at least a trillion calls per year and predicts that conversational voice agents will handle a majority within five years. His scale calculation is conditional: at a trillion interactions, a 1% error rate yields 10 billion incidents annually. The forecast is not established by that arithmetic, but the calculation explains why even a seemingly small error rate becomes consequential at large volume.

He reports that Hamming monitors 10,000 agents and sees an error rate closer to 10%. Examples include finding a policy while skipping eligibility or verification, applying an unauthorized discount, providing incorrect information, and claiming an appointment was booked when it was not. These failures concern business logic and action completion as well as conversational accuracy. The recording does not give the sampling method, measurement window, or precise error denominator, so the reported rate should be understood as an observation from Hamming's monitored population rather than a universal rate.

An error count alone does not express harm. Sharma places repetition and failures to understand a caller toward the annoyance end of a severity spectrum. At the other end, he offers a hypothetical drive-through order involving a vegan burger and peanut allergies. Mishandling the caller's requirements can create a safety concern, particularly when the workflow is deployed widely. The example teaches teams to distinguish the frequency of a failure from the consequences of that failure in its specific setting.

His second comparison concerns how failures spread. A robbery or vehicle theft generally affects a finite set of people involved in a local event. Voice-agent deployments share centralized prompts and architectures. A single change can therefore alter behavior for millions of downstream users. That shared control creates a large blast radius: the same mistake can recur across many calls, making visibility into incidents essential.

4:204:24
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:20 · section reference included

Build a loop that discovers problems as well as scores them

Sharma introduces a framework borrowed from friends who worked on growth at Facebook. First identify problems in the conversational experience, then prioritize them by frequency and severity. Understand the fix, execute it, check whether it worked and whether it caused regressions elsewhere, and continue monitoring production. The sequence matters because changing the agent is only one step. A useful change needs evidence of improvement and continued observation after deployment.

He separates known problems from emerging behavior and asks how many conversations the team analyzes. Known issues include latency, interruptions, and automatic speech recognition problems that can be tracked over time. Emerging patterns may become apparent only across many conversations. Coverage and discovery answer different questions: evaluating more calls expands observation, while looking beyond existing checks determines whether the team can recognize a problem it had not already named.

Manual listening is his recommended starting point. Specific conversations give a team depth, context, and intuition about the experience. But listening by hand does not scale. Teams commonly turn those observations into a spreadsheet with five or 10 rubrics covering greetings, closing, validation, and core logic, then invest in evaluation tools to compute metrics and scores across more calls. The progression converts detailed human observations into repeatable checks.

Predefined scoring still tends to check consistency against known problems. Sharma argues that it does not, by itself, discover novel insights across conversations. Hamming spends substantial effort on cross-conversation analysis, and he says some of the strongest teams it works with do the same. The unit of discovery becomes a pattern across calls rather than a score on one call. He does not describe a specific algorithm for finding those patterns, so the substantive recommendation is to add this form of analysis alongside individual-call evaluation.

6:576:58
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:57 · section reference included

Prioritize recurrence and consequence together

Sharma organizes priorities around whether a problem is isolated or systematic and whether its impact is low or high. An isolated, low-impact issue deserves less attention, but systematic annoyances such as repetition remain worth solving at scale and during a comparison between competing systems. A high-impact isolated failure should not become chronic. Systematic, high-impact failures are the P0 targets. His example is a financial-services agent intended to freeze credit cards: repeatedly failing to perform that action defeats a critical purpose of the workflow.

9:109:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:10 · section reference included

Test the failure, vary the interaction, and measure real users

Selected presentation frame from I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI at 632 seconds
Test the failure, vary the interaction, and measure real users

Attempting a fix is, in Sharma's view, the simplest and lowest-effort part of the debugging loop. The harder question is whether the changed system actually works. His basic test takes a real failed call, such as the appointment conversation, and replays it five, 10, 20, or 50 times to estimate how often the agent passes that case. Repetition checks whether success is consistent rather than relying on one favorable run.

A stronger test preserves the intent while varying wording, conversational patterns, accents, and style, then adds another intent to the interaction. This expands coverage beyond the exact conversation that exposed the bug. The desired result is confidence that the change improves the system across related situations and produces a net benefit. Passing a repeated original case is useful evidence, but it leaves open whether the fix survives ordinary variation in how callers ask for help.

Some hypotheses are difficult to test in a synthetic setting before deployment. Sharma uses outbound calls as an example: he considers the first five seconds especially important, with vocal quality and the exact opening words affecting the interaction. He argues that real-world A/B testing is necessary to assess those choices because simulations alone cannot supply the relevant results. Synthetic tests exercise controlled cases; live experiments measure how actual recipients respond.

He closes this part by returning to the complete loop: identify, prioritize and size impact, understand and execute the fix, check that it works without breaking something else, and keep monitoring. Verification therefore includes both the targeted improvement and regressions elsewhere. Deployment continues the learning process rather than ending it.

9:5610:00
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:56 · section reference included

When callers deliberately exploit the agent

Selected presentation frame from I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI at 766 seconds
When callers deliberately exploit the agent

Sharma then changes the assumption about the caller. Making an agent useful is already difficult when people earnestly want their problems solved. Deliberately adversarial callers introduce another class of risk: persuading an agent or a human to reveal protected health information or personally identifiable information. His broader warning about increasingly capable automated callers is a threat scenario; the recording does not establish the external capability example he invokes as a demonstrated event.

The pressure to make agents more capable expands this risk. Teams give agents more data and tools and deploy them quickly; Sharma argues that each increase in capability enlarges the surface available to an attacker. Natural, human-sounding voice can also be weaponized against people. His concern therefore covers both agents with access to sensitive systems and humans who may trust a persuasive caller.

Hamming released a red-teaming product in April to investigate how many agents it could break adversarially. Sharma reports that the team can break approximately one in five agents, with testing across financial services, healthcare, and consumer settings. He describes bypassing verification and obtaining data the testers should not have received. These are reported failures of access boundaries, although he supplies neither attack transcripts nor a precise definition of a broken agent. The rate indicates the concern within those tests without establishing an industry-wide prevalence.

11:5111:54
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:51 · section reference included

Combine testing, monitoring, and continuous red teaming

Selected presentation frame from I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI at 851 seconds
Combine testing, monitoring, and continuous red teaming

His defensive program starts with deep testing before deployment. Tests can operate text-to-text or voice-to-voice; he acknowledges advantages and disadvantages to both without detailing them here. This first layer checks whether the agent can serve ordinary people who simply want a problem solved. The recording supports using both testing modes as options, but does not establish that either one alone provides complete coverage.

The second layer is production monitoring through per-call scoring, manual evaluation and listening, and cross-call analysis. Sharma expands the object of monitoring to include the users as well as the agent. Teams need to observe what the agent says and does, while also recognizing callers who try to induce prohibited behavior. That distinction helps separate mistakes during legitimate use from deliberate attempts to exploit the system.

His additional recommendation is 24/7 red teaming, especially where bad interactions can be costly. This makes adversarial testing an ongoing activity alongside ordinary monitoring. He does not specify the operating cost, test volume, or a threshold for sufficient coverage, so continuous testing is presented as a response to high stakes rather than a quantified guarantee of safety.

13:3713:39
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:37 · section reference included

The promise of better calls carries personal exposure

Selected presentation frame from I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI at 925 seconds
The promise of better calls carries personal exposure

Sharma ends by preserving the positive case for voice agents. They could make interactions feel more human than clunky IVR trees, chatbots, or waiting on hold. His warning concerns what happens as deployment grows and bad actors exploit vulnerabilities: he expects incident counts to rise and more people to experience the consequences personally. This is his closing prediction, not a measured future outcome. It explains why he regards voice-agent risk as something callers will encounter directly.

The talk closes with an invitation to work in this area and to discuss whether a deployed agent's architecture and evaluations are configured correctly. Sharma mentions Hamming's substantial token use and offers to speak with attendees afterward. The practical emphasis remains on examining both the system that produces behavior and the evaluations used to judge it.

14:3814:40
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:38 · section reference included

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    My name is Suman Yu and I'm the founder

  3. 0:15

    and CEO of Hamming. And before working

  4. 0:19

    on voice agent reliability and safety, I

  5. 0:22

    worked at a company called Citizen

  6. 0:25

    out of New York. Anybody here use

  7. 0:26

    Citizen app? Awesome. Thank you. Uh, and

  8. 0:31

    at Citizen,

  9. 0:33

    we listened to crime, thousands of hours

  10. 0:36

    of police radio station data and sent

  11. 0:40

    millions of alerts to users in San

  12. 0:42

    Francisco,

  13. 0:44

    New York, LA, uh, Chicago, Baltimore,

  14. 0:48

    and so on.

  15. 0:50

    Some obviously gory and pretty sad. Uh,

  16. 0:54

    but others more funny like a person

  17. 0:56

    stealing bags of ice cream from Safeway

  18. 1:00

    or report of a man hanging off the side

  19. 1:02

    of the house after a woman stole his

  20. 1:04

    ladder.

  21. 1:07

    If I actually take a look at the citizen

  22. 1:09

    app right now for those who are

  23. 1:10

    customers or users, I can see that there

  24. 1:13

    is a man yelling at person. There's

  25. 1:16

    indecent exposure. This is real. This is

  26. 1:18

    real time. This is, you know,

  27. 1:21

    couple hours ago. These are real-time

  28. 1:23

    alerts that we're sending.

  29. 1:27

    Now, voice agents scare me more because

  30. 1:29

    they're finally graduating from demos

  31. 1:31

    and PC's to production. We should be

  32. 1:33

    super excited, but I'm nervous. I'm

  33. 1:35

    personally nervous. Uh, they're talking

  34. 1:37

    to users at a scale that would make Gary

  35. 1:39

    Tan and Polygram proud.

  36. 1:43

    When I got started in voice agent

  37. 1:45

    reliability in early 2024, voice was

  38. 1:48

    just starting to work. It was not quite

  39. 1:50

    good yet, but it was just starting to

  40. 1:52

    work. You would have to pay me a lot of

  41. 1:54

    money for me to stop using, you know,

  42. 1:55

    Aqua voice, Super Whisper, uh, Whisper

  43. 1:58

    Flow, and so on. These products are just

  44. 2:00

    getting super, super good. And a big

  45. 2:02

    reason is because the underlying

  46. 2:03

    infrastructure is getting better, and

  47. 2:05

    the orchestration layer is getting

  48. 2:06

    meaningfully better. It's getting much

  49. 2:09

    faster to build products and voice

  50. 2:10

    experiences that maybe are 60% good in a

  51. 2:14

    pretty short period of time, but the

  52. 2:16

    long tail is still Hey, Gorov. the long

  53. 2:18

    tail is still uh wise away.

  54. 2:21

    I think speech speech models are getting

  55. 2:23

    better. Um teams are experiment

  56. 2:25

    experimenting with hybrid architectures

  57. 2:28

    of combining more voicetovoice

  58. 2:32

    modalities and also cascading stacks to

  59. 2:34

    make the experience reliable but still

  60. 2:36

    pretty low latency.

  61. 2:39

    Things are obviously getting better.

  62. 2:41

    Agents are being connected to calendars,

  63. 2:42

    CRM, HRs, reservation systems, and so

  64. 2:45

    on. Voice agents can now take actions.

  65. 2:50

    However, reliability is still the number

  66. 2:52

    one problem holding back most voice

  67. 2:54

    agent deployments at scale. This is

  68. 2:55

    still the number one problem. This is an

  69. 2:58

    example I found on Twitter pretty

  70. 2:59

    randomly, you know, two weeks ago and a

  71. 3:02

    person is trying to get information for

  72. 3:04

    a tradein and gets absolutely confused

  73. 3:06

    with the information that they're

  74. 3:08

    receiving. Alex now has to correct for

  75. 3:10

    this loss of trust but trying to, you

  76. 3:13

    know, call the person and see what see

  77. 3:15

    what happened and fix the situation. Let

  78. 3:17

    me see if audio works here.

  79. 3:20

    >> Screwed up with another customer. We're

  80. 3:22

    getting it fixed, but I got to call him

  81. 3:24

    and see if I can work it out.

  82. 3:25

    >> I'm like, dude, half the time I'm like,

  83. 3:27

    I don't know if I'm talking to AI. I

  84. 3:28

    don't know if I'm talking to a person.

  85. 3:29

    It was just confusing, but we got there.

  86. 3:32

    >> It probably is AI and human. And

  87. 3:37

    >> so I think voices sound very confident.

  88. 3:39

    They sound very natural, but the

  89. 3:40

    information provided is often, you know,

  90. 3:42

    not correct. That's the biggest problem

  91. 3:44

    here.

  92. 3:46

    This example is more personal. I had

  93. 3:48

    booked an appointment with a physician a

  94. 3:49

    couple weeks ago or I thought I did. I

  95. 3:52

    showed up to the appointment and turns

  96. 3:54

    out I was not actually on the schedule.

  97. 3:57

    So the front desk, you me turned me

  98. 3:58

    away. I wasted 2 hours. For me, this was

  99. 4:01

    a waste of time. But what if this was

  100. 4:03

    actually your parent?

  101. 4:06

    What if this was your grandparent?

  102. 4:08

    What if this appointment was for a

  103. 4:10

    procedure instead of a regular checkup?

  104. 4:13

    The costs for these different

  105. 4:15

    permutations of the same failure mode

  106. 4:16

    can actually be super super high.

  107. 4:20

    Now, let's compare crime to voice

  108. 4:22

    agents. Um, I think observation number

  109. 4:24

    one is crime is actually decreasing over

  110. 4:26

    time. This is a good thing and I hope it

  111. 4:30

    crosses the x- axis at some point you

  112. 4:32

    know in the future.

  113. 4:35

    Voice on the other hand is generally

  114. 4:37

    taking off right we're seeing a pretty

  115. 4:38

    fast takeoff of voice agents being

  116. 4:40

    deployed in production. There's at least

  117. 4:42

    a trillion calls that are done every

  118. 4:44

    single year and majority of these will

  119. 4:46

    be done by conversational voice agents

  120. 4:48

    over the next you know five years. If

  121. 4:50

    you assume a 1% error rate that is still

  122. 4:53

    10 billion incidents per year. That's a

  123. 4:55

    lot.

  124. 4:58

    In practice, we currently monitor 10,000

  125. 5:00

    agents and the error rate is closer to

  126. 5:03

    10% in practice. These range from agents

  127. 5:06

    saying they found the right policy when

  128. 5:08

    they actually skipped the eligibility or

  129. 5:10

    verification steps or applying discounts

  130. 5:12

    when they were not really supposed to,

  131. 5:14

    misharing what the person said,

  132. 5:16

    providing incorrect information, or

  133. 5:18

    claiming they booked an appointment when

  134. 5:19

    they actually did not, just like it

  135. 5:21

    happened for me.

  136. 5:23

    Now, not every single call has an

  137. 5:25

    equally, you know, bad cost. Uh, some

  138. 5:29

    range, you know, in the crime land, some

  139. 5:30

    range from trash fires, which are kind

  140. 5:33

    of funny, annoying, not really hurting

  141. 5:35

    somebody. For a voice equivalent, that

  142. 5:38

    would be annoyances like repetition, um,

  143. 5:40

    or just sort of not quite understanding

  144. 5:42

    what the user is saying. all the way to

  145. 5:44

    safety risks like mass shootings or in

  146. 5:47

    the voice agent equivalent, it would be

  147. 5:49

    um a drive-thru that's deploying um

  148. 5:52

    voice agents at scale like a Taco Bell

  149. 5:54

    or McDonald's and a person orders a

  150. 5:57

    vegan burger with peanut allergies.

  151. 5:59

    If one of those two situations are not

  152. 6:02

    handled correctly, that is definitely a

  153. 6:04

    safety concern at scale.

  154. 6:07

    The other big difference between crime

  155. 6:10

    and and voice agent deployments is is

  156. 6:13

    crime generally tends to be pretty hyper

  157. 6:16

    local,

  158. 6:18

    tends to be very decentralized, right?

  159. 6:19

    Things like robbery or motor vehicle

  160. 6:22

    theft or lararseny. They're impacting a

  161. 6:25

    finite set of individuals that are

  162. 6:27

    involved in that um situation.

  163. 6:32

    On the other hand, voice agents are much

  164. 6:34

    more centralized. a single prompt change

  165. 6:37

    or an architecture change can have

  166. 6:38

    pretty massive implications downstream

  167. 6:41

    for all of the millions of you know

  168. 6:43

    users that are um in the crossfire. So

  169. 6:46

    the blast radius is is quite quite

  170. 6:48

    massive. So the natural question is how

  171. 6:51

    do you make these incidents much more

  172. 6:52

    visible and obvious? That's the kind of

  173. 6:54

    obvious question here.

  174. 6:57

    I'll borrow a framework from a couple of

  175. 6:58

    my friends who were OG growth folks at

  176. 7:01

    Facebook. So step one is to identify

  177. 7:04

    okay what are all the challenges and

  178. 7:05

    problems that um exist in your

  179. 7:08

    conversation experience. Step two is to

  180. 7:10

    prioritize an impact size. There's a

  181. 7:12

    frequency and severity analysis that's

  182. 7:14

    pretty important. Step three is to

  183. 7:16

    understand okay how do we actually fix

  184. 7:18

    this? Step four execute. Step five okay

  185. 7:21

    did my change actually work and did it

  186. 7:23

    cause any regressions somewhere else.

  187. 7:25

    And lastly we continue to monitor in

  188. 7:27

    production.

  189. 7:30

    On the y- axis, I think it's important

  190. 7:32

    to highlight there are known problems

  191. 7:34

    that already exist. Things like turnover

  192. 7:36

    latency, interruptions, um maybe some

  193. 7:39

    ASR problems you're, you know, aware of.

  194. 7:41

    And these are known problems that exist

  195. 7:44

    that the team should track over time. On

  196. 7:47

    the other axis is actually emerging

  197. 7:49

    behavior or patterns that are only

  198. 7:51

    obvious across lots of conversations. Um

  199. 7:55

    on the x- axis, you have coverage just

  200. 7:57

    like insurance. Are you analyzing few

  201. 8:00

    conversations? Are you analyzing many,

  202. 8:02

    many conversations? Most teams will

  203. 8:04

    typically start by listening to calls

  204. 8:06

    manually.

  205. 8:08

    And I think that's the best place to

  206. 8:09

    start. I don't think you should skip

  207. 8:11

    that step. There's a lot of depth and

  208. 8:13

    insights to get by actually listening to

  209. 8:15

    specific conversations and building that

  210. 8:17

    texture that that comes from that

  211. 8:19

    intuition. However, it's obviously not

  212. 8:21

    scalable. So most teams end up having a

  213. 8:24

    spreadsheet of I don't know five or 10

  214. 8:27

    different rubrics around greetings,

  215. 8:29

    closing, validation,

  216. 8:31

    um, core logic and so on. To scale that

  217. 8:35

    up even further, you then end up

  218. 8:36

    investing in some eval product, right?

  219. 8:39

    You might run some element as a judge

  220. 8:41

    and compute classic metrics and also

  221. 8:44

    more more deterministic and stoastic

  222. 8:47

    scoring logic. Um, but there you're

  223. 8:49

    still stuck with checking for

  224. 8:51

    consistency of known problems, but

  225. 8:53

    you're not really discovering novel

  226. 8:54

    insights that are actually happening

  227. 8:56

    across conversations. We're spending a

  228. 8:58

    ton of time on performing cross

  229. 9:01

    conversation analysis, not a pattern on

  230. 9:03

    a single call, but across conversations.

  231. 9:05

    And some of the best teams that we work

  232. 9:07

    with are are doing the same.

  233. 9:10

    Now, to prioritize an impact size, I

  234. 9:12

    think there's problems that are one-off

  235. 9:14

    that are low impact. I mean, who cares?

  236. 9:17

    uh even low impact and systematic

  237. 9:19

    problems in the crime world that would

  238. 9:21

    be a trash fire in a voice aation world

  239. 9:23

    it could be some repetitions the team is

  240. 9:25

    experiencing they're still annoying at

  241. 9:27

    scale and if you are doing a bake off

  242. 9:29

    it's still worth solving for them I

  243. 9:31

    would not ignore these class of problems

  244. 9:33

    oneoff and high impact well hope it

  245. 9:35

    doesn't chronic and I think systematic

  246. 9:38

    and high impact are obviously the P 0

  247. 9:40

    you know target areas um for the team to

  248. 9:42

    solve an example of that

  249. 9:45

    would be in a fins serve capacity

  250. 9:47

    There's a voice agent that um helps

  251. 9:49

    users freeze their credit cards. And if

  252. 9:51

    it doesn't do that, well, that's a

  253. 9:53

    massive fail.

  254. 9:56

    All right. So, understand and execute.

  255. 9:58

    I'm pretty sure everyone's doing this.

  256. 10:00

    Please fix my agent. Uh I think fixing

  257. 10:02

    or rather attempting to make a fix is

  258. 10:05

    the simplest and the lowest effort

  259. 10:08

    component of this debugging pipeline and

  260. 10:11

    loop. Um the next step is all right, I

  261. 10:14

    made a change to my system. How do I

  262. 10:16

    actually know this thing works um for

  263. 10:19

    real? A great way that's naive is to

  264. 10:23

    take a real call, for example, in my

  265. 10:25

    case, I booked an appointment and it

  266. 10:27

    didn't get scheduled and replay that

  267. 10:29

    exact conversation and run that maybe 5,

  268. 10:32

    10, 20, 50 times and see, okay, what is

  269. 10:34

    my probability of passing this type of

  270. 10:37

    issue? A better way is to keep the same

  271. 10:40

    intent but change the wordings, change

  272. 10:44

    the patterns, change the accents, change

  273. 10:46

    the style, add one more intent to the

  274. 10:48

    mix. And that gives teams much more, you

  275. 10:51

    know, better coverage to feel confident

  276. 10:53

    that yes, I actually made a change and

  277. 10:56

    my changes are net positive instead of

  278. 10:58

    net negative.

  279. 11:01

    There are certain fixes and I guess

  280. 11:05

    hypothesis that are very difficult to

  281. 11:07

    test in a pre-eployment synthetic

  282. 11:09

    setting. And so AB testing ends up

  283. 11:11

    being, you know, pretty pretty critical

  284. 11:13

    for those circumstances. For example, if

  285. 11:15

    you have an outbound agent, the first 5

  286. 11:17

    seconds of a conversation tends to be

  287. 11:19

    the most important. And so the vocal

  288. 11:22

    quality um and the specific words you

  289. 11:25

    end up using, they matter the most. And

  290. 11:27

    so AB testing that is the only way in in

  291. 11:29

    kind of real life setting to to get

  292. 11:32

    results. You can't really do it through

  293. 11:34

    simulations alone.

  294. 11:37

    And so there we have the loop. Identify,

  295. 11:39

    prioritize, impact size, understand the

  296. 11:42

    fix, execute, check, make sure it didn't

  297. 11:45

    break anything, and then continue

  298. 11:47

    monitoring.

  299. 11:51

    So I think making voice agents useful is

  300. 11:54

    already hard as it is. even when dealing

  301. 11:58

    with earnest users on the other line,

  302. 12:01

    right? These are people who who just

  303. 12:02

    want their problem solved. They're not

  304. 12:03

    trying to mess with you. These are like

  305. 12:04

    legit normal people.

  306. 12:08

    Now, what happens when mythos learns how

  307. 12:10

    to dial?

  308. 12:13

    So, if it can extract trade secrets and

  309. 12:16

    uh you know, from the NSA, it can

  310. 12:18

    certainly, you know, seduce you into

  311. 12:22

    revealing PHI and PII data as well.

  312. 12:26

    And I think both voice agents and humans

  313. 12:29

    will [clears throat] be targeted here.

  314. 12:31

    Voice agents because there's a pressure

  315. 12:33

    to make these more capable. Give them

  316. 12:37

    access to more data. Give them access to

  317. 12:39

    more tools.

  318. 12:41

    Deploy them quickly.

  319. 12:45

    The more the capability, the bigger the

  320. 12:46

    surface area. This is this is pretty

  321. 12:48

    pretty common sense. And the more the

  322. 12:50

    voice agents become natural and human

  323. 12:54

    sounding, the more humans will be

  324. 12:56

    tricked along the way as well for those

  325. 12:58

    who are weaponizing.

  326. 13:01

    Uh we ship a we shipped a red tipping

  327. 13:03

    product um back in April just to test

  328. 13:05

    out this hypothesis for how many agents

  329. 13:07

    can we actually break from a adversarial

  330. 13:10

    capacity and we can probably break one

  331. 13:12

    in five agents at this point. We've

  332. 13:14

    tested this across financial services,

  333. 13:16

    healthcare, um consumer and so on. We've

  334. 13:19

    bypassed verification. Uh we've

  335. 13:22

    definitely had agents, you know, we've

  336. 13:24

    been able to promject uh several agents

  337. 13:27

    and and gotten data we should not have.

  338. 13:29

    So this is not theoretical. This is

  339. 13:31

    actually a real a real concern.

  340. 13:37

    I think the only real defense against

  341. 13:39

    the dark arts is

  342. 13:41

    step one to invest deeply in

  343. 13:43

    pre-eployment testing. This could be

  344. 13:45

    textto text. This could be voice to

  345. 13:48

    voice. There's pros and cons to both.

  346. 13:50

    Happy to chat offline if folks are

  347. 13:51

    interested. And this is just making sure

  348. 13:54

    you're not self-owning, you know, when

  349. 13:56

    you're talking to real people who just

  350. 13:57

    want to get their problem solved. Step

  351. 13:59

    two is to have a great monitoring system

  352. 14:02

    of all kinds. And I've highlighted, you

  353. 14:04

    know, different flavors of monitoring

  354. 14:06

    per call scoring, manual kind of evals,

  355. 14:10

    you know, listening to conversations and

  356. 14:11

    cross call analysis. And this is helpful

  357. 14:14

    both for monitoring what the agent is

  358. 14:16

    saying and behaving and how it's

  359. 14:17

    actually doing, but also the users. Are

  360. 14:19

    the users being adversarial? Are they

  361. 14:21

    being annoying? Are are they trying to

  362. 14:22

    trick the agent into doing things it's

  363. 14:24

    not supposed to be doing?

  364. 14:26

    And I think our new recommendation now

  365. 14:28

    is to run 24/7 red teaming um for your

  366. 14:31

    agents, especially if you believe the

  367. 14:34

    cost of bad interactions can be can be

  368. 14:36

    rather large.

  369. 14:38

    So, I think voice agents um have this

  370. 14:40

    awesome potential of of making the world

  371. 14:43

    feel much more human compared to

  372. 14:48

    interacting with clunky IVR trees or

  373. 14:50

    chat bots or worse um being stuck on a

  374. 14:53

    on a hold.

  375. 14:55

    And when we think about crime, we often

  376. 14:57

    think of crime happening to somebody

  377. 14:59

    else. You know, crime does not happen to

  378. 15:01

    you typically with voice agents,

  379. 15:04

    especially bad actors. As these agents

  380. 15:07

    are deployed and as bad actors start to

  381. 15:09

    exploit a lot of the vulnerabilities,

  382. 15:11

    the number of incidents is about to kind

  383. 15:13

    of go way way up. And so the reason I

  384. 15:16

    fear voice agents more than crime is

  385. 15:18

    that one of these incidents is going to

  386. 15:20

    impact you. It already did for me.

  387. 15:24

    Awesome. So it's time for me to shill.

  388. 15:25

    Well, we burn a lot of tokens. If you

  389. 15:28

    are interested in working in this space,

  390. 15:31

    please come and talk to us. And if you

  391. 15:33

    are deploying voice agents and want to

  392. 15:35

    validate whether your architectured or

  393. 15:39

    your eval are set up correctly, please

  394. 15:40

    come and talk to us. We'll be outside.

  395. 15:42

    And here's here's my number. Here's my

  396. 15:44

    WhatsApp. Thanks everyone.

  397. 16:02

    >> [music]