← All AI Engineer talks

AI Engineer World's Fair 2026

The Prompt Is Still a Punch Card

Read the talk

The Prompt Is Still a Punch Card

Natural language expanded what computers can understand, but submitting a complete request and waiting still leaves people managing the interaction. Conversational AI needs a different protocol.

From a talk by Ted Johnson

Type, submit, wait, repeat

You type a request into a small box and wait. A cursor blinks. A progress indicator cycles through whimsical activities—hullabalooing, tomfoolering, philosophizing—while the system works. Perhaps the answer is useful. Perhaps you rephrase the request and start again. The remarkable thing is how normal this sequence now feels.

Close-up of a cursor and busy indicator beside a progress bar reading 78%.
“Please wait…” above a progress bar marked 78%.

Ted Johnson, co-founder of JoinIn AI, wants to make that familiarity uncomfortable. After 25 years building enterprise software, collaboration systems and AI interfaces, he experienced ChatGPT as both an extraordinary advance and a disappointment in interaction design. Why should something so capable still require people to learn an unnatural way of communicating?

His framework separates three things that are easily conflated: channel, the medium carrying intent; expression, the range and richness of meaning an interface permits; and protocol, the rules governing the exchange. Improving one does not automatically improve the others. That distinction makes it possible to see what prompting changed—and what it inherited.

0:000:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:00 · section reference included

The keyboard feels natural because we learned it

The keyboard is familiar, ubiquitous and learned. Typing takes lessons and practice, and attempts to improve it have continued for generations. Dvorak and Colemak rearrange the keys to reduce finger movement. Some enthusiasts give each thumb four, eight or ten keys rather than spending both thumbs on a space bar. Even people who love keyboards recognize that their conventional arrangement is not inevitable.

Johnson shows a historical patent drawing, which he dates to about 1860, followed by the Hansen Writing Ball: a reminder that a radically different physical arrangement was possible. The broader point does not depend on a particular layout winning. We inherited an input device shaped by old constraints, learned to accommodate it, and eventually stopped noticing the accommodation.

A channel determines what signals can travel through an interface. A keyboard, microphone, screen, punch card and prompt box each offer different possibilities.

Channel or representationWhat it can carry
TextA sequence of discrete symbols
VoiceWords, timing, pitch and hesitation
DiagramSpatial relationships visible together

These are differences in signal capacity, not guarantees of understanding. A microphone can capture hesitation without a system knowing what that hesitation means. People routinely combine channels; forcing every interaction through one medium discards options we normally use without thinking.

2:072:14
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:07 · section reference included

An ocean of expression inside the same box

With language models, the keyboard and Submit button remained in place, but the range of acceptable input expanded dramatically. This is progress in expression: how much of what a person means the interface lets through. The physical act of typing is unchanged; what those keystrokes can communicate is not.

Assembly offers opcodes. Shells add commands, inputs and flags. Programming languages provide primitives that can be composed into increasingly sophisticated instructions. Each remains a machine-defined vocabulary within which the user must formulate intent. Natural language opens that vocabulary to ordinary requests. The slide places this progression alongside form fields, ending with open language.

Five panels labeled Opcodes, Shell, Primitives, Fields, and Open Language show code, form controls, and a natural-language request.
From Opcodes to Open Language: five forms of computer input.

An ordinary request contains context, nuance and intention that people often do not spell out to one another. Language models make much more of that meaning available to software. Yet, in Johnson’s framing, channels that have persisted for decades—and in some cases much longer—now carry this enormous expressive range through essentially the same narrow interaction. More meaning fits through the box, but the user still has to decide how and when to package it.

4:204:30
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:20 · section reference included

Faster batch is still batch

The remaining constraint is protocol: the shape of the interaction and the rules participants must follow. Johnson describes the recent LLM advance as an explosion in expression, while prompting retains the structure of punch-card batch processing.

The historical workflow made the separation explicit:

  1. Encode the entire request away from the computer.
  2. Carry the card deck to an operator and submit the job.
  3. Wait for the run to finish.
  4. Inspect the printout, find the mistake, repair the deck and resubmit.

The machine engaged with the finished package, not with the person while they were working out what they wanted.

StagePunch-card jobPrompt interaction
PrepareAssemble the deckCompose the request
SubmitDeliver the jobSend the message
WaitAwait the printoutAwait the response
RepairCorrect and resubmitRephrase and resend

Updates and summaries improve visibility, and the wait may shrink from hours or overnight to seconds or minutes. But response speed does not determine when the system is allowed to participate. If participation begins only after the user submits a complete turn, the central batch constraint remains. A voice interface that transcribes speech into a box and submits it preserves that constraint too; changing the input channel alone does not change the protocol.

6:146:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:14 · section reference included

Expertise in assembling the request

Prompt engineering often teaches people how to accommodate this protocol: ask for step-by-step thinking, provide examples, request an expert persona—or avoid one. Paste more context, or less. Communicate through Markdown documents. The advice can be useful, but its contradictions also expose how much work goes into preparing a package the machine will accept.

Overlay lists six suggestions, including thinking step by step, giving examples, asking for expertise, and pasting more or less context.
“Packaging a good batch” lists familiar prompting advice.

That skill resembles knowing how to assemble a punch-card deck so the job succeeds. Johnson is not dismissing prompts, punch cards or command lines; each is a powerful response to particular constraints. His question is whether those constraints still justify making users finish packaging their intent before the machine can help.

A system could instead ask a follow-up, clarify a thought as it develops, or acknowledge missing information. Johnson invokes MIT professor Sherry Turkle’s description of conversation: “Conversation is the most human and humanizing thing we do.” Her original context is human relationships; here, the quotation supplies a standard for the richness of interaction that a submit-and-wait loop leaves out.

The punched-card lineage reaches back to weaving looms, where a pattern could be set in advance and then run through the machine. Johnson uses that lineage as an analogy for inherited interaction rules. The objection is not that natural language is a poor encoding. It is that the LLM is permitted to engage only after the whole pattern has been prepared.

8:048:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:04 · section reference included

The work left outside the model

As reasoning, speech, vision, memory and planning improve, the surrounding interface can still leave the human coordinating everything. The user selects relevant context, remembers which question to ask, chooses the moment, notices ambiguity and repairs the output. A capable model behind a prompt box does not automatically take responsibility for any of those tasks.

When the interaction fails, users may conclude that they were insufficiently specific or simply do not understand AI. Johnson locates the problem elsewhere: the interface asks people to operate a new kind of intelligence through an inherited protocol. The race to expose model capabilities has moved faster than the design of the interaction around them.

10:1010:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:10 · section reference included

Hearing speech is not knowing whom it addresses

Johnson’s co-founder tried a simple interaction with an unnamed frontier company’s speech-to-speech voice mode. He asked when the next Timberwolves game would take place, and it answered quickly. He then pretended Johnson had arrived and said, “Hey, Ted, come on in.” The AI responded, “Sure, I'm here. What's on your mind?”

The failure was not answering the sports question. It was treating speech intended for another person as a new message requiring an AI reply. In this example, a one-message/one-reply protocol did not adequately represent the speaker’s intended recipient. Speaker identity, addressee and the obligation to respond are separate pieces of conversational state.

Johnson also points to GPT-Realtime-2 and reports backchannels such as mm-hmm and right in voice mode. OpenAI dates the API announcement to May 7, 2026, earlier than the late-May release timing he gives; that announcement does not establish the separate Voice trial timing or those particular backchannels. The design direction he highlights is participation during an exchange, rather than merely answering after it.

11:0911:21
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:09 · section reference included

Yielding, listening and following the new thread

The next clips demonstrate NVIDIA PersonaPlex, an external research model rather than JoinIn AI’s system. A person says they are thinking about starting a diet. The model begins recommending simple changes, including more vegetables and fruit. Before it finishes, the person interrupts to say they have signed up for a marathon. The model yields, acknowledges the new information and switches to running preparation: regular long runs, hydration, fueling and recovery.

The key behavior is not just cancellation of audio. The model stops, gives up the floor and follows the changed conversational thread. Johnson describes this as real turn-taking: listening and speaking in real time rather than requiring an isolated, completed input before responding.

A second clip is less task-oriented. Someone invites a conversational partner to visit the city, then works through a thought about spray-painted murals and whether people commission them. Short acknowledgments land within the ongoing speech. The listener can signal attention without converting every pause into a full answer or forcing the speaker to surrender the floor.

Conversational flow and group awareness remain different problems. Backchanneling does not by itself establish who is in the room or whether a remark was addressed to the AI. Johnson introduces JoinIn’s work at that boundary: improving the protocol’s representation of human and group conversation.

12:3312:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:33 · section reference included

A meeting has goals, participants and a floor

The JoinIn demonstration opens with greetings involving Sam and Jordan, then a question about the requirement under discussion. The answer identifies REQ-442, expense approvals. Johnson explains the participation mechanism as utility-driven: the system labels statements as questions, proposals or answers, creates goals to fulfill, and takes a turn only when nobody else is speaking or holding the floor. Silence and an available turn are not necessarily the same thing.

The group wants users to approve requests faster, but the scope is initially loose. A participant asks what kinds of requests they mean. Expense approvals come first, with access requests later. Then someone tells the AI to hold while the people clarify whether they are designing expense approvals specifically or a general approval workflow. They settle on expense approvals for the first release and access requests as future scope; a participant notes that this distinction changes the data model.

A participant next asks the AI to pull the material up for everyone. This request explicitly addresses the assistant, making the recipient easy to identify. Johnson flags recipient tracking as essential because later requests will not include an explicit AI address. The system must follow both the topic and the relationships among utterances.

14:1014:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:10 · section reference included

From discussion to a captured requirement

The participants work through an approval constraint: expenses over $5,000 need a second approver, involving the manager and finance. A large expense therefore cannot be handled as a single tap; below the limit, they accept one-tap approval. The group explicitly agrees on expense approvals with a $5,000 threshold. Johnson identifies this as the moment the AI resolves its scope objective without anyone composing a dedicated prompt or pressing Submit.

Only afterward does someone explicitly ask the AI to capture the decision. Its spoken summary includes expense approvals for the first release, excludes access requests, and lets managers approve or reject an expense directly from a notification. It attributes routing above $5,000 to the finance controls policy. A compact representation of that captured content is:

json

{
  "id": "REQ-442",
  "firstRelease": ["expense approvals"],
  "outOfScope": ["access requests"],
  "notificationActions": ["approve", "reject"],
  "secondApproverAboveUSD": 5000
}

The distinction is between following the discussion to resolve an objective and responding to a later request to capture it. Those are separate conversational events.

A participant then asks to make the threshold 10,000 instead of 5,000, without explicitly addressing the AI. The assistant recognizes a possible revision and asks whether it should update the requirement to that threshold. The participant confirms. This establishes a confirmed change request; the clip does not establish authorization to override the finance policy or show a persisted requirement update.

A confirmed revision is not a demonstrated write

Constructed example: Object IDs, state labels and the JSON field secondApproverAboveUSD are teaching representations, not disclosed implementation fields. Requirement identity, scope, threshold values and confirmation come from the demo.

Participant revision — unchanged
Actually, make the threshold 10,000, not 5,000.

Operation: Recognize the revision, ask whether to update the requirement, and receive the participant's confirmation.

Requirement

Before: Captured requirement
REQ-442
After: After verbal confirmation · Unchanged
REQ-442

First release

Before: Captured requirement
Expense approvals; access requests out of scope
After: After verbal confirmation · Unchanged
Expense approvals; access requests out of scope

Last captured second-approver rule

Before: Captured requirement
Expenses over $5,000
After: After verbal confirmation · Changed
Expenses over $5,000; persisted revision not shown

Confirmed change request

Before: Captured requirement
Not present
After: After verbal confirmation · Added
Set secondApproverAboveUSD to 10000; participant confirmed
The assistant infers the intended revision and seeks confirmation; the demonstrated outcome is a confirmed request.

The conversation then changes topics. Someone asks whether the room is free after the meeting. The AI begins answering, is interrupted by a clarification asking whether it is available until 3:00, and confirms that it is. The assistant handles a side question and an interruption without requiring the meeting to be repackaged as a new, carefully constructed task.

14:2415:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:44 · section reference included

Which burdens can the interface take over?

These examples suggest a broader role for AI: intelligence is also an interface technology. Better models behind prompts, agents and loops are only part of the opportunity. Reasoning, listening and inference can help remove the constraints that previously required people to adapt to a machine’s syntax, forms and timing.

The design question becomes: what are humans still doing only because machines used to be unable to do it? The answer need not be another chat window, voice mode or wall of Markdown. It may be a question, a pause, a sketch, a checklist, a quiet aside or no response at all. An intelligent interface can choose an appropriate channel and moment instead of making the user manage both.

Johnson expects removing these burdens to reduce friction and support adoption. Earlier interfaces—punch cards, typed commands, menus, phone gestures and prompts—each improved how people encode intent, while carrying some constraints into the next era. He describes the resulting work as translation, precision, context and repair taxes.

Putting those taxes down does not require making every interaction voice-based, promising magic or replacing human judgment. It requires computers to become more fluent with people. The deeper design challenge is not simply teaching users how to adopt AI: it is deciding how greater machine understanding should change the interface, so that people no longer have to reshape every thought before a machine can participate.

17:1917:33
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:19 · section reference included

Resources

From the talk

  • NVIDIA PersonaPlexArticle12:33

    Research overview with interruption and backchannel examples, architecture, training details and evaluation context.

  • OpenAI's dated API launch announcement explains realtime reasoning, interruption handling and tool use.

Updates since the talk

  • Current meeting-assistant product with objective tracking, contextual answers and a desktop beta.

Read the complete timestamped transcript
  1. 0:00

    [upbeat music] I'm sure that sometime in the last few hours, most of you did this. You typed a request into a small box to a superintelligence, and then you waited.

  2. 0:16

    You watched the cursor blink, maybe a little throbber cycled through clever gerunds like hullabalooing, tomfoolering, and philosophizing to hide the wait. Maybe it gave you what you wanted, maybe you rephrased it and tried again.

  3. 0:32

    It all felt completely normal. I wanna spend the next twenty minutes making prompting feel unfamiliar and strange again. I'm Ted Johnson, co-founder of JoinIn AI. During my twenty-five-year career building enterprise software, collaboration systems, and AI-enabled interfaces, I've always focused on human interaction.

  4. 0:53

    I've also been following AI for two decades, including back to the far less impressive GPT-1 and 2. And when ChatGPT arrived, I felt two things at once. First, and unsurprisingly, amazement, knowing the world would never be the same, followed by actually surprising disappointment I couldn't shake.

  5. 1:13

    This disappointment turned into an observation that started a company, JoinIn AI, and that I keep coming back to, which is, why do we still have to learn AI? Why does something s- this powerful so often feel unnatural to use?

  6. 1:29

    Here's the path we'll take to answer that. We'll start with the most familiar computer interface and make it strange again. Then I'll give you three key concepts, the channel, the physical transport that carries your intent, expression, the range and richness of meaning the channel can carry, and the protocol, the shape or rules of this exchange.

  7. 1:51

    I'll use those three concepts to show you that the prompt is our present-day punch card. We'll share examples of the ways interfaces could progress, and we'll wrap with some practical advice for AI and human-centered design.

  8. 2:07

    Everyone knows what this is, the keyboard. It's everywhere, and it feels completely normal or natural,

  9. 2:14

    but it isn't. We all had to take lessons. We all had to practice. And I say this as someone who loves keyboards, but it seems that we've been trying to fix them as long as they've been around.

  10. 2:25

    People have tried more efficient layouts like Dvorak or Colemak to save their fingers some work. A more extreme, uh, example, some keyboard enthusiasts refuse to squander their two digits on the space bar, giving them four or eight or ten keys, uh, to press with that efficient thumb of theirs.

  11. 2:44

    And what do we even mean by the keyboard? Here's the patent drawing for the layout we use every day. This patent's from about 1860. And my personal favorite, it just as could have easily been the Hansen Writing Ball, which looks anything but like a way you wanna talk to a superintelligence.

  12. 3:03

    We carry these legacies of an arbitrary input device designed under constraints that haven't existed for a century, and we put it between ourselves and the most capable machines ever built.

  13. 3:13

    Nobody alive chose it. We all inherited it, and then we stopped noticing. So that's the first idea, the channel, the medium an interface gives you to work in. A keyboard is a channel, a microphone, channel, a screen, a punch card, a prompt box, all channels.

  14. 3:30

    And channels matter because each one can physically carry a different kind of signal. For example, text is a stream of discrete symbols. Voice can carry timing, pitch, hesitation, and words.

  15. 3:42

    A diagram can carry spatial relationship all at once. But these are differences in what the medium can transmit, its bandwidth, not the differences in the meaning. And carrying more signal isn't the same as the machine understanding any of it.

  16. 3:58

    That's a separate question. That's the next idea. Humans use all these channels constantly without thinking. We never pick one channel and force everything through it. That would be absurd.

  17. 4:09

    And yet, that's what we ask people to do with machines over and over. Hold on to the word channel because here's the plot twist. With AI, the channel never really changed.

  18. 4:20

    You're still typing into a box, but what you were allowed to push through it was about to. For the first time, the computer channels carry rich, complete human language.

  19. 4:30

    Notice I said what it carries, not what it is. You're still typing into a box. You're still hitting the Submit button. The keyboard didn't change. What improved is the range of what you're permitted to express through it.

  20. 4:43

    There's the second idea, expression, how much of what you actually communicate or mean will go through the interface. That's the second idea, expression, how much of what you actually communicate or mean will the interface let through.

  21. 4:58

    Here's an example of expression progress over time with computers. Starting with assembly, which gave you an instruction set, a few dozen opcodes, then came the sh- uh, commands with the shell, inputs, flags, then modern programming languages gave you primitives that you could compose.

  22. 5:17

    While powerful step by step, each one of these is a fixed vocabulary requiring you to express your intent by choosing from a menu the machine will accept. Natural language blew that menu open.

  23. 5:29

    For the first time, you can say almost anything the way you'd say it to another person. And on the expression one axis, the leap is real and enormous. There's an ocean of meaning in an ordinary human request, context, nuance, intent, all things we've never had to spell out to each other.

  24. 5:47

    For the first time, you can say almost anything the way you'd say it to another person. For the first time, a machine can take it in. So here's what should bother us as engineers and designers, as it's bothered and inspired me, the channels for computers have been the same for fifteen years, some a hundred and eighty years.

  25. 6:05

    Now, with AI, we've poured an ocean of expression into it. So why does it feel like we're still sipping through a straw and struggling to learn how the AI thinks?

  26. 6:14

    Because there's a third idea underpinning the other two, and it's really the one that hasn't kept up, the protocol, the rules you follow, and the shape of the interaction itself.

  27. 6:26

    Channel stayed the same. Expression exploded in the last three years with LLMs, but the protocol, prompting, is the protocol of a punch card, and the punch card's protocol is good old batch.

  28. 6:38

    Here's what punch card batch meant. You sat down, away from the machine, carefully encoded your entire request in advance, carried your deck to the operator, you submitted the job, and then you waited, sometimes hours, sometimes overnight.

  29. 6:54

    Then you read the printout, found one thing that was wrong, fixed it, resubmitted it, and waited again. The machine never engaged with you while you were thinking. It engaged with the finished package after the fact.

  30. 7:08

    Now let's look at the prompt. Assemble the whole request, submit it, wait, read what comes back. Something's off, assemble it again, submit again, and wait. We have to acknowledge that there are features, interactive features improving this.

  31. 7:23

    You can ask for updates. You can ask for summaries of what was done. But in the end, it's still batch with interactive sprinkles. It's the same protocol. We shrank the wait time from overnight to a few seconds or a few minutes, and the speed fooled us into thinking that it had become interactive.

  32. 7:41

    It hasn't. It's still batch. You still package a complete turn before the machine is allowed to participate. We learn tricks, send tips to use code skills or rewrite prompts a certain way to manage this.

  33. 7:54

    And speaking doesn't change it. Your voice just gets transcribed into the box and submitted. Shorter batch is still batch because the protocol is the part that did not advance.

  34. 8:04

    The protocol is the part we've had to learn. We just gave it a flattering name. We call it prompt engineering and treat it like it's a power user skill.

  35. 8:12

    Strip the label off, and it's a set of rules for packaging up good old batch. For example, tell it to think step by step. Give it examples. Ask it to be an expert.

  36. 8:23

    Don't ask it to be an expert. Don't ask it that way. Paste more context. Paste less context. Only talk to it through markdown documents. We trade incantations. We've learned the magic words.

  37. 8:36

    That's the illusion. It feels like mastery, but it's the same sort of mastery a punch card operator had, knowing exactly how to assemble the deck so the job wouldn't fail.

  38. 8:45

    Moreover, we've gotten good at prompting or these black boxes, and, and that's the part that should bother us, not reassure us. None of this means prompts are bad. Punch cards weren't bad.

  39. 8:55

    Command lines aren't bad. They're brilliant solutions for constraints of their time. But that's the whole question. Is batch still the right protocol? Are we still prepackaging our intent for a machine that no longer needs us to?

  40. 9:08

    Because it shouldn't need us anymore. It can ask a follow-up. It can clarify mid-thought. It can notice it's missing something and say so. It should be human conversational. Sherry Turkle of MIT puts it very well: Conversation is the most human and humanizing thing we do.

  41. 9:26

    It's one of humanity's superpowers. The capacity to engage and think is right there, and yet we're still making people submit the deck and wait for the run. Even the punch card inherited its protocol, in fact.

  42. 9:39

    Batch came from the weaving loom. You set the whole pattern in advance, then ran the cloth. The punch card got reused on computers by default. We're still standing at the same moment again.

  43. 9:51

    AI could finally meet us in the middle of a thought, got handed a protocol of a loom. That's what I mean. The prompt is still a punch card, not because of how you encode it, the encoding is powerful and awesome, because when the LLM is allowed to engage only after you've packaged a complete turn and submitted it,

  44. 10:10

    and this is where the mismatch bites. Model capacity is shooting straight up. Reasoning, speech, vision, memory, planning, all curving upwards. The interface protocol, flat. Still a box, still a submit button, still the human doing all the work around the LLM.

  45. 10:26

    The human still decides what context matters, still remembers what to ask, still chooses the timing, still notices the ambiguity, still repairs the output, still has to carefully engineer a prompt.

  46. 10:37

    But the intelligence feels magical. It's the interface that still feels like work. And when it feels like work, when the output's wrong, when the magic words don't land, people blame themselves.

  47. 10:48

    They decide they're bad at this. They're not specific enough. They don't get AI. I want to say as clearly as I can, it is not our fault. We are not bad at using AI.

  48. 10:58

    We are being asked to operate a brand-new kind of intelligence through a protocol of a punch card. The mismatch isn't the user, it's the interface. In the race to enable AI, we shortcut the interface.

  49. 11:09

    Okay, let's make this concrete and familiar. A few weeks ago, my co-founder was using a frontier company's voice mode. These are known as speech-to-speech models.

  50. 11:21

    He asked it a normal question, "When is the next Timberwolves game?" Fine. It answered it quickly. Then he pretended I showed up as if to speak to me and said, "Hey, Ted, come on in."

  51. 11:34

    He wasn't talking to the AI, but these models have no way to know that. So the AI did the only thing a prompt box can do. It took his speech as a turn and answered it.

  52. 11:47

    "Sure, I'm here. What's on your mind?" That's not a good answer, but it's not a dumb model. It answered the first question perfectly, but it's a protocol with exactly one slot, your message, then its reply.

  53. 12:01

    It has no concept of who's speaking, whether the words were even meant for it. And the frontier companies want to make strides as well. OpenAI released GPT Realtime 2 in late May and started trying it for their voice mode as well.

  54. 12:16

    It backchannels now. It goes, "Mm-hmm," and, "Right," the little sounds we make to show we're listening actively. We're seeing the field is converging on the same conclusion we built our company on: the interface has to stop being batch and start participating.

  55. 12:33

    Others are working on real-time conversation as well. This is NVIDIA's PersonaPlex, a research model, not ours. Watch what happens when it gets interrupted.

  56. 12:44

    I've been thinking about starting a diet.

  57. 12:46

    Yeah, starting a diet can feel a bit daunting, but you could keep it simple. Focus on eating more veggies and fruits. Try to-

  58. 12:53

    Oh, before I forget, I signed up for a marathon.

  59. 12:56

    All right. Congrats on signing up for the marathon. That's a big challenge. You've got a lot of time. Focus on building a solid base with regular long runs. Stay hydrated.

  60. 13:05

    Make sure you fuel right before and after, and don't forget to stretch and take care of your feet.

  61. 13:10

    PersonaPlex stops. It yields. It picks the thread back up. That's real turn-taking, listening and speaking at once in real time.

  62. 13:21

    You need to come visit me-

  63. 13:22

    Oh, okay

  64. 13:23

    ... 'cause then we can go into the city.

  65. 13:24

    Okay.

  66. 13:25

    'Cause that's the thing, like-

  67. 13:26

    Yeah

  68. 13:26

    ... there's, like, the random spray paint, but then there's also, like...

  69. 13:30

    I, I'm not su- people must commission them-

  70. 13:33

    Yeah, I think they do

  71. 13:34

    ... like these massive, like, mural spray paint pieces.

  72. 13:36

    Yeah.

  73. 13:36

    Um-

  74. 13:37

    I think they do

  75. 13:38

    And PersonaPlex's backchannels, listens, and lands where a person's would.

  76. 13:44

    Beyond conversational flow, there are lots of challenges, and it's a complex problem. Making listening noises is not really the same as knowing who's in the room. These are s- not trained to tell that, "Hey, Ted," wasn't meant for it,

  77. 14:01

    but we are. We're working on improving the protocol to the models by giving it a better understanding of human and group conversation.

  78. 14:10

    Good afternoon, everyone. Good afternoon, Sam. Hi, Jordan.

  79. 14:17

    Good afternoon. Good to see you both.

  80. 14:20

    Quick one. Which requirement is this? Do we have an ID?

  81. 14:24

    This is [REDACTED:generic_id], expense approvals.

  82. 14:28

    There, it answered a question. It only takes actions based on the utility-driven model, so it creates goals to fulfill as it labels each of the participants' statements as a question, a proposal, an answer, and then only takes a turn when no one else is speaking or holding the floor.

  83. 14:47

    Right. We need users to approve requests faster. Yeah. The approval flow's too slow.

  84. 14:56

    What kind of requests, though?

  85. 14:58

    Expense approvals first, access requests eventually. AI, hold that.

  86. 15:12

    Actually, let's pause. Expense approvals or a general approval workflow? Expense approvals, first release. Access requests are future scope.

  87. 15:25

    That changes the data model. Good to know.

  88. 15:31

    AI, pull that up for everyone.

  89. 15:33

    Tracking, determining who s- the speaker's referring to is critical. In this case, it was easy with a direct reference to the AI, but it'll happen again without a direct a- reference.

  90. 15:44

    Oh, I'd forgotten that was a rule.

  91. 15:46

    So over 5,000 needs a second approver? Yep, manager plus finance.

  92. 15:53

    So a big one can't be a single tap.

  93. 15:56

    Right. Over the limit, it routes to a second approver. Agreed. Under five, one tap's fine.

  94. 16:03

    Works for me.

  95. 16:09

    Okay. Agreed. Expense approvals, 5,000 threshold.

  96. 16:17

    Right there, the AI resolved the scope objective. No one wrote the prompt. No one packaged the turn and hit Submit. The system was in the conversation, following it, understanding, and choosing its moment.

  97. 16:31

    AI, capture that for us.

  98. 16:34

    First release supports expense approvals only. Access requests are out of scope. Managers can approve or reject an expense right from a notification.

  99. 16:43

    And per the finance controls policy, anything over $5,000 routes to a second approver.

  100. 16:59

    Actually, make the threshold 10,000, not 5,000.

  101. 17:03

    Want me to update the requirement to a 10,000 threshold?

  102. 17:06

    Yes.

  103. 17:09

    AI, is this room free after the meeting?

  104. 17:12

    Let me check. The room looks free after this, but-

  105. 17:16

    Until 3:00?

  106. 17:16

    Yes. It's yours until 3:00.

  107. 17:19

    And that's the difference between a smart machine behind the same old prompt and an interface that finally participates. Here's the mindset shift I wanna leave you with: AI is not just an intelligence technology.

  108. 17:33

    It's increasingly becoming an interface technology, and if so, then book-smart models alone are not enough. Stop picturing AI as a smarter machine hiding behind prompts, agents, loops, and a- all the old paradigms.

  109. 17:47

    We have to start seeing intelligence itself as a thing that can finally remove interface constraints and amplify human potential. For 75 years, humans adapted the, to the machine, its syntax, its forms, its timing, its batch.

  110. 18:03

    A system that can reason, listen, infer, adapt should be able to meet us partway, if not all the way. Instead, if AI is for users, then we should obsess about maximizing the interface.

  111. 18:16

    So then the design question changes. It needs to become, what burden are we still putting on humans only because the machine used to be too limited to carry that burden itself?

  112. 18:27

    Ask that question, and the whole interface space opens up. The answer isn't always chat. It isn't always voice and not a wall of markdown. It's definitely not a decade-old set of digital constructs.

  113. 18:41

    The right answer is the affordance humans already use with each other: communication, a question, a pause, a sketch, a checklist, a quiet aside, or saying nothing at all. An interface where timing and modality aren't the human's job anymore, where choosing the right channel at the right moment is done by the AI.

  114. 19:03

    And as a usability person, this is the part that excites me the most. When you take that burden off people, the friction disappears and adoption follows. Computing has mostly been about, to date, improving how humans encode their intent for machines.

  115. 19:19

    The punch card, type a command, click a menu, use your thumb on an iPhone, write a prompt. Every step was progress, and every step carried the old constraint forward into the next era.

  116. 19:30

    A translation tax, a precision tax, context tax, repair tax. AI is our chance to put those down, not by making everything magical, not by making everything voice, not by replacing human judgment, but by making computers, for once, more fluent with us.

  117. 19:49

    Most talks and videos cover how to use or adopt AI. The deeper question is how AI intelligence changes the interface. Human conversation is the most human thing we do because if a machine can finally understand more of what we mean, then we can and should stop reshaping ourselves to be understood by it.

  118. 20:09

    Thank you.