← All speakers

Bio, Work & Ideas

Jesse Hall

Conference affiliation: Livekit

On this page

Jesse Hall is the developer educator behind codeSTACKr. His work spans web-development education, conversational memory, and voice applications, explaining how software can remember a user, act on a request, and keep a conversation moving. He spoke as a staff developer advocate at LiveKit at the AI Engineer World’s Fair 2026.

From web-development education to conversational AI

A self-taught developer, Hall built codeSTACKr around tutorials, videos, and courses that make software development approachable. His earlier writing covers CSS interfaces and web-development roadmaps, helping newcomers navigate a crowded landscape without trying to learn every technology at once. He also created a VS Code course and maintains public tutorial repositories, giving learners code to explore alongside his explanations.

At MongoDB, Hall worked in senior and staff developer advocate roles. His teaching connected JavaScript and React applications to database design and, increasingly, generative AI. His 2024 conversational-memory tutorial used MongoDB and LangChain to extend retrieval-augmented generation beyond isolated questions: stored conversation history supplied context for subsequent exchanges. His React application workshop combined Next.js, LangChain, MongoDB Vector Search, and stored chat history to demonstrate that approach.

His later voice-agent teaching carries this application-building focus into a medium with different demands. A caller needs both a useful result and conversational cues that make waiting, interruptions, and changes of mind manageable.

Memory, handoffs, and useful actions

Hall’s guides explain several complementary parts of a working voice application:

  • Conversational memory: His LiveKit and Supabase guide connects a voice agent to a backend that stores user profiles, memories, knowledge, and session reports. It distinguishes shared knowledge retrieval from personal memory and business operations such as looking up an order. Hybrid vector and full-text search retrieves relevant memories, while authenticated user identity scopes the data. These mechanisms let an application carry useful context across sessions.
  • Context-preserving handoffs: His handoff guide routes callers to specialists through natural-language intent. Before transferring, the agent packages the caller’s request, relevant account information, conversation history, and results already retrieved. A specialist or human can then continue without making the caller explain everything again. Hall also argues that an explicit request to speak to a person should be honored immediately.
  • Conversation during slow reasoning: With Darryn Campbell, Hall co-authored a guide to the talker-reasoner pattern. A fast primary model handles the conversation while a slower model performs demanding analysis in the background; an asynchronous tool delivers the result when ready. Their guidance explains the tradeoff: a second model adds little when one fast model suffices, meaningless speech does not improve a delay, and the primary agent must avoid guessing at an unfinished answer.
  • Voice interfaces that act on screen: Hall’s healthcare-intake tutorial demonstrates an avatar guiding a user through a form while function tools update browser fields through remote procedure calls. Speech, avatar animation, and visible application state become parts of one interaction, with the agent performing a defined task.

Evaluating the conversation as a whole

Hall’s “Latency Is a Budget. Humanlike Is the Goal.” connects these application concerns to the experience of speaking with an agent. He argues that model leaderboards cannot determine the best voice system: developers ship a pipeline, and errors can compound across speech recognition, language generation, and speech synthesis. Evaluation therefore needs to measure the agent as a whole.

He also distinguishes measured latency from perceived latency. Appropriate conversational cues can help a caller understand that work is underway during a tool call, while turn detection determines whether the agent responds at the right moment or cuts the caller off. His hotel-receptionist demonstration brings those concerns together through interruptions and a caller changing his mind. The emphasis fits his broader teaching: memory, transferred context, useful actions, and responsive timing all contribute to whether a conversation works.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Jesse Hall explains why voice agents need conversation-level evaluation, audible latency measurements and asynchronous tools—and how to spend a response-time budget on behavior that feels human.

  • Evaluate the complete voice stack: a transcription error can drive a confident but incorrect downstream action.
    3:42 ↗
  • Measure the wait until audible speech. First-byte and first-token metrics can hide leading silence and the time needed to finish a speakable sentence.
    5:12 ↗
  • Separate slow tools from conversation progress so the agent can remain available while backend work continues.
    7:12 ↗
  • Audio-based turn detection and provisional transcripts can overlap pipeline work. Hall’s practical response targets are within 1.5 seconds, with around 600 milliseconds feeling more human.
    9:57 ↗
  • Conversation benchmarks should score correct outcomes, interruptions, tool behavior, privacy and unnecessary steps—not completion alone.
    12:27 ↗

References