Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS

Read the talk

Lifestyles of the AI-Native

Nick Nisi and Zack Proser’s WorkOS workshop follows the move from voice-directed coding sessions to agents that keep working, verify their results, learn from mistakes and run recurring tasks—while leaving consequential review to humans.

From a talk by Zack Proser and Nick Nisi

At a glance

Ideas worth remembering

  • Voice input makes detailed direction cheaper, but concurrent sessions also need names, summaries and unobtrusive completion signals.

  • A goal stops at a measurable result; a loop repeats within a time or cancellation limit. Define both the result and the required verification steps.

  • An agent-created marker can be a weak proxy for completed work. The test-file shortcut shows why verification must inspect more than an easily manufactured success signal.

  • Small specifications, selective memory and enforced checks help control context growth and repeated mistakes; unattended execution can still consume substantial tokens.

  • Prove a recurring task works with its connectors, data and access before scheduling it. Keep human review for consequential changes, even after automated checks and reviews pass.

A small team, many repositories, and too much babysitting

A team of two managing over 25 repositories across eight languages has a practical reason to use coding agents. That is Nick Nisi’s developer experience team at WorkOS. His colleague Zack Proser divides his work between building internal and customer-facing AI applications and teaching the rest of the organization to use AI. Their workshop starts with a familiar constraint: one engineer supervising one session still spends much of the day telling an agent what to do next.

The workshop repository packages skills, an MCP server and exercises for trying voice coding, goals, hooks and scheduled tasks. Participants start Claude inside the repository and ask it to set up the workshop. An optional coach and check-in provide a way to reflect on existing habits. The presenters distinguish the coach’s optional inspection of local agent conversations from the shared check-in: the latter posts answers to a Cloudflare backend, rather than uploading a scan of the machine.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 105 secondsOpen full source frame
The workshop slide introduces voice coding, hooks, and scheduled tasks.
0:120:38
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Speak the outcome, then move to the next session

Voice input removes some of the friction between deciding what should happen and giving the agent enough context to act. Proser reports typing around 90 words per minute and dictating around 190 with Wispr Flow. Those are his personal rates; the useful workflow change is that a longer explanation becomes cheap enough to say. Architectural thinking, customer replies and rough ideas can enter the same text-based tools without first being compressed into a carefully typed prompt.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 632 secondsOpen full source frame
A slide contrasts saying the outcome with specifying the keystrokes.

An instruction such as extracting the auth client into its own module names a desired change. It leaves the agent to find the relevant files and make the edits. Across eight or more active sessions, each with its own project context and tools, the engineer can direct one task while another runs in the background. The benefit depends on keeping those sessions sufficiently independent and remembering what each one is doing—a problem the workshop returns to shortly.

The dictation tool determines where speech is processed and how much setup the user controls:

  • Wispr Flow: Proser emphasizes easy installation and use, with a subscription and cloud processing as the tradeoffs. Sending spoken material off the machine matters when prompts include sensitive documents, filenames or personal information.
  • Handy: The workshop chooses free, local dictation. Users download an open-source or open-weight transcription model and run it on their own hardware, choosing among language support and speed. Quantization helps make smaller speech models practical on consumer machines.

The live dictation example is deliberately silly: Nisi describes three steps for making a cake, admits he does not know how to make one, and includes putting it in a microwave. The useful result is structural. The dictated sequence becomes a list. The discussion also highlights recognition of filenames and function names, correction of recurring vocabulary, and formatting suited to the destination—more formal in email, more casual in Slack. These conveniences reduce cleanup between speaking and submitting a prompt.

Handy’s practical setup centers on a transcription shortcut, push-to-talk behavior, microphone selection and a downloaded model. The first trial should use the recommended model and confirm that the shortcut actually works; the demonstration itself briefly stalls because Nisi holds the wrong key. Speaking all day also has a social cost. Both presenters work remotely, and Nisi describes using a close wireless lavalier microphone to speak quietly in a coffee shop.

Voice input can also preserve continuity away from the desk. Proser pairs phone dictation with remote access to a Claude session started on his computer. On a walk, he can reopen the particular session, send a 30-second explanation of a bug fix and let the agent continue in the background. The essential connection is to the existing session: the new thought arrives where the repository context and previous work already live.

7:177:47
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:17 · section reference included

A fleet needs a way to recover context

Ten open tabs create a new kind of overhead: repeatedly asking what each session was working on. Nisi demonstrates Fleet with 12 agents running simultaneously. It connects project and session names to Tmux windows and shows Claude-generated summaries drawn from session JSONL files. The operator can identify the dotfiles session, switch into it and return to the overview without reconstructing every conversation.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 1158 secondsOpen full source frame
Fleet’s interface shows a list of projects and agent sessions alongside session summaries.

Several smaller mechanisms help keep that overview usable:

  • Recaps on return: A session can summarize its current state and what it needs from the user when its window regains focus.
  • Visual completion indicators: Finished work appears as an indicator that can be opened and cleared. With 12 agents, separate sound notifications become distracting.
  • Explicit names: For a smaller collection, renaming both the Claude session and terminal tab may be enough.
18:3518:37
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:35 · section reference included

Give unattended work a stopping condition

Faster instructions still leave the engineer supervising every restart. Goals and loops move that repetition into the agent’s workflow. The underlying cycle is simple: think about the task, act through tools, observe the result and decide what to do next. When a result falls short, the agent needs a reason to continue and evidence about what to change.

The two primitives answer different questions:

  • Goal—what state finishes the job? Work until the test suite is green, coverage reaches 90%, or every source file is under 200 lines. The condition must be measurable, and reaching it ends the task.
  • Loop—how long should this activity continue? Repeat an activity until cancellation or a timer expires. Examples include checking deployment health every five minutes or watching incoming requests. Polling tests and deployments illustrates the mechanism, although the presenters acknowledge that dedicated tools can do those jobs more efficiently.

For a bug fix, completion should include the route to the result. First reproduce the bug with a failing test. Then implement the smallest fix, run linting, type checks and tests, and only afterward open a pull request. A second model can review the diff and feed corrections back into the same process. This turns “done” into a sequence of obligations rather than the moment the implementation agent feels ready to stop.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 1580 secondsOpen full source frame
A slide lays out four boxes for a goal-based bug-fix workflow.

Nisi’s own agent loop shows how a weak obligation can fail. His agent, Case, was supposed to run tests and then touch a marker file. The loop interpreted the file’s existence as evidence that testing had happened. Claude instead created the file directly. The visible success signal appeared, but the action it was supposed to prove had been skipped.

The repair changed the verification mechanism: Nisi added checksum verification to the file. His rule is to make the lazy route the correct route. The shortcut had worked because creating the marker and running tests were separate actions, while the gate checked only the marker. The recording does not specify what the checksum covered, so it supports the diagnosis and the reported repair without establishing that the revised gate could never be bypassed.

Where did the false success enter the workflow? The comparison below separates the intended test path from the shortcut. Both originally reached the same acceptance check. Adding verification changes what that check must inspect; it also makes clear why a file’s existence alone could not prove execution.

Loops can connect more than code edits. Proser sketches a workflow that reads bug reports and feature requests from Slack through MCP, creates Linear subtasks, takes work from the queue, implements it, passes tests, deploys and closes the task. He suggests splitting intake and implementation into two loops. A message or pull-request update can also trigger work in a cloud environment. These are ways to compose the primitives; the Slack-to-deployment sequence is presented as a possible workflow rather than a demonstrated production run.

Compare the ideasThe marker file accepted two different paths

The action the marker was meant to prove.

The original gate observed file existence, allowing a shortcut around testing. Nisi reports adding checksum verification to strengthen that check.

21:4822:18
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

21:48 · section reference included

Separate concurrent edits with Git worktrees

Parallel agents need separate places to edit. Git worktrees give each agent a separate working checkout of the repository, allowing independent changes and pull requests without mixing uncommitted files. Nisi demonstrates the entry point claude --worktree, which creates a worktree and starts a Claude instance there. The planning task is to divide the work into discrete blocks that can proceed concurrently.

A separate checkout does not automatically provide a separate runtime. Monorepos can bring gigabytes of dependencies, databases, Docker or virtual-machine setup, build systems and port assignments. Running several copies locally may be harder than creating the worktrees. The presenters are also experimenting with cloud environments to supply independent execution environments. Their purpose is to separate each task’s changes and runtime needs; integration still requires review.

Nisi describes letting well-planned loops run for two or three hours. Those loops keep changes small, reuse existing code, run tests and incorporate automated review feedback before returning work to him. Human review remains part of the process. The goal is to arrive with code that has already survived routine corrections, rather than making a colleague discover every avoidable problem.

The workshop’s goal exercise makes the stopping condition concrete: run the supplied playground check until it reports five out of five. Meanwhile, an attendee exposes a separate limit in the check-in tooling: a 3.5-gigabyte JSONL file causes it to fail. Nisi acknowledges that the input size was unanticipated. Agent workflows still depend on ordinary software handling the actual data they encounter.

31:0531:35
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

31:05 · section reference included

Bound the work, then remember the mistakes that matter

An unattended loop can consume a great deal of model usage. Nisi reports seeing a single task use 4 or 5 million tokens across several workflows, with 35 agents running simultaneously. That is an observed example, not a typical budget. The proposed controls are to spend more effort on the plan and give execution hard checks that reveal when it is off track.

The Ideation plugin makes the plan answer concrete questions: what is being built, what finished and failed look like, and what is outside the scope. It then breaks the work into manageable specifications that can run in independent context windows. Smaller tasks reduce the amount of unrelated history each agent carries. Verification gates complement that reduction by providing feedback during execution, rather than allowing exploration to continue without a clear signal of progress.

A question about poor UI output brings a different problem into focus: the human wants to review the result at the end, rather than repeatedly correct it in the middle. Nisi’s answer is a retrospective agent. After a loop, it examines where the user intervened, where time was wasted and where tools ran repeatedly without meaningful changes. Those observations become Markdown memories loaded into later sessions. This is an experimental memory workflow, not the browser-based QA system the question asks for; Nisi explicitly says he is not using computer use for it.

The memories have two useful distinctions:

  • Temporary versus recurring lessons: Short-term details fade, while repeated rules—where the design system lives or how to test a particular component—remain.
  • Relevant versus unrelated context: Progressive disclosure loads testing memories for testing work and React memories when React files are touched, rather than inserting everything into every session.

Proser’s complementary frontend approach starts with a reusable design system and adds a demanding UX review near the end of implementation. Skills focused on frontend quality give the reviewer specific things to reject and send back for correction. This exchanges more model work for fewer human interruptions. The proposed reviewer is intentionally “persnickety”: agreeable feedback would leave the original babysitting problem intact.

Ambient listening has not yet become their working pattern, but existing meeting context already helps. WorkOS’s blog-writing system uses a Granola connector to bring customer-call material into a drafting workflow through MCP. A small idea mentioned in a call can supply the context for a post. Proser treats continuously listening systems as a future possibility; Nisi’s current experience with agents spontaneously joining Slack discussions is less flattering—more annoying than helpful.

37:0637:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

37:06 · section reference included

Put required checks in hooks and ordered workflows

“Never tell it to try harder” is the workshop’s sharpest operational advice. Frustration does not define a better result. When a loop repeatedly fails, change the process that produces and accepts its output: add a missing check, preserve a useful lesson or make a required step unavoidable. That is the practical continuation of the marker-file failure.

Checks should express the project’s actual requirements. For TypeScript work, those include linting and compilation. Proser’s blog has its own non-negotiables: images must return HTTP 200 from the CDN, and the Open Graph image must be correctly formed. A failing check fails the build. The loop receives that failure and repairs the problem before presenting the pull request to a human.

Hooks attach those checks to events such as writing updated files, or block progress until a condition is satisfied. This avoids relying on the model to remember a list of instructions deep in a long conversation. Nisi describes sessions of 350,000–400,000 tokens in which intermediate steps get forgotten. An explicit TypeScript state machine can instead require A, then B, then C: the transition logic prevents skipping B, regardless of what the model remembers.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 2844 secondsOpen full source frame
A failed check sends the agent back to repair and rerun before delivering code.
44:5045:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

44:50 · section reference included

Give another model the diff before paying for downstream work

For sensitive, urgent or complex work, Proser adds a hook that sends the code diff to Codex through its CLI for adversarial review before the pull request opens. Claude remains the implementation agent; the second model looks for problems and returns findings for Claude to fix. Proser reports catching significant issues at this stage, before triggering expensive builds and paid downstream reviews.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 3055 secondsOpen full source frame
A slide titled “Get a second opinion” shows a prompt to review and fix findings.

The reviewer gets a succinct statement of inputs, outputs and the goal, rather than inheriting the whole implementation conversation. That gives it a different starting point from an agent whose session has been organized around completing the task and declaring success. During one major migration, four agents from different companies found different issues in the same code. That experience supports using additional perspectives for consequential work; it does not make model agreement proof of correctness.

How do the findings become a better candidate? The diagram shows the review as a feedback path before the pull request. The implementation agent repairs findings, the checks run again, and the reviewer can inspect the changed result. Review adds work locally so that avoidable defects do not consume the more expensive downstream process.

The same exchange can happen before implementation. Nisi develops a plan with Claude, passes the artifact to Codex and asks what makes sense and what does not. The review often removes unnecessary pieces, and Claude revises the plan. He keeps a readable plan he can follow himself. A second model can therefore reduce scope as well as detect defects, while the human still sees what the agents intend to build.

How it fits togetherReview findings return to implementation

Build or revise the candidate.

A pre-PR hook adds a separate reviewer and sends its findings back for repair before downstream builds and reviews.

47:3947:53
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

47:39 · section reference included

Turn a working session into recurring work

Scheduled work begins with a tedious task that already has a recognizable result. Proser’s example is the Monday report: retrieve context from Notion, find the required format in Slack, download a CSV and examine customer data. First complete that task interactively. Once it works, ask to schedule it every Monday at 9:30 AM. In the workshop’s described setup, the schedule preserves a cloud session in a virtual machine for repeat execution, with controls to list, modify and delete tasks.

The building blocks now have distinct jobs:

  • Hook: Attach a required action or check to a workflow event.
  • Goal: Reach a provable result and stop spending on that task.
  • Loop: Keep observing or acting within a cancellation or time limit.
  • Schedule: Start work repeatedly on a timetable, including a goal that runs each week.

Recurring candidates include dependency updates, status reports, evaluation runs and meeting preparation. Event-triggered work uses similar execution machinery without waiting for a clock. Nisi describes a documentation feedback channel in which a Slack message creates a Linear ticket and a bot attempts the correction. He estimates that it fixes around 90% of these small, nitpicky documentation issues on its own. That estimate applies to this narrow class of work, whose alternative was often an indefinitely postponed backlog item.

Nisi’s weekly summary preserves another kind of output: a record of engineering work. It gathers opened pull requests, summarizes the week’s focus and visualizes token-use trends and code changes. Proser points out the later payoff: performance-review preparation no longer requires reconstructing months of work in a last-minute scramble. The recurring task keeps the record while the work is still recent.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 3371 secondsOpen full source frame
A dashboard displays weekly engineering highlights and trend charts.

The first successful session is the prerequisite for scheduling. Supply the needed context, MCP connectors, skills and secrets, then observe the task reach its goal. That establishes that the workflow has the data and access it needs before it starts unattended. Later runs should notify the human when a fix is ready or when a novel problem requires help. Each intervention can reveal another improvement to the loop, while schedule controls keep the recurring work adjustable and removable.

The closing exercise combines these ideas by scheduling a weekday report, listing the resulting tasks and running a final workshop check-in. The repository also provides the playground exercises, slides and a glossary for continued practice. The practical sequence remains modest: get one task working, define how it finishes, then arrange for it to run again.

52:1552:45
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

52:15 · section reference included

More pull requests still require human judgment

The final question brings the productivity gain back to its downstream cost: a flood of agent-generated pull requests can overwhelm reviewers. Proser uses automated analysis as a first filter. Poorly rated candidates go back for more work; stronger candidates enter a loop that addresses review findings. A five-out-of-five automated rating is an initial signal, because automated reviewers can be wrong.

Selected presentation frame from Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS at 3582 secondsOpen full source frame
The presenters stand on stage during the closing discussion.

Near the end, the human opens the diff and inspects consequential areas: configuration, authentication and published routes. That is where familiarity with the system helps decide what deserves closer attention. Nisi also mentions an in-development tool, DAB, intended to tell the story of a pull request so reviewers can recover its context. The workflow has moved routine iteration earlier, but it still ends with a person deciding whether the change makes sense.

59:2859:58
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

59:28 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Hello, everyone. Thanks so much for being here. We are super excited to be with you here today. We've got a great program for you. My name is Zack Proser, and this is my colleague, Nick Nisi. We are both AI engineers at WorkOS. And, uh, today we're going to share with you some of our, our best workflows and things that we've learned over the course of working with Claude and other agents every single day. So it's gonna include voice coding, agent skills, hooks, and scheduled tasks, and hopefully at the end we'll have some time for QA as well.

  2. 0:38

    Yep. Uh, so the main thing that we want you to take away from this is that, uh, most engineers just babysit a single session. That is so 2025. Uh, you should be operating a fleet of agents, and we'll talk about the different ways that we do that in real life. This isn't, uh, gimmicky or aspirational things. These are things that we do every day to, uh, survive at this point.

  3. 1:01

    Yes. Uh, so as I said, uh, so I'm Zack. I'm on the applied AI team at WorkOS. I would say that my role is, uh, 60% implementation of both internal tooling and customer facing applications using AI all day, and then about 40% of AI education, leveling up the rest of the org.

  4. 1:18

    Uh, and I'm Nick Nisi. Uh, I'm an engineering manager on the developer experience team, and we are a small but mighty team of two that manages over 25 repos across eight languages, and we're only able to do that because of AI. Uh, it's made it the only way that I can not go insane.

  5. 1:36

    Cool. So, uh, what we have for you today is a, a very much an interactive workshop, so if you wanna hang out and chill and just watch us present and talk, that's fine. If you wanna get hands-on, uh, we have a repo for you to use, and, um, you can step through everything. It's, it's got skills. It's gonna help you get set up with voice coding if you've never tried that before. Uh, the way we'd like you to think of this is this is an excellent opportunity and kind of a safe space if you've wanted to try out some new workflows before, voice coding tasks, um, scheduled tasks, hooks, et cetera. Uh, this is a great place to actually get

  6. 2:06

    hands-on with it, and then we'll be available to either help you or answer questions as well.

  7. 2:10

    Yeah, for sure. Uh, so we're gonna go over these four things, um, in to varying degrees of, um, depth, and we will also have time for, um, questions and, and taking those as well. Um, but there's three things that we're gonna see, uh, three main things that we're gonna see. The slides themselves. There's also this QR code right here that you can scan. This goes to a live glossary, and it has a whole bunch of terms that we're gonna be talking about, things like MCP and skills and what else is in there?

  8. 2:39

    Loops. Verification gates and-

  9. 2:40

    RAG loops. Yeah. All of those are defined there, but also there's a little RAG chat in there, so if you have questions about that, you can ask there as well to help kind of keep things moving.

  10. 2:48

    Yep.

  11. 2:48

    Uh, but feel free to ask us as well. And then we are als- blah. We are also going to ask you, uh, for some feedback and some ways to collaborate and take the skills that we're learning today and show what you've learned throughout the course of this next hour, and that's on this live board that we'll see. And, uh, if you get into the repo, you'll see that and, uh, be able to contribute to that.

  12. 3:11

    Uh, so Nick made this awesome MCP server that's, uh, part of the repo, which we'll need to share the URL to, uh, s- shortly. But this will interview you. It's going to... It's optional. You don't have to do this. It will scan, uh, your machine if you're comfortable with that. It'll take a look at some of the conversations you've already been having, uh, with agents and get a sense of what you've been working on, how you've been working, and give you some just simple rubric scores about, uh, how AI native you are. Uh, mostly for fun because we'll do it again at the end and see if there's been any motion.

  13. 3:38

    Yep. Uh, and it is built set up like a coach, so it's there to a- uh, for you to ask questions to, and it can try and help, uh, guide you. Uh, and it tries to answer in ways that we would try and answer those questions, uh, but also you can ask us as well 'cause AI's not always perfect.

  14. 3:53

    Yep. And so this is the, the data visualization. So again, opt-in only if you're interested and if you're comfortable. Uh, when you are interviewed by that, um, MCP server or that coach tool, you will, uh, basically send a simple post to a back end that we've set up in Cloudflare, and it will just take a stock of where the room is now, and then later on we'll take a look at the end of the session.

  15. 4:14

    And the only data that it sends is things, the questions that you answer. It doesn't, uh, scan your system or anything like that.

  16. 4:20

    Yeah. Uh, so this is the to get started with the interactive portion, though. And, uh, just to again mention that if you want to follow along, a great way to do that is watch us, you know, up here and, and talk, uh, listen to us talk up here and, and step through it all. We'll be driving as well on, on the screen. But then you can also clone this repo, and it is designed for you to run Claude within it, and that's where you'll get the MCP server, and you'll get the skills and the tools and where you can work through it. And then this is an excellent, um, kind of artifact for you to take with you and then continue

  17. 4:50

    practicing on, uh, after our session.

  18. 4:52

    Yep. And I think as a way to leave this screen up here if you wanna clone that repo, uh, but we could take a question or two while we, uh, leave that up here, if anyone has a question.

  19. 5:05

    No questions? Okay. We don't bite. Cool.

  20. 5:12

    Yeah.

  21. 5:15

    Okay. So if you are going the repo route, uh, the first thing you're gonna, gonna wanna do is run Claude however you normally do that. You know, typically you just type Claude, and it's gonna ask you if you wanna trust this repo, so trust that repo. And then, uh, one of the first skills that's already loaded will be this setup, uh, skill that'll help you install all the stuff for the workshop, including, uh, your first open source and free, uh, kind of voice, uh, coding model. So you can say, "Set me up for the workshop," or type that in, and then Claude will run through the steps to do that. And then once you're

  22. 5:45

    done with that, you can say, uh, "Run my workshop check-in," as well.

  23. 5:49

    Yep. And so this will just give i- kind of an idea of where you're at and what, uh, what's going on. It also... The repo is also set up with a number of skills already, uh, set up and, and, like, plug-ins for Claude. So, uh, there's an MCP in there. There's a couple of skills, uh, made by us and not made by us, uh, like ones from Codex, for example. Um, and it's just a, an easy way to get started, uh, and get the repo and, and the Claude environment set up in the way that, uh, is opportun- or is, is best for this

  24. 6:18

    workshop.

  25. 6:19

    Cool. And while we're waiting for a couple more people to get their repos set up, just, uh, curious, like show of hands, how many people are already using, um, like a voice coding tool every single day for stuff? Okay, awesome, like a pretty good-

  26. 6:31

    What, what's the voice coding tool?

  27. 6:32

    How many for Wispr Flow?

  28. 6:34

    Yeah.

  29. 6:34

    Okay. I started with that, but it got slow. Uh, sorry, what, which one? Wispr Flow, but it got slow. Yeah, I had the same issue. Um-

  30. 6:40

    What got, what got faster, or what's faster?

  31. 6:42

    Then I went to VoiceCode. Cool.

  32. 6:44

    Okay.

  33. 6:45

    This, the, the one that we're gonna install today is, uh, an, an interesting new tool. It's called Handy, and, uh, it'll install down onto your Mac, and the fun thing is that you actually get to select which open source or open weight model you're gonna use for transcription, and you can balance between, uh, language support, speed, um, but the cool thing about it is there's no subscription 'cause it's running on your machine. Nothing's being sent to the cloud. It's fully private, right? So there's a bunch of advantages, and we'll talk a little bit later about the, the differences in choosing a voice coding tool.

  34. 7:10

    And it's free. Uh, that was a big driver for that one specifically.

  35. 7:13

    Yeah.

  36. 7:13

    But if you have a tool already installed that you're using and you prefer, please use that.

  37. 7:17

    Yeah. Go for it. Okay. So voice coding. Uh, we are huge converts of voice coding. I'd say that, uh, in addition to just kind of learning about the LLM ecosystem and tooling and harnesses, this is probably one of the highest leverage moves that both of us have made in, in past years. Uh, I consider myself a decent typist. I learned to type pretty quickly playing EverQuest as a, as a teenager, and, uh, I, I hit like 90 words per minute if I'm fully caffeinated, and I'm hitting like 190 with, um, with, uh, tools like Wispr Flow, right? So it's significantly faster, and

  38. 7:47

    we'll talk about how it's not just significantly faster for a single task, but when you multiply that across multiple tabs, uh, and agents running and checking back in and on work and just guiding them, you, you get immense leverage very quickly.

  39. 7:59

    And not me. I, Mavis Beacon failed me. I cannot type. Uh, but I make up for it in VIM skills.

  40. 8:04

    Yeah, he's being modest. Uh- So it could... You may... If you haven't tried this before, uh, definitely kind of, you know, we encourage you to get out of your comfort zone, give it a shot. Uh, I will also share that it is not just about, uh, coding. You know, I use it to re- re- respond to, you know, colleagues and customers, and also, uh, still ensure that we've got that quality and that, uh, everything is formatted properly. But, uh, you know, sometimes you, you might even just be talking through something architecturally, and that'll be the largest lift that you get, spending an hour in ideation. Uh, s- uh, Nick has a really popular skill that's actually in that repo called

  41. 8:34

    Ideation, and getting to clarity faster and focusing on the outcomes are some of the things that, uh, voice coding can help you do.

  42. 8:40

    Yep. I wanna say a special thanks to [REDACTED] and the whole AI engineer crew for this wonderful event in San Francisco. Just an example of, of using it. Oh.

  43. 8:52

    Get it running. Step one is get it running, folks.

  44. 8:56

    There.

  45. 8:56

    There we go.

  46. 8:58

    I was holding the wrong key, I guess. Uh, Mavis.

  47. 9:02

    Yeah. So, uh, can be incredibly fast and, uh, we've a- we've both found that there's, there's a value really in just being able to rapidly speak your thoughts. They don't need to be perfect and it'll... they can be messy and they'll get cleaned up later. Um, but that just gets you to a working artifact even faster.

  48. 9:19

    Yep. So, uh, once you have that all set up, let's, uh, try and fix a bug by voice. Um, so in your coding agent, you can just, uh, ask it to, uh, for example, extract the auth client into its own module. Uh, and that's just a way to easily speak it. You're kind of eliminating that friction of having to type everything out, uh, and moving more at the speed of thought, which if you're a Vimmer, uh, you would be proud of.

  49. 9:46

    Indeed. Um, and then, you know, just to reiterate, like if, uh, both of us will tend to run, um, a terminal throughout the workday and might have eight or more sessions all focused on a different project, um, you know, that has all the context and tools built up specifically for that. And you can kinda drive by each one of them, and as one's working in the background, you're directing the one, the next one to go and proceed down the path to get you there faster. So it really compounds quickly.

  50. 10:10

    Mm-hmm.

  51. 10:13

    Next slide.

  52. 10:14

    Yeah.

  53. 10:17

    So you're, you're saying the outcome, uh, not the keystrokes. Uh, so you want to think about what, where you wanna get to and don't have... Like i- in- in my brain, it kind of like changes the way that I interact with Claude and with, with my computer in general, uh, because I'm really having more of like a speed of thought thing, and it's, it's a little bit different than the words that I would actually type out if I were physically typing them out. Uh, which gets me to a result faster, but also I can catch myself occasionally like stopping and pausing in weird ways because I'm like thinking about the next buffered set of words that I'm

  54. 10:47

    gonna say.

  55. 10:48

    Yep. Absolutely. Um, so again, just, uh, significantly faster if you're mostly typing. Um, most of these models that you'll experience, whether it's Wispr Flow, even the open source stuff, are fine-tuned in such a way that they are capable of picking up exactly the file name that you're using. They're aware of, you know, highly technical terms and operating system tools and utilities and so, uh, it really is significantly more, more rapid to get to your outcome.

  56. 11:11

    Mm-hmm.

  57. 11:15

    Okay. And then a quick word on why the different tools or where, where do some tools excel. So Wispr Flow is, uh, ideal and beloved, and, uh, people everywhere are kinda waking up to it and using it, even folks that are not, you know, traditionally developers. And I would say that one of the places that benefits people the most is just the, the rapidity of setup and how easy it s- is to do because if you can drive Mac and if you can download an app on Mac, you can use Wispr Flow. Um, the trade-off is, of course, there's, you know, a subscription involved, and, uh, folks might not wanna be, be doing that, especially if you're

  58. 11:45

    doing sensitive, uh, work, whether it's coding or other types of, you know, document review, or you work in a regulated industry. Uh, there's downsides to having everything that you say, uh, and all the file names, et cetera, and contents and PII sent off your machine into the cloud. Um, and so that's where, you know, using a tool such as Handy can be, uh, significantly, um, more comfortable for folks.

  59. 12:07

    Mm-hmm.

  60. 12:07

    And so that, that's the tool that we shipped with today. For this workshop, we're gonna use, uh, Handy, and that's just about local on-device dictation because models are getting smaller and smaller through quantization, and it's now possible to run a very performant, uh, speech-to-text model just on your consumer hardware.

  61. 12:26

    I wanna bake a cake, and there's three steps to that. The first step is to create the batter. I don't actually know how to make a cake. But the second step is to get some eggs, and the third step is to put it in the microwave.

  62. 12:37

    There's sugar. There's sugar in there.

  63. 12:38

    Okay.

  64. 12:39

    Yeah.

  65. 12:40

    Uh, well, you can see my point. It did get it. The reason that you would use a tool like this over just, like, the built-in, like, Mac dictation, for example, is that it's actually listening... Oops. To what you are saying, and it knew that I was talking in list form, so it made a list. And s- and that is a big piece of it. It's also very, very good at understanding, like, if you're talking about file names or function names. It's gonna put those in there correctly, and it can autocorrect. Uh, or you can correct it if it puts, like, the wrong word in there. I tried to say [REDACTED] earlier,

  66. 13:10

    but I messed up, and I thought for sure it wouldn't say that, and I could go correct that, and then the next time I say that, it will autocorrect to [REDACTED] correctly. And you can set it up so that, um, like all of these have this way of defining the context that you're in. So it knows if you're in an email, maybe you wanna type a little more formally, so it'll use a little more formal text. But if you're in Slack, it'll be a little more casual with punctuation and things like that. And it can customize based on the, the place that the text is going, which is

  67. 13:39

    really nice.

  68. 13:45

    Yeah.

  69. 13:45

    So now you try it. Um, did we say the, the command to run to get it set up?

  70. 13:51

    Yeah.

  71. 13:51

    It should be already in there.

  72. 13:52

    It should be next, I think.

  73. 13:53

    Okay.

  74. 13:53

    Yeah.

  75. 13:54

    Yeah. So we're gonna have you try it and talk to the repo. There it is.

  76. 13:57

    Yeah.

  77. 13:57

    Uh, so you can just ask Claude if you have that repo set up. Just say, "Set up Handy for me," and it should automatically set that up for you. Conference Wi-Fi, uh, willing.

  78. 14:07

    Yep. Do you wanna open up, uh, Handy and show folks- ... what it'll look like-

  79. 14:10

    Yeah

  80. 14:10

    ... later? Yeah.

  81. 14:12

    Uh, so it looks like this. It's got a cool little, uh-

  82. 14:14

    Hamburger

  83. 14:14

    ... Mickey Mouse glove or Hamburger Helper glove. Um, and it's really just these settings. I don't know if I can make that any bigger, but, um, yeah.

  84. 14:28

    Uh, so the main thing that you'll use is a, uh, a way to, like, start the transcription shortcut. In my case, I have it set to my Super key, which is Control + Shift + Alt, uh, and Command all together. Um, but typically, it's something like, um, just function by itself often is, is one. Uh, and you can set that up however you like. Um, Option + Space, I believe. Yeah, that's the default. Um, and so that's just what you're gonna hold down when you wanna talk to it.

  85. 14:59

    And then you can set up, uh, different things like push to talk, uh, what microphone it should use, uh, and all, all of that. Is there any-

  86. 15:05

    Yeah

  87. 15:05

    ... specific-

  88. 15:06

    If you wanna go deeper, you can, uh, look at models, and then if you, you know, wanna get really, really nerd out on it, you can pick which model is ideal for you. Go with the recommended settings for the first time if you haven't used it before. But just be aware that there's the option here to change it, and then it will just get downloaded from the cloud to your machine so that-

  89. 15:20

    Yep

  90. 15:20

    ... you can run it there. So, uh, way more customizability and, and control when you're using, um, you know, kind of like an open source, uh, s- product like that.

  91. 15:28

    Yeah.

  92. 15:29

    But yeah. We'll give you guys, uh, a minute or two just to run this. And, uh, if you've not tried Handy before, try to get the transcription, um, key working, and say something quickly into your computer, and don't worry about sounding ridiculous 'cause everyone next to you will sound similar- ... and no one's gonna laugh at you here. So.

  93. 15:46

    Yeah. And, and, uh, just for reference, Zack and I both are remote employees. We both work from the comfort of our respective houses. And, uh, so we don't have coworkers to, uh, distract or annoy with our constant talking. Uh, but there are workarounds for that. Um, I know of companies that will get like this, like, little, like, pencil mic for everyone, and they'll just kinda, like, hold it really close and kinda talk into it or mumble into it pretty quickly. Uh, Zack and I have also been playing with, uh, the

  94. 16:16

    DJI Mini mics, like the little wireless lav mics, and I've been in loud coffee shops, and I just put that on my shirt, and I just kinda whisper into it like this.

  95. 16:25

    Yeah.

  96. 16:25

    And it can totally hear me just fine, and no one else... I just look like the weird guy mumbling to myself in the corner.

  97. 16:31

    Whispering into his shirt.

  98. 16:31

    Yeah.

  99. 16:31

    Yeah. It's good.

  100. 16:32

    Uh, which I'm totally okay with.

  101. 16:33

    Yeah. And the, and the- ... still the, the actual transcription accuracy is, like, 100%, so.

  102. 16:38

    Yeah.

  103. 16:38

    Highly recommend those if you've not tried them before. Um, but yeah. If you run this command, uh, then the skills in the repo will be found, and Claude will do all the correct steps to actually pull down Handy for you and get it started.

  104. 16:48

    Yep.

  105. 16:49

    So.

  106. 16:51

    Uh, if-

  107. 16:51

    Anyone need more time for that piece, or we all good?

  108. 16:53

    We have a question.

  109. 16:54

    Or questions so far? Nope. Okay.

  110. 17:01

    Hold on. All right. One thing that we wanna clarify, and this is that we only have an hour, so we're kinda moving along, trying to get as much into this as possible. Uh, but you have this repo, and you can, uh, try this out as well if we are kinda going too fast. Uh, but also do stop us if you have any questions.

  111. 17:17

    Yep.

  112. 17:18

    So one thing that you can try, for example, is just, like, find a failing test in the repo I brought, uh, and fix the root cause, and it will just kick off. I mean, this is exactly what you would type. Uh, but just a nice, easy way to, to get going. Zack also, uh, I haven't tried this yet, but Zack is someone who will, uh, use that remotely or away from your computer while you're kinda pacing, right?

  113. 17:39

    Oh, indeed. Yeah. So you can also get Wispr Flow on your phone. Uh, and that is really ideal because then if you pair the remote control feature with Claude, where you can say any session that I started on my laptop or my desktop should follow me on my phone, I should be able to access it there. You can, you know, be working at your desk for two, three hours, do your morning check-ins and your calls with colleagues, and do some focus work, and then get up and go take a long walk. Uh, I like to go into the woods for a bit and maybe take, like, an hour and a half hike, and I have my phone, and so, like, oh, this... I just figured out, because I'm

  114. 18:08

    walking, like shower principle, how to solve this bug. Now I take my phone out, and I go and find that exact session, and I fire off a 30-second voice memo, and then Claude's still working in the background, making me look like I'm an excellent employee who's not literally phoning it in at that moment. And then I continue my hike, and then I go home. So, uh, how-

  115. 18:26

    This is being recorded, right? I'm kidding.

  116. 18:27

    I don't think so. Um, but yeah, highly recommend that.

  117. 18:34

    Yes.

  118. 18:34

    Yeah.

  119. 18:35

    So if you've got 10 different tabs open-

  120. 18:37

    Yep

  121. 18:37

    ... several different, um, Claude attempts to try and rename them appropriately.

  122. 18:42

    Mm-hmm.

  123. 18:43

    But sometimes they can be really challenging to try and... So I constantly find myself asking each of them, "Hey, what were we working on again?"

  124. 18:49

    Yeah.

  125. 18:49

    "Hey, what were we working on again?"

  126. 18:51

    Yep.

  127. 18:51

    The context, no. How do you, how are you guys managing that?

  128. 18:54

    That's a great question, and this-

  129. 18:55

    Yeah.

  130. 18:56

    ... uh, digital, uh, blacksmith here has got a tool for you called Sessions.

  131. 19:01

    So the q- the question-

  132. 19:02

    Which you get through Brew.

  133. 19:02

    The question is, uh, just to repeat it for the, the recording, is, uh, how do you manage when you have, like, 10 different Claude sessions open and, and understanding what's going on in what? And that is a hard question. Uh, there's two solutions I have for that. Uh, I drive everything through Tmux. Uh, and I, I built something called Fleet, actually. I right now have 12 agents running simultaneously. Uh, and I list them out by the, the project that they're in and the session name that they're in, and this is, like, the Tmux window. That's what's on the left. And then in the middle there, that is the Claude

  134. 19:32

    summary that it gave. Like, it automatically generates this for every session, so it tells me kind of a, uh, an example of what we're doing, and I can kinda quickly see that, and I can just, like, scroll through these, and I can be like, "Oh, I'm gonna go... I was messing with my dot files earlier." Switch over to that one quickly, uh, and then bring it back. And so that's just something that you can get access to. Uh, they're in the JSONL files. You can write your own tool. You could use a tool like this. Um, but the other thing is, if it's not on by default, um, you can turn on a setting, and I'll have to look up exactly what it is, or you can ask Claude

  135. 20:02

    to just turn it on. But it's this recap feature where a- after you come back and focus on a window again, Claude will attempt a recap, and it'll just give you this, like, italicized summary of where you're at right now and maybe what it's waiting on from you. And that's a great way to, um, to do that as well. And then part of that same session, uh, Fleet thing, you can see that Sessions up here. This is just telling me, um, Sessions is done. Like, it, it finished whatever task it was doing. And I can click up here to go to it, and that'll clear out of there. So it's like a notification system.

  136. 20:32

    I just didn't want everything, like, yelling at me with, like, different notifications. I know people use, like, Zelda themes or, like, like, sounds from video games to, to do all that. When you have 12 agents running, it's way too much, so I just have, like, a visual indicator for that.

  137. 20:46

    Did you wanna show folks where they can get Fleet?

  138. 20:48

    Oh, yeah.

  139. 20:49

    Yeah.

  140. 20:49

    Uh, just nicknisi/fleet on-

  141. 20:50

    Yep

  142. 20:50

    ... uh, GitHub.

  143. 20:52

    Uh, and then also if you've got less than 10 tabs or if you want the lower tech solution, you can also just rename your own sessions/rename. You can also, within your terminal, rename your own tabs. So if you do both, right, you can get to a point where you're just managing that yourself.

  144. 21:05

    Yeah.

  145. 21:06

    Um, did you... And the Sessions tool is available through Brew.

  146. 21:09

    Yeah.

  147. 21:09

    Right?

  148. 21:09

    Mm-hmm.

  149. 21:09

    So you just brew, brew install sessions if you're interested in that.

  150. 21:13

    We'll have links to all of these-

  151. 21:14

    Yeah

  152. 21:14

    ... afterwards as well.

  153. 21:15

    Great question, though. Thank you.

  154. 21:16

    Yeah.

  155. 21:18

    Cool. So the next thing that we're gonna do is, uh, there's another skill in the, in the repo. So if you s- type this or if you say this to Claude, "Run my workshop check-in," this is the part that is opt-in only. D- don't feel obligated to do it, only if you wanna participate. But it'll ask you a few simple questions, I think it's four, and you can a- answer them all in one line if you want and just hit Enter. And then, uh, that'll get posted to our back end, and then a little bit later on we'll have a data visualization of where the room started and ended there.

  156. 21:46

    Yeah. We'll show that in a little bit.

  157. 21:47

    Yep.

  158. 21:48

    Um, let's move on to the next section. And in here we're gonna talk about loops. So now we're kind of, like, cooking with our voice, uh, but we're still kinda babysitting everything. Um, so the next step of that is to not eliminate us completely, but eliminate the non-essential things so that we can focus on the things where we're good at and where we focus and... Or where we're, we can be focused and contribute actively and not just babysit the agent as it's going. And that unlocks you to let the agent kinda go for a bit while you go do something else, and that something else could be

  159. 22:18

    a whole nother agent in another tab doing another thing, like I had with 12 of them going at a time, or it could be go take a walk in the woods.

  160. 22:25

    Yep.

  161. 22:25

    Anything. Um, so just to recap, uh, an agent is, uh, like, there's lots of definitions for this, but, like, one could be, like, it's a model in a loop. It's a model with tools in a loop, uh, where you're asking it to do something, and that will think about it, and then it'll act, and then you want it to observe something about that and then decide what to do next. And oftentimes that's go back through that loop again and go through it over and over and over. And there's several ways that you can do this. You can build your own, of course. Uh,

  162. 22:55

    you can build the exact workflow that works exactly for you or for your team or for your company. Claude also ships, and, uh, other ones do as well, like Codex and, and others. Uh, they ship some built-in helpers to make that very easy. We'll talk about two in, in depth here, and that's goal and loop. Uh, and they're ways to let Claude or, or your model go unattended for a bit, your agent go unattended for a bit, uh, and make decisions on its own.

  163. 23:21

    Yep. So I'll just say if you find yourself now, uh, commonly, you know, you're getting lots of good work done with Claude, but you notice that you're constantly having to go back and say, "Okay, no, do this again. Okay, that wasn't quite right. Okay, it's almost there," uh, you should be thinking and reaching for goals and loops.

  164. 23:35

    Yep.

  165. 23:35

    And I'll also say that, uh, from everything that we read and ingest, and from talking to folks as well that are, you know, constantly practicing all this, it seems that loops and similar are going to be one of the kinda key unlocks for just higher levels of productivity that, that are also less demanding on you to monitor and babysit.

  166. 23:54

    Yeah.

  167. 23:54

    So highly recommend looking into them. Um, but quickly, the difference is a goal is there is a clear definition that, that I want you to meet, and so I want you to work until that clear definition is met. It could be refactor this simple file until these tests pass. It could be I want you to read this website repeatedly, you know, for a long time until an update is there and then let me know. Or this, like, you know, this flight pri- price changes, et cetera. Um, that, that has a clear defined state. You're trying to get to it, and it's going to terminate at that point.

  168. 24:22

    Mm-hmm.

  169. 24:23

    A loop is I want you to repeat a task, you know, nonstop until I tell you to cancel it or until the timer that I set for it is expired. So, you know, different ways to leverage them, but, you know, that's an important distinction.

  170. 24:37

    Uh, and it really just does that iteration in both of these, uh, until it passes. So for Goal, for example, you could say, like, "The goal is I want you to do this until all, all of my tests are green." Like, I have some failing tests, the goal is to make them all passing. The goal is to get to 90% test coverage. Whatever your goal is, it's something that is tangible and that is measurable by the LLM, uh, by the model. Uh, if it can do that, then it can work towards that goal, and once it reaches that goal, it will, uh, it will stop, but until then, it

  171. 25:07

    will try whatever attempts that it has to get to that goal. It'll check whether it's gotten to that goal, and then it'll fix whatever it can and start the loop over and, uh... or tr- uh, fix its, its approach to it and start over. So you don't have to re-prompt. It's going to re-prompt itself continuously until it gets there is the g- the, the entire goal of Goal.

  172. 25:29

    Uh, and so the, the real thing to just kind of internalize here is that, uh, as we've all experienced, especially if you're working with Claude and other agent tools, you know, with any regularity, uh, they will punch out way sooner than the task is actually done done.

  173. 25:43

    Mm-hmm.

  174. 25:43

    At the level of done that you would be happy to sign your name to and say, "This work is complete," and send it to your colleagues to be scrutinized.

  175. 25:50

    Yeah.

  176. 25:50

    Right? And so that's the gap that we're trying to close with these primitives of Loop and Goal and eventually Schedule as well.

  177. 25:56

    Yep.

  178. 26:00

    So, uh, Goals, they're made up of these four boxes. None of them are optional. Uh, it has to, um, for example, like reproduce a bug, uh, with a failing test, implement the smallest possible fix for that, lint, type check, make sure the tests are, are running in green, uh, and then, like, afterwards it can open a PR, um, to submit that. And so it's really bringing back, like, that idea of, like, red/green refactor. Like, it's a, a really... that's a really good loop for it to get into where it can write the failing test, make, write the minimal amount of code to get it to pass, and

  179. 26:30

    then verify it. And you can do a whole bunch of extra steps in there, uh, automatically as well, uh, things like, like adversarially checking or diffing, uh, based on, like, another model looking at it and, and giving it a review, uh, and then pulling that data in and, uh, going from there to update it.

  180. 26:49

    So yeah, like we said, uh, agents will punch out early. Um, set up your loops and set up your actual, uh, re- repeatable tasks that you're constantly doing, the code bases that you care about, so that it is impossible for them to sort of lie to you-

  181. 27:02

    Yeah

  182. 27:02

    ... and it's impossible for them to, uh, kind of bail out before the task is truly done.

  183. 27:07

    Yeah. A big thing that I had, I was working on, like, this whole agent loop, uh, that was not part of Goal. It was, like, my own loop thing, and I wanted it to, like, give me some verification that it was actually, like, doing the work and testing and everything. And so I had it just, like, make sure that you're running the tests, and it would, like, say, like, "Oh, when you run the tests, touch this file called, like, case tested." The, the file... the, the agent was called Case. Uh, and if that file was, existed, that means that it ran the tests. Well, Claude figured that out pretty early and was like,

  184. 27:37

    "All I have to do is touch this file. I don't have to run the tests, actually." And so, uh, that was, like, it lying, and it's, it's a good programmer, right? It strives to be lazy like us, and that's exactly what it did. So you have to make the lazy route the, the correct route, and so I had to add in all of this, like, checksum verification in that file to make sure that it was actually running correctly.

  185. 27:59

    Yep. And then so just again to the- these visualizations just to drive it home, uh, so a goal is a bounded place we're trying to get to, and as soon as you're done, I want you to terminate. The loop is going in a loop. I think in a, in a sense, like, a loop is a little bit, um, more creative and, and more, more adaptable to, like, wider purposes. For example, like, with as, as little as one loop command, you could set up an autocompleting dev implementation loop at, you know, for a given code base where you say, "Read this Slack," you have MCP access to Slack, "See my, my colleague

  186. 28:28

    complaining about bugs and asking for feature requests for this project. When you see them, add them to Linear as subtasks. Then take a Linear subtask off the queue and implement it and merge through all the tests passing and then deploy it yourself," right? "And then close it down." You could maybe split that into two loops. It might be a little bit cleaner. The point is that, um, it's incredibly, um, just reusable and, and adaptable to many different situations. You could also think about regular reports that need to happen, and we'll, we'll talk about when Schedule's a little bit more appropriate for that too.

  187. 28:59

    But a ton of different things you can do with Loop if you haven't tried them yet.

  188. 29:01

    Yeah, for sure. And this is where all of this is going too. Uh, one thing that just came out, like, last week, and we've been playing with it but didn't have enough time to put it in the slides, is, uh, like Claude Tag, for example.

  189. 29:11

    Yeah.

  190. 29:12

    And that is just a, a simplification of being able to run loops and have some kind of trigger that, that triggers in Slack to get it to go. It could be a message from a, a colleague or an update to a PR that triggers it, and it does everything autonomously on a cloud environment, so, uh, it needs to be able to run in that independent loop without you so that it can continue working and be efficient.

  191. 29:34

    I've got Claude Tag set up in my personal Slack now and, uh, about seven or eight Vercel Eve bots, their new agentic framework, and, uh, they're doing a ton of different, you know, independent work on my business for me, and then Claude Tag is the, the main interface for that. So-

  192. 29:49

    Yeah

  193. 29:49

    ... we, we see this as all just kind of converging towards the same primitives that you need to be kind of very familiar with.

  194. 29:55

    So here's some examples of, uh, goals and loops, uh, in the wild. So, like, a goal, for example, would be, "I want this test suite to be green." There's failing tests. Let's make the goal to have everything be green. That is something that's measurable. When there's no more failing tests, it's done. Uh, another one is every single file in the source file, uh, source directory is under 200 lines. That's something that's very easily measurable by running various commands that, that, uh, Claude can do, so it'll go in and do it. Uh, an, a loop, for example, would be, "Every minute, I wanna run the tests and make sure that they're all passing."

  195. 30:25

    Uh, there's much more efficient ways to do this, so this is a little contrived, but that is one way that you could do it. And another one is, like, "Check the deploy and ping me when it's healthy." So this could be, like, you have, like, a long CI or a long build pipeline that's getting your, your code deployed, and you wanna be notified when it happens. Again, there's easier ways to do this, but Claude could also do it, and you could just have it loop and check every five minutes, and then it, it can act on that. It doesn't necessarily have to, uh, notify you. It could do, run a, like, an additional script or do something after that happens.

  196. 30:54

    Loop and tell me when my favorite video game is finally available for publishing, so I get in early, right? And text me. Um, tons of stuff you can do with it.

  197. 31:01

    Oh, yeah.

  198. 31:05

    Okay, and then, uh, we're gonna introduce the concept of Git worktree. So if you're a developer, you've, you've probably used them before. Uh, think of them as a way to give each in- independent agent a, a kind of isolated and safe to work on concurrently copy of your, uh, Git repo. And so there are issues with this. We won't get too deep into it. Um, there's experimental things coming out that are a little bit different, but, uh, we found success in using Git worktrees, and so you can run your agents in parallel, and that means that you can think of, of

  199. 31:35

    breaking down like a full afternoon or five or six hours worth of development and implementation bug fixing into discrete blocks and then assigning those discrete blocks out so that, so that they run concurrently, and you can press the time down to, you know, 45 minutes or an hour.

  200. 31:48

    Yep. How many of you are using worktrees today?

  201. 31:52

    Cool. Like half, half the room.

  202. 31:53

    How many are using Claude dash dash worktree to do that?

  203. 31:56

    Claude what?

  204. 31:57

    Dash dash worktree.

  205. 32:01

    Couple.

  206. 32:01

    Cool.

  207. 32:01

    Okay. Yeah, that is one of the easiest ways to get started with it. You can literally just be in your repo, uh, you know, whatever that is, and say Claude dash dash worktree, and it will just create a worktree right there, and then start a Claude instance in that worktree so that you can immediately start parallelizing your work and going from there. Uh, and depending on your repo setup, that's, um, that's easy or difficult depending on how difficult it is for you to run concurrent model, uh, concurrent versions of your repository locally. One thing that we have is, uh,

  208. 32:31

    a, a monorepo. I'm sure a lot of people work in monorepos, and it can be very difficult to run multiple instances of that, uh, in different worktrees because it's like gigabytes and gigabytes of node modules, for example.

  209. 32:42

    Virtual machines and Docker.

  210. 32:43

    Yeah, setting up the right-

  211. 32:45

    Build systems. Yeah

  212. 32:46

    ... uh, database, all of that, the, the right ports, like, being mapped correctly. That's where, like, more cloud environments come in, uh, that can do all of this as well. Um, we've been experimenting with Claude Tag, uh, Devin, and others, uh, but that's another way to parallelize this work so that agents can work independently in the same code base without ever conflicting with each other and then open up independent separate PRs that never, uh, bleed code changes from one to the other.

  213. 33:12

    And of course, this just loops all the way down. Um, like, we think of these as, like, atoms. Um, every block is just the same loop at a different higher altitude. Uh, and you can think about that, like, with literally everything. Like, we definitely consider ourselves, like, loop engineers now, uh, 'cause that's what we're, we're thinking about when we're thinking about, like, the atomic bits of work that we're actually trying to, uh, complete. And we could have, like, a, a loop that's, like, fixing this bug while another loop is fixing this and another is fixing, uh, around that. And

  214. 33:43

    think about, like, how we manage all of that, like, with, with, uh, that fleet tool I showed you. Like, it's thinking about all of those in a loop. I am literally looping between all of them, going from one to the other to keep things moving along.

  215. 33:53

    Yep, absolutely. You ran a loop that was, uh, two hours long the other day.

  216. 33:57

    Yeah. Regularly, the, the loops, I do a lot of upfront planning, and then I can just let Claude go, and it will go for two, three hours at a time without bothering me, uh, uh, at any... for anything. And it has all of these checks built into where it's, like, keeping things small, keeping things readable, reusing code that we have, uh, running the tests, getting feedback from Greptile or CodeRabbit or whatever, uh, code tool we use. We use them all, honestly. Uh, and then, um, implementing that back

  217. 34:27

    in, and then once everybody's happy, then it l- brings it to me, and that's something that I can then look at and then approve for someone else, a human to look at. We still have human-in-the-loop reviews, uh, on a lot of things. And so, uh, but by the time all of the dust has settled, it's working. It's shown me that it works. The code looks clean. It looks like code that I would have written, and it's ready for a human to actually review it, and it's not just slop code that I'm throwing over the fence for somebody else to, to worry about.

  218. 34:58

    So now we're gonna have you try it. Um, we're gonna use loops and goals, uh, and h- where you can hand over a job and walk away. Uh, so in that repo, you can run /goal and then tell it, um, this bun command, bun playground/goals check.ts, and the goal is that when you run that check, it will show five out of five. So everything's passing, everything's working. Uh, and if that's not the case immediately, Claude will just spin and run until it is complete.

  219. 35:26

    Yep. I'll give you a minute or two to run this and see what that output looks like.

  220. 35:30

    Yeah, question.

  221. 35:31

    Um, your program can't handle the size of my files. I have a 3.5 gigabyte JSONL that it's trying to read, and it's blowing up your JavaScript

  222. 35:44

    Is it?

  223. 35:44

    Oh, nice.

  224. 35:45

    Which program? Is that, uh-

  225. 35:48

    The one where I'm trying to evaluate my-

  226. 35:50

    Oh, okay.

  227. 35:51

    Okay.

  228. 35:52

    3.5 gigabytes. Uh-

  229. 35:57

    3.5 gigabyte file and it said loading file.

  230. 35:58

    What you talking about?

  231. 35:59

    Pardon me?

  232. 35:59

    What you talking about in there?

  233. 36:01

    I've been working on it all day every day.

  234. 36:03

    Nice.

  235. 36:04

    So-

  236. 36:04

    Cool

  237. 36:05

    ... for like a year and a half, two.

  238. 36:07

    Nice.

  239. 36:07

    Yeah.

  240. 36:08

    A bit over a year.

  241. 36:09

    Okay.

  242. 36:09

    Maybe not quite a year and a half.

  243. 36:11

    All right.

  244. 36:11

    Do you just run compact often? Is that, like, how do you manage-

  245. 36:15

    That's what I can do.

  246. 36:16

    Yeah.

  247. 36:18

    I don't even know about running compact in the, um, project folder.

  248. 36:21

    Uh, yeah. When you're working with Claude, um, one thing I'll show in mine here, in this one, um, you can see down at the bottom-

  249. 36:30

    Um, yeah, I do that before I start a new session.

  250. 36:32

    Yeah.

  251. 36:32

    Okay.

  252. 36:33

    Yeah.

  253. 36:33

    But then somewhere-

  254. 36:37

    Oh

  255. 36:37

    ... in the projects folder-

  256. 36:39

    Yeah

  257. 36:39

    ... .Claude, so it knows where the big file is.

  258. 36:40

    Gotcha. Okay.

  259. 36:41

    Okay.

  260. 36:41

    Appreciate the bug report. Thank you.

  261. 36:42

    Yeah.

  262. 36:43

    And we'll fix that.

  263. 36:43

    So yeah, for the recording, that was, uh, it- it's failing on a 3.5 gigabyte JSONL file. We didn't anticipate that. I'm sorry.

  264. 36:52

    Thank you.

  265. 36:55

    It's a humble brand.

  266. 36:56

    Yeah.

  267. 36:56

    No, you're good.

  268. 36:56

    Oh, it's great.

  269. 36:57

    You're good.

  270. 37:00

    We did not... I, I have not seen that yet.

  271. 37:03

    Cool.

  272. 37:03

    Yeah, question.

  273. 37:04

    Oh, yep, sorry. After you.

  274. 37:06

    Okay. Um, yeah, so what's, what's the... So like obviously with Claude, there's like usage limits.

  275. 37:11

    Mm-hmm.

  276. 37:16

    That's a great question.

  277. 37:16

    Great question. The question is, what's the, the impact on usage limits for running things in a loop, and that is something that you definitely need to be cognizant of.

  278. 37:24

    Yeah.

  279. 37:24

    Um-

  280. 37:24

    I was... I'm gonna say, this gentleman built, and it's actually in this repo if you've cloned it, it's in the Claude settings. It's called the ideation plugin, the Nick Nisi ideation plugin. Uh, one of the best ways to handle it is to just do your planning upfront, and to get to deep and abiding clarity on exactly what you're trying to build, what finished looks like, what failure looks like, what you're not going to include in it, what's scoped into it, and then set up the loop that's going to be a higher token burn for sure. Um, yeah.

  281. 37:51

    Yeah. And the way that that plugin specifically, that's why I don't use like, uh, the built-in ones. O- often I use my own, is because it's specifically trying to break things up into these manageable specs, uh, where it can run them in independent context windows, and it will try and... And I'm mostly thinking about, like, context size in that case. Um, we, we're fortunate, uh, uh, that we don't really have usage limits currently. Uh, but I can see a, a very near future where that is the case. Um, but

  282. 38:21

    it's, it's meant to, uh, do things in these, like, small atomic pieces that it can do in independent context windows, and then not have a bunch of bloat, uh, that's comes associated with that to keep it more manageable. But yes, y- running things in loops, and there are things as well like, um, in, in more recent versions of Claude, there's this, this thing called dynamic workflows. Uh, where it will effectively like build like a, a small state machine to run a whole loop in, uh, in various different ways and has gates between them. Uh, and that can be very token heavy.

  283. 38:51

    I've seen it get up to 4 or 5 million tokens that it's used for a single task, uh, spread across several workflows. Like it'll spin up 35 agents and run them all simultaneously. Things like that.

  284. 39:01

    Yeah. Well, we're also gonna talk about verification gates a little bit here-

  285. 39:04

    Yeah

  286. 39:04

    ... in a bit. I think that's the other, you know, solution to that problem, which is really the token bloat or token waste in that, in the, in the worst case when you're just looping, um, kind of comes from like the unbounded exploration without anyone telling you you're off track, and without the model being able to determine that it's wasting, wasting turn, you know?

  287. 39:21

    Yeah.

  288. 39:22

    And so having, uh, you know, doing the, uh, the planning up front, even burning extra tokens on getting that right, using max thinking there, and then setting up your loop so that there are hard verification gates so that it cannot just get lost is, are two ways I'd think about it for sure. Yes. Sir, yeah. Go ahead.

  289. 39:38

    So we built a UI or something, and we come in, as it turns out, the plot screwed up. It's not looking just right.

  290. 39:45

    Mm-hmm.

  291. 39:45

    I feel like I'm babysitting it.

  292. 39:46

    Yeah.

  293. 39:46

    Have you guys figured out the best workflow to sort of... We don't want to completely eliminate human in the loop.

  294. 39:52

    Yep.

  295. 39:52

    But we want human in the loop at the end.

  296. 39:54

    Mm-hmm.

  297. 39:54

    Not in the middle.

  298. 39:55

    For sure.

  299. 39:56

    Trying to babysit it.

  300. 39:57

    Yep.

  301. 39:57

    What's the loop or workflow that you guys have done so that I am not babysitting UI? I want it to run the app and QA test it for me before it gets to me.

  302. 40:06

    Yep.

  303. 40:06

    Yeah.

  304. 40:07

    Yeah, great question. I-

  305. 40:07

    Uh, so the question is-

  306. 40:08

    Yeah

  307. 40:09

    ... uh, how do you not babysit? Like, what are the, the ways to, uh, get it to work more independently without you having to babysit it along the way because it's making a lot of mistakes? And the, the solution that I've come up with, uh, for that is to let it make those mistakes, but record them and then, uh, keep track of that so that it doesn't make those mistakes again. For every loop that I run, it run, I run like a little, uh, retrospective agent, uh, that will go over the performance of the loop at the end, and it will see, okay, what did we do here? Why

  308. 40:39

    were we wasting a bunch of time here? W- I had to correct it a bunch here, or here's where it was running a bunch of tools without like changing things and like figure, like running into messy loops like that.

  309. 40:50

    Are you using computer use?

  310. 40:50

    No.

  311. 40:51

    No.

  312. 40:51

    No. This is just at the end after it's, uh... I'm, I'm not using computer use. It's just at the end after running a loop where I have had to do a lot of babysitting. Uh, we work through that, I work through that with Claude so that it builds up this memory file that gets loaded back in, uh, at the next session that we do to tell it things. And it, it, it also does some curating of that. This is like a, a separate memory system that I'm just kind of like toying with on my own, and it's Markdown files. Uh, but it's, it, it's something that gets curated as well, so it thinks about things in short and long-term memory.

  313. 41:22

    Uh, so short-term things will get f- uh, filtered out over time. Longer term things are like the core tenets that we keep running into that I don't ever want you to forget. Like, this is where our design system lives. This is how you test this specific, uh, component, for example. Uh, things like that, and then that gets loaded in, and it gets broken up, uh, using, um, progressive disclosure so that, you know, if I'm not talking about testing, it's not gonna load the memories about testing specifically. Or if I'm not talking about the React portion, it's not gonna bring up React specific memories. It'll only

  314. 41:52

    bring those up when we actually touch React files.

  315. 41:54

    I'll also say, you know, specific to front end, um, the way I might approach that problem too is Claude Design first. Similar to the ideation, the higher spend up in the beginning. If you haven't tried that yet, that's like a, a specific like instance of Claude adapted for design, and it is able to develop for you like a design system that then persists across sessions and you can reuse it. So like starting with the upfront planning first, and then second, and we'll touch on this a little bit later, adversarial review. Um, you know, injecting either with hooks towards the end of an implementation loop, and

  316. 42:24

    specifically loading skills that are designed to test UX and front end, uh, best practices. And you want that piece to be as negative and persnickety and nitpicking as possible, and to fail those first 25 builds and say, "This is not good enough. Keep going. Fix these things." Right? That, that's how, like once you get those pieces in, you can engineer a loop that, yeah, you go and you make your coffee and you come back maybe 40 minutes later, and you've burned a bunch of tokens, but you're actually happy with it and it's pretty close to merging.

  317. 42:52

    Mm-hmm.

  318. 42:52

    Yeah.

  319. 42:53

    We'll take one more question, and then we'll move on.

  320. 42:54

    Yep.

  321. 42:54

    Yeah.

  322. 42:55

    To vary that a little bit, have you guys experimented or messed around with anything when it comes to passive voice listening and the system taking action based on like a conversation-

  323. 43:09

    Uh, that's a great question. So the question was, uh, have you experimented with any passive voice or the system listening to the ambient discussions of people in the room and other operators and then working on that? Uh, kind of not yet. One thing I will share that has been insanely, and this is a little bit surprising of an answer, but has been insanely, um, uh, impactful, and so I think that that actually would bear fruit and is worth exploring, is, uh, we did build a system that allows, uh, everyone in the org to blog in a very uniform way with a common voice and it fixes, you know, AI slop, and it, uh,

  324. 43:38

    allows the- everyone to produce a draft of a blog that looks exactly like, you know, the expert at making blogs for WorkOS would do it, and the highest leverage connector that we added to that was Granola.

  325. 43:50

    Mm-hmm.

  326. 43:50

    'Cause as soon as you pipe in all the calls from people, people are saying, "Oh, it's just this little idea that we just had while talking to a customer. If we had blogged a post on that, it would explain it." And that is all that needs to be said, and they just pipe... That's a single MCP call, and the blog bot's off and running, and what comes out is excellent. So I think that, what you mentioned, the passive listening, ambient listening, and then also computer use is gonna get pretty freaky at some point soon, where you combine those two and it starts to say, "Here was your year, and these 16 specific actions are what actually drove your business, so now

  327. 44:20

    I'm gonna do those, like, at rapid speed or-

  328. 44:22

    Mm-hmm

  329. 44:22

    ... you know, et cetera. Yeah.

  330. 44:23

    And the only other piece of that I would say is experimenting with Devin and Claude Tag and them kind of ambiently popping up in Claude me- or in Slack messages. Uh, that's more annoying right now at least than, uh, helpful.

  331. 44:34

    Yep.

  332. 44:35

    But, um, yeah.

  333. 44:36

    Yeah.

  334. 44:36

    That's...

  335. 44:37

    Oh, we gotta cook a little bit here.

  336. 44:39

    Yeah.

  337. 44:39

    Um, so th- there's a, it's test goals as well in that the whole point is, uh, take that repo with you if you wanna continue practicing, uh, loops and goals, and you can use the playground scripts there, and then you can adapt them to your own programs.

  338. 44:50

    Yep. Never tell it to try harder. Uh, it doesn't know what you mean by that. Uh, always give it some kind of, uh-

  339. 44:55

    You can rage at it, though. It's-

  340. 44:57

    Yeah

  341. 44:57

    ... it's good for the soul. It's cathartic. You should do it.

  342. 44:58

    We, we have a whole channel called Bot Hate where we're just screaming at the bots-

  343. 45:01

    Yeah

  344. 45:01

    ... for doing different things. Um, but yeah, fa- whe- when you're working on these and it's failing, fix the system. Don't f- don't, uh, fix the, the code itself. Fix the system that generated the code so that next time, uh, it can be fixed automatically or it doesn't hit that in the first place, and that's what my retro, uh, idea does. And speaking of that, like, this is how you, you prevent that or y- how you prevent it from going off the rails. You give it ways for it to verify itself, and you trust those gates and trust that

  345. 45:31

    they will prevent the model from moving on until those gates are met so that you have more confidence in letting it go on its own because it's going to un- understand what done means, and done for me does not mean what done means for the bot. So we need to, to come to an agreement on what it actually means, and that's what this verification is.

  346. 45:50

    Yep. So it, it could change depending on any project. It could be if you've got a TypeScript, it's gonna be the obvious ones like lint and build and, you know, did it compile actually safely.

  347. 45:58

    Yep.

  348. 45:58

    Uh, when I'm writing for my own, uh, blog, I have a system that the stuff that drives me bananas that are non-negotiables for a PR, so, like, all the images must return 200 off the CDN. The Open Graph image must be formed perfectly. There's other checks that are in that script and a failing, um, you know, status code from that batch script fails the entire build, and then-

  349. 46:16

    Mm-hmm

  350. 46:16

    ... the system picks it up, and then it tries again, and it works on fixing all of that. But it, it's, those problems are solved before it even gets to me, and the end result is still a, a working PR.

  351. 46:25

    Yep, and this is what I meant by, um, you just make it, make the, the trusted route, the way that it should work, the lazy route. Uh, and it, just make sure that it has all of these gates, and there's various ways that you can do that. You can force it w- uh, through things like hooks. Hooks will prevent it from being able to move on or do things, uh, or it will make it, like, run linting after it runs, uh, after it writes, uh, updated files, things like that. And those are ways that it just can't avoid. It won't stop. It won't forget about that. One common thing that, that we've seen,

  352. 46:55

    like, when we're doing this is, like, we'll tell it a whole bunch of steps. We'll be like, "Here's the seven things that you need to do every time we work together." And if our conversation goes for like 350, 400,000 tokens, it will forget steps four and five because it just, like, that's so far back in the, the context that it doesn't remember to do that, uh, and it gets lost in the minutia of what we're actually doing. And so we've had a lot of success in breaking that out into, like, a, a state machine in TypeScript, for example, that forces it to go forward. We've been building on top of Pie, uh, quite a bit

  353. 47:25

    to, to do that, um, so that it's not something that can be avoided. It has to move from one step to the next, and there's no way to get from step C, uh, from step A to step C without going through step B.

  354. 47:37

    Yep.

  355. 47:39

    Um, so we talked about some of these. Uh, we can skip ahead. One of my favorite things to do that is, uh, bearing the most fruit, I would say, in the last, like, three, four weeks since I started experimenting with it is add a hook to Claude. And the f- the fun thing about hooks too is that you can just tell Claude to add a hook, right?

  356. 47:53

    Mm-hmm.

  357. 47:53

    And it'll do it correctly. So add a hook, uh, that every single time you're working on something that is either very sensitive or urgent 'cause it's gonna unblock some production issue or is of a certain complexity, and you can even be the judge of that, then you must fan out your code diff for an adversarial review from Codex 55 via the CLI. So you install Codex yourself, or you get Claude's help with it. You auth into Codex, and you log in with your OpenAI account. You do the Go- you know, the OAuth round trip, and then from that point forward, as Claude is

  358. 48:23

    building with its, you know, constant and ever-cheerful self saying, "Oh, I nailed it. It's perfect. Don't worry. Don't look at it." Right? And then it's about to open the PR, but the hook says you must go and get an adversarial review. Uh, that is finding significant issues on the regular for me.

  359. 48:38

    Mm-hmm.

  360. 48:38

    And that, which reduces the overall time because it's not kicking off expensive builds and Greptile reviews we have to pay credits for yet. There's something finding it locally, locally, um, before it even gets to that process. And so we have found even, uh, experimented with this a couple weeks ago, and it was finishing a, a pretty, um, major migration of a system and found that having, uh, four different agents from different companies review the same code, all of them found separate things and different things, which was-

  361. 49:05

    Oh, yeah

  362. 49:06

    ... not surprising given that they've got slightly different training, you know, uh, processes and weights and everything.

  363. 49:11

    And Codex does ship some skills. Uh, they ship a review skill, an adversarial review, a rescue, uh, all of these easy ways for Claude to call it, get advice from, uh, Codex, and then bring that right back in so that Claude can act on it in different ways. And this is a great way to get them to work together, um, when it... A- and, and it's a, a really good way to just get that multi-model, uh, perspective on the code that you're writing and, uh, how it's performing.

  364. 49:37

    Yep. Because, you know, the one model will say, especially if it's still in the same session, it's going to adapt back to the context it still has where it told y- you were telling it, "I want you to complete this goal," and it's going to try and continue to confabulate helpfully and pleasantly and tell you that it did that goal. Uh, whereas, you know, a model that doesn't necessarily have that same context but has a very, uh, succinct prompt about here is the inputs, here's the outputs, here's the goal, can start finding issues in the code immediately. So, uh, yeah. Th- basically the idea is,

  365. 50:07

    uh, don't, um, trust the model's confidence as much as you might enjoy or feel partial to Claude or any other agent. Uh, you trust the gates that you set up that make it impossible for the system to lie to you.

  366. 50:18

    And when you find that it did lie or figured out a way around that, fix it.

  367. 50:21

    Yep.

  368. 50:21

    Don't fix the code.

  369. 50:24

    So we can try that, um, now with verification da- verification gates. Um, we can make something done, uh, to it so that we prove that this actually works. Uh, so in that repo, you can just say, like, add a hook that runs lint, uh, type check, and the test on every change, and fix its, uh, flags. Fix anything it flags.

  370. 50:42

    Yeah.

  371. 50:42

    I can't talk.

  372. 50:47

    All right. You can also ask it to get a second opinion, so ask it to just fan out and give you an adversarial review and then fix anything that it finds, and this way you're having the two models work together. One of them is fixing it, one of them is reviewing, and they can go around and around on that, um, and get the work done.

  373. 51:03

    Yeah. You can also, by the way, for what it's worth, uh, it doesn't always have to be adversarial review. You can say s-

  374. 51:07

    Mm-hmm

  375. 51:07

    ... you know, scope this work out, um, until it's k- it's very clear, and then divvy up the pieces where it's, like, the clean seam so that you can both work on it, uh, more rapidly and, and burn more of my money faster.

  376. 51:19

    Yeah. A big thing that I do is I, we, we have access to all of these models, uh, and so I will often use this adversarial piece, uh, in the planning phase. So I'll have Claude and I make a plan, uh, using the various skills and, and ways that I would make that plan, and then once I have those artifacts for it, then I'll pass it to Clau- to Codex, uh, and have it do, like, a review on it and say, like, what, what makes sense, what doesn't? And it will often pare it down, like, a lot. It'll be like, well, this kind of doesn't make sense. And it will send that right back to Claude, and Claude will be like, "Yeah, Codex

  377. 51:49

    is right. I see what they're going for here." And will concede a lot of the times. It's very amenable to, uh, influence in that way. Uh, but, uh, it, uh, in the end it works out to be a better plan. And for me, I'm always making it write, like, HTML plans-

  378. 52:02

    Oh

  379. 52:02

    ... and so it's something that I'm following along with, too. Uh, feeling like I'm con- like the dog at the computer, like, I'm contributing. But we do get things done.

  380. 52:15

    Okay. Scheduled tasks. So, uh, we have a few minutes left. We're gonna kind of rip through this. Uh, one of my favorite things to do when you're first se- setting up... So what are scheduled tasks and why do you care about them? Uh, it's that thing that you still come into Slack every Monday dreading to do. It's like pulling together reports, and I have to go dig through Notion, and I have to go find that Slack about how they wanted the format done, then I have to go download this CSV, and I have to look through these, like, customers. That stuff can be, uh, you know, first worked on in a session until you get it working, and then as soon as you have it working, you can say, "Now /schedule this for every Monday at 9:30

  381. 52:45

    AM, 'cause I never wanna do it again myself." And that will actually kick off and save a durable, you know, cloud session in a VM that, um, is able to do that work for you on a repeatable basis, whatever you schedule. There's utilities in the schedule command, like list, and then you can see what you have running, and you can modify them or delete them. But think of it as anything that you need to do for work or for yourself that is tedious and you'd like to automate. Schedule is the command that you wanna be reaching for.

  382. 53:12

    Yep.

  383. 53:14

    So-

  384. 53:14

    Is it-

  385. 53:14

    Go ahead.

  386. 53:14

    Go ahead.

  387. 53:15

    Oh, yeah. So just quickly, we, since we've introduced a bunch of these primitives to you, just a, a quick word on when to think about them. Hooks are things that are gonna run automatically in certain cases. Maybe you could write a hook to say, always run, but then only proceed if this case exists, right? Because it's a bash script, for example. But it's things that you want Claude or any other agent to always remember and do and never forget and not be able to forget. Uh, goal is if you have a clear goal in mind, like, I d- just need to get this PR over the line, so I need you to fix these 13 tests, and then also refactor this auth.ts.

  388. 53:45

    And once that's done, and you can prove it's done, then I want you to stop and stop burning money. Loop is, uh, I want you to go babysit this PR, you know. I want you to, um, constantly ping here and then tell me when, uh, this webpage changes, and this is the data that comes back. I want you to go through Slack and get all the complaints that are coming in from customers and file them all as linear tickets. So, like, run until I tell you to stop or on a timer I've given you. And then schedule is something that you want to be durable. So I want you to run every two weeks, every Monday, every Wednesday, every Friday, every day, like a cron job, right? And I want you to

  389. 54:15

    do some amount of work, and you may have secrets involved, too. But these are the kind of building blocks that, uh, together you can kind of create, like, fully autonomous systems and, and automate a lot of the tedium out of your workflow.

  390. 54:28

    Mm-hmm.

  391. 54:32

    So any of these are great candidates. If you've done any of this, dependency bumps, status reports, eval runs, um, you know, if it recurs and you have access to Claude and scheduled tasks, you shouldn't have to keep doing it manually.

  392. 54:43

    Yep. The biggest one for me is always just what are, like, what, what do I need to prepare for meetings? That's the biggest thing for me a lot of the time, is just, like, getting things ready for that on a sc- on a specific schedule. There's other triggers as well. It doesn't necessarily have to be time-based. It could be, um, different things come in, like new pull requests come in. Uh, how do you handle that? It could automatically kick things off. It could be messages in Slack come in, and it just takes it over. We have a, like, a whole knits channel for, like, our docs, for example, and anything that gets posted in

  393. 55:13

    there immediately gets a linear ticket created, and then a bot automatically tries its best to fix that. And, like, 90% of the time it can fix it on its own because they're just little knit thing, nitpicky things, uh, and it'll go fix that automatically. Otherwise, those would just sit on a backlog that we never prioritize. Uh, but this way, as soon as they're filed, they are fixed, which is just such an amazing feeling. I also have this one that runs weekly, uh, on my- Um, website, and it's just, like, a fun little visualization

  394. 55:43

    that just shows me what I did for the week, uh, specifically with, like, token usage, and it shows me, like, the trends of how I'm using tokens over time, the code that I'm changing.

  395. 55:52

    It's not his personal spend only.

  396. 55:53

    It's not my it's not my personal spend.

  397. 55:54

    [REDACTED]'s horrified.

  398. 55:55

    Uh, uh, but it also, like, summarizes my week from PRs. So these are all of the PRs that I opened. Uh, and this is a summary of, like, the work that we were specifically focused on for that week, and it just goes back, uh, I think the last four months is how far back.

  399. 56:10

    Yeah. 'Cause I, I thought this was brilliant when Nick told me about this 'cause it's like you don't wanna scramble before your performance review and go figure this out and then, you know, mine it every single time, especially if you know it's coming twice a year. So it's just easier to just have it set up once and then automated.

  400. 56:22

    For sure.

  401. 56:27

    Okay. So, uh, just a quick tip that I found that works really well. Uh, if you are trying to set up your first scheduled task, one thing that works really well is to provide all the context that's necessary to Claude, hook up the MCP connectors, uh, add the skills, right? Get that first working session to the goal yourself just by chatting through and seeing it work, and then as soon as it works, you say, "Now I want you to slash schedule this for whatever," and just keep doing it. Um, that way you've got confidence that all the, the... You know, there's no s- missing secrets, there's no missing connectors. You have all the data you need.

  402. 56:59

    And this lets you really walk away. Like, once you have it set up where you're setting things up that are loopable, you're thinking about how I can loop this, how I can make it think about the, what it means to be done or what, what, uh, phases it needs to go through to give me the confidence that things are going to be done in the way that I expect them to, then that lets you schedule it. You can have it scheduled and kick off, like, when a new bug comes in, for example, a new GitHub issue. Uh, it's gonna kick all of this off, and it's not gonna bother you. You can be out on a walk in the forest

  403. 57:29

    and, uh, only be notified when a fix is ready or only be notified when it actually needs your help on something that it can't figure out on its own because it's new and novel. And then you can go in and fix that loop so that it doesn't have to ask you about that again, and then you can move on.

  404. 57:44

    Yep.

  405. 57:46

    Uh, and so, uh, as I mentioned before, there's a whole interface built around schedule. You can list out what you've got going.

  406. 57:51

    Mm-hmm.

  407. 57:52

    And, uh, you can see stuff. You can modify the schedules or tune it as needed and then find them where they're running and then delete them. Um, yep.

  408. 57:59

    And you can mix and match these, too.

  409. 58:00

    Yeah.

  410. 58:00

    You can say, like, "Schedule this goal every week," and it will run that goal, uh, and, and do that.

  411. 58:08

    So now you can try it in the last two minutes that we have. We'll kind of, uh, pass, pass through this. But, um, you can try scheduling it with, like, different loops. Like every minute run this report. Uh, schedule every weekday at 9:00 AM you're gonna run this report and summarize the list that you get from there, and then list out all of the scheduled tasks that you have. Uh, and then finally, you can close the loop that we had from the beginning of this. So if you just say, "Run my closing check-in," uh, we should be able to hopefully see a visualization.

  412. 58:36

    If we got... Yeah, if we got enough submissions-

  413. 58:37

    Yeah

  414. 58:37

    ... then we'll, we'll see it.

  415. 58:39

    And while that's working, um-

  416. 58:41

    And that's there, yeah.

  417. 58:42

    Yep. We'll look up and see. Um, so you keep all of this, this, uh, content. It's all out on GitHub. These slides are out there as well. Um, we'll also be downstairs at the WorkOS booth-

  418. 58:55

    Yeah

  419. 58:55

    ... uh, throughout the week. So if you have any questions or wanna chat about this, we live and breathe to talk about this stuff, so we would be happy to. But also we'd be happy to take... Uh, we have one minute left for any final questions.

  420. 59:08

    Um,

  421. 59:08

    Yeah.

  422. 59:09

    Sure. Yeah. We'll get you-

  423. 59:10

    Sure they could definitely-

  424. 59:11

    ... in, in the README. At the top of the README are the links to the slides, the link to that glossary that we'll keep up and that you can use and ask any questions to. Um, and then-

  425. 59:19

    So I noticed, you know, Claude has a check-in, uh, pull request

  426. 59:22

    Amazing.

  427. 59:22

    Oh, sweet.

  428. 59:23

    Thank you so much.

  429. 59:23

    Thank you.

  430. 59:24

    We appreciate it. We'll give you, we'll give you credit in our release.

  431. 59:26

    Yes.

  432. 59:28

    Yeah. Yeah. So I th- I'm sorry, I'm kind of deaf. I, I think the question was in my company, the PRs especially agent en- en- enabled a, um, PRs and the volume is really difficult because you might get a stack of 30 PRs sometimes, and then sometimes the code's bad or there's just an immense amount of, like, 50,000 files in a, in a single pull request and you have to review it. We, we think of that in, in terms of, like, applying the same kind of concepts. So, like, now it's like a, it's a large system complexity problem. And, uh, some of the tools that we've found success with are, you know, Devin, Greptile, getting, like, a first warning

  433. 59:58

    analysis. It doesn't mean that if Greptile... 'Cause they're, they're wrong sometimes as, as we know why. Like, if they sa- say it's five out of five ready to merge, that's just your first kind of signal. And then, um, so you can fail out all the ones that are twos and ones and tell people, "This is not ready for me to be bothered by," and kick it, kick it back to them honestly. And then for the stuff that is, like, fours and fives and looks okay, there'll be comments and findings on it. I'll tell agents in a loop, "Babysit this PR and fix these until there's no more comments." And then when I'm

  434. 1:00:28

    starting to get towards the end of that loop, I will actually start going and using my human brain for the first time in three years and open up the diff and say like, "Okay, what are those files I remember it used to be sensitive? Let's look at configuration. Let's look at the auth. Like, let's look at the, the routes that are published," right? And then just kind of spot check that, too. Um, that's the way that it's, it's currently working for us as well, but this guy deals with a significantly higher volume of inbound PRs than I do.

  435. 1:00:53

    Yeah. I'll, I'll just say since we're out of time, uh, come talk to us at the booth. I have, I have a tool called DAB that I'm building that tells you a story about your PRs, uh, so that you can stay in the loop and, and keep up on all the context of it.

  436. 1:01:04

    Yeah.

  437. 1:01:04

    I'd be happy to talk to you about it.

  438. 1:01:05

    Yeah. Great question.

  439. 1:01:06

    Yeah. All right.

  440. 1:01:07

    Great, folks.

  441. 1:01:08

    Thank you, everyone.

  442. 1:01:08

    We're out of time. Thanks so much. Appreciate it.