← All speakers

Bio, Work & Ideas

Dru Knox

Conference affiliation: Tessl

On this page

Dru Knox builds products that help people work with complex software, from web-platform documentation to machine-learning tools and coding agents. At Tessl, where he was head of product and design at the time of his 2026 software-factory talk, his work centers on turning developers’ repeated corrections into shared knowledge that improves future agent work.

From web platforms to AI products

Knox’s earlier career included product work at Google and Airtable. At Google, he worked on the web platform and tools for its creators. In 2017, he advocated unified web documentation through MDN, alongside cross-browser testing and infrastructure for identifying inconsistent implementations. Developers needed a dependable account of how APIs worked across browsers, rather than separate explanations scattered among vendors.

In 2018, as a product manager on Google Search, Knox introduced Blog Compass, an Android application for English- and Hindi-speaking bloggers in India. It connected WordPress and Blogger with Analytics, Search Console, and personalized topic suggestions from Google Trends. Bringing those tools together was intended to reduce the administrative work of publishing and leave bloggers more time to write.

At Grammarly, Knox worked as a product manager on machine-learning products for communication. He co-authored an account of modeling reader attention with machine-learning engineer Karun Singh. Their team built an attention heatmap to help writers understand which sentences readers were likely to absorb. To collect measurements at scale, readers navigated an email one sentence at a time while the remaining text was blurred. Reading time served as a proxy for attention; the model weighed a sentence’s informational value against the effort required to read it. This gave writers an explanation they could act on: where important information might be overlooked and how the demands of reading affected its reception.

Knox’s subsequent product work extended into generative AI, including work at Cantina and his own startup, Overworld. He spent a year developing its generative-AI world-building assistant for dungeon masters, authors, and game developers, applying AI to the work of creating fictional worlds. That work preceded his work at Tessl on tools for coding agents.

Turning corrections into shared knowledge

Knox’s Tessl work connects better library knowledge with the larger problem of making agent workflows dependable. His research and product arguments address several parts of that problem:

  • Versioned API knowledge: Knox co-authored Tessl’s coding-agent evaluation research with Maksim Shaposhnikov, Maria Gorinova, and Rob Willoughby. The team tested whether library-specific documentation helped agents use existing abstractions correctly instead of inventing interfaces or rebuilding functionality. In its Claude Code experiments using Sonnet 4.5, documentation tiles improved the average abstraction-adherence score from 59.70 to 80.97, approximately a 35% relative increase. The result concerned correct use of library abstractions on the team’s evaluation tasks; it was not a general measure of software quality.
  • Improvement loops: Knox treats recurring agent mistakes as inputs to an improvement process. His account of Tessl Agent describes using pull requests, session logs, and tickets to identify repeated errors, propose reusable instructions, and establish recurring workflows. A correction that previously required another human review can become a skill available to future tasks. Over time, routine improvements can run on triggers rather than depend on someone remembering to initiate them.
  • Shared, reviewable context: Knox argues for keeping workflow descriptions and team standards in versioned context. A design guideline used by both implementation and review agents illustrates the mechanism: updating the guideline after a review failure helps the next implementation avoid the same mistake. He also describes a Tessl customer-insight automation that lost credibility because colleagues could not understand its decisions. Moving its policy into a reviewable skill gave the team a place to debate and improve the process. Inspectable instructions support both agent behavior and human agreement about what the workflow should do.
  • Portable organizational knowledge: Knox’s open-factory approach puts skills, evaluations, and context in files teams own and can move between agents or providers. These accumulated standards preserve institutional knowledge, allowing a team to change tools while retaining the processes it has developed.

Building a software factory one workflow at a time

Knox’s harness-engineering approach shifts engineering effort toward the systems that guide, coordinate, and improve coding agents. He describes a software factory in terms of autonomy, automation, and quality, with higher quality as a central payoff. That framing makes fewer manual takeovers and less corrective review useful signals alongside the amount of work agents initiate independently.

His preferred starting point is a small workflow whose instructions people can inspect and whose results they trust. Automation follows that agreement; expansion follows experience. The practical unit of progress is a recurring task made more dependable, with its lessons available to the rest of the organization.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Dru Knox explains how cheap development checks, deeper PR review and feedback-driven improvements can turn coding agents into a system engineers build and maintain—one workflow at a time.

  • Autonomy measures how much correction an agent needs; automation measures how much work can proceed without manual verification. Improve both while tracking product quality.
    3:00 ↗
  • Run cheap checks during development, expensive checks on PRs, and a meta loop across logs, reviews and user feedback to improve future attempts.
    6:46 ↗
  • Shared issues, sandboxed agent runs and PR comments make human corrections available for later analysis; reusable skills distribute what the team learns.
    11:53 ↗
  • A recurring task such as a flaky-test hunt can become a skill and an automated workflow, reducing the separate effort required to start each run.
    17:38 ↗