← All speakers

Bio, Work & Ideas

Eric Schwartz

Conference affiliation: Traversal

On this page

Eric Schwartz works on making AI incident investigation useful to engineers responsible for production systems. A product manager at Traversal in 2026, he helped conceive Traversal Workers with the team: agents that follow incident conversations and contribute investigative findings where responders already collaborate.

Schwartz studied at Harvard Business School from 2023 to 2025. His subsequent product work at Traversal connects operational knowledge, alert interpretation, and incident coordination. He has written about Knowledge Bank and Alert Intelligence and co-authored Workers launch announcements with engineer Lyndon Vickrey. These contributions address a practical question: what does an AI investigator need beyond access to telemetry to become useful to a particular engineering team?

Bringing investigation into the incident room

Schwartz’s involvement in conceiving Workers connects his product thinking to the circumstances of incident response. Engineers coordinate through a changing conversation: they propose explanations, share discoveries, rule out possibilities, and decide what to investigate next. An assistant that waits to be invoked depends on someone recognizing when it could help. Scheduled automation depends on someone anticipating the right trigger.

The Workers design, which Schwartz described with Vickrey, puts agents inside that conversation. Workers follow the developing incident and decide when to contribute findings and when to stay quiet. The product’s usefulness therefore depends on the timing and relevance of its participation as well as the quality of its investigation.

Traversal launched Workers in public beta on June 30, 2026. The September release made Incident Workers generally available and introduced Alert Workers in public beta, extending the approach from active incident coordination to incoming alerts.

The conditions for useful autonomy

Schwartz’s writing and public explanations develop several connected positions about AI-assisted operations:

  • Operational memory: His Knowledge Bank explanation treats runbooks, postmortems, and team habits as necessary context for investigation. The system extracts customer-specific lessons from interactions, accepts explicit feedback and manually supplied guidance, and lets engineers inspect or edit what it remembers. Knowledge about tracing requests through a particular middleware system can change how an investigation proceeds; generic debugging advice contributes much less.
  • Alerts need explanations: In his Alert Intelligence writing, Schwartz challenges the habit of dismissing repeated alerts as noise. A recurring signal may represent an unresolved issue, a poorly chosen threshold, or a harmless event that matters only alongside another change. Connecting it to deployment history, neighboring services, and normal behavior gives responders a basis for deciding whether to act.
  • Root cause requires causal reasoning: Schwartz argues that observability can reveal what broke while leaving the harder question of why unanswered. When failures propagate across services, the visible symptom may be several steps removed from its cause. His account of production troubleshooting emphasizes following those dependencies rather than treating the most conspicuous anomaly as an explanation. Faster AI-assisted code production makes this investigative burden increasingly consequential.
  • Autonomy needs preparation and controlled permissions: In a public interview about AI SRE at enterprise scale, Schwartz explains why telemetry must be continuously indexed and compressed before an investigation needs it. He distinguishes the resources appropriate for routine alerts from those warranted by major outages. He also describes action permissions expanding as customers gain confidence, with investigation progressing toward proposed fixes subject to human review.

Schwartz’s contribution brings these concerns together in product behavior: an investigator that remembers how a team operates, explains why a signal matters, and joins the response at a useful moment. The engineering challenge extends beyond producing a plausible diagnosis to giving people findings they can understand and act on.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Eric Schwartz explains Traversal’s path from manual incident response to diagnosis across an enterprise and, eventually, verified fixes. The hard part is connecting a visible failure to a distant cause quickly enough to help the on-call team.

  • Faster code generation can increase troubleshooting work when production grows faster than engineers’ understanding of it.
    1:11 ↗
  • Enterprise root cause analysis requires connecting symptoms across services and teams; a failing checkout API may be several hops from its possible cause.
    4:27 ↗
  • Level four extends diagnosis across an environment. Level five adds fixes and verification, completing the proposed production loop.
    6:25 ↗
  • The Pepsi and American Express examples change the order of work: investigate alerts before asking engineers to triage them, and analyze incidents before paging broadly.
    12:06 ↗
  • Evaluate an AI SRE on data coverage, search cost and load, relationship mapping, autonomous knowledge maintenance, and fast multihop investigation.
    15:12 ↗