← All organizations

AI software development agents

Bismuth

Bismuth developed an AI software developer designed to handle development tasks autonomously. Its foundation, Asimov Agents, was a Python framework for building agent systems around a graph execution engine, model inference and caching. Developers could define task dependencies, let models direct execution, preserve state through snapshots and trace execution with OpenTelemetry.

Co-founders Ian Butler and Nick Gregory ran the business together; Gregory served as CTO. The team also created SM-100, a benchmark for evaluating whether software engineering agents could find and fix bugs in real codebases. It measured discovery without hints, report accuracy, successful fixes and review of bug-introducing changes. Its single-attempt evaluations emphasized whether an agent could produce useful results without repeated runs.

Bismuth wound down operations and offboarded customers as Butler and Gregory joined XBOW. Butler said the business had paying customers, but growth fell short of the founders’ expectations. Its main API and agent implementation, codenamed Daneel, became a read-only public archive in 2025. The framework and benchmark document its engineering work, rather than establishing an ongoing commercial service.

bismuth.sh

2 talks

Newest first

2 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Agents reported thousands of bugs, how many were real? - Ian Butler and Nick Gregory

    Start here to understand how SM-100 uses validated bugs to compare agents on bug discovery and pull-request review.

    Ian Butler · Nick GregoryAI Engineer World's Fair 2025

  2. How to Improve your Vibe Coding — Ian Butler

    Watch this for guidance on checking whether reported bugs were actually fixed and understanding how noisy reports create developer alert fatigue.

    Ian ButlerAI Engineer World's Fair 2025

Messages from the stage

Maintenance demands program understanding

Butler and Gregory argue that reliable bug discovery requires targeted search, program comprehension, and reasoning across files. Their benchmark discussion examines gaps in how coding and security evaluations capture maintenance work.

Make vulnerability searches more specific

Butler recommends reasoning models, OWASP-informed and vulnerability-specific prompts, and careful context management as practical ways to improve coding-agent workflows.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28