Build Dynamic Products, and Stop the AI Sideshow
AI Engineer World's Fair 2025 · 18:10
AI evaluation, observability, and prompt management
Freeplay built a platform for developing and improving AI applications and agents, combining prompt management, evaluations, and production observability. Engineers, product managers, and domain experts could experiment with prompts and models, review outputs, and turn production traffic into test datasets. Online and offline evaluations connected testing with monitoring, helping teams assess changes before deployment and examine behavior in production.
Co-founders Ian Cairns and Eric Ryan previously worked together at Gnip and Twitter; a 2025 interview identified Cairns as CEO and Ryan as CTO. Freeplay supported two ways to manage prompts: teams could let the platform manage versions so non-engineers could change prompts and models without code changes, or retain their codebase as the source of truth and synchronize templates through the API. Platform-managed prompts could be fetched at runtime or bundled during builds.
Freeplay’s self-serve platform reached general availability in 2025, with customers including Chime, Help Scout, and Maze. Its documentation also described enterprise and self-hosting options. Freeplay’s reported planned shutdown was scheduled for May 15, 2026; completion of that shutdown and continued commercial availability remain unconfirmed.
AI Engineer World's Fair 2025 · 18:10
Jeremy Silva · Chris Hernandez
AI Engineer World's Fair 2025 · 12:50
Affiliations reflect their AIE appearances, not necessarily current employment.
Start here for Workday Help and employee self-service examples showing how AI capabilities can mature within enterprise products.
Eliza Cabrera · Jeremy SilvaAI Engineer World's Fair 2025
Learn how iterative evaluations, human review, hallucination monitoring, and feedback loops address the gap between prototypes and production reliability.
Jeremy Silva · Chris HernandezAI Engineer World's Fair 2025
The discussion with Eliza Cabrera connects customer-led discovery to integrated risk planning, aligned roadmaps, and incremental redesign of existing workflows.
The discussion with Chris Hernandez places customer-experience and quality-assurance teams in continuing roles as prompt testers, model reviewers, and AI performance monitors.
Affiliations reflect each recorded session, not necessarily current employment. These are joint discussions with speakers from Workday and Chime.