← All organizations

AI agent simulation and evaluation

Arklex AI; Columbia University

Arklex AI develops tools for agent evaluation, helping teams test conversational agents before deployment. Its ArkSim product generates synthetic users with distinct profiles, goals and knowledge levels, then simulates multi-turn conversations. Developers define scenarios and receive interactive reports with response scores, categorized failures and conversation transcripts. The platform helps identify context loss, tool misuse and policy violations, and lets teams set production-readiness gates.

Founded in 2023 as a spin-off from Columbia University’s Natural Language Processing Lab, Arklex is a company distinct from the university. Its co-founders are Zhou Yu, CEO and a Columbia computer science professor, and Arbit Chen, CTO. Its earlier agent-building work included the Arklex AI Agent Framework, launched in 2025 to combine structured workflows with autonomous reasoning. E-commerce applications used customer and product data to answer shoppers’ questions across chat, email, phone and SMS.

ArkSim fits into developers’ existing testing workflows: it connects through Chat Completions endpoints, the A2A protocol or direct Python imports, and can run in continuous integration with failure thresholds. Teams can run it on their own infrastructure. Its practical focus is generating test conversations as well as scoring them, allowing developers to inspect how an agent behaves through follow-up questions and changing conversational context.

arklex.ai

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Evaluate more than task completion

Yu emphasized realistic multi-agent benchmarks that consider efficiency and security alongside whether agents complete their tasks.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28