← All organizations

Continual learning and AI training infrastructure

Trajectory

Trajectory builds a continual learning platform for companies developing AI products. Its SDK captures traces, corrections, re-prompts and edits so teams can identify failures and direct improvements to models, prompts and agent harnesses. Customers control which data enters training, and model updates pass through their evaluation suites and approval workflows before reaching production.

The company was cofounded by CEO Ronak Malde, formerly a researcher at Windsurf and Google DeepMind, alongside former Apple researcher Arjun Karanam and former Google DeepMind robotics researcher Michael Elabd. Its research extends Self-Distillation Policy Optimization (SDPO) to off-policy data: individual interaction traces that reach training after the model that generated them has changed. This addresses a practical obstacle to learning from asynchronous product usage rather than requiring freshly generated training examples.

Customers include Clay in sales, Decagon in customer support and Harvey in legal AI. Alongside model customization, Trajectory helps customers tune the software harnesses around closed-source models to improve accuracy or reduce operating costs. In August 2026, it raised a reported $40 million in financing led by Sequoia Capital at a $300 million post-money valuation.

www.trajectory.ai

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Token-level teaching versus scalar rewards

Malde compares SFT, DPO, RLHF, and GRPO before presenting on-policy self-distillation as a richer teacher-student alternative to scalar reinforcement-learning rewards.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28