
18:48
Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford
John Ousterhout
AI inference and agentic workloads shift networking needs from bulk throughput toward small-message latency · Tail latency in synchronization exchanges can stall distributed computation and waste GPU capacity · Incast creates receiver-side switch queues that delay short messages and can trigger packet loss