▶ Watch ↗AI Engineer Code 202518:08
Laurent Gil is co-founder and president of Cast AI, which automates cloud infrastructure allocation and has extended that work into AI inference and coding agents. His work centers on the economics of useful computation: reducing wasted capacity and manual resource management so teams can accomplish more with their infrastructure budgets.
Gil’s early career included corporate finance and capital-markets work at Crédit Agricole, with assignments in New York, Tokyo, and Paris. During his MBA at Wharton, he co-founded a boutique investment bank in Brazil focused on Latin American telecommunications transactions. He also co-founded and served as chief financial officer of TAHO, a wireless internet provider in Rio de Janeiro. These ventures involved both financing telecommunications infrastructure and operating a connectivity business.
He later co-founded and led Viewdle, a Ukraine-based machine-learning and computer-vision company. Its video-search technology used facial recognition to index who appeared in footage, frame by frame. That made people within the video searchable rather than relying solely on surrounding text. Google acquired Viewdle in 2012.
Gil then co-founded Zenedge, which supplied web application firewalls and distributed-denial-of-service protection for cloud, on-premises, and hybrid environments. Oracle acquired Zenedge in 2018, and Gil subsequently worked on security product strategy at Oracle Dyn.
His security writing from that period examined a survey associating AI-assisted tools with faster incident detection and response. He also acknowledged that other security software correlated with improvements and that vendors sometimes marketed conventional algorithms as AI. His account distinguished a useful operational result from certainty about what caused it.
Running Zenedge supplied the problem behind Cast AI. Gil, Yuri Frayman, and Leon Kuperman had watched their cloud bill grow as the business scaled. They founded Cast AI to automate infrastructure optimization rather than leave teams continually adjusting resources by hand. Gil served as chief product officer before becoming president.
Cast AI’s initial focus was Kubernetes: selecting machines, placing workloads, and adjusting capacity as applications’ needs changed. Its expansion into AI infrastructure applies that approach to expensive hardware that can remain underused even as demand grows.
Gil’s writing on GPU workloads explains how autoscaling, bin-packing, and GPU sharing address this waste. Autoscaling adjusts capacity to demand; bin-packing places workloads efficiently across machines. Time-slicing lets workloads take turns using a GPU, while NVIDIA’s Multi-Instance GPU divides supported hardware into isolated compute and memory partitions. These mechanisms offer different ways to use existing capacity more effectively before adding hardware.
Gil’s work on AI costs extends resource optimization from the machines running models to the choice of models themselves. Three connected positions shape that work:
▶ Watch ↗AI Engineer Code 202518:08