← All speakers

Bio, Work & Ideas

Laurent Gil

Conference affiliation: Cast AI

On this page

Laurent Gil is co-founder and president of Cast AI, which automates cloud infrastructure allocation and has extended that work into AI inference and coding agents. His work centers on the economics of useful computation: reducing wasted capacity and manual resource management so teams can accomplish more with their infrastructure budgets.

From finance and connectivity to computer vision

Gil’s early career included corporate finance and capital-markets work at Crédit Agricole, with assignments in New York, Tokyo, and Paris. During his MBA at Wharton, he co-founded a boutique investment bank in Brazil focused on Latin American telecommunications transactions. He also co-founded and served as chief financial officer of TAHO, a wireless internet provider in Rio de Janeiro. These ventures involved both financing telecommunications infrastructure and operating a connectivity business.

He later co-founded and led Viewdle, a Ukraine-based machine-learning and computer-vision company. Its video-search technology used facial recognition to index who appeared in footage, frame by frame. That made people within the video searchable rather than relying solely on surrounding text. Google acquired Viewdle in 2012.

Gil then co-founded Zenedge, which supplied web application firewalls and distributed-denial-of-service protection for cloud, on-premises, and hybrid environments. Oracle acquired Zenedge in 2018, and Gil subsequently worked on security product strategy at Oracle Dyn.

His security writing from that period examined a survey associating AI-assisted tools with faster incident detection and response. He also acknowledged that other security software correlated with improvements and that vendors sometimes marketed conventional algorithms as AI. His account distinguished a useful operational result from certainty about what caused it.

Automating cloud resource decisions

Running Zenedge supplied the problem behind Cast AI. Gil, Yuri Frayman, and Leon Kuperman had watched their cloud bill grow as the business scaled. They founded Cast AI to automate infrastructure optimization rather than leave teams continually adjusting resources by hand. Gil served as chief product officer before becoming president.

Cast AI’s initial focus was Kubernetes: selecting machines, placing workloads, and adjusting capacity as applications’ needs changed. Its expansion into AI infrastructure applies that approach to expensive hardware that can remain underused even as demand grows.

Gil’s writing on GPU workloads explains how autoscaling, bin-packing, and GPU sharing address this waste. Autoscaling adjusts capacity to demand; bin-packing places workloads efficiently across machines. Time-slicing lets workloads take turns using a GPU, while NVIDIA’s Multi-Instance GPU divides supported hardware into isolated compute and memory partitions. These mechanisms offer different ways to use existing capacity more effectively before adding hardware.

Measuring AI by completed work

Gil’s work on AI costs extends resource optimization from the machines running models to the choice of models themselves. Three connected positions shape that work:

  • Automatic model selection: Gil argues that developers should describe the result they need while software chooses suitable models for the tasks involved. Kimchi Coding, which reached general availability in July 2026, applies this approach through task-sensitive model routing and output-scoring feedback. Its open-source terminal agent assigns distinct model roles to planning, implementation, review, exploration, and research. Gil advocates measuring cost per completed task: the harness chooses among proprietary and open models based on outcomes, making the cost of useful work the relevant comparison. Kimchi is developed by the Cast AI team, with Žilvinas Urbonas leading its engineering; Gil brings the perspective of the company’s co-founder and president to its approach to AI costs.
  • Token access as developer enablement: In his argument for continuous AI access, Gil objects to cutting developers off midway through reasoning or debugging. He wants optimization to absorb more of the cost-management burden so developers can continue working. He also challenges the decision to stop hiring junior employees because AI can perform entry-level work: organizations still need people to enter the profession and develop into experienced practitioners. His case for cheaper AI includes giving people room to learn as well as produce.
  • AI economics from infrastructure to outcomes: Gil treats token spending as part of a chain: hardware produces tokens, applications consume them, and completed work creates value. His tokenomics framework connects GPU rightsizing, capacity routing, and autoscaling for bursty agent workloads with governance of AI consumption. He serves on the FinOps Foundation governing board and joined the Tokenomics Foundation governing board in 2026. His stated priority there is to connect the infrastructure cost of producing AI with the business outcomes it delivers, bringing that measurement problem into shared industry standards.

Read the topics behind these talks

1 conference talk

Key ideas

Kimchi combines model selection with an outcome-checking coding loop, then moves long-running sessions into remote sandboxes and a shared team board. The aim is to lower the cost of completed work while letting engineers use more tokens.

  • Compare models on the cost of completing the same task at the required quality; token prices alone do not describe that cost.
    3:00 ↗
  • Ferment pairs model selection with milestones, build-and-repair cycles and quality scoring. Its current autonomous delivery stops at staging.
    8:30 ↗
  • Teleport separates the lifetime of an agent run from the engineer’s laptop; Studio makes those remote sessions, plans and review requests accessible to a team.
    12:01 ↗