← All speakers

Bio, Work & Ideas

Žilvinas Urbonas

On this page

Žilvinas Urbonas’s engineering work spans cloud-resource optimization and the systems that help coding agents complete software tasks. A founding engineer and engineering manager at Cast AI, he led engineering for Kimchi at the time of its World’s Fair 2026 presentation. Across these areas, his work addresses how software allocates resources, responds to changing demand, and recovers from interruption.

Matching cloud resources to the work

Urbonas’s background includes more than a decade of full-stack engineering in cloud-native systems. By 2023, his work at Cast AI combined engineering with team leadership.

In 2021, he co-authored a KEDA autoscaling guide with Annie Talvasto and Tom Kerkhove. Their batch-processing example used queue length to determine how many worker pods were needed, ran those workers on spot instances, and removed empty nodes when processing finished. Interrupted jobs could be rescheduled elsewhere. The design connected application demand to infrastructure allocation while accounting for the interruptions that come with cheaper compute.

His subsequent article on GPU workload scheduling addressed a related problem: successfully placing a workload does not guarantee economical use of infrastructure. He explained how node templates specify hardware requirements and exclusions, allowing Cast AI’s autoscaler to select suitable instances and remove GPU capacity after jobs end. The customer example illustrated a team and product accomplishment.

Urbonas also treats cost management as a leadership responsibility. In his writing on engineering teams and cloud spending, he argues that engineers need a reason to care about costs, visibility into what drives them, and tools that handle repetitive optimization. Resource tags make spending attributable, monitoring catches problems early, and automation puts optimization decisions into practice. These mechanisms make cost control part of everyday engineering without requiring constant manual intervention.

Building the system around coding models

Kimchi extends resource allocation into coding-agent workflows. In their joint presentation, Urbonas and Cast AI co-founder Laurent Gil presented an open-source harness that chooses proprietary or open models according to task outcomes. Their organizing measure was cost per completed task: token prices alone do not capture what it costs to get useful work done.

Two capabilities of the Kimchi harness explain how the system supports that approach:

  • Model selection within the workflow: Kimchi separates planning, implementation, review, codebase exploration, and external research into roles. An orchestrator assigns work to suitable models, uses lighter options where appropriate, and escalates difficult tasks. Review receives a fresh context even when the same model handles other stages. This lets model selection respond to the work at each stage instead of imposing one model choice on every operation.
  • Recoverable long-running work: Ferment preserves project goals, phases, steps, decisions, and working memory across sessions. Validated state transitions constrain how work advances, while persisted state lets a project resume after a crash or closed terminal. The system retains both its progress and what remains to be done, supporting coherent continuation after interruption.

This work carries forward concerns visible in Urbonas’s cloud engineering: allocate resources according to demand, automate repetitive decisions, and design recovery into the system. In coding agents, those concerns shape the orchestration and persistent state that support sustained software development.

Read the topics behind these talks

1 conference talk

Key ideas

Kimchi combines model selection with an outcome-checking coding loop, then moves long-running sessions into remote sandboxes and a shared team board. The aim is to lower the cost of completed work while letting engineers use more tokens.

  • Compare models on the cost of completing the same task at the required quality; token prices alone do not describe that cost.
    3:00 ↗
  • Ferment pairs model selection with milestones, build-and-repair cycles and quality scoring. Its current autonomous delivery stops at staging.
    8:30 ↗
  • Teleport separates the lifetime of an agent run from the engineer’s laptop; Studio makes those remote sessions, plans and review requests accessible to a team.
    12:01 ↗