Sheilah Kirui’s work connects GPU-accelerated data processing with the practical demands of serving language models. Her engineering background includes NVIDIA’s RAPIDS organization; she later presented inference work as a developer advocate at Akamai. Across these areas, she explains how software choices, GPU memory and workload patterns affect application performance.
From RAPIDS engineering to model serving
At NVIDIA, Kirui worked on cuDF development and testing, alongside testing and documentation for cloud deployments of RAPIDS libraries. Her May 2024 guide to GPU-accelerated data processing on Databricks explained how developers could bring GPU acceleration into familiar workflows: cuDF for pandas processing, the RAPIDS Accelerator for Apache Spark, and Dask for distributed computing. With cudf.pandas, supported operations run on the GPU while other operations fall back to CPU pandas, allowing developers to retain familiar code.
Her subsequent instruction on self-hosted inference applies that attention to deployment choices to language-model serving. Kirui and Khaja Omer’s joint dedicated-GPU inference workshop covered deploying vLLM and testing concurrent requests. Their work connects serving decisions to observable behavior: requests share GPU capacity, cached attention state consumes memory, and improvements in overall throughput affect individual response times.
When an inference optimization earns its cost
Kirui’s speculative decoding presentation makes a concrete case for workload-aware optimization. A smaller draft model proposes several tokens, and the larger target model verifies them in a single forward pass. Accepting enough draft tokens can accelerate decoding, but hosting the additional model and its KV cache also consumes GPU memory.
Her comparison using vLLM on a single NVIDIA Blackwell GPU showed why the workload matters. Structured output ran 1.6× faster with a high token acceptance rate; creative writing produced much lower acceptance. The result illustrates a conditional benefit rather than a general speedup: spare GPU memory, smaller batches and shorter contexts favor the technique, while long contexts and high concurrency reduce its gains. Kirui’s emphasis is on assessing the requests a system actually serves before committing memory and computation to an optimization.
A small model can draft several tokens before a large model checks them together. Sheilah Kirui’s side-by-side demo explains when that exchange saves time, when it wastes work, and why acceptance rate alone cannot decide.
Speculative decoding uses cheap sequential drafting and one target verification pass to advance several tokens when proposals are accepted.
The structured-output demonstration reported 1.6× faster generation with high acceptance; the creative case showed lower acceptance at high temperature.