Compression at the Edge
Chris Alexiuk · Daniel Han · Asma Beevi · Merve Noyan · Parth Sareen
AI Engineer World's Fair 2026 · 46:01
Accelerated computing, AI infrastructure and graphics
NVIDIA develops accelerated computing hardware and software for AI, scientific computing, graphics and autonomous machines. Its CUDA platform lets developers use GPU parallel processing for scientific simulations and AI model development, while NeMo supports custom generative AI, including speech recognition and synthesis. GeForce RTX serves gamers and creators; Jetson and Isaac help teams develop and deploy robots and edge AI applications across manufacturing, logistics, healthcare and retail.
Founded in 1993 by Jensen Huang, Chris Malachowsky and Curtis Priem, NVIDIA began with a focus on 3D graphics for gaming and multimedia. Huang remains CEO. The company’s ray-tracing research also contributed to RTX hardware, while its DLSS technology uses AI to reconstruct high-resolution images from a fraction of the rendered pixels.
NVIDIA’s products span cloud, data-center, desktop and edge deployment. DGX Spark supports local AI applications, including personal agents that can switch between local and cloud models. For industrial software developers and manufacturers, its simulation tools support physically accurate digital twins for building, training and testing systems before deployment. Omniverse extends this work with camera, lidar and radar simulation that developers can integrate into existing applications.
Chris Alexiuk · Daniel Han · Asma Beevi · Merve Noyan · Parth Sareen
AI Engineer World's Fair 2026 · 46:01
Carter Abdallah · Vincent Weisser · Lucas Atkins · Chris Alexiuk
AI Engineer World's Fair 2026 · 43:21
Nader Khalil · Alex Cheema · Matthew Berman · Ahmad Osman · Joseph Nelson
AI Engineer World's Fair 2026 · 44:29
Walden · Carter · Tanay · Alex Atallah · Nav
AI Engineer World's Fair 2026 · 48:17
AI Engineer World's Fair 2026 · 20:36
AI Engineer Europe 2026 · 10:16
AI Engineer World's Fair 2025 · 16:41
AI Engineer World's Fair 2025 · 20:24
Travis Bartley · Myungjong Kim · Byungjoong · Jaehan
AI Engineer World's Fair 2025 · 16:24
Annika Brundyn · Aastha Jhunjhunwala
AI Engineer World's Fair 2025 · 17:47
AI Engineer World's Fair 2024 · 33:39
Affiliations reflect their AIE appearances, not necessarily current employment.
Start here to understand how request routing and GPU allocation affect the quality, latency, throughput, and cost tradeoff—and why reported gains depend on worker configuration.
Kyle KranenAI Engineer World's Fair 2025
The NVinfo routing case study illustrates when a smaller specialized model can match a larger model's accuracy, with a concrete workflow from feedback collection to redeployment.
Sylendran ArunagiriAI Engineer World's Fair 2025
Read for practical safeguards against familiar infrastructure failures, including short-lived credentials, least privilege, network segmentation, and protection of stored data.
Lovina DmelloAI Engineer World's Fair 2026
NVIDIA's Mark Moyou explains how to size and optimize production LLM inference deployments while controlling GPU costs.
Mark MoyouAI Engineer World's Fair 2024
Mark Moyou connects sequence lengths and KV-cache behavior to deployment sizing, while Kyle Kranen examines separating prefill and decode. Mozhgan Kabiri Chimeh brings these concerns to local hardware through reproducible benchmarks and memory-bandwidth constraints.
Sylendran Arunagiri describes curating ground truth and evaluating fine-tuned models from production interactions. Mitesh Patel addresses a different source of answer quality: combining graph relationships with vector retrieval and improving triplet extraction through data cleaning.
Using exposed Ray clusters, Lovina Dmello examines threats across model, data, supply-chain, and infrastructure layers. Her talk connects access controls and isolation to their latency and throughput costs.
Affiliations reflect each recorded session, not necessarily current employment.