← All organizations

Accelerated computing, AI infrastructure and graphics

NVIDIA

NVIDIA develops accelerated computing hardware and software for AI, scientific computing, graphics and autonomous machines. Its CUDA platform lets developers use GPU parallel processing for scientific simulations and AI model development, while NeMo supports custom generative AI, including speech recognition and synthesis. GeForce RTX serves gamers and creators; Jetson and Isaac help teams develop and deploy robots and edge AI applications across manufacturing, logistics, healthcare and retail.

Founded in 1993 by Jensen Huang, Chris Malachowsky and Curtis Priem, NVIDIA began with a focus on 3D graphics for gaming and multimedia. Huang remains CEO. The company’s ray-tracing research also contributed to RTX hardware, while its DLSS technology uses AI to reconstruct high-resolution images from a fraction of the rendered pixels.

NVIDIA’s products span cloud, data-center, desktop and edge deployment. DGX Spark supports local AI applications, including personal agents that can switch between local and cloud models. For industrial software developers and manufacturers, its simulation tools support physically accurate digital twins for building, training and testing systems before deployment. Omniverse extends this work with camera, lidar and radar simulation that developers can integrate into existing applications.

www.nvidia.com

14 talks

Newest first

19 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Hacking the Inference Pareto Frontier

    Start here to understand how request routing and GPU allocation affect the quality, latency, throughput, and cost tradeoff—and why reported gains depend on worker configuration.

    Kyle KranenAI Engineer World's Fair 2025

  2. Effective AI Agents Need Data Flywheels, Not The Next Biggest LLM – Sylendran Arunagiri, NVIDIA

    The NVinfo routing case study illustrates when a smaller specialized model can match a larger model's accuracy, with a concrete workflow from feedback collection to redeployment.

    Sylendran ArunagiriAI Engineer World's Fair 2025

  3. Your LLM Stack Is a 2008 Database With Better Marketing

    Read for practical safeguards against familiar infrastructure failures, including short-lived credentials, least privilege, network segmentation, and protection of stored data.

    Lovina DmelloAI Engineer World's Fair 2026

  4. Mastering LLM Inference Optimization: From Theory to Cost-Effective Deployment

    NVIDIA's Mark Moyou explains how to size and optimize production LLM inference deployments while controlling GPU costs.

    Mark MoyouAI Engineer World's Fair 2024

Messages from the stage

Performance depends on the workload

Mark Moyou connects sequence lengths and KV-cache behavior to deployment sizing, while Kyle Kranen examines separating prefill and decode. Mozhgan Kabiri Chimeh brings these concerns to local hardware through reproducible benchmarks and memory-bandwidth constraints.

Improve the data and retrieval pipeline

Sylendran Arunagiri describes curating ground truth and evaluating fine-tuned models from production interactions. Mitesh Patel addresses a different source of answer quality: combining graph relationships with vector retrieval and improving triplet extraction through data cleaning.

Security controls have serving costs

Using exposed Ray clusters, Lovina Dmello examines threats across model, data, supply-chain, and infrastructure layers. Her talk connects access controls and isolation to their latency and throughput costs.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-27