← All organizations

Semiconductor and AI infrastructure research

SemiAnalysis

SemiAnalysis provides research, industry models and consulting on semiconductors and AI infrastructure for AI companies, chipmakers, cloud providers and investors. Its AI Cloud TCO Model examines the economics of buying accelerators and selling GPU compute, while its accelerator and datacenter models track chip production and forecast IT power capacity. Networking and wafer fabrication models extend that coverage to interconnects and semiconductor equipment, connecting technical requirements with infrastructure economics.

Founder Dylan Patel remains CEO and chief analyst, having developed SemiAnalysis from a solo venture in 2020. Its public engineering work includes InferenceX, which benchmarks inference across accelerators and serving software with reproducible recipes, logs and results. Its AgentX workload replays the structure of long coding sessions, including subagents and repeated cache reuse, without exposing original user content. This lets engineers examine serving behavior beyond isolated prompts, including how routing and memory management affect performance.

The business combines subscription research with institutional models, retained advisory work and bespoke projects. In August 2026, SemiAnalysis reported operating its benchmark matrix across more than 1,000 chips and approximately 2 MW of compute, and spending over $3 million on the dataset underlying AgentX. It also entered a semiconductor research partnership with Tema ETFs in 2026, with a suite of funds planned.

semianalysis.com

2 talks

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Compute & System Design for Next Generation Frontier Models

    Start here for practical deployment concerns involving Llama 405B, vLLM, TensorRT-LLM, accelerator contention, and service reliability.

    Dylan PatelAI Engineer World's Fair 2024

  2. The Geopolitics of AI Infrastructure

    Use this talk to compare Huawei Ascend and CloudMatrix with NVIDIA Blackwell and NVLink, then follow the discussion into Middle Eastern compute investment.

    Dylan PatelAI Engineer World's Fair 2025

Messages from the stage

Prefill and decoding demand different resources

Patel contrasts compute-intensive prompt prefill with bandwidth-intensive token decoding, discussing continuous batching, disaggregated prefill, and context caching as serving-system considerations.

Compute expansion depends on supply chains and power

The geopolitics talk places accelerator architecture alongside high-bandwidth memory, export-control exposure, and electricity constraints, connecting hardware comparisons to the conditions surrounding infrastructure investment.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28