George He works on the infrastructure that turns complex documents into usable inputs for AI systems. As Head of Engineering & Legal at LlamaIndex in September 2026, he co-authored an account of reliable document orchestration that connects document understanding with practical engineering concerns: recovering interrupted jobs, allocating shared computing resources, and containing unpredictable model behavior. His background spans cloud infrastructure, machine-learning research, financial technology, and intellectual-property law.
From distributed systems to financial workflows
He graduated from Berkeley in 2016 after studying electrical engineering and computer science, then completed a computer-science master’s focused on artificial intelligence at Stanford in 2019. At Berkeley’s AMPLab, he worked on distributed systems and bioinformatics and co-authored research on Mango, a tool for exploring large sequencing datasets. At Stanford’s Partnership in AI-Assisted Care, he developed computer-vision models to identify patient actions from low-resolution depth images. His published résumé documents these research projects alongside his production engineering work.
His collaborative research explored how machine learning could work with constrained inputs and resources. With Sami Oueida and Tucker Ward, he investigated real-time gaze estimation using ordinary webcams, limited computing power, and minimal training requirements. With Daylen Yang and Kelly Shen, he studied transfer learning between Atari games, testing whether reinforcement-learning systems could use knowledge from one game to improve learning in another.
From 2017 to 2019, He worked as a technical lead at Google, leading engineering for a containerized cloud-monitoring product and working on backend architecture and APIs. He subsequently became a senior software engineer at Amazon, where his responsibilities included Prime Video’s international expansion and a new global video-catalog backend. In 2021, he moved to Valon Mortgage, leading its investor-reporting platform and work involving reporting, remittances, and regulatory data requirements. These roles took him from large cloud platforms to a startup processing consequential financial workflows.
He also studied at Harvard Law School, advised early-stage businesses through its entrepreneurship program, and gained experience in intellectual-property litigation. His professional interests include AI policy and improving access to law. By September 2026, his role at LlamaIndex combined engineering and legal responsibilities.
Engineering dependable document AI
He’s technical writing explains how document AI can fail in production and what engineers can do about it. The problems extend beyond extraction accuracy: a slow or pathological document-processing job can consume shared resources, delay other customers, or leave work stranded after an interruption.
Document parsing as distributed work: Complex documents can mix clean text, scanned tables, handwriting, and schematics, requiring different processing paths. In his joint account with Adrian Lyjak, He describes LlamaIndex’s move from increasingly elaborate RabbitMQ orchestration to Temporal. Persisted workflow state lets interrupted work resume, while coordinated permits limit concurrent jobs and access to shared resources. Separating permit-management workers from document-processing workers prevents overloaded processing from delaying permissions and deepening the backlog.
Different failures need different remedies: His analysis of vision-language-model OCR failures distinguishes repetition loops, which consume time and resources, from provider content filters, which can abruptly interrupt legitimate extraction. Execution limits, repetition detection, and whitespace cleanup address runaway generation; explicit error handling and controlled retries address blocked or truncated responses. These safeguards also protect other customers from the effects of one problematic job.
Generated SDKs still require interface design: With Yong Park, He argues that generating client libraries does not remove responsibility for how developers use them. Their account of SDK design and migration explains why method names, types, pagination, exceptions, and language conventions matter independently of the API specification. A change of generator can alter the interface customers depend on even when the server API stays the same.
Search that helps agents investigate
He advocates hybrid retrieval with file navigation: give agents complementary tools and let them choose how to investigate a document collection. His work on LlamaParse Index’s retrieval approach combines semantic search with directory listings, metadata filtering, grep, and direct reading. Search narrows the territory; navigation and reading let the agent inspect surrounding material when a retrieved fragment is incomplete. Parsed text and page images provide different information, especially when tables or visual layouts resist plain-text extraction.
In “Grep or Embeddings? Agentic Search Over Company Documents”, He explains why the right approach depends on the collection. Codebases are often local, text-based, and rich in structural clues. Company documents can be much larger, multimodal, and subject to access permissions and freshness requirements. His position is to give the agent both indexed retrieval and tools for closer inspection, rather than make one search method carry every task.
A demonstration using 135 Alphabet financial filings makes that approach concrete: an agent builds a cash-flow table grounded in the source files. The example connects retrieval to the work it enables—finding relevant documents, inspecting their contents, and assembling an answer readers can trace back to the underlying material.
George He explains how company-document agents can use hybrid retrieval to find a direction, then list, filter, grep and read remote files to build answers grounded in their contents. The hard work includes preserving document layout, keeping permissions current and avoiding repeated ingestion.
Direct file search works well when the corpus is manageable, textual and navigable. Large, mixed-format company collections make pre-indexed retrieval useful for avoiding repeated exploration.
A document-search harness can combine hybrid retrieval with listing, metadata filtering, grep and reading, letting the agent expand or narrow its investigation after the first search.
Prepare structured text and page screenshots together. Preserving table structure and layout matters because downstream models must reason from the representation they receive.
Production freshness includes permissions and metadata as well as content. Ingestion, parsing and indexing are distinct stages with different synchronization work.
Remote grep and bounded reads give an agent familiar file operations without downloading the whole corpus. Connecting a synthesized answer to specific source files makes it easier to check.