← All speakers

Bio, Work & Ideas

Sujee Maniyam

Conference affiliation: Developer Advocate · Nebius · 2026

On this page

Sujee Maniyam is a software engineer, educator, and entrepreneur whose work connects data infrastructure and AI with practical developer education. An AI Developer Advocate at Nebius in 2026, he builds applications and experiments that let developers inspect model behavior, understand performance tradeoffs, and judge tools through use.

From enterprise software to developer education

Maniyam’s early engineering work included Java-based workflow engines at Crossworlds and IBM, followed by web development and distributed data systems. His Hadoop consulting included building an advertising-data warehouse on Amazon’s cloud at Adpredictive, integrating Hadoop with Hitachi Data Systems’ storage infrastructure, and working on WANdisco’s Non-Stop NameNode product. These roles brought together application development and the infrastructure needed to operate large data systems reliably.

In 2013, he co-founded Elephant Scale, serving as principal consultant and providing big-data consulting and training. Teaching became a substantial part of his engineering practice. With Mark Kerzner, he co-authored Hadoop Illuminated and HBase Design Patterns, explaining distributed data processing and practical application design. He also developed Data Analytics with Spark and Hadoop, an O’Reilly course, and a guided machine-learning learning path supported by weekly sessions.

His developer-advocacy work later moved into deep learning and AI applications. He created educational materials for Intel’s BigDL distributed deep-learning framework and worked with MongoDB on vector-search examples and guides in 2023–2024. At the AI Alliance in 2024–2025, he contributed to data preparation and retrieval through upstream Data Prep Kit improvements and runnable examples, and helped create its Office Hours program. He joined Nebius in 2025, focusing on open models, coding agents, and inference through Token Factory developer resources.

Making model behavior tangible

Maniyam connects technical explanation with applications developers can run, inspect, and change:

  • LLM Snake Arena: His browser-based model competition began as a Nebius Token Factory Cookbook demonstration before becoming an independent project. Two models control competing snakes, each moving when its model responds. Latency therefore affects the game directly: a faster response creates another opportunity to move. Logs and latency graphs expose the behavior behind the action, while controls for reasoning, visibility, and movement hints let users explore how the surrounding setup affects results.
  • Practical LLM Evals: His evaluation collection brings together inference benchmarks, model visualizations, and interactive experiments. It complements his interest in how coding models handle repository understanding, debugging, refactoring, and longer tasks—and what completing those tasks costs. Numerical performance measures sit alongside behavior developers encounter while building.
  • Learnable retrieval systems: Maniyam contributed to AllyCat, an AI Alliance open-source website chatbot that makes retrieval-augmented generation understandable through a complete application. It collects website content, cleans and divides it into passages, creates embeddings, and retrieves material to support answers. Local and hosted configurations let developers examine how document processing, retrieval, and model inference fit together.
  • Distinguishing model optimizations: His LEGO-based explanation of fine-tuning, distillation, and quantization separates three often-confused operations: adapting behavior through training, teaching a separate student model, and representing numbers with fewer bits. He also explains the analogy’s limits. Quantization usually preserves the model’s architecture, and reduced numerical precision needs suitable hardware and software to translate into faster execution.

Coding agents with engineering judgment

Maniyam’s enthusiasm for coding agents includes close attention to the work they leave behind. When he migrated his website from Pelican to Hugo using Claude Code, the agent handled much of the conversion and customization. He intervened when it generated an outdated deployment workflow and scattered hardcoded styles across templates. His advice follows from those concrete problems: review generated work, ask for alternatives, and use small commits and branches to keep experimentation reversible.

Community organizing gives his teaching another setting. He founded the Big Data Gurus meetup earlier in his career and is part of the founding and organizing team of Silicon Valley GenAI. His books, courses, developer programs, and open-source applications share a practical teaching method: give people a working starting point, then make its behavior accessible enough to investigate and improve.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Dylan Bristot and Sujee Maniyam explain how Nebius Token Factory connects production data to model improvement, then walk through the hardware, routing, caching and decoding choices behind fast inference.

  • A production model needs a recurring path from inference logs through dataset preparation and post-training back to deployment.
    4:34 ↗
  • Serving the same model on different hardware and engines can produce different performance and cost; engine selection belongs to the model-specific optimization work.
    6:52 ↗
  • Cache-aware routing improves reuse by sending requests toward GPUs that already hold useful cached state.
    13:28 ↗
  • Speculative decoding pairs a fast draft model with a large verifier; custom training uses application traffic to make the draft more useful.
    14:23 ↗
  • Cache offloading preserves reusable work outside GPU memory, while prefill/decode separation assigns different resource demands to separate GPU groups.
    16:53 ↗
  • Quantization requires experiments to balance serving efficiency against model quality loss.
    18:12 ↗

References