← All speakers

Bio, Work & Ideas

Dylan Couzon

Conference affiliation: DevRel Engineer · Qdrant · 2026

On this page

Dylan Couzon connects vector retrieval with practical questions: whether search finds the right product, whether an optimization earns its cost, and whether an AI assistant can remember useful information on hardware its user owns. A developer relations engineer at Qdrant in 2026, he turns those questions into technical guides, application examples, and teaching, including Building AI Assistants with On-Device Memory, developed with DeepLearning.AI and Qdrant.

From solutions engineering to retrieval

Earlier in his career, Couzon worked as a solutions engineer at Branch Metrics, helping customers with complex product use cases. That customer-facing experience preceded his work at Qdrant, where his technical writing and application examples explain how embeddings and retrieval support search, recommendations, and AI memory.

Making retrieval useful

Couzon examines search decisions through their consequences. A larger embedding model can leave an input-data problem untouched; personalization can introduce products the shopper never requested; an evaluation can reward agreement with an older model rather than better results.

  • Hybrid product search: In his account of Qdrant Shopping, Couzon explains a collaboratively built storefront that combines semantic intent with exact keyword matches. A shopper’s request can contain both a precise brand or size and an imprecise need, such as clothing suitable for cold-weather hiking. Dense embeddings and sparse keyword retrieval address different parts of that request. Filters constrain eligible products during search, while personalization reorders relevant candidates. Separating those responsibilities prevents preferences from pulling unrelated items into the results.
  • Better inputs before bigger models: The same storefront exposed a problem in which brand names overwhelmed the garment’s identity. Adding the product category to the embedded text corrected the drift across the models the team tested. Couzon uses this example to argue for inspecting what a system embeds before paying for a larger encoder. He also explains why evaluation methods need scrutiny: a fixed answer set built from an earlier model’s results can penalize useful changes, while a judge examining returned products cannot reveal relevant products the system never retrieved.
  • Retrieval and ranking as separate problems: His search-tuning guide starts by asking whether relevant documents entered the candidate set. Retrieving more candidates will not fix a ranking stage that continues to bury them. He recommends checking silent configuration errors and collecting enough labeled queries to distinguish improvement from noise. Expensive rerankers should earn their added latency against a properly tuned baseline.
  • Cheaper queries over a stable index: Couzon’s explanation of Constella, Qdrant’s research preview, describes how different query encoders can search one set of document embeddings. Smaller encoders learn to match Stella’s embedding space, letting developers change query-time compute without rebuilding the document index. The tradeoff is concrete: Zero pools token vectors cheaply but loses word order; Nano uses a small transformer to preserve contextual relationships. His account connects those design choices to their evaluation and practical use.

Memory that belongs to the user

Couzon’s work on local AI extends his retrieval concerns to ownership. He distinguishes the autonomy of running inference on owned hardware from the continuity supplied by persistent memory: an assistant’s accumulated information can remain dependent on a cloud provider even when its model runs locally.

His framework for AI memory treats the language model as a CPU, its context as RAM, and persistent memory as disk. Memory requires writing, retrieving, and forgetting information. Storing everything in a folder or inserting it all into a prompt leaves the central retrieval question unresolved: which experiences should an assistant bring into context for the task at hand?

Couzon teaches Building AI Assistants with On-Device Memory, developed with DeepLearning.AI and Qdrant. The course makes persistent memory concrete through multimodal information that an assistant can store and search locally. Recognizing a newly encountered object without retraining illustrates the distinction: the assistant can retrieve a remembered encounter rather than require its model weights to absorb every new experience. The accompanying on-device memory course repository gives developers a practical route into that work.

An offline drone demonstration using Qdrant Edge applies the same approach: the drone builds searchable memory with sub-millisecond queries in a reported 15 MB footprint. Couzon also advocates opt-in shared memory, keeping consent part of the design when information moves beyond the device. His work brings retrieval down to the scale of an individual application—and makes the ownership of its accumulated memory an engineering decision.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Dylan Couzon explains how local, searchable memory can give an assistant continuity across sessions—and demonstrates the mechanism with an offline drone-memory application built on Qdrant Edge.

  • Owning inference supplies autonomy; owning persistent memory supplies continuity across conversations.
    4:31 ↗
  • Memory needs write, retrieve and forget operations. Retrieval makes topic selection and changing relevance explicit instead of placing the entire record in every prompt.
    5:19 ↗
  • The offline drone example turns detected pictures and labels into stored embeddings, then retrieves coffee-table sightings with images, timestamps and counts.
    8:25 ↗
  • The proposed memory portability allows changing the reasoning model or device while retaining the same embedding model.
    11:16 ↗
  • Shared memory should be a deliberate extension of separate local stores, with synchronization opt-in and users choosing who receives their memories.
    13:01 ↗

References