AI Engineer Code 2025

Stop Renting Your AI's Memory — Dylan Couzon, Qdrant

Read the talk

Stop Renting Your AI’s Memory

Dylan Couzon explains how local, searchable memory can give an assistant continuity across sessions—and demonstrates the mechanism with an offline drone-memory application built on Qdrant Edge.

From a talk by Dylan Couzon

At a glance

Ideas worth remembering

  • Owning inference supplies autonomy; owning persistent memory supplies continuity across conversations.

  • Memory needs write, retrieve and forget operations. Retrieval makes topic selection and changing relevance explicit instead of placing the entire record in every prompt.

  • The offline drone example turns detected pictures and labels into stored embeddings, then retrieves coffee-table sightings with images, timestamps and counts.

  • The proposed memory portability allows changing the reasoning model or device while retaining the same embedding model.

  • Shared memory should be a deliberate extension of separate local stores, with synchronization opt-in and users choosing who receives their memories.

A capable model still starts as a stranger

What happens to every preference, correction and dead end you teach an assistant when the conversation ends? Dylan Couzon, a developer relations engineer at Qdrant, opens with that practical frustration. An assistant can reason well during a session yet fail to carry the experience into the next one. His phrase for this combination is “geniuses with no long-term memory.” The missing capability is a persistent record that the assistant can use again.

Local inference makes that omission more conspicuous. Couzon describes a machine costing under $2,500 as capable of running roughly the previous year’s frontier, while explicitly stopping short of claiming the absolute frontier. This is his characterization of capability, rather than a benchmark comparison. Even accepting it, putting a capable model on a desk does not give it knowledge of its owner. It remains a “brilliant stranger” until preferences and experience survive between conversations.

That creates a second ownership question beyond where inference runs. Compute, model access, the software that runs an agent and the accumulated personal record can all depend on a provider. The last dependency grows with use: years of corrections and context make an assistant useful in ways that a fresh copy of its model cannot reproduce. Couzon frames the prospect of superintelligence around who will control that accumulated record.

0:121:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Inference gives autonomy; memory gives continuity

Rented inference introduces several dependencies that become especially visible when software runs long agent loops:

  • Access: a provider can withdraw access to a hosted model. Locally held weights remove that particular remote access dependency.
  • Version stability: providers can retire versions or throttle service. A local copy of the weights stays fixed until its owner changes it.
  • Economics: an agent working for hours can consume far more tokens than a short chat. Couzon cites developers using over $5,000 a month of compute through a $200 plan and predicts that such subsidies will disappear; that forecast is an argument for controlling compute costs, rather than a demonstrated pricing outcome.

Owning compute and weights addresses those inference dependencies. Memory solves a different problem: continuity. The next session needs access to what earlier sessions learned. A memory feature supplies that connection without requiring the underlying model to become more capable. An assistant can feel increasingly personal because it retrieves your corrections and preferences, even while its weights stay unchanged.

3:193:49
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:19 · section reference included

Write, retrieve and forget

Memory has three operations: write, retrieve and forget. Writing preserves something worth carrying forward. Retrieval selects what matters for the present task. Forgetting prevents every old observation from retaining equal importance forever. This changes the design question from how much text will fit in a prompt to which stored information should reach the model now.

A folder of Markdown files can hold a record, but dumping its contents into the prompt leaves selection to the model. An indexed retrieval system makes selection explicit. Topic filters can restrict the search; recency and frequency can influence relevance; older information can lose weight as circumstances change. Couzon names these as controls a memory system should support. He does not develop a particular deletion or decay algorithm in this recording.

The CPU, RAM and disk analogy separates the responsibilities. The model performs the computation. The context window holds the information available during the current interaction. Persistent memory keeps a record across interactions. A larger context window expands the temporary workspace; it does not by itself decide what to save, how to find it later or who controls the saved record.

The talk leaves the choice of the best memory architecture open. Its narrower design position is that the persistent record should belong to its user. Local embeddings and vector-search infrastructure provide one way to make that record searchable without placing it in a provider’s data center.

5:195:49
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:19 · section reference included

Put the searchable store inside the application

Qdrant Edge is the implementation used here: a vector-search engine embedded directly in an application. It opens a local store inside the application’s process and writes embeddings alongside payloads. The embeddings support similarity search; the payloads retain the information attached to each memory. Queries can run offline, while each application keeps its own isolated store.

Selected presentation frame from Stop Renting Your AI's Memory — Dylan Couzon, Qdrant at 458 secondsOpen full source frame
A Qdrant slide is titled “Qdrant, embedded in your process.”

The store persists after the application closes. Its folder becomes the object that can follow the user to another device. Couzon also says quantization can fit a million memories into less than a gigabyte, making local storage plausible on a phone or Raspberry Pi. The recording does not specify the vector dimensions, payload sizes or quantization settings behind that capacity claim, so it should be read as a configuration-dependent estimate.

7:257:55
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:25 · section reference included

From a first sighting to a searchable coffee table

The demonstration starts with an empty memory and recorded drone footage of a house. The memory application runs live on Couzon’s laptop, fully offline. YOLO, an object-detection model, identifies objects and produces labels. The application turns the pictures and labels into embeddings, then stores those representations as observations accumulate. The drone example therefore begins with no remembered objects and gains a searchable history as the footage plays.

At the reported point in the demo, the application has recognized 92 different objects and holds about 300 vectors. Multiple representations can belong to the same object; vector count is not object count. A two-dimensional view places similar memories near one another, giving a spatial view of the stored representations. The reported footprint for the memory and Qdrant engine together is 15 MB.

Selecting “coffee table” changes the interface from an accumulating memory to a set of retrieved sightings. Couzon reports that the query returns every coffee table previously seen in less than one millisecond. Each result carries an image, first-seen and last-seen timestamps, a definition and a sighting count. Those fields make the result useful: the application can show both what it remembered and when it encountered it. The timing describes this demo’s retrieval, not the latency of the complete detection-and-embedding pipeline or an arbitrary larger collection.

How does a passing image become something the application can recall? The flow below follows the coffee-table example from footage through detection and storage to retrieval. The important separation is between creating memories and querying existing ones: detection and embedding build the record, while search brings selected observations back with their associated details. The whole demonstrated path runs locally without a network connection.

How it fits togetherHow the application remembers a coffee table

The application starts with an empty memory and processes footage locally.

The write path accumulates observations; the retrieval path returns remembered sightings and their payloads.

7:558:25
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:55 · section reference included

Carry experience between agents, models and devices

Cloud synchronization adds a shared-memory option to the local store. A swarm of drones or robots could contribute observations to a “hive mind,” allowing other systems to benefit when one learns something. Local retrieval provides each device’s own memory; synchronization makes selected experience available beyond that device.

For a chatbot, agent or code assistant, the remembered observations become corrections, preferences and unsuccessful approaches. Retrieving them in a later task can save the user from repeating a correction or the agent from exploring the same dead end. This is the compounding mechanism behind Couzon’s claim that a merely good model can feel increasingly capable: useful experience remains available even when the reasoning model itself has not changed.

Portability has an important condition. The reasoning model and hardware can change while the memory folder moves with the user, but Couzon explicitly requires keeping the same embedding model. The stored representations depend on that model. His portability claim therefore covers changing the reasoner or device while preserving the embedding setup; it does not establish automatic portability across different embedding models.

9:5510:25
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:25 · section reference included

Extend retrieval from objects to a day

The next step extends the drone’s visual memory into everyday questions:

  • Finding an object: “Where did I leave my badge?” asks for a remembered observation associated with a place.
  • Recalling a detail: a daughter’s name or whiteboard contents asks for information encountered earlier.
  • Recovering an association: the song playing when two people met connects an event with something heard.
Selected presentation frame from Stop Renting Your AI's Memory — Dylan Couzon, Qdrant at 733 secondsOpen full source frame
A slide lists questions about a badge, a daughter’s name, whiteboard contents and a song.

Smart glasses supply the proposed route from occasional application memory to a record of what someone sees and hears. The drone demo establishes the object-memory example; the whole-day assistant is an extension of that idea, rather than another demonstrated system. Its usefulness would come from preserving and retrieving experiences that a person would otherwise forget.

This extension raises the stakes of ownership. A record that includes relationships, habits and encounters becomes an index of a person’s life. Couzon warns that AI providers are already building increasingly personal records. His proposed response is to let the user keep and control that index, rather than depend on continued access to a provider’s account.

11:4612:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:46 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Hi, everyone. Thank you for joining me today. I hope everyone is having a great conference so far. Um, so my name is Dylan Couzon. I am a developer relations engineer at Qdrant, and the title of my talk today is The Frontier is Coming Home. So what do I mean by that? Um, somewhere in the next few years, something smarter than all of us is going to exist. We've spent years asking when, but I think when is kind of the boring question. I think the real

  2. 0:42

    one is who it belongs to. Because there are two versions of this. In one, super intelligence l- lives in someone else's data center, and you pay by the token. In the other, it's yours. It learns from your life, your context stays private, and nobody can switch it off, throttle it, or take it away from you. That second version is what I want to talk about today. That's the frontier coming home.

  3. 1:11

    Quick thought experiment. What would you build if your assistant actually remembered everything you ever taught it? Every preference, every correction, every dead end for years. That's not a chatbot any- anymore. That's a second mind. But the, the assistant you're using right now forgets almost all of it the second the conversation ends. Every session starts from zero. We call these same things intelligence, and they are. They're just

  4. 1:41

    geniuses with no long-term memory. This talk is about fixing that, not with a bigger prompt, with actual memory. That genius used to only exist behind someone's API in a data center you'll never see. Now it runs on a-- on hardware that most of us own. Not the absolute frontier, I'm gonna be honest, but shockingly close. A machine under twenty-five hundred dollars today can run what was basically last year's

  5. 2:11

    frontier. And open-weights models are closing on fast. The gap is shrinking every quarter.

  6. 2:19

    You'd expect this to feel like a landmark, uh, the frontier on your desk for the price of a laptop. And yet it doesn't quite land. Because a model on its own is just a bri- brilliant stranger. It's powerful, it-- but it, it isn't yours. What make-- What makes it yours is everything it knows about you, and that part hasn't come home yet. So you can run a frontier class model on your own hardware today, and l- yet look

  7. 2:49

    at the AI stack that most of us are using today. The compute is rented, the models is rented, the harness is rented, and the part that's supposed to become you, the thing that should compound over the years, also rented. If we don't fix that, the most powerful technology any technology any, uh, any of us will ever touch ends up owned by whoever holds the lease, not the people it was built for. So let's look at what renting actually costs

  8. 3:19

    today. Starts where we-- starts with where the model lives because when you rent it, you don't really control it. This shows up in three different ways. Last month, a government order pulled access to two flagship models, Fable and Mythos-5, for every single customer. Now compare that to weights you already-- that are already sitting on your desk. Nobody can reach into your machine and flip those off.

  9. 3:49

    Even when the model, the model stays up, it doesn't stay put. Providers retire versions, and they throttle what's left when the economics get-- gets tight. The model y- you rely on can quietly get worse or vanish. Weights you own never change unless you change them. People also don't wr- don't write prompts anymore. They run agents or loops for hours, and one task can burn thousands of times more tokens than a chat message. Some developers

  10. 4:19

    burn through over five thousand dollars a month of, of compute on a two hundred dollars plan. That's a subsidy, and those subsidies will ultimately disappear.

  11. 4:31

    If you own the compute and the weights, the control comes back to you, but that's still only half the problem. Here's the deeper issue. Owning inference gives you autonomy. Nobody can take the model away. But owning memory gives you continuity. And continuity is exactly what every major lab is trying to package up and sell back to you. Look at last year. Every frontier lab shipped a memory feature, not because the models got smarter, but because they don't carry anything between

  12. 5:00

    sessions. That gap is the actual product. A system that starts simple and keeps learning you will always feel smarter than one that starts brilliant but forget everything about you. That compounding is where intelligence starts to feel personal.

  13. 5:19

    So what actually is memory? It's not a bigger prompt. It's a systems with three verbs: write, retrieve, and forget. Same three th- things your brain actually does. And, you know, there's an obvious pushback. Why not just throw everything in a big folder of Markdown files? And because retrieving the right memory beats dumping everything and praying the model finds it. Retrieval gives you something a prompt cannot.

  14. 5:49

    It is control. You can filter by topic. You can decay by recency and frequency. You can let relevance shift over time, the way human memory actually works. And the infrastructure has been ready for years. HNSW in 2016, um, embeddings have been ru- running locally since 2019, long before any, any model could really use them at scale. The memory side of super intelligence was n-

  15. 6:19

    never the hard part. We were just pointing it out at ourselves.

  16. 6:25

    So today, the memory that actually knows you lands in one of two places. Either it's stuck in an unindexed file you can't really query, or it's in someone else's cloud. One terms, one terms of service change away from re- being reshaped or moved. Carpe- Karpathy framed it in a way that's really stuck with me. It's the model is the CPU, the memor- uh, the context window is the RAM, and the memory is the disk. RAM is fast, but forgets,

  17. 6:55

    forgets the moment the session ends. The disk is what remembers you across every conversation. And right, right now, for almost every AI assistant on, on Earth, the, the disk is in someone else's data center. So, you know, today I'm not going to argue which memory architecture wins or is the best one. That's definitely a different talk. But the disk, the parts that holds the actual record of you, should be yours fully and forever.

  18. 7:25

    So this is the piece that I work on, and it's just one way to do it. Um, a vec- I work on a vector search engine that embeds directly into your app. It opens look, a local store inside your process. It writes embeddings with payloads, then query offline its sub-millisecond time. Same raw score as Qdrant in the cloud, just running where you are. With quantization, a million memories can fit in less than a gigabyte. Small enough to live on a phone or even a

  19. 7:55

    Raspberry Pi. Close the app, come back a year later, the memory's intact. Every app gets its own isolated sto- store, and that folder is the thing that actually follows you. So I'm gonna do a quick live demo here. Um, so right now, um, I have a video recording of a drone, like, going over a house. So basically, that drone, it's, it first time starting. It's, it's first time out in the world.

  20. 8:25

    It doesn't know anything. And so this, this is actually running live right now, fully offline on my laptop. And I'm using YOLO, which is an object detection model. So it detects the object and creates, um, label for every items it sees here. Then we take, take that picture, we take that label, and we turn that into embeddings to create a live real-time memory for that

  21. 8:55

    drone. And, and that way we can re- the drone can remember everything it has ever seen before. And so here you can see a 2D representation of that vector space with all the memories inside. And memories that are the most similar will, will be the closest to each other. So here you can see that we've already recognized 92 different objects, and we have o- over 3, 300 vectors. So that's 300 different

  22. 9:25

    representations of those 90 objects. And you can see that the entire memory and Qdrant engine, uh, footprints is only 15 me- megabytes. And then what we can do, uh, on those memory? Well, we enable, um, semantic search over those memories. So, you know, if I click here on the, on coffee table, you can see that in less than one millisecond, I was able to pull

  23. 9:55

    up every single coffee table that this drone has seen before. And so, you know, we have the image here, we have the first seen times, timestamp, we have the last seen ti- timestamp, we have a definition of that coffee table, and, and also we can see how many times that, uh, table has been seen. So this is all being run and processed locally. There's no network attachment at all. Um, but we also have a capability

  24. 10:25

    of, uh, cloud sync. So for example, you know, y- if you have, like, a swarm of drones or robots navigating the world, they can have, like, a hive mind in, in the cloud where a, you know, one system, uh, learns something, and they can all learn about it.

  25. 10:46

    And, you know, this is what... This was, like, an example for a drone, but you can apply that to, to everything. Just a, a chatbot, an agent that you called to, a code assistant. And, you know, if you create those memories for just a few weeks, the stranger just disappears. It remembers your corrections, your preferences, the dead ends you had to only hit once. And a model that's merely good but never forgets you and

  26. 11:16

    keeps compounding over the years starts to feel super human in practice. Not because the model changed, but because it never stops learning you. And it's portable. Swap the reasoning model, change the hardware, none of it matters. It all comes with you. If you keep the same embedding model and that folder becomes permanent. A lifetime of context fully yours, moving from device to device. But we can go even further than that. You know, this can

  27. 11:46

    go way past a laptop or a drone. Points the same memory at your whole day. Where did I leave my badge? What was her daughter's name again? What was on that whiteboard? And what wa- what was the song that was playing when we met? Every one of those is just a r- a retrieval query. And, you know, you just watched a drone remember objects. Now extend that to everything you see and everything you hear, and, you know,

  28. 12:16

    that's not science fiction. Those are the smart glasses that are already shipping today and that we have seen today in this room. It's not a smarter you, it's a you that doesn't forget. And right now, none of those memories share, uh, none of th- those devices share a memory that you actually own. They could on your own terms, because frontier models are already building an index of you today. They have your relationships.

  29. 12:45

    They have your habits. They, they have your whole inner life. That index is coming, whether you opt in or not. The only question is whether you own it or if you're renting access to yourself.

  30. 13:01

    And one last id- idea, because owning memory doesn't, doesn't mean keeping it to yourself. Edge can sync parts of your store to the Qdrant, Qdrant Cloud out of the box when you choose to. Picture a family, each wearing glasses that record their own day. Separate memories, but they can pull into a hive mind the whole family owns. Mom's day, Dad's day, your day, searchable together. Now scale that up. Instead of one company's

  31. 13:30

    super intelligence serving billions of identical people, you get thousands of small private ones, each shaped by a life, a family, a team. But that same mechanism cuts both ways. A memory can be shared without... A, a memory share can be shared with consent, or it can be extracted without it. That's exactly why it has to start local, and why sharing should always be opt-in, never the default. You should

  32. 14:00

    really decide who gets your continuity and your memories. And this is where I'm gonna leave you. You know, every- everyone here is going to watch super intelligence show up over the next few years. That's, that part is basically settled. What's not settled is who it belongs to, and that's the fight worth having right now while the arch- while the architecture is still up for grabs. So tonight, don't go home and just try a tool. Tr- take the stranger

  33. 14:30

    home and give it memory. Then picture that memory five years from now. Not a chatbot, a private mind that never forgets you, that nobody can switch off, and that knows, knows you better than any system ever has because you were the only one who ever trained it. That's not a s- that's not a smaller super intelligence. That's the only one worth wanting. All right. That's all for me. Thank you very much.