On Thursday 7 May 2026, after a day at DevWorld, I went to an MLOps Community Amsterdam evening about memory for AI agents. This is a different evening from the AI sovereignty and GPT-NL special I wrote up before. The opening slides described the community as part of 75k members worldwide, with chapters in 35 cities and 150+ events.
The agenda slide listed three talks, a panel and networking:
- 6:25 to 7:00 PM: Ewa Szyszka, Qdrant, a deep dive into memory architecture, retrieval quality and evaluation in production AI systems.
- 7:00 to 7:20 PM: Chris Coutinho, Booking.com, âMemory for humans and agents: Nextcloud + MCP + Qdrantâ.
- 7:40 to 8:00 PM: Aryan Sharma, LangWatch, âYour Agent passes every eval and still breaks in productionâ.
- Then a panel with audience Q&A, and networking.

The agenda slide.
Qdrant: the silent tax on every AI agent
The first talk was titled âThe silent tax on every AI agentâ. The argument on the slides was about tokens. A chart compared tokens needed per correct answer against corpus size on three benchmarks (SWE-bench Lite, LongMemEval and FinanceBench). Long-context with caching and grep-style retrieval climbed with the corpus, while dense-vector and the Qdrant composition conditions stayed roughly flat.

Tokens per correct answer versus corpus size, on three benchmarks.
The speaker also covered what each Qdrant primitive contributes to recall, and what a single agent call would cost with several large models for a 200k-token context and 20k output tokens. These are vendor benchmarks, so treat them as the vendorâs own measurements.
Chunking methods and Recall@10
The slide I found most useful was âThe Effects of Chunking Methodsâ: chunking strategy impact on Recall@10 at a 200k corpus size, per benchmark, against a fixed-window 350-token baseline.

Chunking strategies compared on Recall@10, with a fixed 350-token window as the dashed baseline.
On the slide, function-level AST chunking led on SWE-bench Lite, and Chonkie semantic chunking led LongMemEval and tied for the lead on FinanceBench. Several strategies fell below the fixed-window baseline on a given benchmark. My take: this matches my experience, which is that the best chunker depends on the corpus type, so measure it on your own data before committing. A following slide checked model sensitivity across three embedding families and reported the same corpus-size trend.
Chris Coutinho (Booking.com): memory for humans and agents
The second talk was about Nextcloud, MCP and Qdrant. The slide said âI had everything in Nextcloud. I couldnât find anythingâ, then added semantic search through Qdrant, and then noted that agents generating documents made finding them the next bottleneck. The talk framed memory as four flavours.

Four flavours of memory: working, episodic, semantic, procedural.
- Working: what is in the prompt right now, volatile.
- Episodic: what happened, such as conversations and events.
- Semantic: what is true, such as facts and preferences.
- Procedural: how to do something, such as scripts and playbooks.
The slideâs point was that Nextcloud is already your episodic and semantic memory, so the agent inherits a year of you. A âWhy Qdrantâ slide listed a Rust core with predictable latency, payload filtering for multi-tenancy, and hybrid sparse and dense search.
Cliff versus slope
In a demo called âread the shapeâ, two query results were plotted by rank. A real match showed a cliff between ranks 4 and 5, while no match showed a smooth slope. The slideâs advice: watch the shape, not the top score.

Cliff versus slope: the score distribution says more than the top score.
The stale index and verify-on-read
A war story from preparing the talk: the speaker deleted a note and re-ran the query that had surfaced it. The note came back as the top result with score 0.7, all seven chunks ranked above any other document. In the slideâs words, Nextcloud said the note did not exist, Qdrant said it was the best match in the corpus, and both were telling the truth about their own state.

The stale index: a deleted note still came back as the top result.
The fix, on a later slide, was verify-on-read. Every search verifies sources before returning, which costs one Nextcloud round trip per note result, and webhooks still prune the index in the background. Ghost evictions become an observable signal. The slide reported that a re-test the day before the talk showed the deleted note gone six seconds after deletion. The closing line was the one I wrote down.

âAgents donât need bigger context windows. They need a librarian.â
My take: this is the failure mode I would test first in any RAG system. An index is a cache, and caches need invalidation, especially when deletion is a compliance requirement.
LangWatch: when every eval passes
The third talk, by Aryan Sharma of LangWatch, was about scenario testing for agents. One slide showed a suite of 15 scenarios with a 53% pass rate on day one, captioned âevery red tile is a bug I wouldâve shippedâ, and red-team scenarios such as prompt injection that caught real failures.

A scenario suite where the red tiles are the bugs that would have shipped.
The QR codes in the slides pointed to LangWatchâs open-source scenario repository on GitHub.