At the end of November 2024 I spent two days in a row with the Elastic community in Amsterdam. On Monday 25 November I went to the Elastic Meetup Amsterdam, the evening before ElasticON. It had three talks: Elastic on vector search, Kaufland on running e-commerce search, and Devoteam on observability for large Elastic Cloud Enterprise (ECE) platforms. On Tuesday 26 November I went to ElasticON Amsterdam at the Beurs van Berlage.
This post is a throwback built from the photos I took. Iâve only included what was on the slides and signs. At the end I add my own view on running quantized vector search in production.
Steve Kearns: the secret sauce of faster, better vector search
The first talk was by Steve Kearns, GM Search Solutions at Elastic according to his title slide. The talk was called âThe Secret Sauce of Faster, Better Vector Searchâ. An early slide called Elasticsearch âthe most widely deployed vector databaseâ. Thatâs Elasticâs own claim. The talk then followed vector search through the 8.x releases.
![]()
Steve Kearnsâ title slide at Elastic Meetup Amsterdam, 25 November 2024.
int8 scalar quantization, on by default
The âNext Phase of Scalar Quantizationâ slide had three columns:
- Float32 to int8 in ingest. The slide gave the formula int8 â 127 Ă (float32 â min) / (max â min) and listed lower cost, improved performance and ânegligible ranking impactâ.
- âŠon by default. About three quarters less index size and cost, better query latency, better ingest performance, and all of it âdone within Elasticsearch (no effort)â.
flatandint8_flatindex types for small sets after a filter. When a filter leaves only a few documents, brute-force vector search works best.flatis âsyntactic sugar for brute forceâ, andint8_flatadds automatic scalar quantization to it.
The next slide compared the trade-offs, using the same rabbit drawing at each precision. float32 keeps high recall and precision but needs all its vectors in RAM. int8 gives 4Ă RAM savings with good recall and moderate oversampling, and an arrow labelled it the âElasticsearch 8.14+ defaultâ. int4 gives 8Ă savings, but recall drops and oversampling becomes necessary. A single bit per dimension gives 32Ă savings with âbadâ recall and precision.
![]()
float32, int8, int4 and bit: recall, precision, oversampling and RAM savings, with int8 as the 8.14+ default.
Elasticâs documentation agrees. The Elasticsearch 8.14 release blog and the dense_vector field reference say that new dense_vector float fields use int8_hnsw by default from 8.14. Indices created on earlier versions keep plain hnsw.
Better Binary Quantization: 32Ă less RAM
Next came a slide titled âGood, but could it be better?â with a chart: memory in MB for 500k vectors of 1,024 dimensions. float32 needed 2,133.5 MB, int7 548.3, int4 292.3 and bit 66.7. The slide asked â100M vectors? Only 12GB!?!â and answered âToo good to be true!â
That set up Better Binary Quantization (BBQ). The slide said âBBQ: 32X RAM savings. Faster & more accurate than Product Quantizationâ. In the rabbit drawings, BBQ looked much closer to the original than plain bit did.
![]()
From scalar quantization to BBQ: 32Ă RAM savings, according to the slide.
The following slides asked the obvious questions in turn: âBut, how long to index?â, âOK, but recall?â and âOK, but what about larger scale? Does it still work?â. The answers were benchmark charts on E5-Small and Cohere v3 embeddings. The summary slide, âMore than a cool nameâ, compared BBQ with product quantization (PQ):
- CohereV3, 30M vectors. PQ took 13Ă longer to quantize than BBQ. Brute-force search at 98%+ recall was 3Ă slower with PQ.
- E5-Small, 500k vectors. PQ took 56Ă longer to quantize and searched about 2Ă slower.
- The two takeaways on the slide were â10â50x faster quantization speedsâ and â2x or better query speedsâ.
![]()
BBQ against product quantization: quantization time and brute-force search at 98%+ recall.
Elasticâs Elasticsearch 8.16 release blog, published on 12 November 2024, introduced BBQ, so the feature was about two weeks old at the time of the meetup. The blog describes up to 32Ă compression and âover 95%â less memory. The Search Labs post on BBQ in Lucene and Elasticsearch explains how it works: the stored vectors are single bits, the queries are quantized to int4, and the method builds on the RaBitQ research. Since then, the default has moved again. Elasticâs 9.1 post made BBQ the default quantization for dense vectors of 384 dimensions or more.
Hardware, retrievers and ES|QL
The âHardware Optimizationâ slide covered three things:
- Panama-accelerated hardware instructions, with about 30% better query latency and indexing performance âif hardware supports itâ.
- Fused multiply-add for dot products.
- A vector distance function optimised with SIMD: 3â6Ă faster, in native code, and for int8 on ARM AArch64 processors âunder certain conditionsâ.
Then came retrievers and RRF, which the slide marked GA. The diagram had three stages. A standard semantic query and a knn query each return a result set. Reciprocal rank fusion (rrf) combines and ranks them. A text_similarity_reranker then reranks the combined list semantically. The example query was âwhy are retrievers fun?â.
![]()
Retrievers: initial result sets, RRF fusion, then a semantic reranker, all in one request.
The talk ended with the ES|QL roadmap in three groups:
- Search, embeddings and RAG: full-text search, reranking, inference and multi-step querying.
- Time series, metrics and observability: time series data streams (TSDS) and a metrics command.
- Aggregations that keep the data:
STATSandINLINESTATS.
Kaufland: e-commerce search and the road to Elasticsearch 8
The second talk came from Kaufland e-commerce. The opening slide described it as an online B2C marketplace with 11k sellers, more than 800 employees and five storefronts (de, cz, sk, pl, at). The speakerâs name wasnât on the slides I photographed, so I wonât name them.
The architecture slide had three services: an Indexer, a Monolith and a Searcher. Around them were Kafka, BigQuery and Elasticsearch. The indexer reads from Kafka and BigQuery and writes to Elasticsearch, and the monolith and the searcher read from Elasticsearch.
![]()
Kauflandâs search architecture: an indexer, a monolith and a searcher around Kafka, BigQuery and Elasticsearch.
The talk went through the phases of building relevance:
- Analysis. They used âsearch learnerâ data on how users interacted with search terms, and looked at shorter queries by asking âwhat are the important parts of this query?â.
- Search. A
filterremoves invalid products and handles shorter queries. A term-centricmustquery is multiplied withfunction_score, usingfield_value_factorfor product rank andscript_scorefor search learner rank. - FTS vs. signals. The slide said the two are âtough to balance with a single queryâ and that signals are strong in the bigger storefronts. It also asked: âDo you really need FTS scoring, or just boolean matching?â
Then there was a quiz on minimum_should_match.
The second half was the journey to Elasticsearch 8. The timeline went from 7.12.1 in January to 7.17.22 in March and then 8.14.1 in May. On the way they hit:
- Node caches filling up, which pushed query latency up. A cron job to clean the caches worked around an unfixed bug in 7.17, but it introduced another bug that needed restarts.
- Unreleased shard locks. A disk ran full because other shards were being moved around, and the locked shards couldnât be moved away. The fix was the 8.x upgrade.
- Less scripting. âCheck if scripting is needed, if not: remove!â
- Unexpected refreshes.
refresh_intervalwas set to 5m, yet there were dozens of refreshes within five minutes. The reason: the Update API calls a realtime GET, and when the translog location isnât available, that GET triggers a refresh. - Setting
eager_global_ordinals, and a range query vs. term query difference that âseems to have been fixed in 8.xâ. - Tracking total hits. On a hot dataset with
must/should/filterclauses and 11k results,track_total_hits: falsetook 400 ms andtrack_total_hits: truetook 200 ms. The slideâs verdict was âWTF?â, and the explanation was dynamic pruning not working as expected withrank_feature.
![]()
Turning track_total_hits off made this query twice as slow: dynamic pruning was not doing what they expected.
The last technical slides covered switching off a microservice to drop one gRPC network hop. They also showed a separate cluster upgraded from 7.9 to 7.17, after which there was âmuch less GC, CPUâ and some nodes could be removed.
The ElasticON agenda the next day had a Community Track session called âFrom Keyword Search to Data Science: Search evolution at Kaufland e-commerceâ, so the story continued there.
Devoteam: observability at scale with ECE
The third talk was âObservability at Scale â With ECE managed elasticsearch deploymentsâ from Devoteam. Patrick Broek, Lead Consultant, presented the observability part. The title slide listed a second Devoteam consultant as co-presenter, and a second presenter took over for the automation part.
![]()
Devoteamâs âObservability at Scaleâ with ECE-managed Elasticsearch deployments.
The talk started with the ECE high-level architecture. The control plane holds the Cloud UI, the Admin API, ZooKeeper plus the Director, and the Constructor. Proxies sit behind a user-supplied load balancer, and allocators run the deployments, with separate ES Admin and ES Logging clusters. Next came a platform summary of capacity and role distribution across zones, and the question âHow do you build observability on ECE?â. The answer combined the ECE platform APIs (/api/v1/platform/infrastructure/allocators and /api/v1/deployments) with GET _nodes/stats filtered to data-set size, roles and names. The result was a view of allocators per tier and zone.
The âWhat can we gain from thisâ slide listed:
- insight into hundreds of deployments plus ECE itself
- alerting at the right level
- lifecycle management
- capacity planning
- KPI dashboards and reporting
- âmuch moreâ: unassigned shards, Logstash error monitoring and infrastructure patch monitoring
The automation section listed three key requirements: scalable, centralised control with low effort, and source code management. The tools slide shows how they put it together: Ansible in the middle, connected to Terraform, PostgreSQL and git, with Terraform pointing at Elastic and Budibase on top of PostgreSQL.
![]()
Management and Automation: Ansible at the centre, with Terraform, PostgreSQL, git, Budibase and Elastic around it.
My take: this is the right shape for a large ECE estate, with Ansible and Terraform rather than clicks in a UI. When you have hundreds of deployments, the platform only stays manageable if its state lives in git and a tool applies it.
ElasticON Amsterdam at the Beurs van Berlage
The next day ElasticON took over the Beurs van Berlage. According to the agenda board, the day ran from 07:30 to 18:45 in the Grote Zaal, Effectenbeurszaal, Administratiezaal and Beursfoyer. Some of the sessions on the board:
- the welcome keynote, âSupercharge <anything> with Search AIâ
- a customer keynote from Volvo Cars
- âBuild trusted, production-ready generative AI applications on Elasticsearchâ
- âExploring Re-Ranking Techniques for E-commerce Searchâ
- âThe Future of Elasticsearch Queries: Exploring ES|QL and Whatâs Nextâ
- the Kaufland session mentioned above
The lunch break had lightning talks: âActioning AI: Elastic Insights with Tines Workbenchâ, âHybrid Geospatial RAG with Elastic and Amazon Bedrockâ, and an ElastiFlow lightning talk.
![]()
Lightning Talks over Lunch Break at ElasticON Amsterdam, in one of the Beurs van Berlage halls.
In the expo area there was a Customer Advocacy corner in the Elastic Acceleration Zone, which promised an ELK(Y) plushie to the first 50 customers who completed a review. There were also screens for the Elastic Excellence Awards. The sponsor wall thanked AWS, Tines, ElastiFlow and Kangaroot, among others.
![]()
Me at the ElasticON sponsor wall.
My take: quantized vector search in production
My take: the most useful thing about this meetup was that it treated quantization as a capacity-planning lever, not a research topic. Quantization is now on by default: int8 since 8.14, and BBQ for 384+ dimensions since 9.1. So many teams already run quantized vectors without ever having chosen to. Thatâs fine, as long as you know it and measure it.
What Iâd do on a production cluster:
- Measure recall on your own queries, not just on public benchmarks. Keep a small labelled set from real traffic and compare float32 against the quantized index before and after every mapping change.
- Tune oversampling and rescoring on purpose. The trade-off slide said it plainly: lower precision needs more oversampling. The
dense_vectordocs expose this asrescore_vector.oversample. Treat it like any other latency/quality knob and track it in your dashboards. - Plan RAM and disk separately. The RAM savings are real: the slideâs 500k Ă 1,024 example went from about 2.1 GB to under 70 MB with bits. If you rescore, though, you still keep the full float vectors on disk.
- Use
flatorint8_flatwhere filters already shrink the candidate set. Kearnsâ slide made that point, and it applies directly to multi-tenant search: when a tenant filter leaves only a few thousand documents, an HNSW graph isnât worth it. - Upgrade, then re-index. According to the docs, the defaults depend on the version an index was created on. Upgrading the cluster, as Kaufland did on its way to 8.14.1, doesnât change the quantization of an existing index.
Related
- ElasticON Amsterdam 2025: Forge the Future Recap
- AI Native Netherlands at Elastic: Agents, MCP, Workflows
- Elastic Agent Builder MCP: Tools and Workflows, No LLM
- SRE NL at Elastic: OpenTelemetry and Your Brainâs Biases
- Vector Databases on Kubernetes: Qdrant vs Milvus vs pgvector
- Loki vs Elasticsearch 2026: Log Aggregation Comparison