LLM Inference at tech events
18 write-ups · 12 events · 2024–2026
This page collects talks and demos about serving LLMs, from vLLM and distributed inference to KServe and hardware choices. The events I attend keep returning to a few themes: serving open models on Kubernetes, how frontier model providers think about cost and latency, and what operators need to run inference reliably next to everything else on the cluster.
2026
25 Sept 2026 · Builders & Brews: Hack Edition (Nebius x NVIDIA Global AI Hackathon, Amsterdam) · Amsterdam
Builders & Brews Amsterdam: Nebius x NVIDIA AI HackathonAAIF Community Amsterdam's Builders & Brews: Hack Edition: the hackathon rules, Token Factory, Tavily, Kilo Code and a fast agent recipe before 30 October.
29 Jul 2026 · KubeCon + CloudNativeCon Japan 2026 · Yokohama
Tuning Kubernetes for AI: the Real Trade-offs from KubeCon Japan 2026From KubeCon Japan 2026: why AI on Kubernetes needs tuning, not just more GPUs — LLMD caching/routing and the real GPU, memory and electricity cost.
2 Jun 2026 · AI on the Amstel (June 2026): frontier models panel · Amsterdam
AI on the Amstel: DeepMind, NVIDIA & Mistral PanelRecap of the June 2026 AI on the Amstel meetup at VU Amsterdam — a panel with Google DeepMind, NVIDIA, and Mistral on how frontier models are built.
June 2026 · Red Hat Tech Day Netherlands 2026
Red Hat AI Model-as-a-Service with llm-dHow Red Hat's llm-d transforms LLM inference into a composable Kubernetes-native architecture: disaggregated serving, smart autoscaling, and MaaS.
June 2026 · Red Hat Tech Day Netherlands 2026 · Bunnik
Red Hat Tech Day Netherlands 2026: Harness EngineeringInside Red Hat Tech Day Netherlands 2026: 30 speakers and 34 sessions on AgentOps, agentic AI, Quarkus, LangChain4j, and hybrid cloud inference.
June 2026 · Red Hat Tech Day Netherlands 2026
vLLM Inference Optimizations on Red Hat OpenShift AIDeep dive into vLLM inference optimizations: KV cache, continuous batching, quantization, and distributed inference with Tensor Parallelism.
12 May 2026 · Red Hat Summit 2026 · Atlanta
Summit 2026 Keynote: Metal to Agents and TokensNotes from the Red Hat Summit 2026 day one keynote: open models, token economics, the metal to agents stack and an NVIDIA conversation on agent governance.
11 May 2026 · Red Hat Summit 2026 · Atlanta
llm-d at Red Hat Summit 2026: KV-Cache Aware Routing for vLLMRed Hat presented llm-d at Summit 2026: cache-aware load balancing for vLLM. Cold 4.3s vs warm 0.6s (7x faster), $0.30 vs $3.00 per 1M tokens with caching.
11 May 2026 · Red Hat Summit 2026 · Atlanta
Upgrading to OpenShift AI 3.x at Red Hat Summit 2026RHOAI 3.3, 3.4 and 3.5 release channels, support windows, supported configurations and a 5-step migration from OpenShift AI 2.x to 3.x.
11 May 2026 · Red Hat Summit 2026 · Atlanta
Pete Cheslock: From RAG to Roll-Your-Own Models in One YearPete Cheslock on the one-year enterprise AI leap from RAG pilots to self-hosted inference and model training on vLLM and llm-d at Red Hat Summit 2026.
11 May 2026 · Red Hat Summit 2026 · Atlanta
llm-d Maintainer Sally O'Malley on Enterprise AI Inferencellm-d maintainer Sally O'Malley on vendor-neutral, Kubernetes-native distributed inference and what enterprise AI adoption looks like from inside the project.
7 May 2026 · DevWorld Conference 2026 · Amsterdam
DevWorld 2026: The Next Generation of AI Is Powered byA DevWorld 2026 keynote argued that the next generation of AI depends on infrastructure, not models — six pillars from latency to 100K req/s scale.
2025
23 Oct 2025 · Dutch Cloud Native & AI Community Group meetups (23 Oct and 15 Dec 2025) · Amsterdam
Dutch Cloud Native & AI: Spegel and Dash0 Meetups 2025Philip Laine on Spegel, the stateless P2P image mirror, then Dash0's December meetup: Kasper Borg Nissen on OpenTelemetry and a Rust Kafka alternative.
28 Aug 2025 · AI_dev Europe 2025: Open Source GenAI & ML Summit · Amsterdam
AI_dev Europe 2025 Amsterdam: CERN, Cerebras and LanceDBAI_dev Europe 2025 at RAI Amsterdam: CERN's MLOps platform, five Cerebras inference lessons, LanceDB's multimodal lakehouse and Neo4j graph agents.
10 Apr 2025 · KubeCon EU Recap and special guests from Nutanix and AWS (Dutch Cloud Native & AI Community Group) · Hoofddorp
Dutch Cloud Native Recap: Nutanix NKP and AWS TritonDutch Cloud Native & AI meetup, Hoofddorp, 10 April 2025: air-gapped Kubernetes bootstrapping with Nutanix NKP, then Triton on EKS with Karpenter from AWS.
3 Apr 2025 · KubeCon + CloudNativeCon Europe 2025 · London
Benchmarking GPU Workloads on Kubernetes with TritonA KubeCon London 2025 talk on benchmarking AI and GPU workloads in Kubernetes: Triton, GenAI-Perf, fmperf and a time-slicing versus MPS comparison.
2024
13 Nov 2024 · OpenSearch Meetup Amsterdam · Amsterdam
OpenSearch Meetup Amsterdam 2024: Vectors and EmbeddingsOpenSearch meetup at AWS Amsterdam, Nov 2024: local LLMs and hybrid search, Bedrock, vector tuning tips, and Zeta Alpha on InPars and e5-mistral.
25 Sept 2024 · AI Tinkerers Amsterdam (September and November 2024 editions) · Amsterdam
AI Tinkerers Amsterdam 2024: Groq, Fiberplane and NebiusTwo AI Tinkerers Amsterdam demo nights from autumn 2024: Fiberplane's AI API testing, Groq-powered apps, an AI-native React compiler and Nebius AI Studio.