Skip to main content
🎓 Claude Code Masterclass Learn AI-assisted development on Udemy — plus the companion book on Leanpub & Amazon. Start Learning

LLM Inference at tech events

18 write-ups · 12 events · 2024–2026

This page collects talks and demos about serving LLMs, from vLLM and distributed inference to KServe and hardware choices. The events I attend keep returning to a few themes: serving open models on Kubernetes, how frontier model providers think about cost and latency, and what operators need to run inference reliably next to everything else on the cluster.

2026

2025

  • 23 Oct 2025 · Dutch Cloud Native & AI Community Group meetups (23 Oct and 15 Dec 2025) · Amsterdam

    Dutch Cloud Native & AI: Spegel and Dash0 Meetups 2025

    Philip Laine on Spegel, the stateless P2P image mirror, then Dash0's December meetup: Kasper Borg Nissen on OpenTelemetry and a Rust Kafka alternative.

  • 28 Aug 2025 · AI_dev Europe 2025: Open Source GenAI & ML Summit · Amsterdam

    AI_dev Europe 2025 Amsterdam: CERN, Cerebras and LanceDB

    AI_dev Europe 2025 at RAI Amsterdam: CERN's MLOps platform, five Cerebras inference lessons, LanceDB's multimodal lakehouse and Neo4j graph agents.

  • 10 Apr 2025 · KubeCon EU Recap and special guests from Nutanix and AWS (Dutch Cloud Native & AI Community Group) · Hoofddorp

    Dutch Cloud Native Recap: Nutanix NKP and AWS Triton

    Dutch Cloud Native & AI meetup, Hoofddorp, 10 April 2025: air-gapped Kubernetes bootstrapping with Nutanix NKP, then Triton on EKS with Karpenter from AWS.

  • 3 Apr 2025 · KubeCon + CloudNativeCon Europe 2025 · London

    Benchmarking GPU Workloads on Kubernetes with Triton

    A KubeCon London 2025 talk on benchmarking AI and GPU workloads in Kubernetes: Triton, GenAI-Perf, fmperf and a time-slicing versus MPS comparison.

2024

Companies on this topic

Related topics

Free 30-min Production AI consultation

Book Now