Wednesday 2 April 2025 was the first main-conference day of KubeCon + CloudNativeCon Europe at ExCeL London. This post follows my photos from that day in two parts: a few keynote slides I thought were worth keeping, and one breakout talk from Fastly that was mostly a postmortem and a set of rules for changing production safely. The other days are linked from the KubeCon Europe 2025 hub post.
Keynote slides
The keynote hall was one very wide screen with the London skyline artwork in the middle and a slide on each side. I did not note speaker names for the keynote slides, so I describe the content only.
The first slides were about the CNCF turning ten. Under “Cloud Native Today: A Global Community”, one slide says that over the past ten years 275,000+ contributors made 3,350,000+ code commits, 1,200,000+ pull requests and 18,800,000+ contributions, across 200+ projects and 190+ countries. A footnote says Kubernetes trails Linux, which is 33 years old, as the highest velocity project, and points to the CNCF velocity data.

Ten years of CNCF in numbers, from the keynote.
A later slide used the Linux kernel as a comparison. “Linux Developer Community” gives 76,496 changes accepted, 8.7 changes per hour, 40 changes per day in stable trees, 13 CVEs assigned per day, and a new release every 8 to 9 weeks. The slide is marked 2024.

Linux kernel development in 2024, as shown in the keynote.
Three more slides were relevant to my work:
- Platform engineering certification. The slide “What’s New in #PlatformEngineering Education” says the Cloud Native Platform Engineering Associate (CNPA) is available for purchase and that development of the Cloud Native Platform Engineer (CNPE) certification kicks off soon. All attendees got 40% off Cloud Native training, certifications and bundles. The CNPA page at the Linux Foundation describes the exam.
- Sovereign cloud. A slide announced NeoNephos, with the mission “Build a sovereign cloud-edge continuum for Europe.” and member logos including SAP, T-Systems and STACKIT. The NeoNephos website describes it as a project of Linux Foundation Europe, funded by the European Union.
- LLM evaluations. A slide titled “Evaluations: foundations for quality” separates evals (defined inputs plus a scoring function with pass/fail criteria, a countable number of evals, predictable inputs) from observability (always-on capture of live use, unpredictable inputs). Both evolve with your code. The closing slide of that talk pointed to honeycomb.io.

CNPA is available; CNPE is in development.

The NeoNephos announcement slide.

Evals and observability side by side.
Lessons from architecting the highest-scale operational systems
The breakout talk was “Lessons Learned From Architecting the Highest-scale Operational Systems in the World” by Artur Bergman, Founder and CTO of Fastly, in the Platform Engineering track on the official schedule (11:15 to 11:45 BST). The title slide shows his name and role. The schedule description says the session covers lessons from testing the limits of vendor systems and guidance on when to build versus buy.

The opening slide of the Fastly talk.
One slide collected quotes from open source infrastructure teams who rely on Fastly. A director of infrastructure at the Python Software Foundation said Fastly “single handedly made it possible” for them to provide their quality of service. A director of technology at the Rust Foundation said it is hard to overstate how valuable Fastly has been. The Kubernetes infra SIG co-chair and technical lead quoted on the slide said they were confident Fastly would help maintain the sustainability of Kubernetes without scalability worries.

Quotes from three open source foundations on the slide.
A kernel lockup postmortem
The part I found most useful was an incident write-up shown as plain text. It describes the primary issue as a race condition between a kernel thread managing the shrinker_rwsem (a semaphore used for memory management in the kernel) and other processes reading cgroup information from /proc. The race could lead to a deadlock in which the kernel thread got stuck and the machine locked up.
The slide then lists three contributing factors:
- A kernel bug. The kernel version on the affected nodes had a bug in the shrinker_rwsem implementation that made the race more likely.
- Duplicate systemd timers. A bug in a configuration management cookbook deployed duplicate timers. They spawned processes that read cgroup information heavily, which increased the number of readers competing for the semaphore and made the race worse.
- A shared memory leak. A change in one service increased shared memory usage on the nodes. It did not cause the lockups directly but probably added memory pressure and made the shrinker more active.

The postmortem slide: one race condition and three contributing factors.
My take: this is the useful shape for a postmortem. A single root cause is rare. Usually one latent bug, one configuration mistake and one unrelated change line up, and fixing only the obvious one leaves the others to cause the next incident.
Production safety rules
The next slide, “Production safety rules”, is the one I would copy into an engineering handbook. It lists five rules, which I reproduce from the slide:
- The 1-2-24 Blast Radius Rule. All code and configuration changes must be rolled out to at least 1% and at most 2% of the fleet for initial canary testing. They must be stable for 24 hours before going to the first quadrant.
- Code and configuration must be tested with all of their dependencies.
- All new or modified software features must have a rapid rollback mechanism, for example feature flags.
- No Fast Rolls. Code and configuration must not be fast-rolled for instant deployment. Exceptions may be granted if compensating controls limit the blast radius of any failure, or if the operation has to be atomic.
- Every known instance of a crash in canary or production must trigger an investigation to identify contributing factors. Eliminating all causes of crashes must be resolved above all other priorities.

Production safety rules, as shown on the slide.
The same idea appeared at KubeCon two days later in a different form, when New Relic described cell-based architecture as a way to limit blast radius.
The final rule: abstractions
The closing slide is titled “Final rule: abstractions are key”. It says that if you do not understand something, you fix it at the wrong level of abstraction; that abstractions are a source of complexity; and that abstractions cannot hide physical limits. The postmortem above fits: the lockup lived in a kernel semaphore that no platform abstraction on top of it could hide.

“Abstractions cannot hide physical limits.”
I also covered the other side of this argument on day one, in the Platform Engineering Day post on abstraction debt.
