Skip to main content
📬 Get weekly Production AI insights Practical notes on Kubernetes, AI infrastructure and platform engineering. No spam. Subscribe free
Luca Berton in the main hall of EDGECASE 2025 in Amersfoort, with a sponsor banner and a large LED wall showing a galaxy behind him
Conferences

EDGECASE 2025 Amersfoort: Talk, Booths and Interviews

EDGECASE 2025 in Amersfoort: Justin Garrison's Disney+ story and booth interviews with Chainguard, Dell, Veeam, Sysdig, Sidero, HCS and TrueFullstaq.

LB
Luca Berton
· 9 min read

On 23 September 2025 I was at EDGECASE 2025, “The Multiverse Saga”, the TrueFullstaq flagship Kubernetes event, held at the Rijtuigenloods in Amersfoort (official event page). I already wrote a short “last year” section inside the EDGECASE 2026 preview. This post is the longer version: one talk I filmed and the booth interviews I recorded for my show that day.

The venue is a former railway carriage shed, so the expo floor sat between old wagons and steel trusses, under a banner for each sponsor.

Top of the EDGECASE 2025 expo floor in the old carriage shed, with a hanging Edgecase partner banner for AWS reading The Multiverse Saga and a Sysdig banner further down the hall

Banners for the sponsors hung from the roof trusses all along the hall.

The talk: Justin Garrison on building Disney+

Justin Garrison, Head of Product at Sidero Labs (his own site lists the role), gave the talk I filmed. He opened by saying he hates making slides, so he had exactly one, and that the talk was a story from his past rather than a product pitch. The event page lists his session as the one about building Disney+.

The audio of the second half of my recording is too poor to quote reliably, so I only retell the parts I could follow.

The setup. He joined in December 2018 to run cloud infrastructure for a streaming launch. His team was four people when he arrived, and the launch date was November, so about eleven months. He said he had never run production systems at that scale, and that his manager believed he could do it before he did.

Start with what the team expects to break. His first move was to ask the people running the systems what would break on launch day. Elasticsearch was the answer from nearly everyone, and cluster provisioning was the other: at that time clusters were created by a single Ruby script on a developer’s machine. He split the work with a colleague: one took Elasticsearch, the other took cluster management. They looked at declarative cluster management with Cluster API, but the existing clusters were not Kubernetes, so they ended up driving CloudFormation from a new tool that took a YAML description.

The first big rollout went wrong. Small test clusters went well, so in the first week of August they tried a large one. He pointed out one detail from the room: a Lambda function has a 15 minute timeout, which is not enough to swap a cluster of that size, so the process had to move to step functions. Even then, roughly 60% of the time nodes came up healthy but workloads failed, and he did not know which of the two big changes (the new cluster tooling or the Elasticsearch replacement) had caused it.

Three days in the observability console. He spent three days correlating workloads, timestamps, nodes and logs and found the cause: log shipping. Every log line from a container was written to disk and then read back by the log agent, so the disks ran out of IOPS as soon as busy applications started logging. The fixes on the table were more expensive storage or less logging.

Testing with the yes command. To test a replacement logging path he generated load with the Unix yes command. He told the room it was the second time he had taken something down with it, and that this time he used up Disney’s whole logging allowance for about twenty minutes, so nobody else could ship logs from any cluster for the rest of that day.

Advice from his manager. The line he called the best advice he had in infrastructure was about estimates: when the business asks what something costs or how long it takes, give an accurate number, even a large one, rather than an optimistic one. I could not hear every word of that passage, so this is my paraphrase of the point.

My take: this is a good example of changing two things at once. The talk’s lesson is the one I repeat to platform teams: change one layer per release, keep a rollback, and instrument before you scale. The IOPS story is also a reminder that logging is a workload, and a heavy one.

Booth interviews

I recorded short clips at several booths. Product claims below are what the interviewees said, checked against the vendor pages I link.

Sidero: Talos Linux, GPUs and TalosCon

After the talk I asked Justin about AI workloads on Talos. His answer: at the bottom layer it is about accelerators and drivers, and anyone who has installed NVIDIA drivers on a traditional Linux knows the pain of the giant run files. Talos uses an immutable OS where the driver comes as a system extension, built and signed together with the kernel modules, so the combination is known to work. The goal, he said, is that the accelerator works on day one, even if there is still plenty to learn about AI itself. The Talos documentation confirms that kernel modules must be signed and shipped as extensions built with the kernel.

He also said he finds Dynamic Resource Allocation (DRA) in Kubernetes the most interesting change, because accelerators can be described more generically: GPUs today, network adapters or storage tomorrow. He invited people to TalosCon, which he described as an “engineer fest” more than a marketing event, and I went: see my TalosCon 2025 recap.

Chainguard: zero CVEs against a Docker pull

I did two clips at the Chainguard booth, with two different team members (I use roles only). The first described the pitch: pull an open source image from Docker, scan it, find CVEs you have to fix first, whereas Chainguard images are meant to arrive without CVEs. According to the interviewee, all images and libraries come with automatically generated SBOMs, and there are 52 free starter images.

The second clip was the demo: a bar chart comparing the CVEs found by a scan of a JDK image pulled from Docker, 153, against none for the Chainguard image. The explanation was that Chainguard rebuilds the images from source every day. Their site describes the images as hardened and updated daily.

My take: the 153 versus 0 number is a vendor demo and depends on the image and scanner, but the model is real: you outsource base image patching and pay for it. For a small team with no security staff that can be cheaper than doing it yourself; for a large one it is a procurement question.

Dell Technologies: private cloud with your own hypervisor licence

A Dell representative showed Dell Private Cloud and the Dell Automation Platform. The idea: bring your own hypervisor or cloud OS licence (the slide named VMware and Red Hat OpenShift, with more to come) and run it on Dell compute, networking and storage with a cloud-style operating model. His line on AI and sovereignty: “bring the AI to your data”. Dell’s private cloud page describes the same model.

Luca Berton in front of a Dell booth screen showing a high level architecture: Dell Automation Platform (SaaS or on-premises) above Dell Private Cloud on-premises, with VMware and Red Hat OpenShift boxes over rows of servers

The architecture slide on the Dell booth: the automation platform on top, the private cloud software and servers below.

Veeam: Kasten for Kubernetes data protection

A Veeam representative explained that Veeam acquired Kasten about five years ago, and that Kasten is a Kubernetes-native tool for backup, disaster recovery and moving applications across clusters and distributions. I asked about recovery objectives. The answer: they aim for near-zero RTO, depending on configuration, with backups as often as every five minutes for critical apps and copies sent off-site. Backups can be made immutable against ransomware, which matches the Veeam Kasten page. She also said she was glad to see how many companies at the event now run stateful applications in production.

My take: the fact that stateful workloads on Kubernetes are now normal is the real news. Backups are no longer the “later” item; restore tests across clusters are the part to rehearse.

Sysdig: agentic vulnerability management

Emanuela Zaccone of Sysdig (listed as a speaker on the event page, on AI security) talked about Sysdig Sage, their AI analyst on top of Sysdig Secure. It lets a team investigate runtime events in conversation, summarises threats, correlates related events and suggests remediation. She then described an agentic approach to vulnerability management that was about to ship. Her claim, which she said she had validated: it can cut around 90% of the noise and save up to 80 hours per week. That is a vendor figure, and I could not verify it. Sysdig’s own announcement covers the agentic platform.

My take: prioritisation is the right place to put AI in security, because the backlog is the pain. Still, ask how the ranking is explained, and test it on your own cluster before you believe the percentages.

HCS: Open Platform Experience

An HCS event manager told me about the HCS Open Platform Experience on 19 November 2025 in Amsterdam: platform engineering and AI talks, two workshops, and sponsors including Red Hat, Portworx and GitLab. HCS, he said, stands for “Helping Clients Succeed”. Tickets were sold at hope.amsterdam.

TrueFullstaq: the hosts

Two TrueFullstaq people, a principal consultant and the head of technology, explained the theme: “The Multiverse Saga” stands for the different timelines of cloud native, and EDGECASE is about running Kubernetes at the edge, for example a talk about Kubernetes on tractors, which matches the precision farming session on the agenda. They showed hand-made shoes with the Kubernetes helm logo that the company gives to people who get certified, and a pocket-sized Kubernetes cluster that any employee can take home as a home lab. Their service range runs from project start to a fully managed service.

A few minutes with a CNCF veteran

By the Sidero booth I met a long-time CNCF ambassador who started in the community when there were four projects. His view on AI: separate the infrastructure that supports it from the companies riding the hype, expect the result in three to five years to look different from today, and remember that the bottleneck has moved back to bus width and interconnects. On creativity, he said models produce noise, and the human “hey, wait a minute” moment is what matters.

Stage of the main hall at EDGECASE 2025: a wall-sized LED screen showing a galaxy and a smaller screen with the session title, before the next talk

The main stage before the next session, with the galaxy visuals that matched the theme.

What I took away

The sponsors formed a stack: Talos for the OS, Chainguard for images, Sysdig for runtime security, Veeam for data protection, Dell for hardware, and AWS and Grafana Labs alongside. The common thread in the interviews was operations at the edge of what teams can handle: drivers, CVEs, restores and noise. If you want the 2026 edition, see the EDGECASE 2026 preview.

Free 30-min Production AI consultation

Book Now