Between mid 2024 and September 2026 I wrote up more than 40 events that touched platform engineering: KubeCon in London, Amsterdam and Yokohama, PlatformCon in London, a long run of Platform Engineering Amsterdam meetups, vendor days, and a few evenings where I was on the stage myself. Read back in one sitting, what repeats matters more than any single talk. Some ideas arrived in 2024 and were still argued about in 2026. Others changed shape once AI workloads landed on the same platforms.
This is the synthesis, organised by theme. Every claim comes from one of my event posts, linked inline, with the speakersâ hedges kept. Where I add my own opinion, I say so.
1. A platform is a product, and most teams do not behave like it
The first time I heard âplatform as a productâ on a slide was at the Kubernetes 10th birthday meetup in Amsterdam in June 2024. A platform team at RTL walked through âThe journey to a platform productâ: start from shadow IT, then build the team, then adoption. They onboarded the most motivated consumer first (Videoland, in their telling) and migrated the others in stages. That is still how I see most internal platforms grow.
By February 2025 a talk on organising a platform team put the same idea in two lines on a slide: find a backer, find a champion. My note then was that a platform born as shadow IT only survives if someone with budget cares and at least one product team will say publicly that it helps.
The 2026 versions were sharper. At Platform Engineering Day in Amsterdam the âGolden Path as a Productâ session summarised itself as four ideas: start small, community first, defeat bureaucracy, solve shared pain. The one I see skipped is the last. A golden path that fixes a problem only the platform team has is a roadmap item, not a product.
A week earlier at Platform Engineering Amsterdam, Lian Li argued that product thinking gets you started and inner source thinking gets you adoption. Her slide called the failure mode âTechnically Impressive Ghost Townsâ and cited three figures from a Google Cloud blog: 55% of organisations have adopted platform engineering, 85% report developers rely on the platform to succeed, and only 27% have successfully integrated best practices.
The people side kept showing up. At the June 2025 BBQ meetup a deck about a startup platform in a regulated business, speaker unidentified, said tech at a certain point is easy and people are hard. At HOPE 2025, a session built on a Dutch government IT organisation that had run a platform for six years gave adoption and governance as much room on its maturity canvas as capabilities.
What changed: 2024 was about getting a platform team to exist. By 2026 the question was why a platform that exists is not used. Takeaway: write down who your backer and your champion are, and name the shared pain your first golden path removes.
2. Abstractions are the product, and they accrue debt
No theme repeated as cleanly as this one. On 1 April 2025 at Platform Engineering Day in London, Atulpriya Sharma of InfraCloud described the âAbstraction Debt Trapâ. The platform team builds a database provisioning abstraction, it works for ten teams, and then a mobile payments team asks for specific PostgreSQL extensions. The wrong answer on his slide adds a configuration parameter for every request. The right one offers several interface levels, with the top level including escape hatches such as direct PostgreSQL parameters, backed by policy-based guardrails.
Two days later, in the LinkedIn talk, the migration lessons slide said: do not give raw Kubernetes to your customers, invest in abstractions, and invest in guardrails to prevent user errors. And the Fastly talk closed on the opposite warning, âabstractions cannot hide physical limitsâ, after a postmortem where a kernel semaphore race sat below anything a platform layer could hide.
So the same week argued for abstractions and against trusting them blindly. I read that as consistent: build them, tier them, and know what lives underneath.
Later examples were concrete about what a good abstraction looks like:
- At the kro and Config Connector meetup in February 2026, a slide titled âOffer a service instead of a DIY-kitâ showed one five-line custom resource expanding into an IAM policy, a Bigtable instance and two app profiles. Each teamâs resources were provisioned by impersonating that teamâs service account.
- At Saxoâs keynote at KubeCon Amsterdam, the âUpdating Kafka Topic Access Control Listâ slide showed a four-step manual flow repeated for Dev, Test, Simulation and Live, followed by a service blueprint covering Kafka, identity, metrics, databases and DNS.
- At the Azure Platform Engineering meetup in Utrecht in 2024, Rabobank described three teams and 8 platform engineers serving 1,208 development and 1,002 production subscriptions, so they shipped managed templates through Template Specs. Every template had tests that ran three times a week, and a failure opened a work item. Their listed challenges included the balance between freedom and standardisation and the shift from example to product.
One warning was about starting at the wrong end. In the 2026 Backend-First IDP session, âThe Portal-First Trapâ set a polished Backstage catalog next to the pile of mismatched code snippets behind it. My take stayed the same as in 2024: the portal is a thin layer over APIs and automation, so build those first.
Takeaway: treat every abstraction as an API with versions, tiers and a test suite. Decide where the escape hatch is before the first awkward request arrives.
3. Measurement is the weakest spot, everywhere
At the September 2025 Platform Engineering Amsterdam meetup a speaker showed a grid adapted from the CNCF Platform Engineering Maturity Model, with percentages from survey sources cited on the slide. In the measurement row âWe do not measureâ was the largest answer at 44.67%, ahead of DORA metrics at 37.3%. Adoption was led by extrinsic push at 35.8% over intrinsic pull at 28.4%. The speaker was candid that his own measurement was not where he wanted it to be, which matched his slide.
Good examples were small and cheap:
- A 2024 Grafana and Friends talk used Backstage to track which services had adopted the OpenTelemetry SDK. The catalogue already knows every service, so âwho is missing instrumentationâ becomes a query instead of a survey.
- Rabobank showed usage numbers for its templates: 2.879K total template spec runs, 22 templates, 37 projects and 37 subscriptions using them.
- At a Datadog user group in September 2025 a speaker responsible for roughly ten classifieds businesses built a dashboard from data he already had. One of its outputs was a golden path index, and my note was that you cannot run a golden path programme until you can see which teams and repositories belong to which capability.
- At PlatformCon London in June 2026, Kief Morris of Thoughtworks showed deliverability and operability âsensorsâ: DORA metrics as lagging sensors for deliverability, and mutation testing and fuzz testing as leading ones.
One thing the sources do not settle is how many organisations have a platform team. Lian Liâs slide said 55% had adopted platform engineering. The CNCF press conference slide at KubeCon Amsterdam 2026 said 28% of organisations have a dedicated platform engineering team. These are different definitions from different surveys, so I would not put them on one chart.
Takeaway: DORA metrics plus a short developer survey beat nothing, and ânothingâ was the most common answer I saw. Track adoption from the catalogue you already have.
4. Scale means blast radius, and fewer moving parts
Two KubeCon London 2025 talks answered the same question in opposite ways. LinkedIn runs a bare-metal fleet of a few very large clusters, about 5,000 nodes each, with machine pools and etcd tuning around them. New Relic reported 280+ clusters of 300 to 500 nodes each, organised as cells across AWS, Azure and GCP. These sources do not agree on cluster size, and I do not think they need to. Each fit its situation. What they share is the aim to keep failures small.
Fastlyâs â1-2-24 Blast Radius Ruleâ said changes go to at least 1% and at most 2% of the fleet for canary and must be stable for 24 hours before going wider. The same limit-the-damage thinking appeared in rollout tooling in 2026. At ArgoCon in Amsterdam, âThe $10,000 ArgoCD Mistakeâ reported mean sync time falling from 600 s to 60 s after reducing polling, using webhooks and ApplicationSets, and sharding the repo server.
The other half of scale is subtraction. At Cloud Native Rejekts in Amsterdam in March 2026, Steve Wade of Platform Fix told me he is brought in âto delete things to improveâ, with fewer Terraform modules and less cloud native tooling. His point was that the CNCF landscape is both a strength and a problem, since several teams in one company often want different tools for the same goal, and âNot every tool is right for every organization.â My note: deleting a component only counts once its alerts and runbooks go too.
Takeaway: pick a cluster strategy on purpose, cap the blast radius of every change, and keep a written reason next to every tool you run.
5. GPUs turned multi-tenancy from a nice-to-have into the main job
This is the theme where I was a speaker, so I will mark my own claims clearly. At Red Hat Summit 2026 in Atlanta I presented multi-tenant GPU platforms on OpenShift AI. The core of it: contention on shared GPUs is inevitable, and the question is whether it is deterministic or chaotic. My answer was per-tenant GPU caps, explicit priority classes (P0 training, P1 serving, P2 batch, P3 interactive), a stated preemption posture, and a GitOps bootstrap bundle per tenant with namespace, RBAC, network policy and quotas. No tickets. The April 2026 Xebia meetup talk went deeper into day-2 operations.
What made this more than my own story was hearing a company at a very different scale reach a similar design. In the KubeCon Japan keynotes, Preferred Networks compared four isolation options from namespaces to dedicated clusters and marked namespace isolation and dedicated nodes as âOur Choiceâ, a hybrid. Then came what you build once you choose namespaces: hierarchical tenants through a forked HNC, network and admission policy, per-tenant controller scope, and âHRQ + Kueue Quotaâ so that âWithout Limits, One Tenant Can Take Almost Everythingâ does not happen.
SoftBank added a capacity argument in three minutes. Its âReality of AI Agentâ slide showed an agent starting several pods for parallel inference, set against âInfinite Apps, Finite Computeâ. My take: if your quotas assume one request means one pod, agents will break that assumption first.
A DevWorld 2026 keynote argued that infrastructure, not model size, decides what ships, and my Skopje keynote abstract made the platform case for GPUs, governance and cost.
Takeaway: the GPU is the easy part. Decide fairness rules, preemption and tenant bootstrap in code before the second team arrives.
6. Agents are the new tenants, and the platform story shifted
In 2025 AI on the platform was a question. At SREday Amsterdam in November, CotĂŠâs talk listed what platform teams will probably own: hosting models, application frameworks, registries for plugins such as MCP and agent skills, and cost control including AI FinOps. It was still open who chooses and updates models, operations or a data scientist type of role.
By PlatformCon London on 23 June 2026 the sentence on every stage was some form of âthe age of AI runs on platform engineeringâ. The more useful details came from specific sessions:
- A morning session on the agentic IDP said âNot that much changesâ. It split work into agent paths, deterministic paths and hybrid paths where pipeline output feeds the model in a loop. The reference architecture kept the resource plane and added tool security and tool observability. Running dozens or thousands of agents needs a governance plane for identity, security and observability.
- Cloudsmith called LLMs on developer workstations shadow infrastructure and proposed pulling models through a governed registry.
- OpenBaoâs founder advised short-lived, fine-grained tokens scoped per sub-task.
- Nirmata said the enforcement layer stays deterministic while agents work above it on policy generation and compliance reporting.
- Kief Morris described the engineerâs job as building the system that builds the software, with a human on the loop instead of in it.
A March 2026 talk from Coder landed in the same place from the tooling side. After a benchmark table of six agent setups on nine coding tasks, the speaker, Michael, said clever harness tricks probably will not matter long term, and what lasts is the work around the model: benchmarks, sandboxing, audit trails and a human who can see results. My note was that this list is a platform teamâs job description. The kro evening in February made the practical point: an agent is only as safe as the tools it calls, so build the kro layer first and put the agent on top.
One tension remains between sources. The morning agentic IDP session said âNot that much changesâ; my closing note from the same day, in the PlatformCon recap, said AI raises the bar on fundamentals. I think both are true: the shape stays, the bar goes up. The closing panelâs answer on what to build next was to get the fundamentals solid enough that agentic anything is safe on top.
Takeaway: treat every agent as a tenant with an identity, a quota, an audit trail and a registry it pulls from.
7. Cost is a platform feature, and it starts with visibility
Cost talks were the most numerical of the three years. At Cloud Optimization 2025 in March, ABN AMROâs FinOps timeline began in 2019 with show-back and forecasting, added commitments, and in 2023 turned a one-person practice into a team. In 2024 it added automation and GreenOps. The order, according to the slides, was visibility, then commitments, then a team. Wehkampâs cost bot, Kostunrix, taught the other lesson: an alert that says what changed, for which team, with a link, gets read.
The platform can also enforce cost. At the February 2025 meetup, Nirmata showed Kyverno generating a VPA for every workload and a cleanup policy for PodDisruptionBudgets that allow zero disruptions. At PlatformCon 2026 Nirmata described cost management as a growing second use case for policy as code.
A Platform Engineering Amsterdam talk in November 2025 went deeper on storage. According to the speaker, a platform inherited over 2,000 TB in S3 across about 400 million files, and after deduplication, Intelligent-Tiering and removing about a petabyte of hidden noncurrent versions the bill was about $90,000 a month lower. My takeaway was two defaults for any S3 module: a noncurrent-version expiration rule wherever versioning is on, and Intelligent-Tiering where access is unpredictable. Observability has its own bill. The 2024 Tempo migration reported one fifth of the previous providerâs cost at about 6 TB of raw data a day.
Takeaway: start with show-back, make alerts readable, and push the boring savings into your default modules.
What Iâd do on Monday
- Name your backer and your champion, and write down the shared pain your next golden path removes.
- Turn your most requested abstraction into tiers with one documented escape hatch and one policy that governs it.
- Pick two adoption numbers from your catalogue and add DORA metrics plus a short developer survey.
- Write your blast radius rule for platform changes: canary percentage, soak time, rollback.
- List your tools, give each a written reason, and retire one with its alerts and runbooks.
- If you share GPUs, write the fairness rules: tenant caps, priority classes and preemption, in Git.
- Give agents their own identities with short-lived credentials, and route models through a registry.
- Start show-back before you buy commitments, and add noncurrent-version expiry to your storage module.
Sources
2024
- Kubernetes 10th Birthday Meetup in Amsterdam
- Grafana Tempo, OpenTelemetry and Backstage: Meetup 2024
- Azure Meetup Utrecht 2024: Verified Modules and Private DNS
2025
- Platform Engineering Amsterdam Meetup, February 2025
- ABN AMROâs FinOps Journey at Xebia Cloud Optimization 2025
- Platform Engineering Day London 2025: Abstraction Debt
- Kubernetes at Scale: LinkedIn and New Relic, KubeCon 2025
- KubeCon London 2025: Kernel Lockup and Safety Rules
- Platform Engineering BBQ Amsterdam 2025: Observability
- Platform Engineering Amsterdam Meetups, Autumn 2025
- Datadog User Group Amsterdam 2025: Spec-Driven Development
- SREday Amsterdam 2025 Q4: Observability, MCP, ClickStack
- Agents in Production at AI House and HOPE 2025
2026
- kro and Config Connector at a Xebia Meetup in Amsterdam
- Platform Engineering MeetUp Amsterdam: Human Intelligence
- Steve Wade (Platform Fix) at Rejekts 2026: Fewer Tools
- KubeCon EU 2026 Co-located Day: Talks and Slides
- KubeCon Europe 2026 Keynotes and Sessions: On the Slides
- Platform Engineering Meetup NL: Multi-Tenant
- GPUs Take Flight: Multi-Tenant Platform Engineering at Red Hat Summit
- PlatformCon London 2026: The AI Era Runs on Platforms
- Your IDP Is the Foundation for Agentic AI
- Kief Morris on AI Agents and Being âHuman on the Loopâ
- Managing AI Agents at Platform Scale: Cloudsmithâs Take
- OpenBao Founder on Secrets Management for AI Agents
- Kyverno Graduates CNCF, and Nirmata Adds AI Governance
- GitOps and Platform Engineering at KubeCon Japan 2026
- KubeCon Japan 2026 Keynotes: AI Platforms at Scale
- DevWorld 2026: The Next Generation of AI Is Powered by
- AI Tech Summit Skopje 2026: From AI Demo to Production
