Skip to main content
📬 Get weekly Production AI insights Practical notes on Kubernetes, AI infrastructure and platform engineering. No spam. Subscribe free
Busra Koken's title slide, Your Systems Don't Learn from Incidents, You do, at the SRE NL meetup in Amsterdam
DevOps

SRE NL Amsterdam 2025: Learning from Incidents

SRE NL at Datadog Amsterdam, October 2025: Busra Koken on why people, not systems, learn from incidents, then SRE culture and why engineers skip postmortems.

LB
Luca Berton
· 5 min read

On Thursday 30 October 2025, after a day at ElasticON Amsterdam, I went to the evening meetup of Site Reliability Engineering NL (SRE NL). It was hosted at the Datadog office in Amsterdam: the Datadog neon sign was on the wall behind the coffee bar, and the “Special thanks to” slide credited Datadog and the Datadog User Group.

There were three talks, and they fitted together well. The first was about how people learn from incidents, the second about changing an organisation’s culture towards SRE, and the third about why engineers avoid the postmortem, the one meeting where that learning is supposed to happen.

This is a throwback built from my photos, so I’ve stuck to what was on the slides.

Luca Berton taking a selfie at the SRE NL meetup at the Datadog office in Amsterdam, with the audience on yellow chairs and the opening slide on two screens

Before the start: SRE NL at the Datadog office, with the “this is fine” opening slide on both screens.

Community news: 2K members

Robin opened the evening, according to the agenda slide. The community slides came first:

  • “2K!”: SRE NL had passed 2,000 members. The next slide was the Meetup “Total and active members” chart, rising from November 2022 to just over 2,000 in September 2025, with a hand-drawn “Meetup community driven” label at the start.
  • The organisers: “The current organizers driving the community” showed eight people, with logos from Booking, bol., ING, Xebia and mollie under their photos.
  • Contribute? Reach out on meetup.com, find past slides at sre-nl.github.io/slides, talk to the organisers, and send in topics to present.

The SRE NL organiser on stage next to a 2K! slide over a collage of past meetup photos

“2K!”: SRE NL passes 2,000 members.

The SRE NL agenda slide: opening by Robin, Busra on learning from incidents, a ten-minute break, Andrea on SRE culture, Line on postmortems, then drinks

The plan for the evening: three talks, a break and drinks, ending around 21:00 to 21:15.

Busra Koken: your systems don’t learn from incidents, you do

The first talk was “Your Systems Don’t Learn from Incidents — You do” by Busra Koken of Humans in Systems. Her intro slide described her as the founder of Humans in Systems, a coach and mentor, and a “Friendly Consultant who helps organizations navigate incidents with confidence”. It also said: “I also paint. A lot.”

Busra Koken's title slide, Your Systems Don't Learn from Incidents, You do, with Humans in Systems underneath

The title says it all: learning is something people do, not something a system does.

She defined the job as “Engineering in complex systems for our business to continuously meet users’ needs.” The next slide explained why those systems are sociotechnically complex:

  • users that can do all kinds of expected and unexpected things
  • developers that push code and config every few minutes
  • software that runs on other people’s servers

The conclusion was in bold: “So really, failures often emerge in the interactions of many parts, not inside a single component.”

Busra Koken's I am Busra slide: founder of Humans in Systems, coach and mentor, friendly consultant who helps organisations navigate incidents with confidence

The sociotechnically complex systems slide: users, developers pushing code and config every few minutes, software on other people's servers, failures emerging in the interactions of many parts

Busra’s intro and the slide behind the talk’s argument.

My take: this is the right starting point for any incident review. If a failure comes from how parts interact, then the review can’t stop at “component X broke”. It has to cover what people knew, what they saw on their dashboards, and why the action they took made sense at the time. The system doesn’t learn from any of that. The people who run it do.

Andrea: SRE culture and why buy-in is hard

After the break, Andrea spoke on “SRE culture and why it’s so hard to get buy-in in larger organisations”. She started with change management: “What do change models say about changing culture?” and “What is suggested to do when changing a culture?”

The slide showed several well-known models side by side. I recognised Lewin’s change model (unfreeze, change, refreeze), the McKinsey 7-S model, ADKAR, Kotter’s 8-step model and the KĂŒbler-Ross change curve. The slide’s point was that most models come down to a similar set of steps.

Andrea's slide asking What do change models say about changing culture, with diagrams of Lewin, McKinsey 7-S, ADKAR, Kotter and KĂŒbler-Ross change models

Change models first, SRE second.

Then came “So, What about SRE Culture?”. On the left were the 7 principles: simplicity, embracing risk, eliminating toil, monitoring, automation, release engineering, and measure (SLI/O/A). On the right was a hand-drawn SRE culture patterns sketch:

  • work to improve the future state (from reactive to proactive work)
  • embrace risk
  • have an experimenter’s mindset (test your hypothesis)
  • empower developers: from passive safety to active policing, “aka guardrails”
  • deploy fast and often for reliable software
  • eliminate toil: get feedback and automate where possible
  • understand the wider system

Andrea presenting the So, What about SRE Culture slide with the seven SRE principles and a hand-drawn SRE culture patterns sketch

The 7 principles and a hand-drawn map of SRE culture patterns.

My take: in a large organisation, “empower developers” and “guardrails” belong in the same sentence. Platform teams win buy-in when the safe path is also the easy path. Writing an SLO policy and asking teams to follow it doesn’t do that on its own.

Line: why engineers don’t like postmortems

The last talk, by Line, was “Why doing a postmortem should be one of your favorite activities as an engineer!”. The slide I photographed gave two reasons engineers avoid them: postmortems are seen as an administrative task, and there are priority issues.

Next to those were two example write-ups that show the problem:

“We faced bad performance on our instances. Looking at the dashboard we saw that the memory usage was high. A rolling restart of the system solved the problems.”

“Third party software delivered an update and after deploying to production it failed. Root cause is at third party.”

Line presenting the Why don't engineers like doing postmortems slide, with two example postmortem quotes about a rolling restart and a third-party update

Two postmortems that close the ticket but teach nobody anything.

Both examples describe what happened, and neither explains why. Nobody asks why memory grew, why the update went straight to production, or what would catch it next time. That links straight back to Busra’s talk: if the write-up stops at “a restart fixed it” or “the vendor’s fault”, the people involved haven’t learned anything either.

What I took home

Put together, the three talks made one argument: reliability is a people problem with technical symptoms.

  • Review the interactions, not just the component. Busra’s sociotechnical slide is a good checklist for any review: users, deploys, and dependencies you don’t control.
  • Treat culture change as change management. Andrea’s mix of change models and SRE principles is a reminder that rolling out SLOs is an organisational project, not a tooling one.
  • Make postmortems useful, so people want to write them. If they’re paperwork, you get Line’s two examples. If they produce a real follow-up that a team cares about, people will show up.
#SRE NL #Site Reliability Engineering #Incident Management #Learning from Incidents #Postmortems #SRE Culture #Change Management #Datadog #Amsterdam #Meetup
Share:

Want to operate this yourself, in production?

Take the free AI Platform Engineer Readiness Scorecard to see which skills transfer — then build a production-shaped AI platform in the 4-week Bootcamp.

Take the Scorecard →
Luca Berton — The Production AI Expert, Docker Captain

Luca Berton

The Production AI Expert · Docker Captain · KubeCon Speaker

15+ years in enterprise infrastructure. Author of 8 technical books, creator of Ansible Pilot (1M+ YouTube views, 648K site users). Former Red Hat engineer. Speaker at KubeCon EU 2026 and Red Hat Summit 2026.

Free 30-min Production AI consultation

Book Now