Skip to main content
📬 Get weekly Production AI insights Practical notes on Kubernetes, AI infrastructure and platform engineering. No spam. Subscribe free
A full lecture hall at CfgMgmtCamp 2025 in Ghent with a slide showing Chef's history from OpsCode in 2008 to Progress Chef in 2020
Conferences

CfgMgmtCamp 2025 in Ghent: The Talks I Saw

My notes from the CfgMgmtCamp 2025 talks in Ghent: an AI model leaderboard, Chef history, Salt internals, SOPS for Ansible secrets, ARA and tox-ansible.

LB
Luca Berton
¡ 16 min read

From 3 to 5 February 2025 I was at Config Management Camp on the HOGENT Schoonmeersen campus in Ghent. My companion post has the video interviews and my own session. This one covers the talks I sat in on, rebuilt from the slides I photographed. I’ve named speakers only where their name was on the slide. Where my phone videos caught what was said on stage, I’ve added that too.

The mornings in the main auditorium were a run of short talks, one after another. My photo timestamps show a new speaker every five or ten minutes, so some of the sections below are brief.

Monday 3 February

Overlook InfraTech: “Our business ethos”

One of the first slides I photographed on Monday, in a packed main hall, came from Overlook InfraTech and set out how the company plans to work:

  • The fork they’re supporting, and any products built on it, are open source “from top to bottom”.
  • They won’t sell access to products. They’ll enable success with them.
  • As a public benefit company, the ethos is written into the business structure and bylaws.
  • Fiduciary duty includes community benefit, not just shareholder returns.

The fork is OpenVox. According to Vox Pupuli’s OpenVox page, it started in late 2024 as a mirror of Puppet packages from Overlook InfraTech, after Perforce stopped public packaging. It’s now maintained by the community under Vox Pupuli as a “soft-fork”, which aims to stay downstream-compatible with Puppet for as long as it can.

In the Q&A, the speaker recalled a relicensing years earlier. Clients’ company policies had stopped them using the software under its old licence, and he said those policies made no sense legally or logically. He wanted more discussion of how open source projects can protect themselves from licence changes. He pointed the room to Dawn Foster of CHAOSS, who had covered this at FOSDEM two days earlier on the panel “Forked Communities: Project Re-licensing and Community Impact”.

The Puppet fork came up again in my interview with Martin Alfke, which is in the companion post.

An AI LeaderBoard (What did AI ever do for me?)

Karanbir Singh asked a simple question: can current models write configuration code that actually works?

  • The contestants: DeepSeek r1, o3-mini, Mistral Small 3 and a local fine-tune of Llama 3.3 (LoRA). The slide gave a release date for each, all from the two weeks before the talk.
  • The targets: Ansible, Puppet, Terraform and Bash.
  • The method: five cycles of ask for recommendations, secure it, improve performance, improve efficiency, optimise. The models were also asked to assess each other’s code. Context got lost between rounds, so the original prompts had to be fed back in with the generated code.
  • The result: “Over all results start to get very similar [except Bash code]”. The local fine-tuned model “seems to benefit immensely from the topically focused skills”. On the “Does it run?” slide, the generated content all deployed, “but there are gaps”.

The closing slide gave honourable mentions to Claude 3.5, Phi-4, Granite 3.1 and Llama 3.3, under the line “Models today are good”.

Karanbir’s demo went further than the slides: he wired a model into a git post-receive hook and asked it to explain the impact of each change on the infrastructure defined in the repo. His point on stage was that a model with only a few minutes on the task, and nothing but its own context, could already flag a significant change, because it knew why the manifest had been written that way.

Karanbir Singh presenting the AI LeaderBoard scores at CfgMgmtCamp 2025, with bar charts comparing four models on Ansible, Puppet, Terraform and Bash

“Scoring the results of the Spanish Inquisition”: four models, four languages, and Bash at the bottom of every chart.

My take: the result that matters most to me is the fine-tuned local model holding its own against the big hosted ones. A small model trained on your own modules and conventions is a realistic option for platform teams who can’t send their infrastructure code to an external API.

Relics of Evolution: Chef’s history

A history of Chef in a few colourful slides:

  • 2008: OpsCode was founded because its creator “was not satisfied with available tools, so he created one!” OpsCode became Chef in 2013, and Progress acquired Chef in 2020.
  • Then and now: from managing server configurations with cookbooks to policy as code, automated compliance and security across cloud and on-prem.
  • Habitat, introduced in 2016, for application portability and lifecycle management, independent of the underlying infrastructure.
  • An honest slide titled “We aren’t idempotent and sometimes need adhoc changes”: Chef Push Jobs (introduced in 2014, later deprecated), “Users Hack Jobs” (engineers make it happen however they can) and Chef Courier.

The talk closed on three themes: start with config (infrastructure as code), shift security (compliance as code) and manage state (consistent system environments).

On stage, the speaker opened the DevOps part with a slide quoting Prince, “A strong spirit transcends rules”, and said it held for both Chef and DevOps. DevOps has kept evolving with people’s needs, and Chef is also shifting towards a platform engineering mindset. The speaker also plugged a conversational Chef workshop on Wednesday at 1 p.m., open to people who had never used Chef, with labs on Test Kitchen, cookbooks, agentless and InSpec.

The OpsCode, Chef and Progress Chef timeline on screen in a full CfgMgmtCamp 2025 lecture hall in Ghent

OpsCode in 2008, Chef in 2013, Progress Chef in 2020.

Three quick ones

  • Vox Pupuli: an update covering what it is, who’s behind it, money and a three-year recap. One slide listed the services they pay for: inwx.de for domains, Hetzner for DNS and tuta.com for shared email accounts.
  • JJ Asghar: “You say you can exit vim, have you ever tried exiting ed instead?”, with slides of real ed commands, such as printing the whole buffer.
  • Jan Bundesmann (ATIX): “GitLab unplugged: Real-Life Tips and Tales”.

Salt performance work

A Salt talk that started with the design idea: “Salt wants you to think of configuration management as a type of distributed computing”, and “privileged nodes are a liability” that set the scale limit through the inefficiency they add. Then came the concrete changes:

  • MWorker queueing: the MWorker processes moved from round-robin to a single queue, giving about 10% more MWorker throughput. The slide credits Sze Chuen Tan as author and links saltstack/salt#67215 and the ZeroMQ guide’s load-balancing broker pattern.
  • Caching Jinja import_*: Jinja rendering in Python isn’t fast, and {%- import_* %} calls are a convenient point to memoise. The slide also admitted that “caching in salt could be much better in coverage and robustness”.

Salt talk slide at CfgMgmtCamp 2025 on MWorker queueing, reporting about 10 percent more MWorker throughput and linking the Salt pull request

A single queue in place of round-robin: about 10% more throughput.

State of the Ansible Community 2025

In the afternoon I moved to a smaller classroom for the Ansible talks. John “gundalow” Barker, Ansible Community & Partner Engineering Manager, gave the yearly community update:

  • 2024 events included FOSDEM and CfgMgmtCamp, DevConf.cz, Red Hat Summit and AnsibleFest, Ansible London and DevOpsDays, and “38 dedicated Ansible Meetups in 13 cities”.
  • The docs are being consolidated under the Ansible namespace on Read the Docs, which brings PR previews, cross-project search and a native version switcher. Anyone interested can join the DaWGs weekly docs meeting.
  • He showed forum stats, open forum polls on whether to merge PRs, and survey results on what people use to write Ansible content and which development tools they know (Molecule among them).
  • He closed on Community and Partner Engineering working together as a “force multiplier”.

Secrets in Ansible

The speaker introduced himself as an Ansible user since about 2015, a contributor since 2016/2017, a collection co-maintainer since 2020 and a member of the Ansible Community Steering Committee. The talk walked through the options for secrets:

  • Plain vars files (secrets.yml) and Ansible Vault.
  • Lookup plugins for external providers: community.hashi_vault.hashi_vault, community.general.bitwarden, community.general.keyring.
  • SOPS: it started as Mozilla’s internal “Secret OPerationS” tool, went quiet in 2021–2022 and moved to the CNCF in 2023. It handles structured data (YAML, JSON, INI, ENV), allows per-file access control through identities and lets you edit encrypted files interactively. The pitch for using it in Ansible: secrets stay as files in the repository, identities use AGE/PGP public keys, and there’s tooling for re-encryption and re-keying.
  • The community.sops collection: a vars plugin for encrypted host/group vars, a lookup and a filter to decrypt, an action to load encrypted vars files (“as facts, not as variables”), an idempotent module to encrypt, and a role and playbook to install SOPS.

The comparison slide summed it up: Ansible Vault “just works”, SOPS has broad support, and client-server solutions go through a lookup plugin. The wish list was SOPS-aware replacements for host_vars/group_vars directories, vars_files and ansible.builtin.include_vars.

A speaker presenting the community.sops collection features in a classroom at CfgMgmtCamp 2025 in Ghent

What community.sops gives you: vars plugin, lookup, filter, action, module, role and playbook.

My take: for GitOps-style repositories I prefer SOPS to Ansible Vault. Per-recipient public keys and re-keying make it much easier to give a CI runner or a new teammate access without sharing one vault password. I compare the options in my Ansible Vault tutorial.

Making compliance bearable: Ansible Lockdown

The last Ansible talk I photographed on Monday covered the Ansible Lockdown project (on GitHub and Ansible Galaxy):

  • Open-source security and compliance automation, written as Ansible roles, that remediates systems towards industry-recognised baselines. The slide counted “500,000+ downloads from galaxy alone”.
  • Auditing uses Goss (Go Server Spec): a single, self-contained Go binary of about 14 MB with no extra dependencies. That makes it good for air-gapped environments. It runs on the host rather than across the network.
  • Speed: 700 tests in under two minutes on a clean OS build, with sub-10-second runs seen. The YAML config is driven by Ansible and Jinja2.

Tuesday 4 February

Pkl

The morning started with Pkl (pkl-lang.org), pitched as “Configuration that is Programmable, Scalable, and Safe”, which can “generate any static configuration format” from data defined once in Pkl.

The live demo ran in an IDE, where Pkl packages are downloaded and cached locally so the documentation, types and source are a click away instead of on a website. The speaker defined a Prometheus deployment that builds on a Prometheus package, with the prometheus.conf entry of a ConfigMap filled from the package’s rendered text output (prometheus.output.text). “I don’t have to think about how many layers of escaping I need,” the speaker said, and of doing the same by hand: “If you have to do this in YAML, you’re fairly depressed by the end of it.”

ARA Ansible for the teams

Kirill Satarin, Principal Software Engineer at Red Hat (with the footnote “presenting personal opinion”), made the case for ARA under the tagline “Most useful development tool for Ansible content that is rarely used is even better for the teams”:

  • The audience: people who use Ansible and write content (collections, roles, plugins, modules), and want to use ARA on their own and with colleagues.
  • With ARA you don’t need debug tasks in your playbooks, roles and tasks.
  • You don’t need to read badly formatted console output.
  • Everything is persisted, easy to search and shareable.

On stage he priced a shared setup on Azure. The smallest PostgreSQL flexible server came to about 16–17 a month and the compute for both containers to about 70, within the slide’s figure of less than 100 EUR a month. That buys a persistent history of every run, wherever it’s started from, once five environment variables are set. The downsides he named: every task sends data to the server, so playbooks with many tasks take noticeably longer, and someone has to pay for it. His summary was that everyone who uses Ansible or writes Ansible content should use ARA. Locally it takes less than five minutes to set up, and running it “globally” for the team is where it pays off.

Kirill Satarin presenting ARA Ansible for the teams in the main auditorium at CfgMgmtCamp 2025 in Ghent

ARA: record every playbook run so the whole team can search it.

My take: I agree that ARA is underused. Once a team has a shared ARA server, “can you paste me the output?” becomes a link, and debugging a failed run stops depending on whoever happened to have the terminal open.

Positive psychology for DevOps teams

James Freeman brought positive psychology to a config management crowd. He started with the glass half empty or half full, moved the question on from “What’s wrong?”, and showed Martin Seligman’s PERMA model: positive emotions, engagement, positive relationships, meaning, accomplishment. His “Relationships” slide tied it back to Ansible: contribution and a strong community. (He also talks about this in his interview in the companion post.)

In the talk he added hedonic adaptation: the new phone that makes you happy feels normal two weeks later. He also cited the Losada “positivity ratio”, the claim that companies with fewer than three positive interactions for every negative one are statistically likely to go bankrupt, and he warned that you can also be too positive and naive. (Worth knowing: the maths behind that ratio was challenged in 2013, and the modelling in the original paper was withdrawn.) He closed on the spirit of contribution in open source, and on the line “Coding is a creative activity as far as I’m concerned.”

Salt at Cloudflare: “Have your waffle and eat it too”

A talk credited on its thank-you slide to the Cloudflare Platform Configuration team:

  • Backstory: the speaker joined SaltStack in 2014 and asked about Salt’s DSL. The original developers said Salt doesn’t have one (“Users can BYO DSL”), while users said “jinja|yaml is the DSL”.
  • The config shift: production keeps drifting towards many sources of authority (SoAs), such as Vault and NetBox.
  • The finding: there was “almost no variation in realized configs”. Users don’t need the full systems × configs permutation space, so they already had O(config) delivery complexity without knowing it, because the tooling hid it.
  • Open problems: bringing production SoAs into change management and into the dev environment, health-mediated deployment for config changes with OpenSLO, and the rendered config format and delivery pipeline.

Cloudflare Platform Configuration team slide at CfgMgmtCamp 2025 titled Have your waffle and eat it too, about maintaining O of config delivery complexity

James Freeman presenting the PERMA model of positive psychology in the main auditorium at CfgMgmtCamp 2025

Left: Cloudflare found almost no variation in realised configs. Right: James Freeman on PERMA.

My take: the waffle slide is the most useful one I photographed all week. Most fleets have far fewer real configurations than their templating suggests. Rendering and diffing the actual output, rather than reasoning about every input combination, is how I’d approach config change review on any platform.

Two more short talks

  • Asfaload: a proposal for an easy-to-use, multi-signature file authentication system for downloads from the internet, where “only the URL of the file to be downloaded is needed”.
  • “Let Red be Red and Green be Green”: DĂĄna Doherty, Software Engineer at Puppet by Perforce, on transient CI errors. GitHub doesn’t let you automatically rerun entire workflows, so the team built a GitHub Actions Workflow Restarter and showed the test workflow behind it.

Ansible and CyberArk Conjur

In a smaller room, James Freeman showed how to integrate Ansible with CyberArk Conjur (his slides carry his copyright):

  • Prep: install the cyberark.conjur collection, fetch the Conjur server certificate and create a host factory token, then grant the Ansible control node its own host identity.
  • What he learned: you can’t keep the Conjur CLI logged in as a user on the control node once it’s integrated, because the lookup uses the host identity, not yours.
  • The lookup plugin currently fails to verify the Conjur server’s TLS certificate if you use Let’s Encrypt/ACME. His workarounds: turn off TLS verification (“not good”), export CONJUR_CERT_FILE pointing at the system CA bundle, or extract the whole CA chain into /etc/conjur.pem.

Two spoken asides filled in the slides. In Conjur, an account isn’t a user. It’s the top-level object in the hierarchy, holding the variables, policies and users, so naming his first one after himself turned out to be a bit odd. And the role stores the downloaded CA certificate under /etc for you, so it only needs to sit somewhere temporary first. In a lab you can turn validation off, and even work with IP addresses if DNS isn’t set up.

You’re Doing Ansible Roles All Wrong

This talk opened with “The State of Ansible Roles”:

  • Roles started as part of playbook projects, with basic support from Ansible Galaxy and GitHub.
  • Collections arrived in Ansible 2.8 (December 2018), bundling modules, plugins and roles under namespaces, yet standalone roles are still widely used.
  • Roles see extensive use in large deployments (“hundreds, even thousands”), but they lag behind modules and plugins in functionality.

tox-ansible for CI

The last Ansible talk I photographed on Tuesday was about testing, under the heading “Taking CI/CD testing to the next level (Option 3)”:

  • tox-ansible isn’t only for local testing. It can also generate GitHub test matrix jobs.
  • You no longer need to update GitHub workflows for new Ansible or Python versions.
  • You control the test matrix locally.

A speaker presenting tox-ansible as a way to generate GitHub test matrix jobs at CfgMgmtCamp 2025 in Ghent

tox-ansible: one test matrix, locally and in GitHub Actions.

Wednesday 5 February

On the third day I joined a hands-on InstructLab session. The chalkboard pointed to instructlab.ai and ibm.com/granite for the IBM Granite models, and the screen showed an ollama list with DeepSeek-R1, several Granite variants and an InstructLab Granite 7B lab model. Later slides covered Chroma, “the open-source AI application database”, and an nvidia-smi view of an NVIDIA A100-SXM4-80GB instance. A slide titled “DeepSeek-R1 Release” showed how recent all of it was: the same model had been a contestant in Monday’s AI leaderboard.

The presenter ran the models locally with Ollama. He noted that its commands feel familiar if you know Docker, because it comes from people who worked on Docker. His list included DeepSeek-R1, Granite 3.0 and 3.1 and Granite Code, IBM’s coding model. For chat he preferred AnythingLLM to Open WebUI, because its API was easier to script. On his 64 GB Mac M3 Pro he usually runs 8B models and sometimes 14B, while the full 671B DeepSeek-R1 needs a machine costing around $100,000. The hard part, he said, is finding a practical use case. His was scoring text: he gave the model examples labelled 100% and 50% sales pitch, and it returned a graded score with its reasoning, which he collected into a CSV.

On fine-tuning versus RAG he described an “accordion problem”. Stable facts belong in the model through the InstructLab workflow, at the cost of more GPU time and hardware. Facts that change often belong in RAG. His live demo asked about the 2024 US budget, which the model couldn’t answer because its training ended in April 2024. Once he had embedded the budget PDF on the laptop, it answered with citations and match percentages.

What I took away

Two threads ran through the three days. AI showed up in the configuration management programme, from model leaderboards to InstructLab on a GPU. Meanwhile the established tools (Chef, Salt, Puppet and Ansible) were doing the unglamorous work: queueing, caching, secrets, CI retries, test matrices.

My take: that mix is the right one. AI-generated playbooks still need the same testing, secrets handling and change review as hand-written ones, and those were exactly the areas the Ansible and Salt talks covered.

Free 30-min Production AI consultation

Book Now