Skip to main content
📬 Get weekly Production AI insights Practical notes on Kubernetes, AI infrastructure and platform engineering. No spam. Subscribe free
A speaker beside a Space shifting slide showing the Sailfish Operator, KEDA and a grid intensity formula
Platform Engineering

Carbon-Aware HPC with KEDA: Sustainability Meetup 2024

A July 2024 Amsterdam sustainability meetup: Sailfish-HPC on Kubernetes with KEDA, carbon-aware space shifting, and cutting Azure emissions.

LB
Luca Berton
¡ 3 min read

On 4 July 2024 I went to a sustainability meetup in Amsterdam, held in a venue with an “A LAB” sign on the stage. Two talks are in my photos: an open-source high performance computing (HPC) framework that schedules work by carbon intensity, and a practical session on reducing cloud emissions in Azure. I don’t name the organiser, since it isn’t on the slides I photographed. This is a throwback built from those slides.

Two presenters on a small stage in front of a projected slide, with the A LAB sign beside the stage

The stage at the start of the HPC talk.

Sailfish-HPC: HPC on Kubernetes with KEDA

An “About us” slide introduced the first speakers as Lisette van Leeuwen and Zeid Adabel, both software engineers. Their talk was about Sailfish, an HPC framework built from Kubernetes components such as KEDA and a Knative gateway, as their design slide showed. The demo slide said “It’s open-source!”.

The “Framework Design” slide listed the principles: open-source based, bring your own container, managed operators for components, Kubernetes for orchestration, and an infrastructure layer that “does not matter”. Another slide showed the flow: a Knative gateway API takes a job from the user, a job queue holds tasks, managers split them and workers pick them up.

Framework Design slide: bring your own container, KEDA and other operators, Kubernetes orchestration

Framework Design: open source, bring your own container, operators on Kubernetes.

KEDA is the Kubernetes event-driven autoscaler, a CNCF graduated project. Its site highlights scale-to-zero and sustainability. That fits an HPC queue: workers scale with the number of waiting tasks instead of idling.

Carbon intensity and space shifting

The talk used Green Software Foundation material. The Green Software Foundation is a Linux Foundation nonprofit that publishes standards such as the Software Carbon Intensity metric. The slides built the argument in steps:

  • A slide on global data-centre electricity consumption showed a rise from 460 TWh in 2022 to 1000 TWh in 2026.
  • “Quantifying carbon emissions in cloud” combined energy consumption with carbon intensity to give emissions.
  • A world map showed carbon intensity of electricity grids.

Slide on global carbon emissions with data centre electricity consumption of 460 TWh in 2022 and 1000 TWh in 2026

Global electricity consumption of data centres: 460 TWh in 2022, 1000 TWh in 2026.

Then came the Sailfish Operator and space shifting. The slide showed the rule MIN(grid_intensity[EU_West], grid_intensity[EU_North]). In other words, run the work in the region whose grid is cleaner right now. A “Sustainable Development” slide said the platform “deploys across time and space”, meaning the same operator can also shift when jobs run. A later screenshot showed queries mixing grid carbon intensity with Azure region availability.

Space shifting slide with the Sailfish Operator, KEDA and a minimum of grid intensity across two EU regions

Space shifting: pick the region with the lower grid intensity.

Demo slide with the sailfish-hpc logo and the text It's open-source

The demo slide for sailfish-hpc.

My take: location shifting only works for jobs that tolerate moving, such as batch and HPC runs with no data-residency limits. Latency-bound services can’t take part. I wrote about the scheduler side in Carbon-aware Kubernetes scheduling.

Reducing carbon emissions in Azure

The second talk came from the founder of CloudXcellence, per the bio slide. It started with the simplest lever: “Pick your region wisely”, with a slide comparing the CO2 of different Azure regions.

Slide asking what the simplest way is to reduce any cloud emission, answer: pick your region wisely

“What is the simplest way to reduce any cloud emission? Pick your region wisely.”

Other slides covered:

  • Azure’s carbon optimisation view keeps only a 12-month sliding window of scope 1, 2 and 3 emissions, while preview APIs give up to five years.
  • Top Azure emissions on one slide: storage, virtual machines, Databricks/SQL, logging and Azure Data Factory.
  • Storage tips: move non-production from GRS to LRS and archive after 180+ days.
  • Gartner’s prediction, quoted on a slide, that 75% of organisations will have a data-centre infrastructure sustainability programme by 2027.

A FinOps versus GreenOps slide quoted that FinOps gets better deals on the number of machines but can’t say whether those machines are needed. The overlap was right-sizing and clean-up.

Venn diagram slide of FinOps practices versus GreenOps practices

FinOps vs GreenOps: savings plans and reserved instances on one side, low-carbon regions on the other.

The take-aways slide was short: check your own footprint, start sustainability groups within your company, follow meetups, and learn to write green software, for example in hackathons.

Free 30-min Production AI consultation

Book Now