Skip to main content
🚀 Taking AI from prototype to production? Find the architecture, GPU, security and governance gaps before they become incidents. Get a Production AI Readiness Assessment
Slide showing the NVIDIA DRA driver for GPUs sharing one GPU among two containers with a ResourceClaimTemplate
AI

KubeCon London 2025: DRA GPU Sharing and CERN Kubeflow

From the AI co-located day at KubeCon London 2025: NVIDIA DRA driver examples for GPU sharing and MIG, and how CERN runs ML challenges on Kubeflow.

LB
Luca Berton
· 6 min read

On Tuesday 1 April 2025, the co-located day before the main KubeCon + CloudNativeCon Europe programme at ExCeL London, I spent the afternoon in the AI-focused sessions. Two stand out when I go back through my photos: a talk on Dynamic Resource Allocation (DRA) with NVIDIA’s GPU driver, and a Kubeflow Summit talk on how CERN runs machine-learning challenges on Kubernetes. They connect, because CERN’s roadmap slide lists DRA as a next step.

Kubeflow at CERN: ML challenges as pipelines

The Kubeflow Summit session was “Streamlining Competitive Data Science at CERN: Running ML Challenges With Kubeflow”, presented by Raulian-Ionut Chiorescu and Hannes Hansen of CERN. The title slide carries their names and the date, 1 April 2025, London.

The title slide Running ML Challenges With Kubeflow from Kubeflow Summit Europe, with two CERN presenter names and the date 1 April 2025

The opening slide at Kubeflow Summit Europe.

The slides lay out the idea. Particle-physics work at CERN uses machine learning for tasks such as track finding and fitting, reconstruction of particle positions and anomaly detection. A challenge platform gives the community a central entry point, promotes collaboration across teams, tasks and datasets, and lowers the barrier to taking part. Participants submit code as a container image, and the platform scores it. The “Why on-premise?” slide gives three reasons: resource availability and control, a shared pool of resources with fair boundary conditions for all participants, and full control over sensitive datasets and code.

A slide titled MLOps at CERN listing Kubeflow for notebooks, distributed training, hyperparameter optimisation, serving and pipelines, integrated with CERN internal tools and OIDC

“MLOps at CERN”: Kubeflow as the main service for ML workloads, integrated with internal storage and OIDC.

The “MLOps at CERN” slide says Kubeflow is the main service for running ML workloads at CERN: interactive notebooks, distributed training, hyperparameter optimisation, serving and pipelines. It is integrated with CERN’s internal tools, namely the storage that hosts the challenge data and authentication and authorisation through OIDC. A following slide adds that Kubeflow Pipelines orchestrates the challenge runs and that Kyverno validates and mutates resources, for example to override home directories in Jupyter notebooks.

I also have a recording of the audio from the live demo and part of the platform description. As I heard it, the platform is heterogeneous: nodes with many CPU cores, a few T4 GPUs for notebooks, and V100, A100 and H100 GPUs, plus AMD GPUs and possibly some ARM devices. Each user gets an automatically created personal profile with a quota of one GPU and 30 GB of block storage. Team profiles with higher quotas are available on request. The speakers described deploying Kubeflow with GitOps and Argo and managing secrets with Vault. They also described a custom controller that injects and refreshes the OIDC tokens needed to mount the storage system as the notebook home directory, and sidecar restart policies that pod defaults did not yet support, which they hoped to contribute upstream.

The demo created a challenge from a name, a downloader image for external data, a scoring image and a Markdown description. The example challenge was the classic iris classification problem. A participant then submits a solution by naming a container image and the command to run, and the platform starts a pipeline: download and prepare the dataset, run the user code, run scoring on the prediction and submit the score to the leaderboard. In the demo the score was plain accuracy.

A CERN roadmap slide listing automated team profiles, better GPU availability and sharing with MPS and DRA, Kueue integration and a dedicated team serving entrypoints

The roadmap slide: automated team profiles, MPS added to existing MIG setups, DRA, MultiKueue and a dedicated team for entrypoints.

The roadmap slide ties this to the rest of the afternoon. It lists automated management of team profiles, improved GPU availability and resource sharing, with improved NVIDIA sharing policies (adding MPS to existing MIG setups) and Dynamic Resource Allocation “working closely with upstream Kubernetes/NVIDIA”, integration with Kueue for better job scheduling, MultiKueue for external resources, and a dedicated team serving entrypoints. For background on the sharing modes, see my GPU sharing guide for MIG, MPS and time-slicing.

A short stop at the edge

In the same Kubeflow Summit track, the schedule lists “Using Training-Operator To Schedule Distributed Edge-Cloud Collaborative AI Applications” by Bincheng Wang of Huawei and Ming Tang of DaoCloud, at 15:55 BST. The slides I photographed describe the KubeEdge project as designed for edge-cloud collaboration and the first graduated edge computing project in the CNCF, which matches the KubeEdge website. I do not have enough from that talk to say more.

Dynamic Resource Allocation with NVIDIA’s GPU driver

The DRA talk was on the Cloud Native + Kubernetes AI Day stage; the slide template carries that logo. The title slide, headed “people & governance”, lists John Belamaric of Google, Patrick Ohly of Intel and Kevin Klues of NVIDIA, and says the Working Group Device Management was formed in 2024. I could not match the session to a schedule entry, so I describe it from the slides only. Two presenters were on stage and I cannot say which of the listed people spoke.

The opening Dynamic Resource Allocation slide naming John Belamaric of Google, Patrick Ohly of Intel and Kevin Klues of NVIDIA, with Working Group Device Management formed in 2024

The DRA slide listing the working group’s people.

The Kubernetes documentation describes DRA as a way to request and share devices such as accelerators among Pods, similar to dynamic volume provisioning. It involves four API objects: ResourceClaim, ResourceClaimTemplate, DeviceClass and ResourceSlice. The slides follow the same model.

A slide titled Dynamic Resource Allocation introducing ResourceClaim and ResourceClaimTemplate as new primitives for requesting devices

“New primitives for requesting devices”: ResourceClaim and ResourceClaimTemplate.

The next three slides are examples for NVIDIA’s DRA driver for GPUs.

Example 1: share one GPU among two containers. A ResourceClaimTemplate named one-gpu-rct requests the gpu.nvidia.com resource class. A Pod with two containers refers to the same claim, so both containers get the same GPU. The footnote says it works in Kubernetes 1.32.0 with NVIDIA DRA Driver for GPUs 25.3.0-rc.2.

Slide showing a ResourceClaimTemplate for one GPU and a Pod whose two containers share the claim

Example 1: one GPU shared by two containers through a single claim.

Example 3: request a dynamically provisioned MIG device. A ResourceClaim asks for a “mig-4g-20gb” device of class mig.nvidia.com and selects it with a CEL expression on the device’s profile attribute. The slide annotates it as powerful for maximising cluster-global resource utilisation, and the footnote refers to Kubernetes 1.33.0. NVIDIA’s MIG documentation explains the partitioning behind it, and I wrote about the operator side in MIG partitioning with the NVIDIA GPU Operator.

Slide showing a ResourceClaim requesting a MIG device with a CEL selector on the GPU profile

Example 3: a MIG device requested by profile, with a CEL selector.

Example 4: a Multi-Node NVLink group through a ComputeDomain. A ComputeDomain object with two nodes refers to a ResourceClaimTemplate, and the slide notes that the driver automatically provisions, just in time, a securely isolated IMEX channel that the workload pods can use.

Slide showing a ComputeDomain resource for a Multi-Node NVLink connected GPU group and the Job that uses it

Example 4: a ComputeDomain for a Multi-Node NVLink connected GPU group.

The live demo for this example was running in the room, and my audio of it matches the slide. As I heard it, on the second node the import-from-shareable-handle call succeeded, which showed that the IMEX channel had been created, and the daemon log confirmed the memory exchange. The speaker said all of that was handled by the compute domain logic of the DRA driver.

A note on versions

These slides date from April 2025, when the API group was still in beta (the examples show resource.k8s.io v1beta1 and v1beta2). DRA has since graduated to stable in Kubernetes v1.34. If you copy anything from the slides, check the field names against the docs for your cluster version, and check the driver README for which features are supported at the moment. When I looked, the README described ComputeDomains as supported and GPU allocation as still experimental.

My take: DRA is the piece that moves GPU sharing from node-level configuration to something a workload can ask for, and the Triton benchmark session the next day was a useful reality check on what sharing costs. That one is in the Triton GPU sharing post.

Free 30-min Production AI consultation

Book Now