On Tuesday 1 April 2025, the co-located day before the main KubeCon + CloudNativeCon Europe programme at ExCeL London, I spent the afternoon in the AI-focused sessions. Two stand out when I go back through my photos: a talk on Dynamic Resource Allocation (DRA) with NVIDIA’s GPU driver, and a Kubeflow Summit talk on how CERN runs machine-learning challenges on Kubernetes. They connect, because CERN’s roadmap slide lists DRA as a next step.
Kubeflow at CERN: ML challenges as pipelines
The Kubeflow Summit session was “Streamlining Competitive Data Science at CERN: Running ML Challenges With Kubeflow”, presented by Raulian-Ionut Chiorescu and Hannes Hansen of CERN. The title slide carries their names and the date, 1 April 2025, London.

The opening slide at Kubeflow Summit Europe.
The slides lay out the idea. Particle-physics work at CERN uses machine learning for tasks such as track finding and fitting, reconstruction of particle positions and anomaly detection. A challenge platform gives the community a central entry point, promotes collaboration across teams, tasks and datasets, and lowers the barrier to taking part. Participants submit code as a container image, and the platform scores it. The “Why on-premise?” slide gives three reasons: resource availability and control, a shared pool of resources with fair boundary conditions for all participants, and full control over sensitive datasets and code.

“MLOps at CERN”: Kubeflow as the main service for ML workloads, integrated with internal storage and OIDC.
The “MLOps at CERN” slide says Kubeflow is the main service for running ML workloads at CERN: interactive notebooks, distributed training, hyperparameter optimisation, serving and pipelines. It is integrated with CERN’s internal tools, namely the storage that hosts the challenge data and authentication and authorisation through OIDC. A following slide adds that Kubeflow Pipelines orchestrates the challenge runs and that Kyverno validates and mutates resources, for example to override home directories in Jupyter notebooks.
I also have a recording of the audio from the live demo and part of the platform description. As I heard it, the platform is heterogeneous: nodes with many CPU cores, a few T4 GPUs for notebooks, and V100, A100 and H100 GPUs, plus AMD GPUs and possibly some ARM devices. Each user gets an automatically created personal profile with a quota of one GPU and 30 GB of block storage. Team profiles with higher quotas are available on request. The speakers described deploying Kubeflow with GitOps and Argo and managing secrets with Vault. They also described a custom controller that injects and refreshes the OIDC tokens needed to mount the storage system as the notebook home directory, and sidecar restart policies that pod defaults did not yet support, which they hoped to contribute upstream.
The demo created a challenge from a name, a downloader image for external data, a scoring image and a Markdown description. The example challenge was the classic iris classification problem. A participant then submits a solution by naming a container image and the command to run, and the platform starts a pipeline: download and prepare the dataset, run the user code, run scoring on the prediction and submit the score to the leaderboard. In the demo the score was plain accuracy.

The roadmap slide: automated team profiles, MPS added to existing MIG setups, DRA, MultiKueue and a dedicated team for entrypoints.
The roadmap slide ties this to the rest of the afternoon. It lists automated management of team profiles, improved GPU availability and resource sharing, with improved NVIDIA sharing policies (adding MPS to existing MIG setups) and Dynamic Resource Allocation “working closely with upstream Kubernetes/NVIDIA”, integration with Kueue for better job scheduling, MultiKueue for external resources, and a dedicated team serving entrypoints. For background on the sharing modes, see my GPU sharing guide for MIG, MPS and time-slicing.
A short stop at the edge
In the same Kubeflow Summit track, the schedule lists “Using Training-Operator To Schedule Distributed Edge-Cloud Collaborative AI Applications” by Bincheng Wang of Huawei and Ming Tang of DaoCloud, at 15:55 BST. The slides I photographed describe the KubeEdge project as designed for edge-cloud collaboration and the first graduated edge computing project in the CNCF, which matches the KubeEdge website. I do not have enough from that talk to say more.
Dynamic Resource Allocation with NVIDIA’s GPU driver
The DRA talk was on the Cloud Native + Kubernetes AI Day stage; the slide template carries that logo. The title slide, headed “people & governance”, lists John Belamaric of Google, Patrick Ohly of Intel and Kevin Klues of NVIDIA, and says the Working Group Device Management was formed in 2024. I could not match the session to a schedule entry, so I describe it from the slides only. Two presenters were on stage and I cannot say which of the listed people spoke.

The DRA slide listing the working group’s people.
The Kubernetes documentation describes DRA as a way to request and share devices such as accelerators among Pods, similar to dynamic volume provisioning. It involves four API objects: ResourceClaim, ResourceClaimTemplate, DeviceClass and ResourceSlice. The slides follow the same model.

“New primitives for requesting devices”: ResourceClaim and ResourceClaimTemplate.
The next three slides are examples for NVIDIA’s DRA driver for GPUs.
Example 1: share one GPU among two containers. A ResourceClaimTemplate named one-gpu-rct requests the gpu.nvidia.com resource class. A Pod with two containers refers to the same claim, so both containers get the same GPU. The footnote says it works in Kubernetes 1.32.0 with NVIDIA DRA Driver for GPUs 25.3.0-rc.2.

Example 1: one GPU shared by two containers through a single claim.
Example 3: request a dynamically provisioned MIG device. A ResourceClaim asks for a “mig-4g-20gb” device of class mig.nvidia.com and selects it with a CEL expression on the device’s profile attribute. The slide annotates it as powerful for maximising cluster-global resource utilisation, and the footnote refers to Kubernetes 1.33.0. NVIDIA’s MIG documentation explains the partitioning behind it, and I wrote about the operator side in MIG partitioning with the NVIDIA GPU Operator.

Example 3: a MIG device requested by profile, with a CEL selector.
Example 4: a Multi-Node NVLink group through a ComputeDomain. A ComputeDomain object with two nodes refers to a ResourceClaimTemplate, and the slide notes that the driver automatically provisions, just in time, a securely isolated IMEX channel that the workload pods can use.

Example 4: a ComputeDomain for a Multi-Node NVLink connected GPU group.
The live demo for this example was running in the room, and my audio of it matches the slide. As I heard it, on the second node the import-from-shareable-handle call succeeded, which showed that the IMEX channel had been created, and the daemon log confirmed the memory exchange. The speaker said all of that was handled by the compute domain logic of the DRA driver.
A note on versions
These slides date from April 2025, when the API group was still in beta (the examples show resource.k8s.io v1beta1 and v1beta2). DRA has since graduated to stable in Kubernetes v1.34. If you copy anything from the slides, check the field names against the docs for your cluster version, and check the driver README for which features are supported at the moment. When I looked, the README described ComputeDomains as supported and GPU allocation as still experimental.
My take: DRA is the piece that moves GPU sharing from node-level configuration to something a workload can ask for, and the Triton benchmark session the next day was a useful reality check on what sharing costs. That one is in the Triton GPU sharing post.
