Skip to main content
🚀 Taking AI from prototype to production? Find the architecture, GPU, security and governance gaps before they become incidents. Get a Production AI Readiness Assessment
A Preferred Networks speaker on the KubeCon Japan 2026 main stage in front of a Looking to the Next Decade slide reading 2016, 2026 and 20XX
Conferences

KubeCon Japan 2026 Keynotes: AI Platforms at Scale

The KubeCon Japan 2026 day 1 keynotes from Yokohama: SoftBank's infinite agents, Fujitsu's GPU-centric Kubernetes, PFN multi-tenancy and Hyundai's Argo CD.

LB
Luca Berton
¡ 9 min read

On Wednesday 29 July 2026, the first full day of KubeCon + CloudNativeCon Japan 2026, I was in the Main Hall at PACIFICO Yokohama for the morning keynotes. I was there as a media partner, so I took photos and notes rather than presenting. In my media-partner preview I said these keynotes were the ones to watch. This post covers what the speakers actually put on screen.

The opening keynote by Chris Aniszczyk and Jonathan Bryce, “Keep Cloud Native Moving: Building Japan’s Platform for Open Innovation”, set the theme: according to its abstract, AI is redefining what comes next for cloud native infrastructure. The keynotes after it were short, between 3 and 10 minutes each according to the official schedule. Each one showed a company running AI or a large fleet on Kubernetes. I’ve grouped them below in the order they ran. Speaker names are written as they appeared on the title slides.

SoftBank: “Infinite Agents, Finite Kubernetes”

The first keynote after the opening was “Infinite Agents, Finite Kubernetes” by Mohammad Mikal Bin Amrul Halim Gan, Platform Engineer, SoftBank Corp. It lasted three minutes and was built around one tension: agents generate effectively unlimited demand, but data centres have a fixed amount of compute, space and power.

The SoftBank keynote title slide Infinite Agents, Finite Kubernetes on the KubeCon Japan 2026 main stage, with the speaker in front of the screen

The title slide, with the Yokohama skyline artwork used across this year’s keynotes.

The slides made the argument in four steps:

  • “Reality of AI Agent.” An LLM drives an agent executor, which calls the Kubernetes API and starts a set of agent-executed pods. The slide said: “The AI agent spins up multiple pods to run parallel inferences and determine the optimal answer.”
  • “Infinite Apps, Finite Compute.” A row of data centres (DC1, DC2 … DCn) with “N ≠ ∞” next to them.
  • “Technical Problems: Managing Tradeoff.” The slide listed Security, Noisy Neighbor and Resource Efficiency, and plotted resource efficiency against isolation: as isolation goes up, efficiency drops.
  • “Finite Hardware, Infinite Capability.” Workspaces spread across three data centres, drawn over a map. The slide before it, “Optimizing Workload”, showed workspaces (WS1 to WS3) being rearranged across two data centres.

SoftBank slide Reality of AI Agent showing an LLM, an agent executor and the Kubernetes API spinning up agent-executed pods

SoftBank slide Technical Problems: Managing Tradeoff with security, noisy neighbor and resource efficiency, and a curve of resource efficiency falling as isolation rises

Left: an agent turns one request into many pods. Right: the trade-off between isolation and efficiency.

SoftBank slide Finite Hardware, Infinite Capability showing workspaces spread over three data centres drawn on a map

“Finite Hardware, Infinite Capability”: spreading workspaces across data centres.

My take: the slide that matters is “Reality of AI Agent”. Platform teams plan capacity per service. An agent that fans out into parallel pods for every request looks more like a batch scheduler than a web service. If your quotas and admission control assume one request means one pod, agents will break that assumption first.

Fujitsu: GPU-centric infrastructure for AI workloads

Next came “The Next Evolution of Kubernetes: GPU-Centric Infrastructure for AI Workloads” by Takao Indoh, Director, Fujitsu Limited. The session abstract says Kubernetes is moving from CPU-based cloud infrastructure to GPU-centric infrastructure. It names GPU shortages and electricity costs as the pressures. It adds that with agentic AI, one request can trigger many downstream tasks, which makes compute demand hard to predict. Fujitsu’s proposed answer is Composable Disaggregated Infrastructure.

The Fujitsu keynote title slide The Next Evolution of Kubernetes: GPU-Centric Infrastructure for AI Workloads by Takao Indoh, Director, Fujitsu Limited

Fujitsu slide Driving the AI Platform with OSS reading Fujitsu will provide trusted technology through innovation enabled by OSS technologies

The Fujitsu keynote: the title, and the closing “Driving the AI Platform with OSS” slide.

The keynote ended on “Driving the AI Platform with OSS”, with the line “Fujitsu will provide trusted technology through innovation enabled by OSS technologies”. Composable disaggregated infrastructure is also the idea behind CoHDI (Composable Hardware in Disaggregated Infrastructure), a CNCF sandbox project. CoHDI came up again at the CNCF press briefing later that day; see my CNCF Japan briefing post.

Preferred Networks: a multi-tenant AI platform with CNCF projects

The longest of the company keynotes, at ten minutes, was “Building a Multi-Tenant AI Platform with the CNCF Ecosystem” by Aya Igarashi (@Ladicle), Preferred Networks, Inc. It was the most useful one for anyone building shared GPU platforms, because it explained specific design decisions.

The talk opened with “The Power of Community and Extensibility”: operators and extensibility combined with the ecosystem to assemble your own platform. Next came PFN’s stack: the slide “PFN’s AI Platform Built on Kubernetes” layered AI chips, the computing infrastructure (PFCP), generative AI foundation models, and solutions and products. “Why We Choose Multi-Tenancy” made the cost argument: accelerators are expensive, so tenants share them.

Preferred Networks slide Balancing Cost and Isolation comparing namespace isolation, virtual cluster, dedicated node and dedicated cluster, with Our Choice above namespace isolation and dedicated node

“Balancing Cost and Isolation”: four options from low cost to stronger isolation, with “Our Choice” marked above two of them.

The core slide was “Balancing Cost and Isolation”. It compared four options, from namespace isolation with a shared control plane, through a virtual cluster and a dedicated node, to a dedicated cluster. The bottom axis ran from lower cost to higher cost and stronger isolation, labelled API isolation, kernel isolation and physical isolation. PFN marked namespace isolation and dedicated nodes as “Our Choice”: a hybrid, not one model for everyone.

The following slides covered what you have to build yourself once you choose namespaces:

  • “Managing Hierarchical Tenants.” A cluster split into Tenant A and Tenant B, each owning project namespaces such as tenant-b--project-3. These are managed with HNC (Hierarchical Namespace), which the slide said is “maintained via PFN fork”. It linked both the retired upstream repository and PFN’s fork.
  • “Securing Every Layer.” Network policies at the network layer, admission policy enforcement at the API layer, and a per-tenant watch scope for controllers, so Tenant A is blocked from reaching Tenant B.
  • “Fairness in a Shared Environment.” “Without Limits, One Tenant Can Take Almost Everything”. “With HRQ + Kueue Quota, Shares Stay Fair”, and “Quota Can Flex With Business Priority”.
  • “Maximizing Cost Efficiency.” Two ideas side by side. Start All Together: a partial start leaves pending pods sitting idle and wasting resources, so start all of a job at once. Bin-Packing: in a fragmented cluster a pending pod “can’t fit”, while in a packed cluster it fits.

Preferred Networks slide Managing Hierarchical Tenants showing tenants A and B with project namespaces managed by HNC

Preferred Networks slide Fairness in a Shared Environment showing that with HRQ plus Kueue quota the shares between tenants A, B and C stay fair

Hierarchical tenants with HNC, and fair sharing with HRQ plus Kueue.

Preferred Networks slide Maximizing Cost Efficiency comparing partial start with all-at-once start, and fragmented with packed bin-packing, with the speaker on stage

Gang start and bin-packing: two ways idle accelerators waste money.

The talk ended with “Looking to the Next Decade”: 2016 as Stateless Apps, 2026 as AI Foundation and 20XX as The Next Shift, shown with a question mark.

Preferred Networks slide Looking to the Next Decade with a timeline of 2016 stateless apps, 2026 AI foundation and 20XX the next shift

Ten years from stateless apps to an AI foundation.

My take: this matches what I see with customers who share GPUs between teams. Picking namespaces is the easy part. The real work is everything the PFN slides listed afterwards: hierarchy, admission policy, controller scoping, quota and gang scheduling. I covered the same trade-off from the hallway interviews in Security, Isolation & Sovereign AI on Kubernetes.

Subaru and LY Corporation

Two keynotes followed that I cover in more depth elsewhere.

Subaru (Ryoji Kobayashi, DevOps Engineer) presented “How Subaru Accelerated AI Model Development for Next-Generation EyeSight with Kubernetes”. The sched abstract describes cutting container image pull time from about three hours to about three minutes. Subaru won this year’s case study contest. See Cloud Native in the Real World: Optics, Subaru & Uber.

LY Corporation (Shota Yoshimura, Senior Platform Engineer) gave “From 5 to 1,300+ Clusters: Declarative Scaling for Private Kubernetes”. The story started in 2016 with VMs on OpenStack. Back then, provisioning often took more than a week, nodes stayed unpatched for long periods, and node failures meant on-call pages “often at 3 a.m.” The fix was to use Kubernetes as the infrastructure control plane, “inspired by Kubernetes’ own Deployment, applied to VMs”.

LY Corporation slide Ten Years Later 2026 showing cluster growth from 5 to more than 1,300 clusters operated by 15 engineers, with one platform engineer operating 90 clusters

“Scale grew without linear team growth”: 1,300+ clusters, 40,000+ nodes and about 1M containers, run by 15 engineers.

The “Ten Years Later — 2026” slide showed 1,300+ clusters, 40,000+ nodes and about 1M containers, operated by 15 engineers. Its punchline: “one platform engineer operates 90 clusters”.

Hyundai AutoEver: one Argo CD, 5,000+ applications

The last keynote before the closing remarks was “Out of the Box, at Multi-Region Scale: How Hyundai Scales Its Platform” by Jaewoo Choi, DevOps Engineer, Hyundai AutoEver. His title slide also described him as an Argo CD maintainer.

The Hyundai AutoEver keynote title slide Out of the Box, at Multi-Region Scale: How Hyundai Scales Its Platform by Jaewoo Choi

Three minutes on running a global platform with the CNCF tools as they come.

“Our Own Private Cloud, Built on Open Source” described HKS (hCloud K8s) as the infrastructure, with Argo CD and Harbor as the delivery platform. Then came the numbers:

  • “One Argo CD”: 5000+ applications and 300+ clusters on a single instance, kept healthy with sharded controllers, jitter tuning and processor tuning. The sched abstract adds that the number of applications grows by hundreds every month.
  • “Multi-Region Harbor”: a South Korea region and a multi-region Harbor connected server to server with pull/push replication, a Harbor proxy and policy-based controls. On the slide, a direct cross-region pull was labelled “10x slower”. The abstract mentions 60k+ artifacts in 5k+ repositories, replicated across three regions to balance data-residency rules against performance.
  • “Tune What You Already Have — Five Nines”: 99.999%, “Use what Argo CD, Harbor, and other CNCF projects already provide”, “Keep tuning as we grow”. The slide ended: “That’s what ‘Out of the Box’ means to us”.

Hyundai AutoEver slide One Argo CD showing 5000+ applications and 300+ clusters handled with sharded controllers, jitter tuning and processor tuning

Hyundai AutoEver slide Multi-Region Harbor showing replication, a Harbor proxy and policy-based controls between a South Korea region and a multi-region Harbor

One Argo CD for 5,000+ applications, and Harbor replicated across regions.

Hyundai AutoEver slide Tune What You Already Have, Five Nines, 99.999 percent, use what Argo CD, Harbor and other CNCF projects already provide

“Out of the Box”: tune before you build.

My take: I liked this one most because it pushes back on the usual advice. Teams often split Argo CD into many instances once it gets slow. Hyundai kept a single instance and tuned it: sharding, jitter and processor settings. Before you build a custom multi-instance control plane, check whether the knobs Argo CD already has are enough.

What tied the keynotes together

Every keynote used Kubernetes as the control plane for something bigger than containers: agent fan-out at SoftBank, GPUs as composable hardware at Fujitsu, shared accelerators at PFN, VMs and 1,300+ clusters at LY, and a 5,000-application delivery platform at Hyundai. None of them claimed to have built a new platform from scratch. They took CNCF and Kubernetes projects (Kueue, HNC, Argo CD, Harbor, custom resources and controllers) and tuned them. That is also how I’d advise most platform teams to approach AI infrastructure in 2026.

The same afternoon CNCF held a press briefing with the Japan numbers and project news. That’s in my CNCF Japan briefing post.

Free 30-min Production AI consultation

Book Now