The Tuesday morning keynote at Red Hat Summit 2026 in Atlanta (12 May) was titled “The next platform is choice” on Red Hat’s recording, and the official line-up lists Matt Hicks, Ashesh Badani and Chris Wright as presenters. I filmed several short clips from the audience. The stage slides did not show speaker names, so below I describe what was said and shown by segment and do not attribute individual lines to named people. Numbers are the speakers’ claims on stage; the official announcement of Red Hat AI 3.4 is the better reference for product details.
The opening: the thing that cannot break
The first speaker addressed a room full of people responsible for systems that “cannot break, cannot go down, cannot be wrong”, from financial transactions to healthcare networks. The framing was that new initiatives, organisational friction and a critical legacy system all hit at the same moment, and that this combination is the villain in the room. A warning followed that neither extreme works: moving so slowly that you never catch up, or so fast that you build on a foundation you cannot extend.
Red Hat’s own AI journey

The “Red Hat’s AI journey” slide opens the internal case study.
A segment on Red Hat’s own use of AI described an internal agent application that grew from about ten agents to almost two hundred. The speaker walked through how the team replaced frontier-model calls layer by layer with smaller open models running on its own infrastructure: document search first, then hallucination detection, then safety management (controlling what the system would answer and commit to), and finally planning, where an agent decides which other agents to ask. Knowledge and even code had been contributed to the deep research agent by teams across the company, including legal, inside sales and operations.

The “Self-hosted models” slide: control AI costs, control AI security, control corporate data.
My take: this is the pattern I see in customer platforms too. Start with a frontier model to prove the idea, then move the narrow, high-volume steps (classification, safety checks, retrieval) to smaller models you can host and measure.
Three speed examples and token economics
A later speaker said the pace of AI change is the point to plan for. Three examples were given, with the gap between a closed breakthrough and an open answer shrinking each time: eight months, five months, then two months for open agents. The speaker pointed to agentic coding tools taking off in late November, and named OpenClaw as the open source agent project that became the fastest growing open source project ever two months later. The advice was to build flexibility into the AI strategy.
The economics argument went like this, as stated on stage:
- Per-token prices fall 75 to 90 percent a year, but token consumption can rise over 500 percent a year.
- Reasoning models were described as consuming 10 to 20 times more tokens than standard models, and agents roughly another 5 times, because they plan, call tools and loop.
- So the organisations that win are the ones that move from token consumer to token provider, running self-hosted models for the jobs where that makes sense.
Metal to agents
The core of the keynote was a layered stack, called “metal to agents”, running from accelerators to agents, on any accelerator and any cloud. As presented:
- AI infrastructure: RHEL as the Linux foundation, OpenShift as the Kubernetes platform for VMs, containers and AI workloads, with network isolation and GPU sharing so expensive accelerators are not stranded between inference calls.
- Inference services: vLLM, which Red Hat described itself as the leading contributor to, plus llm-d for distributed, KV-cache-aware routing. The stated one-year improvement was three times more output tokens and ten times faster time to first token; a paper with more detail was mentioned. For the routing details see my notes on llm-d KV-cache routing.
- Model services: models as a service through an AI gateway, with token quotas, credentials, priorities and rate limits in one control plane, and a Validated Models programme described as more than 20 collections and nearly 700 variants checked on vLLM.
- Data services: retrieval, fine-tuning and integrated evaluation.
- Agent services: the “bring your own agents” idea. Whatever the team uses (the speaker mentioned Claude Code and Codex among the tools staff are already downloading), the platform should give every agent a verified identity, a versioned lifecycle, MCP tool access and tracing, turning bring your own agents into agents as a service.
A conversation with NVIDIA

The NVIDIA conversation on the keynote stage.
The NVIDIA guest described AI infrastructure as a five-layer cake: energy, chips, infrastructure, software infrastructure, then models and agents, all of which must be optimised together to convert power into tokens. The discussion then moved to security. The two companies are collaborating on OpenShell, a sandboxed, zero-trust runtime for agents, and the guest described it as a governance layer with fine-grained policy, comparable to the role-based access control we give employees. Confidential computing in the Red Hat AI Factory with NVIDIA was also mentioned, so that frontier model providers could run on infrastructure the customer owns. The guest also said that all 40,000 employees had been given access to a first agent. For the detail of OpenShell and the confidential computing stack, I covered the session in Red Hat and NVIDIA AI Factory at Summit 2026.
What I take away
- Plan for model churn. Platform choices that lock you to one model or one agent framework will age quickly.
- Treat tokens like a budget. Routing, caching and right-sized models are cost controls, not optimisations for later.
- Agent governance is infrastructure work: identity, policy, sandboxing and traces belong in the platform, not in each agent.