Skip to main content
🤖 Running agents for a team, not just yourself? Get an independent review of identity, secrets, failover, observability and governance. Assess your agent platform
Slide on hierarchical control for agent fleets: control tower, fleet and agent layers with owners and capabilities
AI

AI Agents in Production: What 100+ Event Talks Taught Me

What I learned about running AI agents in production from 100+ event write-ups, 2025 to 2026: workflows, memory, evals, AgentOps, commerce and approvals.

LB
Luca Berton
¡ 14 min read

I have published 109 event write-ups tagged AI agents, from a February 2025 demo night in Amsterdam to a CTO breakfast on 1 October 2026. I attended or covered almost all of them. I spoke at three: an OpenClaw demo evening in March 2026 and keynotes in Skopje and for Clarkson Hyde Global in September 2026.

This post is the synthesis, grouped by theme. Every claim comes from a write-up I already published, and vendor figures stay labelled as such. Security has its own piece, MCP and AI Agent Security: What 50+ Talks Taught Me, so I touch it only briefly.

1. The demo-to-production gap is mostly about people

In July 2025, at the MLOps Community agent tour, the orq.ai co-founder said the agents really in production were in typical areas such as sales and lead research. Sensitive sectors, he named pension funds and insurers, were not shipping yet. Others agreed the hype was ahead of reality (write-up).

In November 2025 the Just Eat Takeaway.com CTO, Mert Öztekin, gave the talk that reframed this for me. His slide called the gap “Cultural Lag”: technology moves like a sports car, culture like a saloon. The company’s enterprise chatbot dashboard showed 69.44% overall adoption, alongside a volunteer AI ambassador programme (HOPE and Agents in Production). My own take from that talk: if performance reviews still reward the old way of working, the dashboard plateaus.

The 2026 numbers are harder to reconcile. At AGNTCon in September, the AAIF Member Pulse Survey (n=186) reported 81% of respondents with agents in production today, 51% at production scale (recap). In May, the Red Hat Summit AgentOps lab cited Capgemini 2025: 82% of tech executives planned to integrate agents within one to three years, but under 10% had production-grade deployments (lab). The posts do not reconcile them. My reading is that the survey sampled people who already build agents.

What the sources agree on is how to start. A Zoku talk in January 2026 listed the pattern one speaker saw working with clients: a high-friction workflow with a measurable KPI, controls where risk lives, quality as engineering, and design for adoption (Zoku). My own Clarkson Hyde keynote said the same for accountants: pick one narrow, high-volume workflow, with document reconciliation the easiest (keynote). My Skopje keynote argued that a model working in a demo is not yet a product, and a platform underneath is what makes it one (Skopje).

Takeaway: treat adoption as a measured change programme with an owner, not a tooling rollout.

2. Workflow first, agent second, and watch the bill

The most useful advice came early. In July 2025 Rafael V. Pierre, describing an agent for a data platform, said not to use agents unless the use case justifies it, because workflows are more predictable. He noted that agents abstract complexity “at the expense of tokens” (write-up).

In January 2026 an Elastic architect drew the same ladder as architecture. Stage one is data-driven search, stage two is RAG, where the slide said the prompt is the only output control, and stage three is agents with MCP tools, where output is “guided by the agent and tools”. (Elastic).

The disagreement is about multi-agent designs. The AAIF survey said 60% of respondents run multi-agent systems. Yet in March 2026 Orq.ai told The Future of Product audience that, with the same model, a multi-agent version was 2.5x more expensive and worse on their evaluator (Future of Product).

In June 2026 the Red Hat Tech Day keynote called the harness the operating system: the LLM is the CPU, the context window is RAM, and the harness decides what the model sees. It listed ten harness components, from the agentic loop to observability and evals (Tech Day). At PlatformCon London, one session split work into agent paths, deterministic paths and hybrid paths, where deterministic output feeds back into the model in a loop (PlatformCon). At the July 2026 docomo session the principle was that the LLM is the reasoning engine, not the executor, and agents coordinated through Kubernetes CRDs rather than a message broker (Japan).

On cost, the May 2026 Red Hat Summit keynote said, as stated on stage, that per-token prices fall 75 to 90 percent a year while consumption can rise over 500 percent. Reasoning models were said to use 10 to 20 times more tokens, agents about 5 times more again. Red Hat’s own agent app grew from about ten agents to almost two hundred, moving calls from frontier models to smaller open ones layer by layer (keynote). The hardware side is in GPU Infrastructure and LLM Inference: Lessons from Conference Talks.

Takeaway: start with a deterministic workflow, add one agent where it earns its place, and price the multi-agent version before you build it.

3. Context and memory are the real bottleneck

In November 2025 a Redis talk titled “Not all context is good context” argued that bigger inputs hurt performance and budget, then demoed semantic caching (HOPE and Agents in Production). In March 2025 Stephen Chin of Neo4j said a graph can be both retriever and memory for agentic systems (interview).

By 2026 memory had become a talk topic of its own. At OpenClaw Builders in February, a memory demo showed three kinds of recall from Markdown files the agent wrote: factual, multi-hop and temporal. My take there was that multi-hop is the hard one (OpenClaw Builders). In May, an MLOps evening covered four flavours of memory (working, episodic, semantic, procedural). One Qdrant slide, vendor benchmark, showed the best chunker depends on the corpus. The Booking.com speaker described a stale index: a deleted note still ranked first, and the fix was verify-on-read. His line: “Agents don’t need bigger context windows. They need a librarian.” (MLOps)

Tool vendors moved the same way. Cursor described dynamic context discovery in February 2026, where large tool outputs go to files the agent reads selectively. Cursor’s own test reportedly cut MCP tool token use by almost 47% (Cursor). At the 1 October 2026 CTO breakfast a slide put “Data + meaning” as the gap, with 7% of enterprises saying their data is ready for AI (CTO Network). A Stream speaker’s slide, sourced from Mintlify, showed .md requests at 54.4% and read it as most documentation traffic now coming from machines (AI Builders).

Takeaway: invest in retrieval quality, invalidation and your data’s meaning before buying a bigger context window.

4. Evaluation grew from prompts to trajectories

In February 2025 orq.ai opened a demo with the $1 Chevy Tahoe screenshot and ran airline questions through four prompt and model variants with JSON-schema and tone evaluators. My take: a prompt change is a code change (Tinkerers). By July 2025 the same company described evals as a testing pyramid, with programmatic checks at the base, an LLM judge in the middle and humans on top. The pitch included reusing evaluators online as guardrails (MLOps).

In 2026 the talks asked harder questions. A January slide asked for a “Definition of Done” for AI output (Zoku). In March the Orq.ai speaker added “Do not forget to test your evals too”. In May LangWatch showed a 15-scenario suite with a 53% pass rate on day one, “every red tile is a bug I would’ve shipped” (MLOps).

Others pushed on feedback loops. At the Claude Code meetup in February, one speaker’s advice was to give Claude a way to verify its work, because prompting alone won’t make a non-deterministic system deterministic (Claude Code). Another speaker that month said the hardware does not change, your feedback loop does (10x). In July the docomo talk gated autonomy on evals for tool usage, grounding, goal success and resilience, aiming at a higher autonomy level for known failure patterns (Japan).

The sharpest evidence came in September 2026. Dexter Horthy’s keynote, titled “There’s no Dark Factory Without Better Verifiers”, cited SlopCodeBench, whose paper reports the best agent passing 14.8% of checkpoints (Day 2). A Red Hat demo made the same point: an agent can write code that meets the spec and still add no meaningful tests (Summit day 2).

Takeaway: one green run tells you little. Keep a fixed scenario suite, score tool use and long-run quality, and test the evaluator.

5. Observability is now the AgentOps conversation

In 2025 this was about tracing your own agent. In May 2026 it became about every agent in the delivery system. At SRE NL, Lewis Isaac of Coralogix said “We see production in detail. We barely see the agent that helped build it.” He listed five questions (cost, quality, flow, security, outcome) and said “A session is a trace. Every action is a span.” He also showed three coding agents naming token attributes differently, so a schema gate helps (SRE NL).

The Red Hat lab the same month made the cost of not tracing concrete. A single 200 OK hid more than twelve tool calls and several LLM invocations in an MLflow trace, and the lab named six metrics exposed through Prometheus (lab). PlatformCon London listed agent observation and tool observability as new layers of the platform (PlatformCon).

Takeaway: emit one trace per agent session with identity, model, tokens, cost and outcome, and use OpenTelemetry so it lands next to your other telemetry.

6. Agents work best through the platforms you already run

The strongest 2026 demos added no new infrastructure. At Red Hat Tech Day in June, the Automation Orchestrator preview put it as “AI isn’t improvising against production infrastructure, it’s acting through AAP”. An agent found affected hosts via MCP and proposed a plan. A human approval step, with a timeout that fails the workflow, came before deterministic remediation of CVE-2024-6387 on 12 hosts (Orchestrator). A second demo had OpenClaw choose a pre-approved playbook through the AAP MCP server, behind an OPA policy check and a ServiceNow change record (CVE demo).

The counterexample was staged. At Summit, an agent SSH’d into production and rebooted servers with no ticket and no window, simply because it had credentials (Summit day 2).

The same pattern appeared elsewhere. The docomo slides claimed “$0 new infrastructure” for agents built on etcd, RBAC and CRDs (Japan). At AWS Summit Amsterdam, Mambu described adding agents to a multi-tenant banking SaaS on Bedrock AgentCore, where an incoming tenant token plus an agent identity yields a workload token (AWS). At PlatformCon, one session argued “Not that much changes” (PlatformCon), and the day’s message was that the age of AI runs on platform engineering (PlatformCon day). I cover that side in Platform Engineering: Lessons from Conference Talks.

Takeaway: expose a small set of pre-approved, idempotent actions to the agent and keep your ticketing, policy and audit systems in charge.

7. Commerce and payments: the backend owns the state

In March 2026 an AI House panel with people from Prosus, Adyen, Visa and ABN AMRO started from identity: who is this agent, who authorised it, what are its limits, how do you revoke it. It also called the demo-to-production gap wide (wallets). In February an AI House evening had closed with a commerce panel (Stripe, JET, iFood). My take was that Cowork plugins, ClawHub skills and Cursor background agents all package know-how as versioned units, which moves the hard problems to permissions, review and audit (Code to Commerce).

In May, Stripe’s DevWorld keynote described the Universal Commerce Protocol, a Shared Payment Token and a Machine Payment Protocol built on HTTP 402. Its 12.3% conversion with AI chat against 3.1% without is Stripe’s own figure (DevWorld). The engineering half of the session is the part I would copy. The agent loop had a hard MAX_ITERATIONS cap. The backend owned the checkout state and the model only read a status such as ready_for_payment. A slide said “the prompt IS the ethics policy”. My view: a prompt states the policy, the backend enforces it (UCP demo).

The sources pull in different directions: the keynote said “Your checkout flow is already obsolete”, while the panel said the gap to enterprise adoption is still significant.

Takeaway: give a spending agent its own identity, a scoped budget, a step cap and server-side state.

8. Humans on the loop, with approvals where actions are irreversible

Sources disagree on how much oversight. In June 2026 Kief Morris argued for “on the loop”: in-the-loop review degrades into rubber-stamping, so humans should build and tune the system that builds the software. His slides added spec-driven development and leading and lagging sensors (Kief Morris).

Others kept gates. A RoachFest speaker said a database-migration agent should pause at ambiguous decisions such as UUID versus string, not guess (RoachFest). The CTO breakfast put a human at “Approve”, “the one step that doesn’t scale yet”, and noted personal-agent products differ in what they “ask you first” (CTO Network). Dexter Horthy stamped a lights-off factory “NO THANKS” (Day 2). Patrick Debois said at a Power Platform meetup he does not believe in auto-deploy to production with this technology (Power Platform). Summit’s trust model ran from approved single actions to supervised workflows, then autonomy for proven systems.

My reconciliation: the sources agree on irreversible actions and ambiguous decisions, and differ on routine steps. At the OpenClaw demo evening, where I also demoed, I liked another builder’s email-triage flow with a “notify for approval” branch (OpenClaw demo). On security, the September 2026 preday (Preday) said to treat agent skills like a supply chain, and the pillar linked above covers it.

Takeaway: design approval touchpoints for irreversible or ambiguous steps, and measure how long your reviewers take.

What I’d do on Monday

  1. Pick one high-friction workflow with a KPI and an owner, and measure adoption from week one.
  2. Write the workflow as deterministic steps first. Add an agent only where a step needs judgement.
  3. Build a fixed scenario suite with a “Definition of Done” for the output, run it before every release, and test your evaluator.
  4. Emit one OpenTelemetry trace per agent session with identity, model, tokens, cost and outcome.
  5. Test retrieval after a deletion, and measure chunking on your own data.
  6. Route agent actions through your existing platform (AAP, CRDs, ITSM, policy) and expose only pre-approved actions.
  7. Write down which actions need approval, set timeouts that fail safe, and review how long approvals wait.

Sources

Posts cited above, grouped by year.

2025

2026

Free 30-min Production AI consultation

Book Now