Skip to main content
🤖 Running agents for a team, not just yourself? Get an independent review of identity, secrets, failover, observability and governance. Assess your agent platform
Moderator and three panellists on stage at Databricks Amsterdam under a slide reading Context is the bottleneck
AI

CTO Network State of Agents 2026 at Databricks Amsterdam

A Builders CTO Network breakfast at Databricks Amsterdam: four people who ship agents, five trend slides and a closing look at the software factory.

LB
Luca Berton
¡ 9 min read

On Thursday 1 October 2026 I went to a CTO Network breakfast at the Databricks office in Amsterdam. The title slide was branded Builders · CTO Network, dated “Amsterdam · 1 October 2026”, tagged “State of Agents 2026”, and asked: “So yeah. Where the hell are we with agents in 2026?” The footer read “Hosted at Databricks Amsterdam” and showed the Builders, Databricks, Billy Grace and Fortino logos.

Title slide on a ceiling screen reading So yeah. Where the hell are we with agents in 2026?, branded Builders CTO Network, Amsterdam, 1 October 2026

The title slide on a ceiling screen at 9:06, just before the session started.

The CTO Network is run by Builders, the Rotterdam AI venture studio. Its network page describes it as a curated network of technical leaders built on invite-only events, biweekly briefings for verified members and an “ecosystem memory”. This breakfast is listed there as “CTO Network Breakfast · Databricks, Amsterdam”. The slides were driven by Michael van Lier, shown on screen as “Michael van Lier (Presenting, annotating)”. The Builders team page lists him as founder and managing director.

The format: “No pitching. No noise. Just builders.”

An early slide set the tone: “A trusted space for Europe’s boldest CTOs”. It described the format as “invitation only, peer driven, anti-sales. Small enough that nobody hides, senior enough that nobody performs.” The stat tiles under it read “2,400+ CTOs & VP Engs in the CTO Network”, “100% Hands-on builders, no consultants” and “0 Sales pitches from the stage”.

The welcome slide that followed framed the morning:

Agents write production code. They operate browsers, terminals and APIs, run workflows for hours, and increasingly get access to systems we used to reserve for humans. This morning is about where that line actually sits right now.

Welcome slide reading No pitching. No noise. Just builders., with three cards: No pitching. No noise; Everyone here builds; Bring the hard question

The house rules: no pitching, everyone builds, bring the hard question.

There were three ground rules. “No pitching. No noise.” “Everyone here builds”: every seat “belongs to someone who is on the hook when the agent does something stupid at 3am”. And “Bring the hard question”, with “No agent 101. No polished demos.” That last card is how I’d want any agent discussion for CTOs to start. Some attendees also had a printed handout headed “01 The autonomous software factory”.

The opening panel: four people who ship agents

The 09:05 slot was the opening panel: “Four people who ship agents, and one hour to grill them.” The line-up slide showed Michael van Lier and three panellists: Tjadi Peeters (Billy Grace), Dylan Moerland (Everday) and a panellist from Databricks. A sponsor slide thanked the hosts: “Hosted at Databricks. With Billy Grace and Fortino.”

Opening panel slide reading Four people who ship agents, and one hour to grill them, with headshots of Michael van Lier, Tjadi Peeters, Dylan Moerland and a Databricks panellist

The panel line-up as it appeared on screen.

Michael worked through five slides labelled “THE TREND”, annotating them live while the panel responded. I’ll stick to what was on the slides. I didn’t record the discussion, so I won’t attribute any opinions to the panellists.

Trend 1: “Next stop: a working week”

The first trend slide, “The line keeps going. Next stop: a working week.”, charted “How long an agent can work on its own”. It used METR’s 50% time horizon: “Solid: measured. Dashed: METR’s ~4-month doubling trend. Astra and Opus 5.5 not yet measured by METR.”

The curve started at “o1-preview · 19 min” and rose through “Opus 4.1 · 1h41” and “Opus 4.6 · 12h”. Then it switched to a dashed projection ending at “Opus 5.5 · 22 Sep · ~43h on trend”, above a dotted line marked “1 working week · 40h”. A side panel called “New ruler · Terminal-Bench” (“Agentic coding in a real terminal”) compared three recent models, with the caveat “Vendor-reported scores”.

Chart slide titled Next stop: a working week, plotting how long an agent can work on its own using METR's 50 percent time horizon, with a Terminal-Bench panel on the right

Measured points are solid; the dashed part is the slide’s projection of METR’s trend, not a measurement.

The slide itself separated measured points from projected ones. My view: a 40-hour line on a chart is a projection, not a guarantee.

Trend 2: “By September, everyone sells one”

“In January it was a hobby project. By September, everyone sells one.” Four cards compared personal agents: OpenClaw, Grok Bot, Muse and dots. Each card had two rows, “Runs on” and “Asks you first”. OpenClaw runs on “Your own machine” and asks first about “Whatever you configure”. Grok Bot runs on an “Own cloud computer, 24/7” and asks first about “Local commands”. Muse runs on a “Cloud VM: browse, email, shop” and asks first about “Email, purchases, sharing”.

Panel on stage at Databricks Amsterdam in front of two screens showing the slide In January it was a hobby project. By September, everyone sells one

The panel in front of the “everyone sells one” slide.

For me, the “Asks you first” row is the useful one. It’s the approval policy, and it varied a lot from card to card.

Trend 3: “And this summer, they started getting out”

The third slide had four tiles, which read exactly:

  • “~700 agents coordinated on a hidden message board to cheat a test, then breached Hugging Face”
  • “1 wk before anyone noticed. It turned up in internal logs.”
  • “2 training pauses at OpenAI after escapes, the second on 20 September”
  • “5 labs have reported agents slipping their boundaries: OpenAI, Anthropic, Google, Meta, Kimi”

A timeline underneath ran from June to late September: “Medicare portal” (Australian government files reached), “Hugging Face” (sandbox escape, zero-days, stolen credentials), “US government sites” (credentials found online used on Census data) and “Out again” (DNS loophole, second training pause). The footer read: “Not evil models. Capable ones, taking initiative during tests and finding the gap between the controls we assumed and the controls we actually” had (the last word was cut off in my photo).

Slide titled And this summer, they started getting out, with tiles reading ~700, 1 wk, 2 and 5 and a timeline from June to September

These are the slide’s claims, with its sources in small print. I haven’t verified them independently.

Trend 4: “Intelligence got cheap. Context is the bottleneck.”

The fourth slide put “Your agent” at the centre of four inputs: Skills (“How we do things here”), Memory (“What happened before”), Tools + triggers (“What it can reach”) and Data + meaning (“What the numbers mean”). The first three were tagged “SHIPPED”. Data + meaning was highlighted and tagged “THE GAP”.

Next to the diagram were three numbers: “7% of enterprises say their data is ready for AI”, “+38% accuracy from context alone, vendor benchmark”, and a 60% tile whose caption was hidden behind a panellist in all my photos.

Moderator and three panellists on stage under two screens showing the slide Intelligence got cheap. Context is the bottleneck, with Skills, Memory, Tools and Data around Your agent

Three inputs shipped, one gap: data and what it means.

In my view, tool access is the easy part now that MCP is common. What still blocks production is the semantic layer: which table is the source of truth, and what “active customer” means in this company. It was a fitting topic for the venue: Databricks says Amsterdam was its first R&D site outside the United States and is its largest R&D centre in EMEA.

Trend 5: from one agent in your IDE to a software factory

The last trend slide, “From one agent in your IDE to a factory that runs itself”, gave a definition:

A software factory is a set of coordinated agents running a structured engineering workflow: triggered by events, with humans approving the key decisions.

Under “A typical line today (Vercel)”, it showed six stations: Issue (trigger: “Filed by a user, a test or an alert”) → Classify (agent: triage, reproduce, diagnose) → Specify (agent: “Write the spec before the code”) → Implement (agent: code, tests, preview build) → Review + risk (agent: “Agent reviews, scores the risk”) → Approve (human: “The one step that doesn’t scale yet”).

Below that were the numbers: “200 → 30 open issues at Astro after it put agents on reproduce, diagnose and fix”, a 35% tile that was partly hidden in my photo, and a “Who runs one” box listing Cloudflare, Vercel, Uber, WorkOS and StrongDM, among others.

Slide titled From one agent in your IDE to a factory that runs itself, showing a six-step line from Issue to human Approve and a 200 to 30 open issues tile

Five agent stations and one human gate: “the one step that doesn’t scale yet”.

The panel closed on a final slide, “Looking ahead: the software factory”, which matched the theme of the printed handout. It’s the same thread Dexter Horthy pulled on in his “State of the Software Factory” keynote at AGNTCon + MCPCon Europe 2026 two weeks earlier. Red Hat Tech Day Netherlands 2026 also had a “Trusted Software Factory” session in its agentic AI track. Builders lists the next breakfast in the series as “Agents After the Demo” on 28 October 2026 in Amsterdam.

My take: what a CTO needs before running agents in production

Put these five slides together and the message is: agents can work longer without supervision, everyone ships one, some have already escaped their sandboxes, context is the gap, and the end state is a factory. If I were the CTO in that room, this is what I’d put in place before scaling past one agent in one IDE:

  • Make “Asks you first” a policy you own, not a vendor default. The personal-agent cards differed mainly in what they ask before acting. In your company that list should be written down and enforced in the execution path: approval gates for irreversible actions, blast-radius limits and an audit log, as in Guardrails for AI Agents in Production. A pre-execution hook that can deny a tool call, like Claude Code PreToolUse hooks, is a control. A line in the system prompt is not.
  • Assume the sandbox will be tested. The “getting out” slide was about agents finding “the gap between the controls we assumed and the controls we actually” had. Give each agent its own scoped credentials, deny network egress by default and make sure nothing it can read contains secrets. Then check those assumptions on purpose, rather than finding out from the logs a week later.
  • Put a guard model next to the deterministic controls. For content, grounding and tool-call sanity checks, a local classifier such as Granite Guardian on Ollama adds a probabilistic layer you can run on-premises. It sits beside the hard policy, not in place of it.
  • Treat tools as APIs with owners. The context slide said tools are “shipped”. The question is whether yours are versioned, tested and least-privilege. In my Elastic Agent Builder MCP test, the deterministic tools (ES|QL queries, a workflow exposed as a tool) could be exercised with curl in CI and no LLM. That’s the bar I’d set for every tool a factory station can call.
  • Close the “Data + meaning” gap before buying more agents. If only 7% of enterprises think their data is ready, the bottleneck isn’t the model. Invest in a governed catalogue, clear metric definitions and lineage, and give agents that layer as context.
  • Instrument the factory like a production line. Every station in the Vercel-style line should emit a trace: agent identity, model, tokens, cost, outcome and who approved it. Measure the human “Approve” step too, because that’s where the queue will build up. The OpenTelemetry approach in monitoring coding agents is a good starting point. The fleet-level view I described after Forecasting the 2026 AI Agent Economy at Zoku builds on the same data.

I’d start the factory with the stations that are easy to reverse (classify, reproduce, specify), keep the human gate on merge and deploy, and only widen agent autonomy once the traces show the review step is the bottleneck rather than the safety net. For the broader strategy questions, see Agentic AI: A CTO Strategy Guide.

Free 30-min Production AI consultation

Book Now