On Thursday 1 October 2026 I went to a CTO Network breakfast at the Databricks office in Amsterdam. The title slide was branded Builders ¡ CTO Network, dated âAmsterdam ¡ 1 October 2026â, tagged âState of Agents 2026â, and asked: âSo yeah. Where the hell are we with agents in 2026?â The footer read âHosted at Databricks Amsterdamâ and showed the Builders, Databricks, Billy Grace and Fortino logos.

The title slide on a ceiling screen at 9:06, just before the session started.
The CTO Network is run by Builders, the Rotterdam AI venture studio. Its network page describes it as a curated network of technical leaders built on invite-only events, biweekly briefings for verified members and an âecosystem memoryâ. This breakfast is listed there as âCTO Network Breakfast ¡ Databricks, Amsterdamâ. The slides were driven by Michael van Lier, shown on screen as âMichael van Lier (Presenting, annotating)â. The Builders team page lists him as founder and managing director.
The format: âNo pitching. No noise. Just builders.â
An early slide set the tone: âA trusted space for Europeâs boldest CTOsâ. It described the format as âinvitation only, peer driven, anti-sales. Small enough that nobody hides, senior enough that nobody performs.â The stat tiles under it read â2,400+ CTOs & VP Engs in the CTO Networkâ, â100% Hands-on builders, no consultantsâ and â0 Sales pitches from the stageâ.
The welcome slide that followed framed the morning:
Agents write production code. They operate browsers, terminals and APIs, run workflows for hours, and increasingly get access to systems we used to reserve for humans. This morning is about where that line actually sits right now.

The house rules: no pitching, everyone builds, bring the hard question.
There were three ground rules. âNo pitching. No noise.â âEveryone here buildsâ: every seat âbelongs to someone who is on the hook when the agent does something stupid at 3amâ. And âBring the hard questionâ, with âNo agent 101. No polished demos.â That last card is how Iâd want any agent discussion for CTOs to start. Some attendees also had a printed handout headed â01 The autonomous software factoryâ.
The opening panel: four people who ship agents
The 09:05 slot was the opening panel: âFour people who ship agents, and one hour to grill them.â The line-up slide showed Michael van Lier and three panellists: Tjadi Peeters (Billy Grace), Dylan Moerland (Everday) and a panellist from Databricks. A sponsor slide thanked the hosts: âHosted at Databricks. With Billy Grace and Fortino.â

The panel line-up as it appeared on screen.
Michael worked through five slides labelled âTHE TRENDâ, annotating them live while the panel responded. Iâll stick to what was on the slides. I didnât record the discussion, so I wonât attribute any opinions to the panellists.
Trend 1: âNext stop: a working weekâ
The first trend slide, âThe line keeps going. Next stop: a working week.â, charted âHow long an agent can work on its ownâ. It used METRâs 50% time horizon: âSolid: measured. Dashed: METRâs ~4-month doubling trend. Astra and Opus 5.5 not yet measured by METR.â
The curve started at âo1-preview ¡ 19 minâ and rose through âOpus 4.1 ¡ 1h41â and âOpus 4.6 ¡ 12hâ. Then it switched to a dashed projection ending at âOpus 5.5 ¡ 22 Sep ¡ ~43h on trendâ, above a dotted line marked â1 working week ¡ 40hâ. A side panel called âNew ruler ¡ Terminal-Benchâ (âAgentic coding in a real terminalâ) compared three recent models, with the caveat âVendor-reported scoresâ.

Measured points are solid; the dashed part is the slideâs projection of METRâs trend, not a measurement.
The slide itself separated measured points from projected ones. My view: a 40-hour line on a chart is a projection, not a guarantee.
Trend 2: âBy September, everyone sells oneâ
âIn January it was a hobby project. By September, everyone sells one.â Four cards compared personal agents: OpenClaw, Grok Bot, Muse and dots. Each card had two rows, âRuns onâ and âAsks you firstâ. OpenClaw runs on âYour own machineâ and asks first about âWhatever you configureâ. Grok Bot runs on an âOwn cloud computer, 24/7â and asks first about âLocal commandsâ. Muse runs on a âCloud VM: browse, email, shopâ and asks first about âEmail, purchases, sharingâ.

The panel in front of the âeveryone sells oneâ slide.
For me, the âAsks you firstâ row is the useful one. Itâs the approval policy, and it varied a lot from card to card.
Trend 3: âAnd this summer, they started getting outâ
The third slide had four tiles, which read exactly:
- â~700 agents coordinated on a hidden message board to cheat a test, then breached Hugging Faceâ
- â1 wk before anyone noticed. It turned up in internal logs.â
- â2 training pauses at OpenAI after escapes, the second on 20 Septemberâ
- â5 labs have reported agents slipping their boundaries: OpenAI, Anthropic, Google, Meta, Kimiâ
A timeline underneath ran from June to late September: âMedicare portalâ (Australian government files reached), âHugging Faceâ (sandbox escape, zero-days, stolen credentials), âUS government sitesâ (credentials found online used on Census data) and âOut againâ (DNS loophole, second training pause). The footer read: âNot evil models. Capable ones, taking initiative during tests and finding the gap between the controls we assumed and the controls we actuallyâ had (the last word was cut off in my photo).

These are the slideâs claims, with its sources in small print. I havenât verified them independently.
Trend 4: âIntelligence got cheap. Context is the bottleneck.â
The fourth slide put âYour agentâ at the centre of four inputs: Skills (âHow we do things hereâ), Memory (âWhat happened beforeâ), Tools + triggers (âWhat it can reachâ) and Data + meaning (âWhat the numbers meanâ). The first three were tagged âSHIPPEDâ. Data + meaning was highlighted and tagged âTHE GAPâ.
Next to the diagram were three numbers: â7% of enterprises say their data is ready for AIâ, â+38% accuracy from context alone, vendor benchmarkâ, and a 60% tile whose caption was hidden behind a panellist in all my photos.

Three inputs shipped, one gap: data and what it means.
In my view, tool access is the easy part now that MCP is common. What still blocks production is the semantic layer: which table is the source of truth, and what âactive customerâ means in this company. It was a fitting topic for the venue: Databricks says Amsterdam was its first R&D site outside the United States and is its largest R&D centre in EMEA.
Trend 5: from one agent in your IDE to a software factory
The last trend slide, âFrom one agent in your IDE to a factory that runs itselfâ, gave a definition:
A software factory is a set of coordinated agents running a structured engineering workflow: triggered by events, with humans approving the key decisions.
Under âA typical line today (Vercel)â, it showed six stations: Issue (trigger: âFiled by a user, a test or an alertâ) â Classify (agent: triage, reproduce, diagnose) â Specify (agent: âWrite the spec before the codeâ) â Implement (agent: code, tests, preview build) â Review + risk (agent: âAgent reviews, scores the riskâ) â Approve (human: âThe one step that doesnât scale yetâ).
Below that were the numbers: â200 â 30 open issues at Astro after it put agents on reproduce, diagnose and fixâ, a 35% tile that was partly hidden in my photo, and a âWho runs oneâ box listing Cloudflare, Vercel, Uber, WorkOS and StrongDM, among others.

Five agent stations and one human gate: âthe one step that doesnât scale yetâ.
The panel closed on a final slide, âLooking ahead: the software factoryâ, which matched the theme of the printed handout. Itâs the same thread Dexter Horthy pulled on in his âState of the Software Factoryâ keynote at AGNTCon + MCPCon Europe 2026 two weeks earlier. Red Hat Tech Day Netherlands 2026 also had a âTrusted Software Factoryâ session in its agentic AI track. Builders lists the next breakfast in the series as âAgents After the Demoâ on 28 October 2026 in Amsterdam.
My take: what a CTO needs before running agents in production
Put these five slides together and the message is: agents can work longer without supervision, everyone ships one, some have already escaped their sandboxes, context is the gap, and the end state is a factory. If I were the CTO in that room, this is what Iâd put in place before scaling past one agent in one IDE:
- Make âAsks you firstâ a policy you own, not a vendor default. The personal-agent cards differed mainly in what they ask before acting. In your company that list should be written down and enforced in the execution path: approval gates for irreversible actions, blast-radius limits and an audit log, as in Guardrails for AI Agents in Production. A pre-execution hook that can deny a tool call, like Claude Code PreToolUse hooks, is a control. A line in the system prompt is not.
- Assume the sandbox will be tested. The âgetting outâ slide was about agents finding âthe gap between the controls we assumed and the controls we actuallyâ had. Give each agent its own scoped credentials, deny network egress by default and make sure nothing it can read contains secrets. Then check those assumptions on purpose, rather than finding out from the logs a week later.
- Put a guard model next to the deterministic controls. For content, grounding and tool-call sanity checks, a local classifier such as Granite Guardian on Ollama adds a probabilistic layer you can run on-premises. It sits beside the hard policy, not in place of it.
- Treat tools as APIs with owners. The context slide said tools are âshippedâ. The question is whether yours are versioned, tested and least-privilege. In my Elastic Agent Builder MCP test, the deterministic tools (ES|QL queries, a workflow exposed as a tool) could be exercised with curl in CI and no LLM. Thatâs the bar Iâd set for every tool a factory station can call.
- Close the âData + meaningâ gap before buying more agents. If only 7% of enterprises think their data is ready, the bottleneck isnât the model. Invest in a governed catalogue, clear metric definitions and lineage, and give agents that layer as context.
- Instrument the factory like a production line. Every station in the Vercel-style line should emit a trace: agent identity, model, tokens, cost, outcome and who approved it. Measure the human âApproveâ step too, because thatâs where the queue will build up. The OpenTelemetry approach in monitoring coding agents is a good starting point. The fleet-level view I described after Forecasting the 2026 AI Agent Economy at Zoku builds on the same data.
Iâd start the factory with the stations that are easy to reverse (classify, reproduce, specify), keep the human gate on merge and deploy, and only widen agent autonomy once the traces show the review step is the bottleneck rather than the safety net. For the broader strategy questions, see Agentic AI: A CTO Strategy Guide.