On Wednesday 9 July 2025 the MLOps Community stopped in Amsterdam with its AI Agent World Tour, hosted at ABN AMROâs headquarters. According to the event page, the format was practitioner-driven: engineers who run agents in production, unscripted conversations, and a panel at the end. The sponsors on the list were JetBrains and orq.ai, with Sirach Ventures as a partner. ABN AMRO was thanked as host on stage.
This post is a throwback. I filmed the talks and photographed the slides from my seat, so what follows paraphrases what was said, and I only quote a number when it was clear on a slide or in the recording. Where a speaker was selling a product, I say so.

The lecture hall just before the opening. The slide on screen invites people to join the community Slack.
The opening: why a tour about agents
An organiser opened with a few slides on the community. MLOps Community is a global online network with local in-person groups (the host mentioned more than 40), and it publishes podcasts, reading groups and workshops. The World Tour is a series of events in different cities about putting AI agents into production, and the Amsterdam stop was described as the sixth. The pitch for the format was simple: there is a lot of talk about agents from people who have not shipped one, and too little âthis works, this doesnâtâ. The agenda: three talks, a break, a fourth talk, then a panel.
Rafael V. Pierre: Lumina, an assistant for a data platform
The first talk was by Rafael V. Pierre, listed on the event page as Principal GenAI Engineer at weet.ai and introduced on stage as a former colleague of the host at the bank. He presented Lumina, an agent project at ABN AMROâs analytics technology teams.
The context he gave: the teams run several shared platforms for data engineering, ML and GenAI, with hundreds of users, plus frameworks that spare users from writing ETL or prompt code from scratch. The cost is onboarding. New users face a steep learning curve, support people spend time answering the same questions, and everyone has to learn the compliance rules and infrastructure-as-code conventions. The stated goal was user satisfaction.
The first step had been Binder, a documentation-as-code portal: you write docstrings and comments in the frameworks, and the documentation is rendered and hosted centrally, with changes reviewed as pull requests and history in Git. Lumina is the next step: instead of clicking through documentation cards, users ask questions, a kind of internal Perplexity for developers.
His advice on agents was the most useful part for me. Unless the use case justifies it, he said, do not use agents: workflows are more predictable and give more control. Agents earn their place when a use case is complex enough that coding every guardrail and edge case by hand would take longer than letting a model decide. He also had no love for the crowded framework landscape and called the OpenAI Agents SDK refreshing by comparison. A slide titled âNot all fun and gamesâŚâ made the trade-off plain.

âAgents abstract away some of the complexity of workflows⌠but that comes at the expense of tokensâ, and notebooks are not engineering.

The Lumina roadmap: plan, build a âminimum lovable productâ, test and evaluate, then run workshops with teams being onboarded. The slide headline reads âPoC Hell is behind us!â
orq.ai co-founder: the agent control tower
The second talk was The Agent Control Tower, by Sohrab Hosseini, co-founder of orq.ai (listed as such on the event page). He said up front that it is a concept, not a finished product, and that he works for a vendor with an LLM platform, so treat the specifics accordingly.
His argument: many prototypes never reach production not because the software fails, but because people do not feel in control. They worry about what the agent outputs, which tools and databases it touches, data residency and privacy, and cost. He gave an anecdote from a prospect, an e-commerce company with annual revenue of 750 million, whose shopping agent worked well in testing until they extrapolated its usage and got a bill of about 500 million a year. Treat it as a single story, but it shows why cost belongs in the design from day one.
He then walked through a lifecycle that will be familiar to anyone from DevOps:
- Build: pick a framework, write instructions and tools, wire up MCP servers, manage context and memory.
- Experiment: run offline experiments on curated or synthetic datasets, add automated evaluators so humans only review exceptions, and feed production edge cases back into the dataset, like regression tests. Role-play with simulated personas is the more advanced step.
- Deploy: canary releases, staged rollouts by user group or market, A/B tests, and routing different models to different markets or clients.
- Run: tracing the whole agent trajectory, and reusing the same evaluators online as guardrails that can block input or output.

Lifecycle management: product managers, domain experts, security and data officers and engineers all have a role in getting an agent to production.
He also compared evals to the testing pyramid: cheap programmatic checks at the base, LLM-as-judge in the middle, and human review at the top, which is the most expensive and should be minimised but never reaches zero. Agents should escalate to a human when they are stuck, he said, like a driverless car handing over to a control centre, and implicit feedback (a user copying or deleting the draft) is a better signal than thumbs up and down.

The evaluation pyramid, with cost increasing towards the top.
My take: this lines up with how I would set up LLM quality gates in CI. If you want a hands-on version of the same idea, see my posts on LLM regression testing with promptfoo and LLM observability in production.
Pebbling: a protocol for agents to talk securely
The third talk came from Raahul Dutta, listed as founder of Pebbling.ai. It was clearly a product and open-source-protocol pitch, so read it as the projectâs own claims. His starting point: today an agent in one company cannot easily and securely talk to an agent in another, and there is no trust or negotiation layer, whatever framework each is built on.

Pebblingâs vision slide: a decentralised protocol stack that gives agents identity, memory, trust and coordination.
According to the talk, Pebbling is a peer-to-peer, open-source protocol where each agent gets a decentralised identifier, and conversations are encrypted between the agents. The slides named three further parts: Hibiscus (a federated registry for discovery), Sheldon (a certificate authority) and Imagine, aimed at multi-party scenarios such as drug discovery. He positioned MCP as the way an agent reaches tools, and Pebbling as the way one agent reaches another. He argued it was more secure and more developer-friendly than Googleâs A2A, and said negotiation and micro-payments were planned. The project was in beta at the time. In Q&A, an attendee asked what happens if a connected agent turns out to be malicious; the answer was a feedback and rating loop in the registry.
I would want to see a threat model and independent review before trusting any new agent identity layer with money or regulated data. Still, the problem is real; I covered the standards side in A2A and the Agentic AI Foundation.
Vladislav Tankov (JetBrains): Koog, AI agents everywhere
After a break, Vladislav Tankov, Director of AI at JetBrains, introduced Koog, JetBrainsâ open-source Kotlin framework for AI agents (repository). He started with a short history of JetBrains: Kotlin since 2016, IDEs, and a pivot to AI from 2022-2023.
Koogâs selling points, per the slides and talk: LLM provider support (OpenAI, Anthropic, Google, Ollama, OpenRouter and others), MCP integration, embeddings, memory, tracing, streaming, history compression, parallel tool calling, structured output, caching, and graph-based workflows with sub-graphs, where you model an agent as a strategy graph with nodes and edges. His example was a banking-style agent that classifies a request, then routes it to a âtransfer moneyâ or âanalyze transactionsâ branch. He mentioned that the framework was built quickly, in what he called conference-driven development.

A Koog agent as a graph: classify the request, branch, finish. The code on the right defines nodes and edges.
The main argument of the talk was the âwhyâ. Because Koog is built on Kotlin Multiplatform, the same agent code can run on desktop, in the browser and on mobile. His claim was that agents will increasingly move to the edge: models of useful size already run locally via llama.cpp, ONNX and MLX, Chrome has a Prompt API for Gemini Nano, and Apple and Android are shipping on-device models. Small local models can handle the background work (collecting context, embedding, summarising) so the big cloud model is called less. He said cloud inference costs add up quickly at JetBrains scale, and that moving suitable work to devices could save a large share of costs, which he put at â80 plus percentâ as an estimate, not a measured result. His closing line was that edge devices are the cost-effective future of AI. There was no time for questions, so he asked people to find him after the panel.
The panel
The final session was a panel with the speakers. Some takeaways:
- What is an agent? Opinions differed on purpose. Vladislavâs definition: a system that interacts with an environment, mostly through tools, and gets feedback from it. Another panelist pointed to the way Anthropic separates âagenticâ systems (an LLM makes at least some decisions) from agents (more autonomy, tool use, memory). One speaker said âagentâ is mostly a business and marketing word. Several noted that non-technical people call anything with an LLM an agent.
- Cost. One panelist warned that agents are expensive, and told of a colleague whose multi-agent test cost 125 dollars for a single call. His advice: if a small classical model such as logistic regression solves 99 percent of the problem, use it.
- Is 2025 the year of agents? The orq.ai co-founder said the agents that are really in production are in typical areas such as sales and lead research, while more sensitive sectors, naming pension funds and insurers, were not shipping yet. Others agreed the hype was ahead of reality and compared it to the Gartner hype cycleâs plateau, and the JetBrains speaker noted that vendors had been relabelling AI products as agents for a while.
Takeaways
- Start with a workflow, and use an agent only when the complexity justifies it. This came from the first speaker and from a panelist alike.
- Evals, staged rollouts and tracing are the same discipline as software delivery, applied to a non-deterministic component.
- Cost is a design constraint, not an afterthought: tokens, judge models and human review all need a budget.
- Interoperability between agents is still unsolved, and several vendors, including the one in this room, are racing to define it.