Skip to main content
🤖 Running agents for a team, not just yourself? Get an independent review of identity, secrets, failover, observability and governance. Assess your agent platform
Speaker on stage next to a slide titled Agentic OCR Key Steps at the LlamaIndex Data Agent Builders meetup in Amsterdam
AI

LlamaIndex Data Agent Builders Meetup, Amsterdam 2025

LlamaIndex's Data Agent Builders meetup in Amsterdam, 14 July 2025: GUI-agent datasets, what makes an agent, MCP tools and an agentic OCR pipeline.

LB
Luca Berton
¡ 6 min read

On 14 July 2025 I spent an evening at the Data Agent Builders meetup in Amsterdam, organised by LlamaIndex. The closing slide read “Data Agent Builders / Meetup Amsterdam / LlamaIndex” and carried a link to the run-llama/llama_index repository. A flipchart by the door said Hotel Arena, and the room was full for the three talks I photographed. Here is what the slides covered, plus a fourth talk that I only recorded.

A talk on GUI agents and their data

The first talk I photographed was about agents that operate graphical interfaces. One slide was titled “What Makes a Good Dataset for GUI Agents”, and another showed GroundUI, with a screenshot of a desktop-style application and a citation to the AgentStudio paper (“AgentStudio: A Toolkit for Building General Virtual Agents”). I did not catch the speaker’s name on a slide, so I leave it out.

Speaker presenting a slide titled What Makes a Good Dataset for GUI Agents to a full room at the LlamaIndex meetup in Amsterdam

The room during the GUI-agents talk, with a flipchart reading Hotel Arena beside the screen.

Speaker gesturing next to a screen showing a slide titled GroundUI

A GroundUI slide: training and evaluating agents that click through real interfaces needs screenshots and annotations of real interfaces.

What makes an agent

The second talk started from basics. The slide “What’s makes an agent” (sic) drew a loop from a query, through an agent and a retrieval step, to a response, and listed the ingredients: reasoning, a query or task planning layer, a tool interface for the external environment, reflection, memory for personalisation, and multi-turn conversation.

Speaker at a lectern next to a slide listing reasoning, task planning, tool interface, reflection, memory and multi-turn as what makes an agent

The agent anatomy slide: reasoning, planning, tools, reflection, memory and multi-turn.

LlamaIndex’s own documentation describes an agent in similar terms, as a semi-autonomous piece of software powered by an LLM that picks tools and decides when a task is done. The same page covers FunctionAgent and AgentWorkflow for building multi-agent systems.

Bring MCP tools to your agent workflows

A section headed “Extra: Bring MCP Tools to your Agent Workflows” followed. The diagram put an Agent on top of MCP, on top of LlamaCloud, with Extract Agent and Vector Index below it, and sources and sinks such as GitHub, an invoice extractor, a CV extractor, Google Drive and S3. A corner of the slide was about running the LlamaCloud MCP server.

Wide shot of the room with a speaker at the front and a slide titled Extra: Bring MCP Tools to your Agent Workflows

The MCP slide, with an agent calling LlamaCloud services through the Model Context Protocol.

The LlamaCloud docs describe LlamaParse as the document-processing platform from LlamaIndex, with products for parsing, extraction, classification, splitting and indexing, and an MCP server for agents. My take: exposing document services as MCP tools is the sensible way to let an agent reach them without bespoke glue code. Do review which tools an agent can call, though, because MCP widens what a prompt can trigger.

Agentic OCR for ID documents

The last talk was my favourite. It started with “What is OCR?” and a “brief history” slide that ran from an early OCR patent in 1929 and Emanuel Goldberg’s system in 1931, via commercial machines in the 1950s, Kurzweil’s omni-font OCR in the 1970s and Tesseract, to LSTM-based Tesseract 4.0 in 2018 and, in the 2020s, multimodal large language models.

Speaker next to a slide titled A brief history listing OCR milestones from 1929 to the 2020s

“A brief history” of OCR, ending with multimodal LLMs.

The “Agentic OCR: Key Steps” slide then described the pipeline:

  • OCR Agent: orchestrates the OCR actions (recognise, analyse, validate).
  • Multimodal LLM: the core engine that extracts text and visual data.
  • Validation Agent: checks output quality and adherence to business rules, and decides the next step.
  • Iterative Loop: failed validation can trigger retries, specially in human-in-the-loop flows.
  • Agent Handover: successful results pass to the next system or agent.

Speaker beside a slide titled Agentic OCR Key Steps with a flowchart of OCR agent, multimodal LLM and validation agent

The agentic OCR pipeline: an extraction step, a validation step and a retry loop.

The “Challenges” slide was candid about the risks:

  • Prompt engineering: agentic systems are only as good as their prompts.
  • Evals and observability: agents add a whole new layer of metrics.
  • Injection attacks: guardrails must mitigate injection via document contents.
  • Document size: proprietary models limit image size and throughput.
  • Racial bias: for ID documents, guardrails must ensure fairness.
  • Language bias: a performance gap between mainstream and niche languages can affect accuracy and fairness.

Speaker next to a slide titled Agentic OCR: Challenges listing prompt engineering, evals, injection attacks, document size, racial bias and language bias

The challenges slide for agentic OCR on identity documents.

Computer vision at Schiphol: opening the turnaround black box

After the OCR talk I recorded one more talk, and this one has no slide photos, so everything below comes from my audio recording and I did not catch the speaker’s name. The speaker said they were part of the first corporate scale-up at Schiphol and wanted to share what the team does, how, and what it learned, so that others might start their own computer vision projects.

The context figures were a useful reminder of scale: last year the airport handled 66 million passengers, with 201 direct destinations, roughly half a million flight movements a year and six runways. Running an airport is like managing a small city where every operation depends on the others. The top challenge named was delays, and the other was capacity: more people fly, but building another runway is expensive, so the aim was to extend capacity without changing the infrastructure.

The core problem was the aircraft turnaround, the flight preparation process of fuelling, unloading luggage, boarding and so on. A few years earlier the airport, airlines and ground handlers had all seen that 30% to 40% of delays traced back to inefficiencies in the turnaround, but they had no way to measure it beyond a person taking notes. Six years ago a data science and engineering lab was set up, and the question became whether the cameras already on the ramps could turn images into detections: “opening the turnaround black box”. The setup was two cameras per ramp, filming the aircraft from all angles, across roughly 100 ramps.

The most transferable idea was the data-centric AI approach. Instead of chasing the latest model, the team paid attention to the data. When a model fails, they want to know why, so they invested in model explainability. They also select a diverse dataset using metadata (winter versus summer operations, lighting conditions, different equipment on each ramp) and use the model’s own probabilities to pick informative samples. One example was a ramp being renovated, where the model reported an error. The main “community model” makes more than 70 detections, such as the aircraft arriving, the stopping position, and the power cable being connected.

Detections were only the first step. They feed data analysts who look for deficiencies in turnaround history, and real-time views for handlers. The detections also feed machine learning models that predict when a flight’s turnaround will finish, and the screen explains why, for example “the fuel is still active”, so the handler can act. A mobile app was planned for the future, and the speaker joked about a smartwatch app. The reported results were a 2.5% to 5% increase in airport capacity, because gate planners can allocate resources better with full visibility, and “one source of truth” that ended arguments between handlers and the airport about whose delay it was.

My take: the lesson that travels furthest is not the cameras but the sequence of working with the people who use the output (the handlers), investing in data quality before model size, and agreeing a single measured version of events before optimising it.

What I took away

My take: the validation agent plus a retry loop is the part worth copying. A single LLM call over a scanned document gives you text; a second step that checks it against business rules gives you something you can put in a workflow. The injection point on the challenges slide matters too: a document is untrusted input, so what the extraction agent can do next should be limited, and every run should be traced.

Free 30-min Production AI consultation

Book Now