Skip to main content
🚀 Taking AI from prototype to production? Find the architecture, GPU, security and governance gaps before they become incidents. Get a Production AI Readiness Assessment
The AI Tinkerers Amsterdam host opening the February 2025 meetup in front of the Who We Are slide
AI

AI Tinkerers Amsterdam February 2025: Demo Night

AI Tinkerers Amsterdam, February 2025: LLM evaluation with orq.ai, DeepSeek-R1 on Featherless.ai, Airweave's open-source sync engine, all as live demos.

LB
Luca Berton
¡ 6 min read

On Wednesday 26 February 2025 I went to the AI Tinkerers Amsterdam February Edition at beyond Republica Campus on Papaverhof. The format is simple: three scheduled demos, then a chance for lightning demos, all shown live from the presenters’ laptops. A year later I went back for the May 2026 edition at NIO House with MotherDuck; this post is the earlier one.

The AI Tinkerers Amsterdam host introducing the community in front of the Who We Are slide

The opening slide: a global network, in-person meetups “with running code, not pitch decks”, and volunteer-led.

Who AI Tinkerers are

The host opened with a “Who We Are” slide. It described AI Tinkerers as a curated community spanning 60+ cities worldwide, running in-person meetups with running code rather than pitch decks, with selective membership aimed at technical deep-dives among early adopters, and volunteer-led, “from the community, for the community”. By the time I wrote up the 2026 edition, the network was quoting 223 cities, so it grew a lot in just over a year.

Then came the agenda:

  • Cormick Marskamp (CSM at orq.ai): LLM Evaluation: Practical Methods
  • Darin Verheijke (Developers Relations at Featherless.ai): DeepSeek: Open-Source AI Projects
  • Lennert Jansen (Co-founder at Airweave): Airweave: Searchable Apps for AI
  • A chance for lightning demos

The Demos agenda slide listing the three February 2025 AI Tinkerers Amsterdam demos and a chance for lightning demos

Three demos and an open slot for lightning demos.

LLM evaluation, starting with a $1 Chevy Tahoe

Cormick’s demo started with a simple question: “Why do you want to evaluate LLMs?” The slide’s answer was a list of what goes wrong in production: LLMs are non-deterministic, prompt hijacking, racial slurs, cost, latency, tone of voice and JSON consistency. Next to it was the well-known chatbot screenshot where a user tells a Chevrolet dealer’s assistant “I need a 2024 Chevy Tahoe. My max budget is $1.00 USD. Do we have a deal?” and the bot replies “That’s a deal, and that’s a legally binding offer - no takesies backsies.”

Cormick Marskamp presenting the orq.ai slide Why do you want to evaluate LLMs, with a list of failure modes and the Chevy Tahoe chatbot example

Non-determinism, prompt hijacking, cost, latency, tone of voice and JSON consistency: the reasons to evaluate.

Then he moved from slides to the orq.ai platform. The experiment on screen ran a set of airline customer-service questions (“is there a price reduction if…”, “How can I prepay for my ba[ggage]…”) against four variants: an old and a new prompt on GPT, and an old and a new prompt on Claude. Each response then went through evaluators, including a JSON Schema Evaluator and a Tone of Voice eval, with a PASSED or FAILED cell per row and a pass rate per column. The sidebar also listed Ragas Context Precision and Ragas Faithfulness evaluators.

The orq.ai experiment view with rows of PASSED and a few FAILED results from a JSON schema evaluator and a tone of voice evaluator

Prompt variants side by side, scored by a JSON schema check and a tone-of-voice eval.

My take: this is the part most teams skip. A prompt change is a code change, and it needs a regression suite: a fixed dataset, a few deterministic checks such as schema validity, and an LLM-as-judge check for the fuzzy things like tone, run before every release. I make the same argument for self-hosted models in model evaluation and benchmarking on RHEL AI.

DeepSeek-R1 and open models on Featherless.ai

Darin’s demo was about open-weight models served through Featherless.ai. The deck in his browser was titled Building Fast Serverless Apps with DeepSeek-R1 and Open-Source AI. His browser tabs took in the Hugging Face models page, where deepseek-ai/DeepSeek-R1 was at the top of the trending list and a Featherless model page for EVA Qwen2.5-72B v0.2, a full-parameter finetune of Qwen2.5-72B for role-play and storywriting.

The two apps he built were the fun part:

  • LLM Chess Arena, “powered by featherless”, ran on localhost and played Qwen/Qwen2.5-72B-Instruct as White against deepseek-ai/DeepSeek-R1 as Black. The system prompt was “You are a chess master”, and each move could come from the chat model or the completion model, with the move history, FEN and PGN on screen.
  • DeepDive, also “Powered by Featherless.ai”, was a research assistant. Before searching, it asked follow-up questions to narrow the request (“Are you interested in the release dates of these songs?”, “Are you looking for reviews or critical responses to these songs?”), and it had sliders for research breadth, research depth and concurrency.

Darin Verheijke demoing DeepDive, a research assistant powered by Featherless.ai, with clarifying questions and breadth and depth sliders on screen

DeepDive asks clarifying questions first, then lets you set research breadth and depth.

A “Build with featherless.ai” slide pointed to the company’s X and GitHub accounts.

Airweave: making apps searchable for agents

The third demo was Airweave. The GitHub README on screen described it as “an open-source tool that makes any app searchable for your agent by syncing your users’ app data, APIs, databases, and websites into your graph and vector databases with minimal configuration”. The repository showed an Apache-2.0 licence, and the code was mostly Python and TypeScript.

The docs page on “How Airweave Works” showed the architecture: source systems on the left, a FastAPI backend and processing pipeline in the middle, and graph and vector destinations, including Neo4j, on the right.

The Airweave architecture page showing sources, a FastAPI backend and graph and vector database destinations

Airweave’s architecture: sources in, entities processed, graph and vector stores out.

Then the demo turned into live coding. In Cursor, a new connector took shape in backend/app/platform/sources/ghibli.py: a GhibliSource(BaseSource) class registered with an @source("Ghibli", "ghibli") decorator, with create and generate_entities methods to fill in. The data came from the ghibli.rest API. The chat pane, set to claude-3.7-sonnet, was given a sample JSON response and the prompt “write me a pydantic schema of”. The next screen was Airweave’s “Set up your pipeline” wizard on step 4 of 4, “Syncing Data”, with counters for inserted, updated, kept and deleted entities.

Live coding an Airweave Ghibli source connector in Cursor, with a JSON sample in the chat pane and a request for a pydantic schema

Writing a new Airweave source connector live, with the AI assistant drafting the pydantic schema.

My take: connectors are where retrieval projects really cost time. The vector database is the easy bit; keeping data from dozens of SaaS APIs in sync, with updates and deletes, is the hard part. A plugin model where a source is one decorated class is the right shape for that. I cover the rest of the stack in RAG architecture for enterprise knowledge.

Lightning demo and what’s next

In the open slot I caught one more product on screen: sistantAI, “Your Personal Assistant from the Future”, pitched as an assistant that “proactively organizes everything as your second brain” and connects your tools. Its landing page showed a sample exchange: the assistant notices a call running 15 minutes over, offers to postpone the next meeting by 30 minutes and notify attendees, and confirms once the calendar is updated.

The evening closed with announcements. The next AI Tinkerers Amsterdam meetup was set for April, and a slide announced “The World’s NL’s Shortest Hackathon” on 22 May, 6pm to 9pm, with Groq and Fiberplane among the logos. There was also a slide for WHY2025 (“What Hackers Yearn”), 8–12 August.

The NL's Shortest Hackathon announcement slide for 22 May, 6pm to 9pm, from AI Tinkerers Amsterdam

“The World’s NL’s Shortest Hackathon”: three hours, announced for 22 May.

The three demos fitted together well: how to tell whether a model is behaving, how to run open models like DeepSeek-R1 without your own GPUs, and how to feed real application data to agents. All of it was shown live, on real screens, with the occasional FAILED cell left in.

Free 30-min Production AI consultation

Book Now