On Wednesday 26 February 2025 I went to the AI Tinkerers Amsterdam February Edition at beyond Republica Campus on Papaverhof. The format is simple: three scheduled demos, then a chance for lightning demos, all shown live from the presentersâ laptops. A year later I went back for the May 2026 edition at NIO House with MotherDuck; this post is the earlier one.

The opening slide: a global network, in-person meetups âwith running code, not pitch decksâ, and volunteer-led.
Who AI Tinkerers are
The host opened with a âWho We Areâ slide. It described AI Tinkerers as a curated community spanning 60+ cities worldwide, running in-person meetups with running code rather than pitch decks, with selective membership aimed at technical deep-dives among early adopters, and volunteer-led, âfrom the community, for the communityâ. By the time I wrote up the 2026 edition, the network was quoting 223 cities, so it grew a lot in just over a year.
Then came the agenda:
- Cormick Marskamp (CSM at orq.ai): LLM Evaluation: Practical Methods
- Darin Verheijke (Developers Relations at Featherless.ai): DeepSeek: Open-Source AI Projects
- Lennert Jansen (Co-founder at Airweave): Airweave: Searchable Apps for AI
- A chance for lightning demos

Three demos and an open slot for lightning demos.
LLM evaluation, starting with a $1 Chevy Tahoe
Cormickâs demo started with a simple question: âWhy do you want to evaluate LLMs?â The slideâs answer was a list of what goes wrong in production: LLMs are non-deterministic, prompt hijacking, racial slurs, cost, latency, tone of voice and JSON consistency. Next to it was the well-known chatbot screenshot where a user tells a Chevrolet dealerâs assistant âI need a 2024 Chevy Tahoe. My max budget is $1.00 USD. Do we have a deal?â and the bot replies âThatâs a deal, and thatâs a legally binding offer - no takesies backsies.â

Non-determinism, prompt hijacking, cost, latency, tone of voice and JSON consistency: the reasons to evaluate.
Then he moved from slides to the orq.ai platform. The experiment on screen ran a set of airline customer-service questions (âis there a price reduction ifâŚâ, âHow can I prepay for my ba[ggage]âŚâ) against four variants: an old and a new prompt on GPT, and an old and a new prompt on Claude. Each response then went through evaluators, including a JSON Schema Evaluator and a Tone of Voice eval, with a PASSED or FAILED cell per row and a pass rate per column. The sidebar also listed Ragas Context Precision and Ragas Faithfulness evaluators.

Prompt variants side by side, scored by a JSON schema check and a tone-of-voice eval.
My take: this is the part most teams skip. A prompt change is a code change, and it needs a regression suite: a fixed dataset, a few deterministic checks such as schema validity, and an LLM-as-judge check for the fuzzy things like tone, run before every release. I make the same argument for self-hosted models in model evaluation and benchmarking on RHEL AI.
DeepSeek-R1 and open models on Featherless.ai
Darinâs demo was about open-weight models served through Featherless.ai. The deck in his browser was titled Building Fast Serverless Apps with DeepSeek-R1 and Open-Source AI. His browser tabs took in the Hugging Face models page, where deepseek-ai/DeepSeek-R1 was at the top of the trending list and a Featherless model page for EVA Qwen2.5-72B v0.2, a full-parameter finetune of Qwen2.5-72B for role-play and storywriting.
The two apps he built were the fun part:
- LLM Chess Arena, âpowered by featherlessâ, ran on localhost and played Qwen/Qwen2.5-72B-Instruct as White against deepseek-ai/DeepSeek-R1 as Black. The system prompt was âYou are a chess masterâ, and each move could come from the chat model or the completion model, with the move history, FEN and PGN on screen.
- DeepDive, also âPowered by Featherless.aiâ, was a research assistant. Before searching, it asked follow-up questions to narrow the request (âAre you interested in the release dates of these songs?â, âAre you looking for reviews or critical responses to these songs?â), and it had sliders for research breadth, research depth and concurrency.

DeepDive asks clarifying questions first, then lets you set research breadth and depth.
A âBuild with featherless.aiâ slide pointed to the companyâs X and GitHub accounts.
Airweave: making apps searchable for agents
The third demo was Airweave. The GitHub README on screen described it as âan open-source tool that makes any app searchable for your agent by syncing your usersâ app data, APIs, databases, and websites into your graph and vector databases with minimal configurationâ. The repository showed an Apache-2.0 licence, and the code was mostly Python and TypeScript.
The docs page on âHow Airweave Worksâ showed the architecture: source systems on the left, a FastAPI backend and processing pipeline in the middle, and graph and vector destinations, including Neo4j, on the right.

Airweaveâs architecture: sources in, entities processed, graph and vector stores out.
Then the demo turned into live coding. In Cursor, a new connector took shape in backend/app/platform/sources/ghibli.py: a GhibliSource(BaseSource) class registered with an @source("Ghibli", "ghibli") decorator, with create and generate_entities methods to fill in. The data came from the ghibli.rest API. The chat pane, set to claude-3.7-sonnet, was given a sample JSON response and the prompt âwrite me a pydantic schema ofâ. The next screen was Airweaveâs âSet up your pipelineâ wizard on step 4 of 4, âSyncing Dataâ, with counters for inserted, updated, kept and deleted entities.

Writing a new Airweave source connector live, with the AI assistant drafting the pydantic schema.
My take: connectors are where retrieval projects really cost time. The vector database is the easy bit; keeping data from dozens of SaaS APIs in sync, with updates and deletes, is the hard part. A plugin model where a source is one decorated class is the right shape for that. I cover the rest of the stack in RAG architecture for enterprise knowledge.
Lightning demo and whatâs next
In the open slot I caught one more product on screen: sistantAI, âYour Personal Assistant from the Futureâ, pitched as an assistant that âproactively organizes everything as your second brainâ and connects your tools. Its landing page showed a sample exchange: the assistant notices a call running 15 minutes over, offers to postpone the next meeting by 30 minutes and notify attendees, and confirms once the calendar is updated.
The evening closed with announcements. The next AI Tinkerers Amsterdam meetup was set for April, and a slide announced âThe Worldâs NLâs Shortest Hackathonâ on 22 May, 6pm to 9pm, with Groq and Fiberplane among the logos. There was also a slide for WHY2025 (âWhat Hackers Yearnâ), 8â12 August.

âThe Worldâs NLâs Shortest Hackathonâ: three hours, announced for 22 May.
The three demos fitted together well: how to tell whether a model is behaving, how to run open models like DeepSeek-R1 without your own GPUs, and how to feed real application data to agents. All of it was shown live, on real screens, with the occasional FAILED cell left in.
