On Thursday 26 September 2024 I went to the Google Cloud User Group Amsterdam, September Edition, hosted at the ML6 office. The sponsor slide showed ML6, Wiz, Xebia, HashiCorp and Google Cloud. There were two talks, and they made a good pair: a bank explaining how it uses machine learning against fraud in production, and an AI consultancy explaining how LLM applications get attacked and what guardrails are for.

The room at ML6 before the start, title slide already up.
The agenda slide listed:
- Thomas Vrancken, Machine Learning Engineer, ML6: How to build robust and safe LLM applications using guardrails and evaluation
- Ali el Hassouni, Head of Data, Bunq: Project Finn - Bunq’s journey to GenAI-driven operational excellence
- A quiz and wrap-up, then drinks
On the night the order was swapped: bunq went first, then ML6.
bunq: AI-driven operational excellence
Ali el Hassouni’s title slide read “bunq’s journey to AI-driven operational excellence”. bunq is the Dutch neobank, and most of the talk was about one problem: transaction monitoring, meaning spotting fraudulent payments among millions of legitimate ones. An early slide listed the bank’s AI applications (large language models, automated customer support), and the next set the context with a diagram of how transaction monitoring works.

bunq’s talk opened the evening.
Why rules are not enough
The “Impact of Financial Fraud” slide gave the share of people who had experienced financial fraud in the past two years: 62% in France, 56% in Germany, 55% in Spain and 46% in Poland. The slide added that, in the literature, financial institutions are widely seen as responsible for enabling such schemes. So the bank has a reason to care beyond its own losses.
The traditional answer is a rule engine, and the next slide listed its five problems:
- High maintenance: rules are static, and fraud changes fast.
- False positives: rules introduce a trade-off in effectiveness.
- Limited applicability: each rule only covers a narrow category of fraudulent behaviour.
- Labour intensive: creating and maintaining monitoring rules takes a lot of work.
- Rapid obsolescence: many rules are easy for smart fraudsters to reverse-engineer.

The case against rule-based transaction monitoring, in five arrows.
Known unknowns and synthetic fraud
The “Case study: Automated Fraud Detection” slide used a Johari window to frame what a model knows. A known known is a confident and correct prediction. An unknown unknown is a confident but wrong one. In between are known unknowns: patterns or categories of fraud that people know about but the model hasn’t learned yet.
That gap is where generative AI came in. The “Generative AI: How do we generate synthetic fraudulent samples” slide showed a GAN-style loop. The training data was only the fraudulent samples from the labelled observations. A generator tries to mimic them (“think of a band of counterfeiters”), a discriminator tries to tell real from generated (“think of the police”), and backpropagation feeds the decisions back so the generator improves.

Generator as counterfeiter, discriminator as police: synthetic fraud samples.
Generated samples are not all useful, so the next slides were about filtering. A “Rejection sampling” table showed synthetic transactions with columns for amount, sender and receiver country, international, high payment (>100) and payment type, with some values marked in red where a sample broke the natural rules. Then came two scores for ranking what survives:
- A uniqueness score: how unique is a feature value compared to the average value?
- A perceptibility score: what do human experts pay attention to? To avoid one person’s bias, bunq asked eight transaction-monitoring experts to rate every feature the model uses from 0 (not important at all to their decision) to 10 (extremely important). The perceptibility score is based on the inverse of the average rating per feature.

Turning expert attention into a number per feature.
I find the inverse the clever part. A synthetic fraud sample that differs from normal traffic in features the experts rarely look at is exactly the case a human reviewer would miss.
The pipeline and the 16%
The “Machine Learning Pipeline” slide, subtitled “To improve the robustness to never-seen-before fraudulent transactions”, put it together in five steps:
- Create synthetic fraudulent transactions.
- Select only the transactions that follow the natural rules.
- Rank and filter the most interesting transactions.
- Test whether these transactions outperform random transactions.
- Train a weakly supervised model using these extra transactions.
The headline number on the same slide: a 16% reduction in losses incurred by fraud.

Five steps from synthetic fraud to a weakly supervised model.
Later, a “Real-world implications” table broke this down for 1 million processed transactions. The supervised model had 734 false negatives and 69,755 false positives. The weakly supervised model had 524 false negatives and 79,720 false positives. So it missed fewer frauds and raised more alerts. A second table then priced that trade-off at three different cost ratios between a missed fraud and a false alarm. The weakly supervised model came out ahead each time, with a loss difference of 6%, 11% and 16%.
That’s the honest way to present a fraud model: the 16% depends on how expensive a missed fraud is compared with an analyst checking a false alarm.
Explainability, support and ethics
The infrastructure slide, “Explainable AI infrastructure: High-level blueprint”, showed explainers as first-class components next to the models: model and explainer training, versioning and sign-off, deployment and serving of both models and explainers, audit trails and compliance checks, monitoring of drift with alerting, and an investigation step.
The last part of the talk went beyond fraud:
- User support: draft acceptance to cut agents’ time, automated ticket resolution, and what the slide called an intelligent support ecosystem. The example on screen was an answer to a user who had lost their bunq card.
- Automated marketing: personalised ad campaigns adjusted in real time and automated split testing.
- Ethical AI: data privacy, bias mitigation, transparency, explicit user consent before each use of confidential user data, and accountability through regular audits.
- Learnings: define clear objectives, ensure data quality and availability, build a feedback loop, design human-AI collaboration so people can validate AI suggestions, and comply with data-privacy laws and sector regulations.
ML6: building trustworthy LLM applications with guardrails
After the break, Thomas Vrancken’s title slide read “Building trustworthy LLM applications. Using guardrails.”, with “Meetup September 2024” and the ML6 logo. He opened with a Sundar Pichai quote, that AI is more profound than fire or electricity and that companies that fail to invest in it will fall behind. Then he moved straight to how that goes wrong.

ML6’s talk: guardrails, and how they get bypassed.
The framing slide: “Today we’ll talk about LLM safety mechanisms (guardrails) and how they are being bypassed (e.g. jailbreaking)”. The first example was the well-known “grandmother napalm” jailbreak, where the model is asked to pretend to be the user’s deceased grandmother, a chemical engineer at a napalm production factory. Then came a slide of headlines under “It is very common practice these days”: a prankster tricking a GM chatbot into agreeing to sell a $76,000 Chevy Tahoe for $1, an airline held liable for its chatbot giving a passenger bad advice, Anthropic warning that LLM guardrails fall to a simple “many-shot jailbreaking” attack, and “5 Easy Ways to Get Around ChatGPT Security Filters”.

“It is very common practice these days”: the headlines that make the case for guardrails.
The agenda slide, “A few key topics…”, promised four questions: how to identify the zones of vulnerability in an LLM application, what the typical types of vulnerability are, how to implement guardrails to prevent them, and what tools exist to build those guardrails. He then set up the reference architecture: Retrieval Augmented Generation (RAG), “a way to let LLMs capitalise on your knowledge database”, adjusting what the LLM knows by giving it access to extra information such as company internal data or third-party paid data.
The last slide I photographed was the “Long list of LLM application vulnerabilities”:
- Personal Identifiable Information (PII) leakage: PII can be present both in the user input and in the context retrieved by RAG.
- Content filtering: specific input or output the system should avoid (violence, hate, inappropriate or sensitive content).
- Prompt injections: attacks against applications built on top of LLMs, injecting malicious input into a trusted base prompt.
- Profanity: offensive, vulgar or inappropriate language in user inputs.
- Jailbreaking: attacks that try to bypass the safety filters built into LLM applications.
- Bias and fairness: unfair prejudice or favouritism towards certain viewpoints or groups in generated content.

Six vulnerability types, and the point where my photos stop.
My photos stop there, so I can’t tell you which guardrail tools he showed in the second half.
My take
The two talks were really about the same thing: you don’t trust a model’s output on its own, you put checks around it and keep humans in the loop. bunq does it with explainers, expert-weighted scores and audit trails around a fraud model. ML6 does it with guardrails around an LLM.
For LLM applications, the vulnerability list above is a good checklist, and I’d sort it into where the check runs. PII, prompt injection and jailbreak detection belong on the input. Content filtering, profanity and PII redaction belong on the output. Retrieval needs its own controls on what the RAG step can read. Since then I’ve tested this pattern locally with a safety model as the classifier in Granite Guardian on Ollama, and written up the agent version, where the risk is actions rather than text, in Guardrails for AI Agents in Production. The lesson from the Chevy Tahoe story still holds: a guardrail is cheap and an apology is expensive.
On bunq’s side, what I’d copy is the reporting: false negatives and false positives at several cost ratios, rather than a single accuracy number.