Skip to main content
🚀 Taking AI from prototype to production? Find the architecture, GPU, security and governance gaps before they become incidents. Get a Production AI Readiness Assessment
Attendees at the Databricks and RevoData Fundamentals Bootcamp in Amsterdam, with the welcome slide on the screen
AI

Databricks Fundamentals Bootcamp with RevoData 2025

A day at the Databricks Fundamentals Bootcamp with RevoData in Amsterdam: lakehouse, Unity Catalog, Lakeflow, Genie and Spark real-time mode.

LB
Luca Berton
· 4 min read

On 22 May 2025 I spent a full day at the Databricks Fundamentals Bootcamp, a joint Databricks and RevoData session held at The Wings in Amsterdam. It followed the Databricks Data + AI World Tour in Amsterdam and was the hands-on, slower-paced counterpart: a small room, laptops open, one platform walked through from storage to AI.

This is a throwback post written from the slides I photographed. The speaker slide listed Jesse Schouten (Revodata), Jasper Vogelzang (Databricks) and Charlotte Blankenberg (Databricks). I am not matching names to individual talks, since the slide only lists the team.

Small training room with attendees at tables facing a screen showing Welcome to Databricks Fundamentals Bootcamp

The welcome slide: “Databricks x RevoData, Welcome to Databricks Fundamentals Bootcamp!”

Open formats first: Delta Lake and Iceberg

The first technical block was about the data estate. A slide called out the problem as a fragmented estate with high costs and proprietary formats. The answer was open table formats: a slide showed Delta Lake and Apache Iceberg side by side, with Databricks and Tabular, and a later slide listed Delta Lake UniForm alongside Unity Catalog, AI models, data sharing, access control and lineage.

Slide showing Delta Lake and Iceberg, with Databricks plus Tabular below

Delta Lake and Iceberg: the open-format story behind the lakehouse.

Warehouses, lakes and the lakehouse

The framing slide contrasted data warehouses (structured tables, BI and SQL analytics, table ACLs) with data lakes (files, logs, text, images and video, feeding data science, ML and streaming). The cost of having both is copying subsets of data between them, with separate governance on each side. The lakehouse was presented as “an open, unified foundation for all your data”.

Slide titled Data Warehouses vs. Data Lakes showing BI and SQL on the warehouse side and data science, ML and streaming on the lake side, with data copied between them

Two systems, two governance models, and data copied between them. That is the problem the lakehouse targets.

Platform architecture

The architecture slide split the platform into a control plane (web app, Unity Catalog with metastore and access control, Mosaic AI, Workflows, Git folders, notebooks, DBSQL) and a data plane with elastic compute (clusters and SQL warehouses) over storage in S3, ADLS or GCS with Delta Lake. A following slide announced that serverless compute is generally available, described as hands-off, auto-optimised compute managed by Databricks.

Infrastructure and Platform slide with the control plane on the left and the data plane with elastic compute and storage on the right

Control plane on the left, data plane with compute and storage on the right.

Governance with Unity Catalog

Governance took a good part of the morning. The overview slide gave three points: unify governance across clouds with fine-grained control based on ANSI SQL, unify data and AI assets in one interface, and unify existing catalogs with no hard migration required. The Databricks docs describe Unity Catalog as the unified governance layer for data and AI, covering access enforcement, lineage tracking and audit logging.

Databricks Unity Catalog overview slide with three columns: unify governance across clouds, unify data and AI assets, unify existing catalogs

The Unity Catalog overview slide.

Related slides covered Delta Sharing and the Databricks Marketplace, clean rooms for mutually approved computation without replicating data, and natural-language data discovery.

Pipelines: Lakeflow and Jobs

The data engineering section started with the building blocks of Databricks Workflows: a job is the unit of orchestration, made of tasks and triggers, with control flow such as for-each loops and a timeline view of task execution for troubleshooting. The LakeFlow slide grouped three layers: Connect for ingest, Pipelines (with Delta Live Tables noted beside it) for transformation, and Jobs for orchestration, all on Unity Catalog and serverless compute. Current wording is in the Lakeflow documentation.

LakeFlow slide with Ingest, Transform and Orchestrate layers mapped to Connect, Pipelines and Jobs

LakeFlow: ingest, transform and orchestrate in one place.

AI/BI Genie and Mosaic AI

On the analytics side, an AI/BI Genie slide described asking questions of your data in natural language, as a self-service route that does not need a data analytics specialist for every question. The Genie documentation shows the product has kept evolving since that day. The AI section listed Mosaic AI: agent framework, model serving and model training, plus a gateway, guardrails, usage tracking and LLM-judge based agent evaluation.

Streaming latency

One slide that stood out in the afternoon compared Spark real-time mode with micro-batch for stateless joins, transformations, simple aggregations, deduplication and stream-static joins with aggregation. Its headline claimed p50 latency in tens of milliseconds and p99 around 100 ms. It is a vendor chart, so I would test it on your own workload, but the gap on the stateful cases is large.

Bar chart slide comparing Spark Real Time Mode with Microbatch latency across five workload types

Spark real-time mode versus micro-batch, as presented.

Next sessions

The closing slides announced the next bootcamp, on Data Warehousing on Databricks on 19 June at The Wings, and a save-the-date for Databricks Data + AI World Tour Amsterdam on 6 November 2025.

Slide inviting attendees to the next bootcamp on Data Warehousing on Databricks on June 19th at The Wings

The next bootcamp: Data Warehousing on Databricks.

My take: for platform teams the useful part of a day like this is the vocabulary. Control plane versus data plane, Unity Catalog as the single governance point, and serverless compute are the three ideas that decide how you design access, networking and cost around Databricks.

Free 30-min Production AI consultation

Book Now