Skip to main content
šŸ“¬ Get weekly Production AI insights Practical notes on Kubernetes, AI infrastructure and platform engineering. No spam. Subscribe free
Luca Berton taking a selfie in the stands of the Johan Cruijff ArenA, with the Databricks stage and agenda screens on the pitch
database

Databricks Data + AI World Tour Amsterdam 2024

A stage on the Ajax pitch and a hands-on Get Started with Databricks for Data Engineering class: compute, Delta Lake, Workflows, DLT and Unity Catalog.

LB
Luca Berton
Ā· 5 min read

On Thursday 14 November 2024 the Databricks Data + AI World Tour came to Amsterdam, and the venue was a football stadium: the Johan Cruijff ArenA, home of Ajax. The main stage stood on the pitch, facing the lower tier of red seats, and my training room had ā€œAJAX for the futureā€ on the wall behind the instructor.

The part of the day this post focuses on is the instructor-led ā€œGet Started with Databricks for Data Engineeringā€ class. Below is what it taught, from the slides I photographed, with links to the current Databricks documentation, because some product names have changed since.

Luca Berton smiling in front of a Databricks DATA+AI World Tour Amsterdam screen

Arrival at the World Tour.

A stage on the pitch

The agenda screens next to the stage listed the day: keynotes part 1, the Benelux Data + AI Awards, keynotes part 2, lunch, a first block of breakouts and trainings, an afternoon break, a second block of breakouts and trainings, a networking reception and the ā€œData After Hoursā€ party.

The Databricks stage and screens set up on the grass of the Johan Cruijff ArenA, seen from the lower stands

The keynote stage on the pitch, with the audience in the lower tier.

Between sessions I stopped for a photo at the Ajax and Eredivisie sponsor wall.

Luca Berton standing with arms crossed in front of the Ajax and Eredivisie sponsor press wall

At the Ajax and Eredivisie sponsor wall.

The training: Get Started with Databricks for Data Engineering

The class ran in Databricks Academy as ā€œDAIWT 2024: Get Started with Databricks for Data Engineeringā€, Amsterdam session, in English, with a hands-on lab environment. The instructor was Tjerk Kostelijk, a senior technical instructor at Databricks according to his introduction slide. The course learning objectives were:

  • describe the available compute options for workloads on the Databricks Data Intelligence Platform;
  • navigate the Databricks Workspace UI;
  • describe the architecture and benefits of Delta Lake;
  • apply various techniques for ingesting data into Delta Lake;
  • the Medallion Architecture;
  • describe how Workflows provide unified orchestration in Databricks.

The instructor at the lectern next to a screen showing the Course Learning Objectives, in a meeting room with AJAX for the future on the wall

The course objectives for the afternoon.

The platform and its two planes

The overview started with the Data Intelligence Platform slide, which grouped AI/BI Dashboards and Genie, Mosaic AI, Databricks SQL and Delta Live Tables on top of a shared base. The infrastructure slide split the platform into a control plane and a data plane. Storage and governance sat on cloud data storage with Unity Catalog and Delta Sharing, and Partner Connect linked out to tools such as Fivetran and Rivery. The slide’s message was ā€œBuilt on an open foundationā€.

Compute: runtimes, SQL warehouses and serverless

The compute slide described two runtime families. The Standard runtime is Apache Spark plus many other components and updates for optimised big data analytics. The Machine Learning runtime adds popular libraries such as TensorFlow, Keras, PyTorch and XGBoost. Under Specialized Compute were SQL warehouses, designed for SQL BI workloads with built-in optimisation for price/performance and enhanced with Databricks Photon. A sample configuration on a later slide used Databricks Runtime 14.3 LTS (Scala 2.12, Spark 3.5.0) on a multi-node cluster.

Slide: Databricks Compute, with Standard and Machine Learning runtimes and SQL Warehouses enhanced with Databricks Photon

Runtimes on the left, SQL warehouses on the right.

The benefits of serverless slide made the case in three columns:

  • Higher user productivity: queries start immediately without waiting for cluster start-up, and more concurrent users are handled by instant scaling.
  • Zero management: no configuration, performance tuning or capacity management, and automatic upgrades and patching.
  • Lower cost: pay for what you use, no idle clusters or over-provisioning, and idle capacity is removed 10 minutes after the last query.

The current serverless compute docs describe the same model: Databricks allocates and manages the compute, rather than you provisioning it in your own cloud account.

Orchestration and ETL

The Orchestration and ETL slide drew a line I found useful. Databricks Workflows is control flow: it orchestrates anything in the platform, including DLT pipelines. Delta Live Tables is data flow: automated data pipelines for Delta Lake. You use DLT to declare how data moves between tables, and Workflows to decide when that pipeline runs and what runs before and after it.

Slide: Orchestration and ETL, comparing Databricks Workflows as control flow with Delta Live Tables as data flow

Workflows for control flow, Delta Live Tables for data flow.

The names have changed since the training. The Databricks docs now call these Lakeflow Jobs and Lakeflow pipelines, but the split between control flow and data flow is the same.

Unity Catalog

The governance section had three Unity Catalog slides. The first compared workspaces before and after Unity Catalog. The second put the metastore at the top, with external storage access, catalogs, query federation and Delta Sharing beneath it. The third showed the object hierarchy, from catalog to schema to tables, views, volumes, functions and models, ending in the three-level name you query with:

SELECT * FROM catalog1.schema1.table1;

Slide: Unity Catalog Overview, showing a Databricks account, metastores, catalogs, schemas and tables, views, volumes, functions and models

The Unity Catalog hierarchy and its three-level namespace.

The Unity Catalog documentation describes the same catalog.schema.object namespace, with every level a securable object you can grant permissions on.

My take: I came at this from the platform side, and Unity Catalog was the part that mattered most to me. A single metastore per region with grants at catalog, schema and table level is what lets one platform team serve many data teams without copying data or duplicating permission systems. It’s the same pattern as a Kubernetes cluster with namespaces and RBAC: one control point, many tenants.

The expo

I also walked the partner expo, with stands including Fivetran and Dataiku, and stopped by the Databricks swag store.

The expo floor seen from above, with Fivetran and Dataiku stands and attendees walking between them

The partner expo from the upper level.

#Databricks #Data + AI World Tour #Amsterdam #Data Engineering #Unity Catalog #Delta Lake #Delta Live Tables #Databricks Workflows #Serverless #Training
Share:
Luca Berton — The Production AI Expert, Docker Captain

Luca Berton

The Production AI Expert Ā· Docker Captain Ā· KubeCon Speaker

15+ years in enterprise infrastructure. Author of 8 technical books, creator of Ansible Pilot (1M+ YouTube views, 648K site users). Former Red Hat engineer. Speaker at KubeCon EU 2026 and Red Hat Summit 2026.

Free 30-min Production AI consultation

Book Now