Highlights
- Full Lakehouse-to-production journey: architecture, Spark, Delta Lake, governance, performance, streaming, pipelines, orchestration, CI/CD
- Every topic paired with a hands-on lab in a live Databricks workspace — not just slides
- Deep, dedicated Delta Lake coverage: ACID transactions, schema evolution, time travel, and optimisation
- Full Unity Catalog governance module: permissions, lineage, external locations, and Microsoft Purview integration
- Real production-engineering content: Structured Streaming, Auto Loader, Lakeflow Declarative Pipelines (DLT), Workflows, and CI/CD with Asset Bundles
- Closes with an end-to-end capstone tying raw files through to a governed, optimised, orchestrated, promoted pipeline
- Maps directly to Microsoft's official DP-3011 curriculum, with extended depth in performance tuning and production operations
- Covers Azure and AI integration, including Azure AI Foundry working with Databricks tables
Course Details
Module 1 — Lakehouse Foundations, Compute & Apache Spark (Day 1, Morning)
- Understand the Lakehouse architecture and how Azure Databricks relates to Apache Spark and the wider Azure ecosystem
- Choose the right compute for a workload: all-purpose clusters vs. job compute vs. serverless, including autoscaling and cluster policies
- Learn how Spark actually executes work: transformations vs. actions, lazy evaluation, and the DAG
- Hands-on labs throughout: tour the workspace, provision compute, and read/filter/aggregate data with the Spark DataFrame API
Module 2 — Delta Lake Deep Dive & Databricks SQL (Day 1, Afternoon)
- Master Delta Lake fundamentals: table anatomy, the transaction log, and how it delivers ACID guarantees
- Create and manage Delta tables in both SQL and PySpark, including managed vs. external tables and concurrency handling
- Apply schema evolution, time travel (VERSION AS OF / TIMESTAMP AS OF), and table optimisation (OPTIMIZE, VACUUM, Z-ORDER)
- Get hands-on with Databricks SQL: the SQL editor, SQL warehouses, and building first visualisations
Module 3 — Unity Catalog Governance & Performance Optimisation (Day 2, Morning)
- Design a governance model using Unity Catalog's metastore → catalog → schema hierarchy, with fine-grained permissions and access control
- Trace data lineage, audit access, and register external locations and storage credentials — including Microsoft Purview integration
- Diagnose and fix slow queries using the Spark UI, query profiles, and partitioning/Z-ORDER/liquid clustering strategies
- Apply advanced Databricks SQL techniques: window functions, materialised views, caching, and Photon acceleration considerations
Module 4 — Ingestion, Pipelines, Orchestration & Production Integration (Day 2, Afternoon)
- Build reliable streaming ingestion with Auto Loader and Structured Streaming, including checkpointing and exactly-once semantics
- Author a Lakeflow Declarative Pipeline (DLT) implementing the Bronze/Silver/Gold medallion architecture with enforced data-quality expectations
- Orchestrate multi-step production workflows with Lakeflow Jobs, including scheduling, dependencies, error handling, and monitoring
- Promote work safely with Git integration, CI/CD via Databricks Asset Bundles, and connect to the wider Azure and AI ecosystem (Azure AI Foundry) — closing with a full end-to-end capstone
Who should attend
- Data Engineers looking to enhance their cloud-based data processing skills
- Data Scientists wanting to leverage Azure Databricks for advanced analytics
- Professionals with experience in data handling who want to scale up to large datasets
- Anyone interested in mastering Apache Spark and Azure Databricks for big data solutions
Feedback
4.8 out of 5 average
" I enjoyed the depth that we covered analytical techniques such as anomaly detection and cluster analysis, whilst improving my knowledge on DAX and KPIs."BC, Performance analyst, Data Analysis with Power BI, April 2021
Watch live client feedback from Data Analytics courses:
“JBI did a great job of customizing their syllabus to suit our business needs and also bringing our team up to speed on the current best practices. ” Brian F, Team Lead, RBS, Data Analysis Course, 20 April 2022