Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Learning Objectives, and Migration Strategy
- Course goals, alignment with participant profiles, and definitions of success criteria
- High-level migration approaches and associated risk assessments
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, Delta Lake overview, and Databricks architectural components
- Key differences between SMP and MPP models and their impact on migration
- Medallion (Bronze→Silver→Gold) architecture design and Unity Catalog introduction
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure to a Databricks notebook
- Mapping temporary tables and cursors to DataFrame transformations
- Validating and comparing results against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization techniques
Day 2 Lab — Incremental Ingestion & Optimization
- Implementation of Auto Loader ingestion and MERGE workflow processes
- Application of OPTIMIZE, Z-ORDER, and VACUUM with result validation
- Measurement of improvements in read/write performance
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL features: window functions, higher-order functions, and JSON/array manipulation
- Interpreting Spark UI, DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
- Query tuning strategies: broadcast joins, hints, caching, and reducing spill issues
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL queries
- Using Spark UI traces to detect and resolve skew and shuffle problems
- Benchmarking before and after changes and documenting tuning procedures
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Modularization techniques, UDFs/pandas UDFs, widgets, and building reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Transforming a procedural ETL script into modular PySpark notebooks
- Incorporating parametrization, unit-style testing, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error handling
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark
Day 5 Lab — Building a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retry mechanisms, and automated validations
- Executing the full pipeline, verifying outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, data lineage, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook development
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and key lessons learned
- Gap analysis, recommendations for follow-up activities, and handoff of training materials
- References, further learning pathways, and support options
Requirements
- A solid grasp of core data engineering principles
- Proficiency in SQL and stored procedures (Synapse / SQL Server)
- Knowledge of ETL orchestration concepts (ADF or equivalent tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers looking to transition from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption and implementation