☁️ Category 05 • Multi-Cloud Lakehouse & Infrastructure

Cloud: AWS, Azure, Amazon S3 & Databricks

Designing scalable multi-cloud architectures. In-depth coverage of AWS (S3, EMR, Glue, Redshift, MSK), Microsoft Azure (ADLS Gen2, Data Factory, Synapse), Databricks Unity Catalog governance, IAM least-privilege security, and cloud cost optimization.

1. Multi-Cloud Data Engineering Strategy

Enterprise Infrastructure

Modern enterprise organizations rarely rely on a single vendor. Understanding the direct equivalencies and architectural tradeoffs between AWS and Microsoft Azure, alongside Databricks as an open lakehouse compute engine, is a defining competency for Senior Data Engineers and Cloud Architects.

AWS Data Stack

Amazon S3, AWS Glue (Catalog + Serverless Spark), Amazon EMR, Amazon Redshift, Amazon MSK (Kafka), and AWS MWAA (Airflow).

Azure Data Stack

Azure Data Lake Storage (ADLS Gen2), Azure Data Factory (ADF), Azure Synapse Analytics, Azure Event Hubs, and Azure Databricks.

Databricks Lakehouse

Decoupled Control Plane (managed SaaS) and Data Plane (in customer VPC/VNet). Delta Lake ACID tables governed by Unity Catalog.

Storage vs Compute Decoupling

Persisting data in low-cost object storage (S3/ADLS) while scaling compute clusters (EMR/Databricks) dynamically on demand.

Topic 5.1

AWS Fundamentals & Managed Big Data Services

AWS Ecosystem

Amazon Web Services powers high-throughput data platforms through a suite of integrated managed services.

AWS Glue (Catalog & ETL)

Hive-compatible metastore, automated schema crawlers, and serverless distributed Spark runtime charged per DPU-hour.

Amazon EMR (Elastic MapReduce)

Managed clusters running PySpark, Trino, Presto, and Flink on EC2 Spot/On-Demand or EMR Serverless.

Amazon Redshift & Redshift Serverless

Massively Parallel Processing (MPP) columnar cloud data warehouse with Spectrum support for querying external S3 tables directly.

Amazon Athena & EventBridge

Serverless interactive SQL queries powered by Presto/Trino over S3 Parquet, triggered automatically via EventBridge on file arrival.

Topic 5.2

Microsoft Azure Fundamentals & Enterprise Integration

Azure Ecosystem

Microsoft Azure dominates enterprise financial and telecom workloads (e.g. XXX XXX engagement), providing seamless integration with Active Directory (Entra ID) and enterprise Power BI.

ADLS Gen2 (Hierarchical Namespace)

Blob storage with POSIX-compliant atomic file directory renames, eliminating expensive recursive file copies during Spark writes.

Azure Data Factory (ADF)

Cloud-based graphical ETL/ELT service for orchestrating batch copy activities, triggering Databricks notebooks, and SSIS package lift-and-shift.

Managed Identities & Azure Key Vault

Passwordless service-to-service authentication between ADF, Databricks, and ADLS without embedding client secrets in code.

Azure Synapse Analytics

Unified enterprise analytics service combining dedicated SQL pools (formerly SQL DW), serverless SQL pools, and Apache Spark.

Topic 5.3

Amazon S3 Deep Dive: Tiers, Lifecycle, Security & Prefix Tuning

Object Storage Master

Amazon S3 is the foundational storage layer of the modern data lake. Mastering S3 requires understanding strong read-after-write consistency, multipart upload mechanics, lifecycle transitions, and high-throughput prefix sharding.

Storage Tiers & Lifecycle Policies

S3 Standard → S3 Standard-IA (after 30 days) → S3 Glacier Flexible (after 90 days) → S3 Deep Archive (after 365 days). Slashes storage bills by 80%.

Prefix Performance Tuning

S3 automatically scales to 3,500 PUT/POST/DELETE and 5,500 GET requests per second per prefix. Partitioning across prefixes avoids HTTP 503 SlowDown throttling.

Multipart Uploads & S3 Transfer Acceleration

Mandatory for files > 100MB. Parallelizes part uploads across network threads, resuming interrupted uploads without restarting.

S3 Object Lock & Immutability

WORM (Write Once, Read Many) compliance mode preventing record deletion or modification even by the root AWS account, satisfying FINRA/SEC rules.

📄 s3_lakehouse_storage.tf
# Terraform Infrastructure as Code for Secure S3 Lakehouse
resource "aws_s3_bucket" "lakehouse_curated" {
  bucket = "enterprise-financial-lakehouse-gold"
}

# Enforce SSE-KMS Customer Managed Key Encryption
resource "aws_s3_bucket_server_side_encryption_configuration" "lake_kms" {
  bucket = aws_s3_bucket.lakehouse_curated.id
  rule {
    apply_server_side_encryption_by_default {
      kms_master_key_id = "arn:aws:kms:eu-west-1:123456789012:key/de-lake-key"
      sse_algorithm     = "aws:kms"
    }
    bucket_key_enabled = true # Reduces KMS API request costs by 99%
  }
}

# Block all public access
resource "aws_s3_bucket_public_access_block" "block_public" {
  bucket                  = aws_s3_bucket.lakehouse_curated.id
  block_public_acls       = true
  block_public_policy     = true
  ignore_public_acls      = true
  restrict_public_buckets = true
}

# Automated Lifecycle transition to Glacier Flexible after 90 days
resource "aws_s3_bucket_lifecycle_configuration" "lake_lifecycle" {
  bucket = aws_s3_bucket.lakehouse_curated.id

  rule {
    id     = "archive-historical-partitions"
    status = "Enabled"

    transition {
      days          = 90
      storage_class = "GLACIER"
    }
  }
}
Topic 5.4

Databricks Architecture, Unity Catalog & Delta Lake

Lakehouse Leader

Databricks is the premier enterprise lakehouse engine, unifying data engineering, SQL BI, and machine learning on top of open cloud object storage.

Control Plane vs Data Plane

The Control Plane hosts the web UI, cluster manager, and job scheduler. The Data Plane resides inside the customer's AWS/Azure VPC where Spark nodes compute.

Unity Catalog Governance

Centralized 3-tier namespace (catalog.schema.table) with cross-workspace governance, automated data lineage, and row/column access policies.

All-Purpose vs Job Clusters

All-purpose clusters serve interactive development (expensive). Job clusters spin up automatically for a scheduled workflow and terminate immediately (50% cheaper).

Databricks Serverless Compute

Instantly provisioned Spark compute managed directly by Databricks, eliminating cluster startup wait times (sub-10 seconds startup) and idle costs.

Topic 5.5

Cloud Cost Optimization & FinOps for Data Platforms

FinOps Strategy

Cloud bills can quickly spiral out of control. Senior data engineers implement automated FinOps practices:

EC2 Spot Instances for Spark Workers

Run 80% of Spark worker nodes on Spot instances (up to 70% discount) while keeping the Spark Driver on On-Demand to prevent cluster termination.

S3 Bucket Keys for KMS

Enabling S3 Bucket Keys caches KMS encryption keys, slashing AWS KMS request fees by 99% when processing millions of small Parquet files.

Auto-Termination Policies

Enforcing mandatory 20-minute auto-termination on all interactive Databricks and EMR clusters when no notebooks are active.

Delta Lake Compaction & VACUUM

Running OPTIMIZE to merge small files and VACUUM to prune historical transaction snapshots older than 7 days, freeing terabytes of S3 storage.

2. Multi-Cloud AWS & Azure Enterprise Architecture

Production Architecture
=========================== AWS DATA LAKEHOUSE =========================== [S3 Bronze: Raw Ingestion] → [Amazon EMR / Serverless Spark] → [S3 Silver: Parquet] ↓ [Amazon MSK (Kafka)] → [AWS Glue Catalog] → [Amazon Redshift / Athena SQL] ========================== AZURE DATA LAKEHOUSE ========================== [ADLS Gen2: Bronze] → [Azure Databricks (Job Cluster)] → [ADLS Gen2: Silver] ↓ [Azure Data Factory] → [Unity Catalog Governance] → [Power BI DirectQuery]
Hands-On Lab

Lab 05: Build a Medallion Lakehouse on AWS S3 & Azure Databricks

Lab Guide
1

Provision S3 Medallion Buckets

Create Bronze (raw JSON), Silver (curated Delta), and Gold (aggregated dimensional) buckets with SSE-KMS encryption.

2

Connect Azure Databricks with Service Principal

Mount ADLS Gen2 / configure S3 bucket IAM role credentials in Databricks Unity Catalog external storage credentials.

3

Execute Bronze to Silver Ingestion

Run an automated PySpark job cleaning customer transactions, applying schema validation, and writing to Delta Lake.

4

Configure Auto-Scaling & Spot Optimization

Configure cluster min 2 / max 8 workers with 100% Spot instances, enabling auto-termination after 15 minutes of inactivity.

3. Cloud Data Engineering Interview Questions

Architect Questions
Q1: How do you design an S3 data lake to prevent HTTP 503 SlowDown throttling errors? ▲

Answer: S3 enforces rate limits of 3,500 PUT/POST/DELETE and 5,500 GET requests per second per prefix. In S3, a "prefix" is any directory path string between delimiters (e.g. bucket/prefix/object.parquet).
Mitigation:
1. Sharded Hash Prefixes: Prepend an MD5 hash or randomized partition key before sequential date strings (e.g. s3://lake/a8f2/2026/10/05/data.parquet).
2. File Compaction: Avoid writing thousands of 1KB micro-files. Compact Spark writes into 128MB–512MB files using repartition() or Delta Lake OPTIMIZE.
3. Exponential Backoff: Configure the AWS SDK / Hadoop S3A filesystem client retry parameters with exponential backoff and jitter.

Q2: What are the architectural differences between AWS EMR and Databricks? ▼

Answer:
• Amazon EMR: Lower runtime cost, open-source Spark binaries, highly customizable with custom bootstrap actions, integrated natively with AWS IAM and Lake Formation. Requires more operational DevOps effort for cluster configuration and tuning.
• Databricks: Proprietary optimized runtime (Photon C++ vectorized engine), proprietary enterprise features (Unity Catalog, Liquid Clustering, Serverless SQL warehouses), collaborative web workspace, faster startup and autoscaling, but incurs an additional DBU (Databricks Unit) license fee on top of underlying cloud VM costs.

4. Related Tracks

Explore Next

Advance into Data Platforms & Storage Engines

Master Delta Lake ACID transactions, Snowflake micro-partitioning, and Kimball dimensional modeling.

🏛️ Explore Data Platforms Track →
Open AWS & Azure interview questions →