DE tool

Spark

Shuffle, skew, OOM diagnosis, structured streaming. Open-ended prompts; your written answer gets senior-reviewer feedback on the coached tier.

Easy

1 scenario
EZ

Stage 4 of 7 takes 80% of the wall clock — find the skew

~10 min

Your nightly ETL joins orders (200M rows) to customers (5M rows) and writes a denormalized fact table. The job has 7 stages. Stages 1–3 and 5–7 finish in under 2 minutes each. Stage 4 takes 38 minutes — 195 of its 200 tasks finish in the first 90 seconds, then 5 tasks crawl for the rest. The business owner wants the job under 15 minutes. Where do you look, and what's your fix?

Medium

2 scenarios
MD

Broadcast join blew the executor heap — works 2 runs in 3

~15 min

A daily Spark job that broadcast-joins a "small" product dimension table into a fact table has run cleanly for 6 months. Last week it started OOMing the executor every third run — non-deterministically, on different nodes each time. The product table was 40MB when the job shipped; it's now 780MB and growing ~30MB/week as new SKUs onboard. Executors are 8GB each. Walk through your diagnostic process, the actual fix, and the guardrails so this class of failure stops surprising you.

Unlock from ₹1,099 →
sparkbroadcast-joinoom
MD

One customer is 30% of the clickstream — pick the join strategy

~18 min

You're joining 5B clickstream events to a 20M-row customer dimension on customer_id. Histogram of events by customer: one enterprise tenant generates 30% of all events; the next 9 together are another 25%; the long tail is the rest. You've tried Spark 3.5 AQE with skew join enabled — it helps, but stage 3 still takes 4× the wall clock of stages 2 and 4. You have one afternoon to get this job under SLA. Walk through your options.

Hard

2 scenarios
HD

Structured streaming reprocesses on restart — duplicate rows in Iceberg

~25 min

A Structured Streaming job reads from Kafka, applies a stateful aggregation over a 15-minute event-time window, and writes to an Iceberg sink. After a 10-minute outage and restart, the downstream analytics team reports ~3% duplicate rows in the affected window. Your team's lead engineer says "checkpoints should handle this" and is annoyed. Walk through what Spark actually guarantees here, why duplicates appeared, and how you'd redesign the sink path so this stops happening — without giving up on event-time correctness.

Unlock from ₹1,099 →
sparkstructured-streamingexactly-once
HD

Daily batch from 45 min to 4 hours — find the regression in 30

~25 min

A production daily batch job has run reliably at ~45 min for 18 months. After last week's release, runtime jumped to 4+ hours and is now blowing through the downstream SLA. The release diff is large — multiple model changes, a config bump, a new MERGE for a slowly-changing dimension. The business owner wants the job back under SLA by EOD. Walk through your forensic process: where to look in 30 minutes, what to rule out, what's most likely.

Unlock from ₹1,099 →
sparkregressionspark-ui