DE tool

Flink

Stateful stream processing, checkpoints. Open-ended prompts; your written answer gets senior-reviewer feedback on the coached tier.

Easy

1 scenario
EZ

Backpressure on the source — find the bottleneck operator

~10 min

A Flink job reads from Kafka, runs a stateless map, a stateful keyed aggregation, and writes to Elasticsearch. The Kafka source operator is showing backpressure 100% (HIGH) in the Flink UI. End-to-end latency, measured at the sink, has grown from ~2s to ~90s over the last hour. The on-call DE just pinged you. Where do you look, in what order, and what's the likely root cause?

Unlock from ₹1,099 →
flinkbackpressureoperator-chains

Medium

2 scenarios
MD

Checkpoint duration went from 30s to 8min — RocksDB at 200GB

~15 min

A Flink keyed-state job has been running for 11 months. Until last month, checkpoints took ~30s and never timed out. Now they take 6–8 minutes and ~5% time out (configured timeout: 10 min). The RocksDB state backend reports ~200GB of state per TaskManager, growing 2GB/week. The job's business logic hasn't changed. Walk through your diagnostic process, the actual fix, and the operational guardrails so checkpoints stay fast as the job ages.

Unlock from ₹1,099 →
flinkcheckpointsrocksdb
MD

Event-time window never fires — one Kafka partition is idle

~15 min

Your Flink job reads from a 12-partition Kafka topic, groups by tenant_id, applies a 5-minute tumbling event-time window, and emits aggregates. For most tenants it works fine. But for 3 tenants the window NEVER fires — their aggregates never make it downstream. Investigating: partitions 4, 7, and 9 have had zero traffic for the last 45 minutes (those tenants went quiet). The other 9 partitions are flowing normally. Walk through what's broken, why, and the right fix.

Unlock from ₹1,099 →
flinkevent-timewatermarks

Hard

2 scenarios
HD

Kafka→Postgres exactly-once — the JDBC sink that wasn't

~25 min

A Flink job reads payment events from Kafka, enriches them via a keyed transformation, and writes them to Postgres. The team configured CheckpointingMode.EXACTLY_ONCE and the default JdbcSink. After a JobManager failover last week, Postgres has ~1,200 duplicate payments in the affected window. Walk through what end-to-end exactly-once actually requires, why this setup didn't deliver it, and how you'd redesign the sink path to make duplicates impossible after a failover.

Unlock from ₹1,099 →
flinkexactly-oncetwo-phase-commit
HD

Rescale 16→32 parallelism with 500GB keyed state, before Black Friday

~22 min

Black Friday is in 6 days. Your sessionization Flink job runs at parallelism 16 with ~500GB of keyed state in RocksDB. Projected traffic is 2.3× normal; load tests at parallelism 16 show backpressure pinning at ~1.8× normal load. You need to rescale to parallelism 32 before peak. The on-call DE has never rescaled a stateful job before. Walk through the plan — what you'd do, what can fail, and the ONE config-time detail that, if wrong, makes the whole rescale impossible.

Unlock from ₹1,099 →
flinksavepointsrescale