Backpressure on the source — find the bottleneck operator
A Flink job reads from Kafka, runs a stateless map, a stateful keyed aggregation, and writes to Elasticsearch. The Kafka source operator is showing backpressure 100% (HIGH) in the Flink UI. End-to-end latency, measured at the sink, has grown from ~2s to ~90s over the last hour. The on-call DE just pinged you. Where do you look, in what order, and what's the likely root cause?