A distributed job computes statistics over a huge column. Each worker processes one partition and reports a partial:
{"n": count, "mean": mean, "m2": m2, "min": min, "max": max}
where m2 is the sum of squared differences from that partition's mean, Σ (x − mean)². A worker that got no rows reports {"n": 0, "mean": None, "m2": None, "min": None, "max": None}.
Write merge_partials(partials) returning the statistics of all the values together, as if computed in one pass:
{"n", "mean", "min", "max", "variance"}
where variance is the population variance (divide by n). If there are no values at all, return {"n": 0, "mean": None, "min": None, "max": None, "variance": None}.
The values can be large (around 1e9) with a small spread, and the result must stay accurate for them.
Python 3.13 in your browser — the standard library plus pandas and numpy; no pip installs.