Why not just average the models?
Because the average believes everything, including the models that are broken. Below are three real Arasense runs: one where averaging costs you 64 percentage points, one where our weighting changes almost nothing and we say so out loud, and one where a single confident number would hide a split ensemble.
One broken model. Sixty-four points.
The observed worst rain day around Mexico City is about 67 mm. One downscaled model in the ensemble carried a baseline of 1,513 mm, which is 22.6 times reality. Both runs below use the same pipeline; only one screens for that.
Every model gets a vote, including BCC-CSM2-MR at 1,513 mm. The corrupted baseline drags the whole ensemble, and the output looks like the region's rainfall collapses by two thirds. Sign this number and a technical reviewer will take it apart.
The extreme-plausibility screen checks every model against the observed 67 mm, excludes the broken one by name, and reports what remains: a genuinely contested signal, published with its 50% agreement attached.
When weighting changes nothing, we say so.
Most vendors would never show you this. Once Bologna's ensemble is screened, our skill weights and a plain equal-weight average land within 0.06 percentage points of each other. We publish that, because the claim we sell is not a magic number. It is the audit trail.
The average prints confidence. We print the doubt.
A plain ensemble average produces one authoritative-looking number for every city below. The agreement column is what it hides: in each of these, credible models genuinely disagree about the direction.
Agreement is the share of trusted-model weight backing the printed direction. Below ~60%, we tell you to treat the number as a question, not an answer. No single average can carry that information.
Which of these numbers would you sign?
Every Arasense engagement produces the right-hand column: screened, weighted, agreement-stamped evidence a technical reviewer can interrogate line by line. Start with one location.