The skeptic's question

Why not just average the models?

Because the average believes everything, including the models that are broken. Below are three real Arasense runs: one where averaging costs you 64 percentage points, one where our weighting changes almost nothing and we say so out loud, and one where a single confident number would hide a split ensemble.

Exhibit A · Mexico City, max 1-day rainfall, 2050

One broken model. Sixty-four points.

The observed worst rain day around Mexico City is about 67 mm. One downscaled model in the ensemble carried a baseline of 1,513 mm, which is 22.6 times reality. Both runs below use the same pipeline; only one screens for that.

UNSCREENED ENSEMBLE
−69.2%
263 mm baseline · a physical impossibility

Every model gets a vote, including BCC-CSM2-MR at 1,513 mm. The corrupted baseline drags the whole ensemble, and the output looks like the region's rainfall collapses by two thirds. Sign this number and a technical reviewer will take it apart.

TRUST-FIRST · SCREENED
−4.6%
71.5 → 68.3 mm · 50% agreement · doubt reported

The extreme-plausibility screen checks every model against the observed 67 mm, excludes the broken one by name, and reports what remains: a genuinely contested signal, published with its 50% agreement attached.

64.6 points of pure artifact, removed not by judgement calls but by a one-parameter physical screen anchored to observations. The excluded model is named in the output, so the correction itself is auditable. See the published Mexico City profile →

Exhibit B · Bologna, 32 trusted models

When weighting changes nothing, we say so.

Most vendors would never show you this. Once Bologna's ensemble is screened, our skill weights and a plain equal-weight average land within 0.06 percentage points of each other. We publish that, because the claim we sell is not a magic number. It is the audit trail.

Equal weight, screened models
+14.31%
32 trusted models, uniform weights
Trust-weighted, same models
+14.25%
same 32 models, Aras Diagram skill weights
What you actually buy. The screening that rejected 2 of 34 models and can catch a 1,513 mm artifact; the named weights and rejections a reviewer can interrogate; the spread (±6.0 mm) and the 78% agreement printed next to the number. Where weighting moves the estimate, that is reported; where it does not, that is reported too. This analysis ships with our paper, not buried in a drawer.

Exhibit C · four cities, one habit

The average prints confidence. We print the doubt.

A plain ensemble average produces one authoritative-looking number for every city below. The agreement column is what it hides: in each of these, credible models genuinely disagree about the direction.

Paris−5.2% 46% Mexico City−4.6% 50% Sydney−2.3% 55% Venice+3.3% 56%

Agreement is the share of trusted-model weight backing the printed direction. Below ~60%, we tell you to treat the number as a question, not an answer. No single average can carry that information.


The offer

Which of these numbers would you sign?

Every Arasense engagement produces the right-hand column: screened, weighted, agreement-stamped evidence a technical reviewer can interrogate line by line. Start with one location.