A combine portal · ten layers, one axis

The Ruler Nobody Declared

Ten layers of this archive do the same honest thing: rebuild a published result under its authors' exact choices, enumerate every other choice that would have been defensible, and report where the published number sits inside the cloud that produces. Each one prints a percentile. Nobody had ever put those percentiles side by side, and doing it changes what they mean. The percentile is not a property of the finding. It is a property of the grid somebody chose to enumerate, and this page hands you the grid.

What is on this page

Three things, none of which a member can say alone

Every number below was read out of the ten member pages by booting each one in a real browser with the shared kit intercepted, so the values are the members' own, produced by their own shipped code on their own shipped data. The first thing this page checks about itself is that its percentiles are bit-identical to the ones those pages print.

published results located inside their own cloud, across ten layers and specifications. Their percentiles run from to .
the largest move a single published number makes when one analytic axis is held at the authors' own choice, in a grid of a thousand cells or more. Eleven of the fourteen move by ten points or more.
of the specifications in the one grid built to model real analysts are more conservative than the median of the twenty-nine real analyst teams it stands in for. All thirty-four of its single-axis slices are.

One · the ruler

Fourteen findings on a single axis

A specification curve answers a question of the form how far out on the range of defensible answers did the published one land? The answer is a percentile, and a percentile is comparable across subjects in a way the underlying quantities are not: an odds ratio, a warming trend, an earnings ratio and a language-split date share no units, but "sixty-ninth percentile of its own multiverse" means the same thing everywhere. So the first thing to do with ten such layers is to line them up.

Each row below is one published result. The filled dot is where its authors' own choice sits inside its own cloud. The pale ticks on the whisker are the same number ranked inside a grid one axis narrower, one tick per axis, computed live in your browser from the members' own specification values. Press the button to add the control.

the authors' own choice the same number, one axis held the same operation on a shuffled grid

Loading the fourteen grids…

Two · the counterfactual, by hand

Hold one choice fixed and watch the number move

Pick a result, then suppose its authors had never thought to vary one of their axes. The value in the cell does not change. The population it is ranked against does, and so the percentile does. This is not a criticism of any member: every one of them declares its grid in full. It is that the grid is a choice, made by one analyst, and the headline percentile inherits it.

LayerThe result locatedDefined cells Own percentileHeld-axis rangeSpan
Loading…

Select a row above.

Three · the undeclared parameter

How wide is a multiverse?

Every grid in this wave is a judgement about what counts as a defensible analysis, and every one of those judgements is an integer nobody argues about. The twelve grids these ten layers run span three orders of magnitude. Two of them evaluate fewer specifications than they declare, because the crossed product contains combinations the page rules out before running them, and the difference between those two numbers is not printed anywhere else.

LayerAxesDeclaredRun DefinedUndefined
Loading…

Four · the control the wave contained and never used

Twenty-nine real analysts against six thousand imagined ones

A specification curve is a model of analyst behaviour. It says: here is what a room full of competent people might have done with this dataset. In exactly one member of this wave, that room was actually observed. Twenty-Nine Answers, One Dataset rebuilds Silberzahn and colleagues' crowdsourced study, in which twenty-nine independent teams analysed one football dataset and reported twenty-nine different answers, and it also enumerates a grid of 6,912 specifications over the same data.

That page asks whether the machine's range contains the humans. It does: twenty-eight of twenty-nine land inside. The question it does not ask is whether the machine is calibrated to them, and the answer is that it is not, in a way one number makes plain.

of the machine's specifications sit below the median of the twenty-nine real teams. If the grid were a fair model of what analysts do, that would be near 50.
is where the machine's own median falls among the humans. The two distributions are offset, not nested.
of the twenty-nine teams land at or above the machine's ninetieth percentile. Four land at or below its tenth.

It is not one axis, and it is not the missing model class

The obvious explanation is that the grid omits an axis the humans varied. It does omit one: seventeen of the twenty-nine teams describe a multilevel, hierarchical, mixed, clustered or Bayesian model, and no option anywhere in the grid's thirty-four declared levels is any of those. That explanation fails anyway. Those seventeen teams have a median odds ratio of and the other twelve have , which is the same number. The prediction was made here, tested here, and is reported as lost.

What is left is a property of the whole grid rather than of any axis in it. Slice the 6,031 defined specifications by every level of every axis, thirty-four slices in all, and take the median of each. Not one of the thirty-four reaches the median human team. The highest is , still short of the humans' . There is no axis you could drop, and none you could add from within the declared set, that closes the gap.

AxisLevelCellsMedian odds ratio Distance below the median human team
Loading…

Five · what this does not say

The limits, stated before anyone has to find them

Six · the check

What stands behind this page

The apparatus is research/the-ruler-nobody-declared/. It is three programs: harvest.mjs boots each of the ten member pages in Chromium with /_kit/multiverse.js intercepted and records every call the page makes into it; derive.mjs reshapes that into the two binaries this page fetches; verify.mjs re-derives every claim above and asserts it. Run it with node research/the-ruler-nobody-declared/verify.mjs.

The load-bearing check is the first one. If the percentiles here were not bit-identical to the ones the member pages print, this portal would be a re-implementation quietly disagreeing with its sources, which is the most likely way a page like this lies. Three of the checks are controls designed to fail if the thing they guard is not really being measured, and two record predictions this work made and lost.