Skip to content

Rewrite forecast-accuracy claims around a defensible benchmark#88

Merged
vahid-ahmadi merged 2 commits into
mainfrom
benchmark-honesty
Jul 24, 2026
Merged

Rewrite forecast-accuracy claims around a defensible benchmark#88
vahid-ahmadi merged 2 commits into
mainfrom
benchmark-honesty

Conversation

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Follows boe-var-model#6. Three claims on the validation page did not survive scrutiny; two of them were flattering and one was unflattering.

1. The CPI "win" was a benchmark artefact

The page said the SVAR "beats the random walk clearly on CPI at every horizon" and has "useful inflation signal". Against a random walk with drift — the textbook naive for a trending log level — CPI goes 0.63 → 0.83 at h=1 and 0.67 → 1.03 at h=8, i.e. no better than naive.

What survives is Bank Rate: the one non-trending series, which improves under the harder benchmark to 0.79 at h=1 (p=0.018). The page now states that is the defensible forecasting claim, not inflation.

2. The GDP "loss" was overstated in the other direction

1.06–1.09 is not statistically distinguishable from a random walk at any horizon (p = 0.38–0.67), and drops to 0.77 excluding six Covid-target origins. Both published; neither preferred. Over-claiming a loss reads as humility but is the same error as over-claiming a win.

3. Coverage was scored backwards

"All fourteen outturns fall inside the 68% bands" was listed as a headline validation statistic. A calibrated 68% band should contain ~9.5 of 14. Containing all of them means the intervals are ~3× wider than a 0.3pp RMSE warrants — over-dispersion, presented as a pass. And the fourteen points are seven quarters × two variables from a single origin, so coverage isn't establishable either way. The page now makes no coverage claim.

4. The multiplier gap was missing from the evidence page

The OBR emulator's impact multiplier for current spending is ~1.0 by construction against the OBR's own published 0.6 — a two-thirds overstatement against the institution being replicated, applying to every spending-side figure. It was on the model page and absent here. Now stated, along with the three exogenous channels (imports, Bank Rate, CPI) that bias upward and the dead dividends channel that biases down — and that they do not cancel.

predictive_validation for boe-svar downgraded moderate → weak, which is what the evidence now supports. 183 integration tests pass.

🤖 Generated with Claude Code

vahid-ahmadi and others added 2 commits July 24, 2026 19:06
The site claimed the SVAR "beats the random walk clearly on CPI at every
horizon" and had "useful inflation signal". A driftless random walk on a
trending log price level forfeits the whole trend as forecast error, so
that comparison is close to uninformative. Against a random walk with
drift the CPI ratio goes from 0.63 to 0.83 at h=1 and from 0.67 to 1.03 at
h=8 -- no better than naive.

The tell was in the original numbers: the model won only on the two
trending price series and tied or lost on the four series where a random
walk is genuinely hard to beat. That is the signature of a benchmark
artefact, not of inflation signal.

What survives is Bank Rate -- the one non-trending series, where no-change
is the right naive -- which improves under the harder benchmark to 0.79 at
h=1 (p=0.018). The page now says that is the defensible forecasting claim.

Also corrects the GDP framing in the other direction: 1.06-1.09 is not
significant at any horizon, and drops to 0.77 excluding six Covid-target
origins. Both figures published, neither preferred.

The AR(1) comparison is shown as unusable rather than as a 3:1 win, since
94.5% of its eight-step squared error for UK GDP comes from one origin.

Predictive validation for boe-svar downgraded moderate -> weak, which is
what the evidence now supports.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two items the persona reviews flagged independently.

Coverage: "all fourteen outturns fall inside the 68% bands" was listed
among the headline validation statistics. A calibrated 68% interval should
contain about 9.5 of 14; containing all of them means the bands are ~3x
wider than a 0.3pp RMSE warrants. The model is under-confident and the
figure was being scored as a win. It now says so -- and adds that the
fourteen points are seven quarters x two variables from a single origin,
so coverage cannot be established from this sample in either direction.
The page now makes no coverage claim.

Multiplier: the OBR emulator's impact multiplier for current spending is
~1.0 by construction against the OBR's own published 0.6 -- a two-thirds
overstatement against the institution being replicated, applying to every
spending-side figure on the page. It was disclosed on the model page and
absent from the evidence page. Also names the three exogenous channels
(imports, Bank Rate, CPI) that bias the estimate upward and the dead
dividends channel that biases the other way, and states that they do not
cancel.

RMSE rounded 0.32pp -> 0.3pp: two significant figures from seven
overlapping observations implied precision that is not there.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Jul 24, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
macromod Building Building Preview Jul 24, 2026 6:08pm

Request Review

@vahid-ahmadi
vahid-ahmadi merged commit 1526cee into main Jul 24, 2026
3 checks passed
@vahid-ahmadi
vahid-ahmadi deleted the benchmark-honesty branch July 24, 2026 18:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant