[trace-mcp]

Preregistration — tool response token cost

This file is retrospective. The measurement ran on 2026-09-05 and this was written the same day, after the numbers were known. It was not preregistered. The bar below binds the next run; it is not evidence about this one.

The reason this measurement exists at all is a failure of exactly the kind preregistration catches. For months the published figure was ~40–50% on average, which descended from a counter that scored every call before the tool ranRAW_COST_ESTIMATES[tool] × 0.15, a constant with zero variance across thousands of calls (TRA-880, #915). Arithmetic presented as measurement, caught by a person reading the code rather than by any process.

Question

Across the call mix a real session actually produces, do trace-mcp tool responses cost fewer tokens than the file reads they replace — and which tools cost more?

Metric

reduction_pct = (Σ calls × baseline_per_call − Σ calls × measured_per_call) / Σ calls × baseline_per_call

Net: tools that cost more than their baseline subtract from the total. The second ratio, credited_reduction_pct, floors each tool’s loss at zero — that is what the corrected in-product counter books, it is higher, and it is never the headline.

measured_per_call is a real o200k_base count of a real tools/call response over stdio, emitted by scripts/bench-response-tokens.ts into response-tokens.json. baseline_per_call is RAW_COST_ESTIMATES imported from src/savings.ts. The join is scripts/gen-response-tokens-data.tsdocs/_data/response_tokens.json, and that file regenerates byte-identically from its inputs or CI fails.

Corpus

Two frozen inputs, both committed:

One machine’s mix, and the surfaces that quote the figure have to say so.

Pass bar

Unadjustable after seeing data. A future run at 22% publishes as MISSED at 22%.

Prediction

We expected the corrected figure to land well below the 40–50% it replaced, and we expected a minority of tools to cost more than the reads they stand in for — the old counter credited those a saving too, so the correction had to move in this direction. We did not predict the size of either.

Control — absent, and that is the finding

There is no measured control arm. The baseline half — what a Read/Grep would have cost instead — is a hand-written table in src/savings.ts, not a measured alternative run. So a miss on this metric cannot be told apart from a mis-calibrated baseline, and neither can a beat.

That limit is why this figure does not lead the storefront: the PR review context benchmark has a real control arm and runs on code we do not own. Every surface quoting the aggregate has to say the baseline is still an estimate, and tests/docs/savings-claims.test.ts fails when one stops.

Building a measured control is the outstanding work on this measurement.

Verdict — MISSED again, and on the other half (TRA-945, 2026-09-05)

The first run of this measurement (TRA-880) missed on coverage: twelve tools carrying 88.4% of recorded calls against a declared 90%. The stated fix was “measuring the tail, not lowering the line”. The tail is measured — twenty-four tools, 97.2% of recorded call volume — and the bar is missed again, on the other half:

21.1% net reduction (32.6% credited) against a declared 25%. The coverage half now passes; the primary half does not.

That is the result, and it is the opposite of what closing a coverage gap was expected to do. The prediction was that the unmeasured 11.6% would move the figure a little in an unknown direction. It moved it 8.2 points down, because the tail held the two most expensive things in the product:

The bar is not moved. It said “unadjustable after seeing data. A future run at 22% publishes as MISSED at 22%”, and this run publishes as MISSED at 21.1%. The figure stays on the storefront with the miss stated next to it, because it is not wrong — it is smaller than we hoped and better supported than what it replaces.

A third number is published alongside for the first time: 19.5%, all-in, with the 1473 no-baseline calls counted on the spend side and nothing on the baseline side. It answers “what does a session cost” where reduction_pct answers “what does a lookup cost”. It is the lowest of the three and the right one to plan a budget against.

The fix this time is not more coverage. It is response shaping on the ten tools that cost more than they replace — one issue each, with the per-tool table as the before number.

Re-measured after the first three were shaped (TRA-952, 2026-09-05)

The three worst ratios — list_projects 10.5x, get_dead_code 4.0x, get_call_graph 3.5x — were shaped, and the whole table was re-measured on the same protocol at the same declared bar. The current data on this page is that run:

21% net reduction (31.5% credited, 19.4% all-in), still MISSED against the declared 25%.

The three tools dropped 65–83% each and 8 of the 22 tools with a baseline still return more tokens than the figure credits them. The headline moved 0.1 points, because those three carry 82 of 18319 recorded calls: worst-ratio-first finds defects, volume-first moves the number, and only the second was ever going to clear a bar. search (4,441 calls, 1.54x) is where that starts.

Measured at trace-mcp 3.18.0 (94dbaf70) on 5 September 2026.