Sample size. 456 questions give a 95% interval about 7 points wide overall; per-question samples of 20 to 28 are wider still, and calibration bins hold a few items each.
One annotation team. The labels come from the authors of the agent kit, with no inter-annotator agreement yet. Shared assumptions could make items easier or harder for ember than real traffic.
Curated, not sampled. Items are short, English, and written for the benchmark rather than drawn from real agent sessions.
Ordinal judgments. Risk and effort levels are judgments; neighbouring levels are often defensible, which is why MAE, RPS, and within-one are reported alongside exact accuracy.
One model, one machine. One pinned model revision in fp16 on MPS; numerics on CPU or CUDA can differ slightly. Repeat runs on the same server are deterministic for identical requests.
Fixed question sets. Questions are weighed jointly, so results hold for these question sets. The routing and effort option sets are templates; a project that replaces them should measure its own.
Model and server only. The benchmark calls the HTTP endpoint directly. It does not measure an agent's decision to consult ember or the quality of the questions it writes.