ember

13Limitations

  • Sample size. 456 questions give a 95% interval about 7 points wide overall; per-question samples of 20 to 28 are wider still, and calibration bins hold a few items each.
  • One annotation team. The labels come from the authors of the agent kit, with no inter-annotator agreement yet. Shared assumptions could make items easier or harder for ember than real traffic.
  • Curated, not sampled. Items are short, English, and written for the benchmark rather than drawn from real agent sessions.
  • Ordinal judgments. Risk and effort levels are judgments; neighbouring levels are often defensible, which is why MAE, RPS, and within-one are reported alongside exact accuracy.
  • One model, one machine. One pinned model revision in fp16 on MPS; numerics on CPU or CUDA can differ slightly. Repeat runs on the same server are deterministic for identical requests.
  • Fixed question sets. Questions are weighed jointly, so results hold for these question sets. The routing and effort option sets are templates; a project that replaces them should measure its own.
  • Model and server only. The benchmark calls the HTTP endpoint directly. It does not measure an agent's decision to consult ember or the quality of the questions it writes.