ember

2Context

ember runs Cloudflare's Clef-Flash decision model [1] locally on Apple Silicon and gives coding agents one Model Context Protocol tool, advise [2]. An agent sends a state (the evidence) and a set of typed questions, and ember returns a probability for every option of every question. It generates no text: it advises, and the agent decides.

Clef is a decision model, not a chat model. A Qwen3.5 backbone with a joint schema head scores the options of every question in one forward pass, so the questions in one call are weighed jointly. ember serves it in fp16 on the Apple GPU (MPS) behind a local HTTP server.

Agents act on ember's answers through thresholds published in ember's agent kit: trust a choice at confidence 0.85 or above, read a noul as yes at P >= 0.80 and no at P <= 0.20, and read a score through its expected value. A threshold is only safe when confidence tracks accuracy, so this benchmark measures calibration and the outcomes of those decision rules, not only accuracy.

TypeAsksReturnsCounts as correct when
noulA yes/no propositionP(true)P >= 0.5 matches the gold label
choiceMutually exclusive, named optionsThe top option, its confidence, and the full distributionThe top option is the gold option
scoreOrdered levelsA distribution over levels and its expected valueThe expected value rounds to the gold level