2Context
ember runs Cloudflare's Clef-Flash decision model [1] locally on Apple Silicon and gives coding agents one Model Context Protocol tool, advise [2]. An agent sends a state (the evidence) and a set of typed questions, and ember returns a probability for every option of every question. It generates no text: it advises, and the agent decides.
Clef is a decision model, not a chat model. A Qwen3.5 backbone with a joint schema head scores the options of every question in one forward pass, so the questions in one call are weighed jointly. ember serves it in fp16 on the Apple GPU (MPS) behind a local HTTP server.
Agents act on ember's answers through thresholds published in ember's agent kit: trust a choice at confidence 0.85 or above, read a noul as yes at P >= 0.80 and no at P <= 0.20, and read a score through its expected value. A threshold is only safe when confidence tracks accuracy, so this benchmark measures calibration and the outcomes of those decision rules, not only accuracy.
| Type | Asks | Returns | Counts as correct when |
|---|---|---|---|
noul | A yes/no proposition | P(true) | P >= 0.5 matches the gold label |
choice | Mutually exclusive, named options | The top option, its confidence, and the full distribution | The top option is the gold option |
score | Ordered levels | A distribution over levels and its expected value | The expected value rounds to the gold level |