A local gut feeling for coding agents. ember runs a decision model (currently
Cloudflare’s Clef-Flash) on your Apple
Silicon Mac and gives agents one MCP tool, advise: describe a situation, ask typed
questions, and get back a calibrated feeling about every option. It’s a little buddy for
judgment calls — it advises; the agent decides.
How it works
A decision model is not a chat model: it takes a state plus a schema of typed
questions and returns one probability per option, with no text generation. So ember
plugs into agents as a tool, while their reasoning stays on their normal LLM:
- The model server (
ember/serving/server.py) loads the model once and stays warm across agent sessions. - The MCP server (
ember/mcp/mcp_server.py) never loads the model; it starts the model server on the first tool call, so the MCP handshake stays instant.
Names
| Thing | Name |
|---|---|
| Product, Python import, repository | ember |
Distribution (uv tool install, PyPI) |
gut, because ember is taken on PyPI |
| CLI | ember (short alias: gut) |
| MCP server / command | ember / ember-mcp |
| Tool | advise — opencode: ember_advise; Claude Code: mcp__ember__advise |
| Playbook skill / MCP resource | ember-advise / ember://guide |
| Environment variables | EMBER_* |
“Clef” always refers to Cloudflare’s upstream model, never to this product.
Verified
On a MacBook Pro M4 Max / 128 GB, torch 2.14.1, transformers 5.18.0, mcp 2.3:
| Step | Result |
|---|---|
| Model load | ~5 s |
| Warm request | ~0.9–1.3 s for ~220–360 input tokens |
| opencode end to end | ✅ the agent reads the instructions, lists the skill, and calls the tool unprompted |
Install (Apple Silicon)
Requirements
- Apple Silicon Mac (M-series) on macOS. Intel Macs and NVIDIA/CUDA are out of scope.
- Unified memory above the model’s size: 32 GB or more for
flash(9B), 64 GB or more forfull(27B). Only 128 GB has been verified. - Disk: about 18 GiB for
flashor 55 GiB forfull, in Hugging Face’s shared cache (~/.cache/huggingface). Config, state, and logs live in~/Library/Application Support/ember. - Python 3.12, managed by uv.
uv tool install --python 3.12 "gut @ git+https://github.com/shapeandshare/ember"
ember model pull # ~18 GB, resumable, disk-space checked
ember doctor # platform, dependencies, model, and server
ember init --opencode # register with opencode: config, plugin, and skill
Keep --python 3.12: uv otherwise picks your newest interpreter, which the pinned
torch/transformers stack is not tested on.
The repository is public, so the install above needs no credentials. To use SSH instead,
install from git+ssh://git@github.com/shapeandshare/ember.
Restart opencode and the agent gains ember_advise. The model server stays lazy —
it starts on the first tool call (or with ember start).
Everything is pinned for reproducibility: flash to the commit verified on MPS (17f0b0a),
full to its release commit (2f3de3d, not yet verified locally), and torch/torchvision to
the tested minor series. Set EMBER_MODEL_DIR to run another weights directory.
Cloud
ember is Apple-Silicon-first. In the cloud, run it on an Apple Silicon host with the memory
above, or on any host with the CPU fallback (EMBER_DEVICE=cpu — float32, roughly twice the
memory, much slower). NVIDIA/CUDA is out of scope. The HTTP server binds to loopback with no
authentication, so exposing it beyond the host needs your own access controls. See
COMPATIBILITY.md and SECURITY.md.
Agent onboarding
Installing the tool is half the job; the other half is making agents want to consult it
at the right moments and read its answers sensibly. ember/agent_kit/ ships that
guidance through every channel each agent actually reads:
| Channel | opencode | Claude Code | Codex CLI | How you get it |
|---|---|---|---|---|
| MCP server instructions (when to consult, how to ask, how to read answers) | ✅ in the system prompt | ✅ (2 KB cap) | — | built in, nothing to do |
ember://guide resource (full playbook) |
✅ via read_mcp_resource |
✅ | — | built in |
ember-advise skill (playbook, loaded on demand) |
✅ | ✅ | ✅ | ember agents install --agent <agent> |
| AGENTS.md / CLAUDE.md policy block | ✅ | ✅ (CLAUDE.md) | ✅ | ember agents show snippet >> AGENTS.md |
ember init --opencode # opencode: config entry, plugin, and skill
ember agents install --agent claude # .claude/skills/ember-advise/SKILL.md
ember agents install --agent codex # .agents/skills/... (opencode reads this too)
ember agents show snippet >> AGENTS.md # then edit the project policy at the end
The skill is a playbook, not a reference card: the decision points worth consulting ember about, copy-paste question sets for intent and readiness, failure triage, change risk, routing, and effort, and starting confidence thresholds calibrated from observed model output (re-measured whenever the pinned model revision changes). The snippet ends with a project policy — edit it to wire ember into your own workflow, e.g. “check change risk before every push”.
Agent-specific notes:
- opencode reads skills from
.opencode/skills,.claude/skills, and.agents/skills, so install one copy per project to avoid duplicate listings. - Claude Code: register the server with
claude mcp add ember -- ember-mcp; the tool appears asmcp__ember__advise. - Codex CLI support for MCP server instructions and resources is unconfirmed, so rely on the skill and the AGENTS.md snippet there.
CLI
gut is a short alias for every command below.
# lifecycle
ember doctor # platform, dependencies, model, and server
ember serve # run the model server in the foreground
ember start | stop | restart | status | logs
ember config path | show
ember uninstall [--purge-models] # also removes global opencode/skill installs
# models
ember model pull [flash|full] # flash = 9B (default), full = 27B
ember model list | path [name] | rm [name]
# opencode and agents
ember init [--opencode] [--global]
ember agents install [--agent opencode|claude|codex] [--global]
ember agents show instructions|skill|snippet
ember mcp # the MCP stdio server agents launch
Usage
Ask the agent in natural language — “Is this bug report urgent, and which team should own
it?” — and it calls ember_advise with a state and typed questions:
{
"model": "clef-flash",
"answers": {
"urgent": { "type": "noul", "noul": 0.967 },
"team": {
"type": "choice",
"choice": "db",
"confidence": 0.9782,
"probabilities": { "db": 0.9782, "frontend": 0.0218 }
}
},
"usage": { "input_tokens": 220, "output_tokens": 0 },
"latency_ms": 930.9
}
Question types: noul (yes/no → P(true)), choice (named options), score (ordered options →
expected score plus legend). The model field echoes the upstream model label.
When the evidence is visual, attach images or video frames inline: images is a list of
data:image/png;base64,... URIs (or {"content_type": "image/png", "base64": "..."}
objects), and videos is a list of videos, each a list of frame refs. Remote URLs and local
paths are rejected — the model server never reads host files or fetches URLs for an agent.
Configuration
Settings resolve as CLI flag > environment variable > config file > default. The config file
is JSON at ember config path (keys model, host, port, device, max_length; a
max_length of 0 means the model’s own maximum).
| Variable | Default | Purpose |
|---|---|---|
EMBER_HOST / EMBER_PORT |
127.0.0.1 / 8765 |
Model server address |
EMBER_DEVICE |
auto |
auto, mps, or cpu |
EMBER_MODEL |
flash |
flash (9B) or full (27B) |
EMBER_MODEL_DIR |
— | Run weights from this directory instead of the pinned cache |
EMBER_MAX_LENGTH |
0 (the model’s maximum: 262144) |
Token cap per request; 0 derives it from the model |
EMBER_MAX_REQUEST_LENGTH |
32768 |
Per-request token cap enforced before inference; 0 disables it (uses the model maximum) |
EMBER_SERVER_URL |
http://127.0.0.1:8765 |
Where the MCP server sends requests |
EMBER_AUTOSTART |
1 |
Let the MCP server start the model server on demand |
EMBER_START_TIMEOUT |
300 |
Seconds to wait for the model server to start |
EMBER_STATE_DIR |
Application Support | Where the pid file and logs live |
Metrics
While the model server is running it exposes Prometheus metrics at
http://127.0.0.1:8765/metrics (the EMBER_HOST/EMBER_PORT address):
| Metric | Type | Meaning |
|---|---|---|
ember_advise_requests_total{status} |
counter | advise requests by HTTP status (200, 422, 503, 500) |
ember_advise_latency_seconds |
histogram | advise request latency by status |
ember_advise_input_tokens_total |
counter | input tokens processed |
ember_advise_output_tokens_total |
counter | output tokens produced |
ember_model_info{model,device,dtype} |
gauge | 1 while a model is loaded |
The endpoint binds to the same address as the rest of the API (loopback by default), so
it is reachable only there unless you change EMBER_HOST. Ember’s metrics live in a
dedicated Prometheus registry, so the standard python_*/process_* collectors are not
included — /metrics shows only the table above.
Development
ember/
cli.py # the `ember` command (alias `gut`)
models.py # pinned model registry + pull/list/rm
serving/ # HTTP model server, MPS runtime, lifecycle, media
server.py # FastAPI: POST /v1/systemone, GET /health (reports pid)
runtime.py # MPS-safe loader (CPU load → .to("mps")) + Engine
process.py # model-server lifecycle (pid file + HTTP health)
media.py # base64 data-URI / {content_type,base64} → PIL decoder
mcp/ # MCP stdio server and wire types
mcp_server.py # MCP stdio server: advise tool, instructions, ember://guide
mcp_types.py # Pydantic wire types (Question, AdviseInput)
cfg/ # configuration and platform paths
config.py # config resolution (CLI flag > env > file > default)
paths.py # platform-aware app dirs (macOS Library, XDG)
opencode/ # opencode integration
opencode_config.py # generates opencode.json
opencode_plugin.py # installs the npm plugin
agent_kit/ # what agents read: instructions, ember-advise skill, AGENTS snippet
api.py # public API: instructions(), skill(), snippet(), install_skill()
packages/opencode-plugin/ # npm-ready opencode plugin source
scripts/ # make-only dev tools: MPS smoke, MCP e2e, site build, provenance/vault audits
evals/eval/ # benchmark harness: run, report, export, snapshot, agent-in-the-loop
tests/ # pytest suite (unit + model-backed, host-isolated)
vault/ # project memory (Obsidian): decisions, discoveries, session logs
.specify/ # spec-kit; memory/constitution.md governs this repo
AGENTS.md CLAUDE.md # guidelines for agents working on this repo
Makefile # contributor lifecycle (wraps the CLI)
make bootstrap # deps + pinned weights + opencode.json + doctor (idempotent)
make start # model server in the background; stop | restart | status | logs
make opencode # this checkout's opencode plugin and skill
opencode.json, .opencode/plugins/ember.js, and .opencode/skills/ember-advise/
embed this clone’s absolute paths or copy packaged files, so they are gitignored — regenerate
them with make init / make opencode after cloning. .opencode/opencode.json is shared and
committed: it registers the vault MCP server that agents use to read and write vault/
(launch opencode from the repository root).
Make targets
| Target | What it does |
|---|---|
make help |
List all targets (default) |
make bootstrap |
From a fresh clone: setup + init + doctor |
make setup / sync / download |
Deps + weights / deps only / pinned weights to .models/ |
make init / make opencode |
ember init / ember init --opencode for this checkout |
make serve / start / stop / restart / status / logs |
Model-server lifecycle via the CLI |
make mcp / make mcp-list |
Run the MCP server / opencode mcp list |
make test / test-fast / test-strict |
Full suite / unit tests only / full suite that fails without weights |
make test-evals |
Calibration eval suite: positive + negative recipe cases (loads model) |
make eval-run |
Run the benchmark dataset against the live server; writes results/ |
make eval-snapshot |
Copy the latest run into benchmark/ for the site to render |
make eval-report |
Render the most recent run as a Markdown table |
make mcp-check / make smoke |
MCP end-to-end check / direct MPS inference |
make compile / make check |
Byte-compile / compile + unit tests |
make ci |
bootstrap + check + test-strict |
make doctor |
ember doctor |
make vault-audit |
Check vault/ notes: frontmatter, tags, wikilinks, code-refs, orphans |
make site / make site-serve |
Build the Pages site into site/_site / preview it at :4000 (needs Docker) |
make release-dry |
Preview the next ember version bump without changes |
make release-ember / make release-plugin |
Trigger a component release workflow on main (bump PR → GitHub Release) |
make clean / make clean-model |
Caches and build output / weights (EMBER_FORCE=1 skips the prompt) |
Make re-syncs the environment automatically when pyproject.toml or uv.lock changes.
Tests
make test # full suite (~30 s, loads the model once)
make test-fast # unit tests only, no model load (~12 s)
make check # compile + test-fast
make test-strict # full suite; fails (not skips) if weights are missing
make ci # bootstrap + check + test-strict
tests/ covers the supported call paths:
GET /healthandPOST /v1/systemoneacrossnoul/choice/score, probability normalization, and request-validation422s- MCP tool discovery,
adviseover stdio, actionable errors (server down, model missing, malformed questions), and autostart on a configured port with pid cleanup - the agent kit: instructions size, skill frontmatter, the
ember://guideresource - the CLI:
agents,init --opencode,status/stopon unused ports,doctor,uninstall - process safety:
stopnever signals a pid that is not an ember server it started
Safe on a host running other opencode instances. The suite binds random free ports (never
8765), keeps state in temporary directories, stops only the processes it started, and never
invokes the opencode CLI or touches global opencode config.
CI (.github/workflows/ci.yml) runs uv sync --locked, make check, uv build, and an
install smoke of the built wheel on a hosted Apple Silicon runner, and audits the workflows
with zizmor. Model-backed tests are
not run in CI: the ~19 GB fp16 model does not fit the available runners (hosted or the
org’s 8 GiB self-hosted VMs), so run make test locally for model-affecting changes.
Coverage is tracked and gated (constitution Article XI). Run make test-cov to see the
full report. The enforced floor (fail_under in pyproject.toml) is the current measured
level and may only increase — currently 71 %. Lowering it requires explicit, recorded
approval per Article XI §11.2.
Benchmark
evals/clef-flash.jsonl is a 264-item benchmark covering the five agent-kit recipes (intent
and readiness, failure triage, change risk, routing, effort and approach) plus four vision
recipes (vision_noul, vision_choice, vision_score, vision_video): 456 scored questions
(184 choice, 152 noul, 88 score, 32 across the four vision recipes), split into
dev (130 items) and test (134). Vision items carry an images or videos field
(base64 data: URIs) alongside state and questions. Each line is
self-contained (state, the recipe’s fixed question set, a gold label for every question, and
a rationale), so other implementations can score it without this harness.
make eval-run # or: ember eval run [--split dev|test] [--category <recipe>]
make eval-report # Markdown report of the most recent run
make eval-export # or: ember eval export; the reviewer bundle described below
ember eval report --compare results/<a>_results.json results/<b>_results.json
ember eval export writes results/<run_id>_report/ for external reviewers:
report.html (one self-contained file with inline charts, light and dark themes, a print
layout, and item filters; it loads nothing from the network), report.md with its
figures/*.svg, and data/ with the run’s results, trace, and the exact dataset scored.
The report covers context, the system under test, benchmark design, metric definitions with
references, results, calibration, the agent kit’s decision rules replayed on every item, a
review card for every miss, limitations, and reproduction pins.
The report gives accuracy with a 95% bootstrap interval, macro-F1, top-label ECE, and Brier
score (choice, noul); ranked probability score and MAE (score); and coverage and accuracy
at the agent kit’s act-on-it thresholds. Each run records the git hash, the engine, and the
dataset’s SHA-256, and --compare warns when two runs scored different item sets.
Gold labels are judged from state alone against the recipe’s option descriptions, and they
are fixed before any run: needs_review is true exactly when risk is Medium or High, and
retry only for flaky failures. A person reviews disagreements; labels are never changed to
match the model. make check validates the dataset (tests/test_eval_benchmark.py).
ember eval reads evals/ and scripts/ from the checkout, so it works only in a clone.
Agent in the loop
The benchmark above scores the model. ember eval agent scores ember the way it is used:
through a coding agent. It gives opencode 54 scripted requests (vague and precise asks,
failing tests, risky and trivial commits, routing and effort questions, and controls), each
in a throwaway git repo, and judges the agent only by what it did (files, commits, tests,
its reply). Each scenario runs under four conditions: none (no ember), mcp (the MCP
server and its instructions), skill (plus the ember-advise skill), and full (plus the
AGENTS.md policy).
ember start # the agent sessions call the warm server
make eval-agent-smoke # 6 scenarios x 4 conditions x 2 models, 1 trial
make eval-agent # or: ember eval agent [--models ...] [--trials 3] [--parallel 8]
ember eval export --agent latest # put the agent results first in the reviewer report
Every session runs opencode run --pure with a private HOME and XDG directories and no TCP
port, so your opencode config, plugins, sessions, and running instances are untouched. The
provider key is read from the environment or opencode’s auth file and passed only to the child
process. The run reports, per model and condition: gold-action accuracy, its paired change
against none, how often the agent consulted ember at decision points (and on controls),
whether it consulted before acting, asked the kit’s recipe questions, passed the evidence, and
followed ember’s answer, plus cost and time per session. It spends real provider credit and is
never part of make check, make test, or CI.
Caveats
- Gated DeltaNet fallback. Qwen3.5’s backbone is a hybrid linear-attention model. The
optimized CUDA kernels (
causal_conv1d,flash-linear-attention) don’t exist for Apple Silicon, so it uses the pure-PyTorch reference path: correct but slower. You’ll see two “falling back” log lines — expected. - fp16 on MPS.
bfloat16works but is emulated and less battle-tested on MPS, so the loader usesfloat16; CPU mode usesfloat32. - Vision works on MPS. Images and video frames run through the same fp16 path as text.
Inline them as base64
data:URIs inimages/videos; remote URLs and local paths are rejected on purpose, so an agent cannot make the server read host files or fetch URLs. Text-only inputs skip the vision tower entirely. - Numerics. MPS can differ slightly from CUDA/CPU. If calibrated probabilities matter,
cross-check with
EMBER_DEVICE=cpu ember restart. device_map={"": "mps"}segfaults with the pinned stack. The loader loads on CPU and then moves the model to MPS — don’t “simplify” that away.
Troubleshooting
- Start with
ember doctor, thenember logs(~/Library/Application Support/ember/logs/server.log). - The agent can’t see the tool:
opencode mcp listshould show✓ ember connected. - “model … is not pulled”: run
ember model pull, or pointEMBER_MODEL_DIRat the weights. Qwen3VLVideoProcessor requires Torchvision→ torchvision is a pinned dependency; runuv sync(or reinstall the tool).- Any single unimplemented MPS op falls back to CPU (
PYTORCH_ENABLE_MPS_FALLBACK=1, set by the runtime).
License
MIT (see LICENSE). Cloudflare’s Clef weights and joint_schema_model.py are Apache-2.0; they
are downloaded from Hugging Face at runtime, not redistributed here.
Brand assets and palette · Provenance and licensing · Project site