3System under test
Two runners exercise the same server. The benchmark runner posts every item to the model server's POST /v1/systemone, the endpoint the MCP tool forwards to, so it measures the model and server end to end. The agent eval runner drives a real coding agent through opencode run --pure, so it also measures whether the agent calls ember at the right moments, asks well, and acts on the answer.
advise call. The dashed routes are the two runners: the agent eval drives a coding agent, and the benchmark posts straight to the model server.Text version
agent eval runner (scripts/run_agent_evals.py)
│ opencode run --pure, one sandbox per session
▼
coding agent ──tools/call advise──► ember-mcp (stdio, mcp_server.py)
│ HTTP POST /v1/systemone
▼
model benchmark runner ──HTTP──► model server (server.py)
(scripts/run_evals.py) │ Engine.advise
▼
Clef-Flash, fp16 on MPS| Fact | Value |
|---|---|
| Model | Cloudflare/clef-flash, 9B parameters |
| Pinned revision | 17f0b0ad64efb65d273590632833508766b2aae6 |
| Device and precision | mps, float16 |
| Context window | 262,144 tokens |
| Host | Apple M4 Max, macOS-26.6.2-arm64-arm-64bit |
| Software | Python 3.12.11, gut 0.1.0, torch 2.14.1, transformers 5.18.0, mcp 2.3.0 |
| Server | http://127.0.0.1:8765 |
| Commit | cadc1ec |
| Run at | 2026-10-03 20:38 UTC |
| Report generated | 2026-10-03 22:55 UTC |
| Dataset | clef-flash.jsonl |
| Dataset SHA-256 | 6a457f6badb7dd53b50a9ce04857d4d01e885d4718f3eb6db9db58658d9f923c |
| Dataset unchanged since the run | yes |
| Request errors | 0 |