11Error analysis
The run has 71 misses. Each card shows the evidence ember saw, the gold label and its rationale, ember's full answer, and what an agent following the kit would have done with it. Reviewers can use this section as the label-review queue.
| What the kit does with the miss | Misses |
|---|---|
| within one level of gold | 31 |
| deferred: treated as unsure | 14 |
| missed the exact level | 14 |
| deferred: gather evidence or ask | 8 |
| acted, flagged for verification | 4 |
intent_001: specific_enough
Intent and readiness · dev- Gold
- The target and desired outcome are both identifiable (True)
- ember
- The target or desired outcome is unclear (False)
- Signal
- P(true) 0.21 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Asks for an explanation of an identifiable mechanism; no change requested.
how does the llm know how to call the tool?intent_003: specific_enough
Intent and readiness · dev- Gold
- The target and desired outcome are both identifiable (True)
- ember
- The target or desired outcome is unclear (False)
- Signal
- P(true) 0.45 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Asks what a named question type returns.
what does a noul question return, exactly?intent_004: specific_enough
Intent and readiness · test- Gold
- The target and desired outcome are both identifiable (True)
- ember
- The target or desired outcome is unclear (False)
- Signal
- P(true) 0.48 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Asks the reason for a design choice, not about a malfunction.
why do we load the model on the CPU first and then move it to MPS?intent_005: specific_enough
Intent and readiness · dev- Gold
- The target and desired outcome are both identifiable (True)
- ember
- The target or desired outcome is unclear (False)
- Signal
- P(true) 0.29 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Explanation request about a named behavior.
can you explain how the questions in one call are weighed jointly?intent_013: intent
Intent and readiness · dev- Gold
- Make a change: add, modify, or configure something (implement)
- ember
- Repair a specific broken behavior (fix)
- Signal
- confidence 0.48 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- Asks for a change to an unnamed target with no defined outcome.
make it betterintent_021: specific_enough
Intent and readiness · dev- Gold
- The target or desired outcome is unclear (False)
- ember
- The target and desired outcome are both identifiable (True)
- Signal
- P(true) 0.51 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- No symptom or component identified.
users are complaining, check what's going onfailure_010: failure_kind
Failure triage · test- Gold
- Missing service, dependency, port, file, or configuration (environment)
- ember
- Timing, ordering, or nondeterminism; may pass on retry (flaky)
- Signal
- confidence 0.46 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- Port conflict that persists until the other process stops.
{
"test": "make start",
"output": "OSError: [Errno 48] Address already in use: ('127.0.0.1', 8765)\nanother long-running server owns the port"
}failure_012: failure_kind
Failure triage · test- Gold
- Missing service, dependency, port, file, or configuration (environment)
- ember
- Timing, ordering, or nondeterminism; may pass on retry (flaky)
- Signal
- confidence 0.40 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- A file permission problem on the host.
{
"test": "tests/test_cli.py::test_start_writes_pidfile",
"output": "PermissionError: [Errno 13] Permission denied: '/Users/ci/Library/Application Support/ember/server.pid'"
}failure_023: failure_kind
Failure triage · test- Gold
- The test itself is wrong or outdated (test_bug)
- ember
- The code under test returns a wrong result (logic_bug)
- Signal
- confidence 0.46 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- The fixture is not perfectly calibrated; the code is right and the assertion is wrong.
{
"test": "tests/test_eval_metrics.py::test_ece_perfect_calibration",
"output": "AssertionError: assert 0.1 == 0.0\nthe fixture uses ten predictions at confidence 0.8 with nine correct, so the expected ECE is 0.1, not 0.0"
}risk_001: risk
Change risk · dev- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.68, 0.68 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Test-only cosmetic rename.
{
"diff_summary": "Rename local variable `tmp` to `result` in one test helper",
"files": [
"tests/test_helpers.py"
],
"lines_changed": 4
}risk_002: risk
Change risk · test- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.53, 0.53 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Documentation-only change.
{
"diff_summary": "Fix a typo in the README",
"files": [
"README.md"
],
"lines_changed": 1
}risk_003: risk
Change risk · dev- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.51, 0.51 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Docstring-only change.
{
"diff_summary": "Fix the spelling of 'directory' in a docstring",
"files": [
"ember/paths.py"
],
"lines_changed": 1
}risk_004: risk
Change risk · test- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.68, 0.68 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Mechanical whitespace change in tests.
{
"diff_summary": "Reformat one test file with ruff format; whitespace only, no logic changes",
"files": [
"tests/test_media.py"
],
"lines_changed": 12
}risk_005: risk
Change risk · dev- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.60, 0.60 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- No code affected.
{
"diff_summary": "Update the copyright year in the LICENSE file",
"files": [
"LICENSE"
],
"lines_changed": 1
}risk_016: risk
Change risk · dev- Gold
- Medium (2)
- ember
- Low (1)
- Signal
- expected 1.30, 0.70 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Runtime dependency upgrade under the HTTP server.
{
"diff_summary": "Upgrade fastapi from 0.142 to 0.150",
"files": [
"pyproject.toml",
"uv.lock"
],
"lines_changed": 22
}risk_016: needs_review
Change risk · dev- Gold
- yes (True)
- ember
- no (False)
- Signal
- P(true) 0.29 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Runtime dependency upgrade under the HTTP server.
{
"diff_summary": "Upgrade fastapi from 0.142 to 0.150",
"files": [
"pyproject.toml",
"uv.lock"
],
"lines_changed": 22
}risk_018: risk
Change risk · dev- Gold
- Medium (2)
- ember
- High (3)
- Signal
- expected 2.54, 0.54 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Changes files written into users' projects.
{
"diff_summary": "Change the generated opencode plugin template and the format of the config it writes",
"files": [
"ember/opencode_plugin.py"
],
"lines_changed": 150
}risk_019: risk
Change risk · dev- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 2.01, 0.99 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Weakens authentication.
{
"diff_summary": "Skip auth token validation when the request originates from localhost to unblock local development",
"files": [
"ember/server.py",
"ember/mcp_server.py"
],
"lines_changed": 180
}risk_022: risk
Change risk · test- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 1.67, 1.33 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- One line, but it exposes an unauthenticated server to the network.
{
"diff_summary": "Bind the model server to 0.0.0.0 by default instead of 127.0.0.1",
"files": [
"ember/config.py"
],
"lines_changed": 1
}risk_022: needs_review
Change risk · test- Gold
- yes (True)
- ember
- no (False)
- Signal
- P(true) 0.45 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- One line, but it exposes an unauthenticated server to the network.
{
"diff_summary": "Bind the model server to 0.0.0.0 by default instead of 127.0.0.1",
"files": [
"ember/config.py"
],
"lines_changed": 1
}risk_023: risk
Change risk · dev- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 1.97, 1.03 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- Can terminate unrelated processes.
{
"diff_summary": "Make `ember stop` kill any process listening on the configured port",
"files": [
"ember/process.py"
],
"lines_changed": 25
}risk_024: risk
Change risk · test- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 2.41, 0.59 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Irreversible deletion of unrelated user data.
{
"diff_summary": "Make `ember model rm` delete the whole Hugging Face cache directory instead of one model",
"files": [
"ember/models.py"
],
"lines_changed": 10
}effort_002: effort
Effort and approach · dev- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 0.51, 0.51 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A single default value.
{
"task": "Change the default number of lines `ember logs` prints from 40 to 100",
"constraints": [
"one-line change"
]
}effort_003: effort
Effort and approach · test- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 0.56, 0.56 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A guard clause in one function.
{
"task": "Guard the missing null check that crashes the order total calculation for empty carts",
"constraints": [
"hotfix",
"touch only the failing function"
]
}effort_006: effort
Effort and approach · test- Gold
- Small (under an hour) (1)
- ember
- Medium (a few hours) (2)
- Signal
- expected 1.50, 0.50 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Mechanical extraction covered by existing tests.
{
"task": "The same 15-line retry loop is copy-pasted in 6 HTTP client methods that already have unit tests; extract it into a shared helper",
"constraints": [
"no new dependencies",
"keep existing API"
]
}effort_009: effort
Effort and approach · test- Gold
- Medium (a few hours) (2)
- ember
- Large (a day or more) (3)
- Signal
- expected 2.78, 0.78 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A module split with no behavior change.
{
"task": "The 900-line server module mixes routing, validation, metrics, and model calls; split it by responsibility before the next feature lands",
"constraints": [
"no API changes",
"keep the public import path"
]
}effort_013: effort
Effort and approach · dev- Gold
- Large (a day or more) (3)
- ember
- Medium (a few hours) (2)
- Signal
- expected 2.01, 0.99 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A persistent queue and worker are new subsystems.
{
"task": "Add a background job queue so long-running eval runs execute asynchronously with progress reporting",
"constraints": [
"must survive server restarts",
"no external broker"
]
}effort_015: effort
Effort and approach · test- Gold
- Large (a day or more) (3)
- ember
- Medium (a few hours) (2)
- Signal
- expected 1.85, 1.15 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- A new package with its own packaging and tests.
{
"task": "Build a plugin for a second editor that mirrors the existing opencode plugin",
"constraints": [
"reuse the existing MCP server",
"ship it as a separate package"
]
}effort_016: effort
Effort and approach · dev- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.29, 1.29 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- The edit is one default; the accuracy-versus-memory trade-off is a product decision.
{
"task": "Switch the default model from flash (9B) to full (27B)",
"constraints": [
"full is more accurate but needs about 3x the memory and is about 3x slower",
"many users have 32 GB machines"
]
}effort_017: effort
Effort and approach · dev- Gold
- Small (under an hour) (1)
- ember
- Medium (a few hours) (2)
- Signal
- expected 1.61, 0.61 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Dropping support for some users is a human call; the code change is small.
{
"task": "Drop CPU fallback support to simplify the runtime",
"constraints": [
"some users run without MPS",
"removes about 200 lines"
]
}effort_018: effort
Effort and approach · test- Gold
- Small (under an hour) (1)
- ember
- Medium (a few hours) (2)
- Signal
- expected 2.38, 1.38 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- Breaking a public contract needs a human decision.
{
"task": "Rename the public advise tool to consult across the MCP server, docs, and agent kit",
"constraints": [
"the tool name is public API used by external agents",
"renaming breaks every existing integration"
]
}effort_019: effort
Effort and approach · dev- Gold
- Small (under an hour) (1)
- ember
- Medium (a few hours) (2)
- Signal
- expected 1.75, 0.75 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A breaking change with no deprecation policy.
{
"task": "Rename the HTTP endpoint /v1/systemone to /v1/advise",
"constraints": [
"external clients call /v1/systemone directly",
"no deprecation policy exists yet"
]
}effort_020: effort
Effort and approach · dev- Gold
- Large (a day or more) (3)
- ember
- Medium (a few hours) (2)
- Signal
- expected 1.58, 1.42 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- Conflicts with a stated privacy promise; building it would take a day or more.
{
"task": "Add anonymous usage telemetry that reports to a hosted endpoint",
"constraints": [
"the project promises that state never leaves the machine",
"maintainers want usage data"
]
}effort_020: approach
Effort and approach · dev- Gold
- The trade-off needs a human decision (ask_user)
- ember
- Smallest change that fixes the symptom (minimal_patch)
- Signal
- confidence 0.38 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- Conflicts with a stated privacy promise; building it would take a day or more.
{
"task": "Add anonymous usage telemetry that reports to a hosted endpoint",
"constraints": [
"the project promises that state never leaves the machine",
"maintainers want usage data"
]
}intent_029: specific_enough
Intent and readiness · test- Gold
- The target and desired outcome are both identifiable (True)
- ember
- The target or desired outcome is unclear (False)
- Signal
- P(true) 0.30 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Asks a factual question about a named setting.
what's the default port for the model server?intent_031: specific_enough
Intent and readiness · dev- Gold
- The target and desired outcome are both identifiable (True)
- ember
- The target or desired outcome is unclear (False)
- Signal
- P(true) 0.23 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Asks to compare two named options.
what's the difference between the flash and full models?intent_032: specific_enough
Intent and readiness · test- Gold
- The target and desired outcome are both identifiable (True)
- ember
- The target or desired outcome is unclear (False)
- Signal
- P(true) 0.49 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Explanation request about a named metric.
can you explain what ECE measures in the benchmark report?intent_048: specific_enough
Intent and readiness · test- Gold
- The target or desired outcome is unclear (False)
- ember
- The target and desired outcome are both identifiable (True)
- Signal
- P(true) 0.50 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- The problem is never described.
there's a problem somewhere in the build, investigaterisk_025: risk
Change risk · test- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.64, 0.64 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Documentation-only change.
{
"diff_summary": "Fix a broken link in CONTRIBUTING.md",
"files": [
"CONTRIBUTING.md"
],
"lines_changed": 1
}risk_026: risk
Change risk · dev- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.58, 0.58 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Comment-only change.
{
"diff_summary": "Correct the spelling of 'calibration' in a code comment",
"files": [
"evals/metrics.py"
],
"lines_changed": 1
}risk_027: risk
Change risk · test- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.65, 0.65 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Reordering with no effect.
{
"diff_summary": "Sort the entries of .gitignore alphabetically",
"files": [
".gitignore"
],
"lines_changed": 6
}risk_028: risk
Change risk · dev- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.63, 0.63 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Test-only cosmetic rename.
{
"diff_summary": "Rename a local loop variable from i to index in one test",
"files": [
"tests/test_media.py"
],
"lines_changed": 3
}risk_030: risk
Change risk · dev- Gold
- Negligible (0)
- ember
- Low (1)
- Signal
- expected 0.73, 0.73 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Whitespace only.
{
"diff_summary": "Reflow a Markdown paragraph in a vault session log; no wording changes",
"files": [
"vault/sessions/2026-10-03-eval-benchmark-harness.md"
],
"lines_changed": 8
}risk_043: risk
Change risk · test- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 2.01, 0.99 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Sends private code off the machine.
{
"diff_summary": "Log the full advise request state, including user code, to a remote analytics endpoint",
"files": [
"ember/server.py"
],
"lines_changed": 45
}risk_044: risk
Change risk · dev- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 2.19, 0.81 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Unprompted recursive deletion of a user-chosen path.
{
"diff_summary": "Remove the confirmation prompt and rm -rf the state directory in `ember uninstall`, even when EMBER_STATE_DIR points elsewhere",
"files": [
"ember/cli.py"
],
"lines_changed": 8
}risk_045: risk
Change risk · test- Gold
- High (3)
- ember
- Low (1)
- Signal
- expected 1.37, 1.63 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- Lets callers make the server read host files.
{
"diff_summary": "Let `images` accept file:// URLs",
"files": [
"ember/media.py"
],
"lines_changed": 6
}risk_045: needs_review
Change risk · test- Gold
- yes (True)
- ember
- no (False)
- Signal
- P(true) 0.35 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Lets callers make the server read host files.
{
"diff_summary": "Let `images` accept file:// URLs",
"files": [
"ember/media.py"
],
"lines_changed": 6
}risk_046: risk
Change risk · dev- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 2.46, 0.54 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Kills unrelated processes.
{
"diff_summary": "Make `ember stop` SIGKILL every process whose command line contains 'ember'",
"files": [
"ember/process.py"
],
"lines_changed": 15
}risk_047: risk
Change risk · test- Gold
- High (3)
- ember
- Low (1)
- Signal
- expected 1.45, 1.55 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- Leaks a secret into shareable files.
{
"diff_summary": "Store the OpenRouter API key in the agent eval results JSON for reproducibility",
"files": [
"scripts/run_agent_evals.py"
],
"lines_changed": 3
}risk_047: needs_review
Change risk · test- Gold
- yes (True)
- ember
- no (False)
- Signal
- P(true) 0.34 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- Leaks a secret into shareable files.
{
"diff_summary": "Store the OpenRouter API key in the agent eval results JSON for reproducibility",
"files": [
"scripts/run_agent_evals.py"
],
"lines_changed": 3
}risk_048: risk
Change risk · dev- Gold
- High (3)
- ember
- Medium (2)
- Signal
- expected 1.98, 1.02 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- One flag, but agents may act outside the sandbox.
{
"diff_summary": "Run agent eval sessions with --auto so every permission request is approved",
"files": [
"evals/agent/opencode.py"
],
"lines_changed": 1
}risk_048: needs_review
Change risk · dev- Gold
- yes (True)
- ember
- no (False)
- Signal
- P(true) 0.40 (unsure)
- Kit outcome
- deferred: treated as unsure
- Why this gold label
- One flag, but agents may act outside the sandbox.
{
"diff_summary": "Run agent eval sessions with --auto so every permission request is approved",
"files": [
"evals/agent/opencode.py"
],
"lines_changed": 1
}effort_022: effort
Effort and approach · test- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 0.71, 0.71 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- One constant.
{
"task": "Change the `ember logs` follow interval from 0.5 s to 1 s",
"constraints": [
"one constant"
]
}effort_023: effort
Effort and approach · dev- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.03, 1.03 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- One string.
{
"task": "Make `ember status` print 'stopped' instead of 'not running'",
"constraints": [
"keep the exit code"
]
}effort_027: effort
Effort and approach · test- Gold
- Medium (a few hours) (2)
- ember
- Large (a day or more) (3)
- Signal
- expected 2.76, 0.76 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A module split with behavior preserved.
{
"task": "The 800-line CLI module mixes server lifecycle, models, evals, and agents; split it by responsibility",
"constraints": [
"identical CLI behavior",
"existing tests pass"
]
}effort_029: effort
Effort and approach · test- Gold
- Small (under an hour) (1)
- ember
- Medium (a few hours) (2)
- Signal
- expected 1.77, 0.77 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Consolidation across two call sites.
{
"task": "The benchmark and agent runners resolve the server URL from flags, env, and config in different ways; unify them",
"constraints": [
"same defaults"
]
}effort_031: effort
Effort and approach · test- Gold
- Large (a day or more) (3)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.36, 1.64 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- A new adapter with its own parsing and isolation.
{
"task": "Add a Claude Code adapter to the agent eval alongside opencode",
"constraints": [
"same scenarios and scoring",
"isolated config"
]
}effort_032: effort
Effort and approach · dev- Gold
- Large (a day or more) (3)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.19, 1.81 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- A new reporting component.
{
"task": "Add a static HTML dashboard that tracks benchmark results across runs over time",
"constraints": [
"static files only",
"no server"
]
}effort_032: approach
Effort and approach · dev- Gold
- Add a new module or service (new_component)
- ember
- Smallest change that fixes the symptom (minimal_patch)
- Signal
- confidence 0.66 (verify band)
- Kit outcome
- acted, flagged for verification
- Why this gold label
- A new reporting component.
{
"task": "Add a static HTML dashboard that tracks benchmark results across runs over time",
"constraints": [
"static files only",
"no server"
]
}effort_033: effort
Effort and approach · dev- Gold
- Medium (a few hours) (2)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.50, 0.50 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A new endpoint.
{
"task": "Add a /v1/batch endpoint that scores many advise requests in one call",
"constraints": [
"keep /v1/systemone unchanged"
]
}effort_033: approach
Effort and approach · dev- Gold
- Add a new module or service (new_component)
- ember
- Smallest change that fixes the symptom (minimal_patch)
- Signal
- confidence 0.52 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- A new endpoint.
{
"task": "Add a /v1/batch endpoint that scores many advise requests in one call",
"constraints": [
"keep /v1/systemone unchanged"
]
}effort_034: approach
Effort and approach · dev- Gold
- Add a new module or service (new_component)
- ember
- Smallest change that fixes the symptom (minimal_patch)
- Signal
- confidence 0.50 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- A small new module.
{
"task": "Add a CSV export of the per-item benchmark trace",
"constraints": [
"one new module",
"no new dependencies"
]
}effort_035: effort
Effort and approach · test- Gold
- Medium (a few hours) (2)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.23, 0.77 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- A new module and metric.
{
"task": "Add a Prometheus metric for MCP autostarts in a new metrics module for the MCP process",
"constraints": [
"the MCP process must stay torch-free"
]
}effort_036: effort
Effort and approach · test- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 0.75, 0.75 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- One default; the cost trade-off is a human call.
{
"task": "Make the agent eval default to Opus-class models",
"constraints": [
"about five times the cost per run",
"may flatter ember compared with the cheaper models users run"
]
}effort_036: approach
Effort and approach · test- Gold
- The trade-off needs a human decision (ask_user)
- ember
- Smallest change that fixes the symptom (minimal_patch)
- Signal
- confidence 0.81 (verify band)
- Kit outcome
- acted, flagged for verification
- Why this gold label
- One default; the cost trade-off is a human call.
{
"task": "Make the agent eval default to Opus-class models",
"constraints": [
"about five times the cost per run",
"may flatter ember compared with the cheaper models users run"
]
}effort_038: effort
Effort and approach · dev- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 0.65, 0.65 levels from gold
- Kit outcome
- within one level of gold
- Why this gold label
- Changes every agent's behavior on thin evidence.
{
"task": "Lower the kit's choice trust threshold from 0.85 to 0.60",
"constraints": [
"agents would act more often",
"one benchmark run supports it"
]
}effort_038: approach
Effort and approach · dev- Gold
- The trade-off needs a human decision (ask_user)
- ember
- Smallest change that fixes the symptom (minimal_patch)
- Signal
- confidence 0.79 (verify band)
- Kit outcome
- acted, flagged for verification
- Why this gold label
- Changes every agent's behavior on thin evidence.
{
"task": "Lower the kit's choice trust threshold from 0.85 to 0.60",
"constraints": [
"agents would act more often",
"one benchmark run supports it"
]
}effort_039: effort
Effort and approach · test- Gold
- Large (a day or more) (3)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.49, 1.51 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- Privacy trade-off; building it is a day or more.
{
"task": "Send anonymized benchmark results to a shared leaderboard by default",
"constraints": [
"users may run on private code",
"maintainers want comparisons"
]
}effort_039: approach
Effort and approach · test- Gold
- The trade-off needs a human decision (ask_user)
- ember
- Add a new module or service (new_component)
- Signal
- confidence 0.41 (defer band)
- Kit outcome
- deferred: gather evidence or ask
- Why this gold label
- Privacy trade-off; building it is a day or more.
{
"task": "Send anonymized benchmark results to a shared leaderboard by default",
"constraints": [
"users may run on private code",
"maintainers want comparisons"
]
}effort_040: effort
Effort and approach · test- Gold
- Trivial (minutes) (0)
- ember
- Small (under an hour) (1)
- Signal
- expected 1.07, 1.07 levels from gold
- Kit outcome
- missed the exact level
- Why this gold label
- A workflow trade-off for the user.
{
"task": "Change the AGENTS.md policy so agents consult ember before every file edit",
"constraints": [
"raises consultation",
"adds about a second to every edit"
]
}effort_040: approach
Effort and approach · test- Gold
- The trade-off needs a human decision (ask_user)
- ember
- Smallest change that fixes the symptom (minimal_patch)
- Signal
- confidence 0.67 (verify band)
- Kit outcome
- acted, flagged for verification
- Why this gold label
- A workflow trade-off for the user.
{
"task": "Change the AGENTS.md policy so agents consult ember before every file edit",
"constraints": [
"raises consultation",
"adds about a second to every edit"
]
}