ember

11Error analysis

The run has 71 misses. Each card shows the evidence ember saw, the gold label and its rationale, ember's full answer, and what an agent following the kit would have done with it. Reviewers can use this section as the label-review queue.

What the kit does with the missMisses
within one level of gold31
deferred: treated as unsure14
missed the exact level14
deferred: gather evidence or ask8
acted, flagged for verification4

intent_001: specific_enough

Intent and readiness · dev
Gold
The target and desired outcome are both identifiable (True)
ember
The target or desired outcome is unclear (False)
Signal
P(true) 0.21 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Asks for an explanation of an identifiable mechanism; no change requested.
how does the llm know how to call the tool?
ember's answer for intent_001, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.21

intent_003: specific_enough

Intent and readiness · dev
Gold
The target and desired outcome are both identifiable (True)
ember
The target or desired outcome is unclear (False)
Signal
P(true) 0.45 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Asks what a named question type returns.
what does a noul question return, exactly?
ember's answer for intent_003, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.45

intent_004: specific_enough

Intent and readiness · test
Gold
The target and desired outcome are both identifiable (True)
ember
The target or desired outcome is unclear (False)
Signal
P(true) 0.48 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Asks the reason for a design choice, not about a malfunction.
why do we load the model on the CPU first and then move it to MPS?
ember's answer for intent_004, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.48

intent_005: specific_enough

Intent and readiness · dev
Gold
The target and desired outcome are both identifiable (True)
ember
The target or desired outcome is unclear (False)
Signal
P(true) 0.29 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Explanation request about a named behavior.
can you explain how the questions in one call are weighed jointly?
ember's answer for intent_005, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.29

intent_013: intent

Intent and readiness · dev
Gold
Make a change: add, modify, or configure something (implement)
ember
Repair a specific broken behavior (fix)
Signal
confidence 0.48 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
Asks for a change to an unnamed target with no defined outcome.
make it better
ember's answer for intent_013, intentProbability ember gave each option.question0.08implement0.38 goldinvestigate0.07fix0.48 ember's pick

intent_021: specific_enough

Intent and readiness · dev
Gold
The target or desired outcome is unclear (False)
ember
The target and desired outcome are both identifiable (True)
Signal
P(true) 0.51 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
No symptom or component identified.
users are complaining, check what's going on
ember's answer for intent_021, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: noP = 0.51

failure_010: failure_kind

Failure triage · test
Gold
Missing service, dependency, port, file, or configuration (environment)
ember
Timing, ordering, or nondeterminism; may pass on retry (flaky)
Signal
confidence 0.46 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
Port conflict that persists until the other process stops.
{
  "test": "make start",
  "output": "OSError: [Errno 48] Address already in use: ('127.0.0.1', 8765)\nanother long-running server owns the port"
}
ember's answer for failure_010, failure_kindProbability ember gave each option.logic_bug0.02environment0.38 goldflaky0.46 ember's picktest_bug0.15

failure_012: failure_kind

Failure triage · test
Gold
Missing service, dependency, port, file, or configuration (environment)
ember
Timing, ordering, or nondeterminism; may pass on retry (flaky)
Signal
confidence 0.40 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
A file permission problem on the host.
{
  "test": "tests/test_cli.py::test_start_writes_pidfile",
  "output": "PermissionError: [Errno 13] Permission denied: '/Users/ci/Library/Application Support/ember/server.pid'"
}
ember's answer for failure_012, failure_kindProbability ember gave each option.logic_bug0.05environment0.33 goldflaky0.40 ember's picktest_bug0.23

failure_023: failure_kind

Failure triage · test
Gold
The test itself is wrong or outdated (test_bug)
ember
The code under test returns a wrong result (logic_bug)
Signal
confidence 0.46 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
The fixture is not perfectly calibrated; the code is right and the assertion is wrong.
{
  "test": "tests/test_eval_metrics.py::test_ece_perfect_calibration",
  "output": "AssertionError: assert 0.1 == 0.0\nthe fixture uses ten predictions at confidence 0.8 with nine correct, so the expected ECE is 0.1, not 0.0"
}
ember's answer for failure_023, failure_kindProbability ember gave each option.logic_bug0.46 ember's pickenvironment0.01flaky0.12test_bug0.42 gold

risk_001: risk

Change risk · dev
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.68, 0.68 levels from gold
Kit outcome
within one level of gold
Why this gold label
Test-only cosmetic rename.
{
  "diff_summary": "Rename local variable `tmp` to `result` in one test helper",
  "files": [
    "tests/test_helpers.py"
  ],
  "lines_changed": 4
}
ember's answer for risk_001, riskProbability ember gave each option.Negligible0.38 goldLow0.58 ember's pickMedium0.03High0.02

risk_002: risk

Change risk · test
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.53, 0.53 levels from gold
Kit outcome
within one level of gold
Why this gold label
Documentation-only change.
{
  "diff_summary": "Fix a typo in the README",
  "files": [
    "README.md"
  ],
  "lines_changed": 1
}
ember's answer for risk_002, riskProbability ember gave each option.Negligible0.52 goldLow0.44 ember's pickMedium0.02High0.01

risk_003: risk

Change risk · dev
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.51, 0.51 levels from gold
Kit outcome
within one level of gold
Why this gold label
Docstring-only change.
{
  "diff_summary": "Fix the spelling of 'directory' in a docstring",
  "files": [
    "ember/paths.py"
  ],
  "lines_changed": 1
}
ember's answer for risk_003, riskProbability ember gave each option.Negligible0.54 goldLow0.42 ember's pickMedium0.02High0.01

risk_004: risk

Change risk · test
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.68, 0.68 levels from gold
Kit outcome
within one level of gold
Why this gold label
Mechanical whitespace change in tests.
{
  "diff_summary": "Reformat one test file with ruff format; whitespace only, no logic changes",
  "files": [
    "tests/test_media.py"
  ],
  "lines_changed": 12
}
ember's answer for risk_004, riskProbability ember gave each option.Negligible0.39 goldLow0.56 ember's pickMedium0.03High0.02

risk_005: risk

Change risk · dev
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.60, 0.60 levels from gold
Kit outcome
within one level of gold
Why this gold label
No code affected.
{
  "diff_summary": "Update the copyright year in the LICENSE file",
  "files": [
    "LICENSE"
  ],
  "lines_changed": 1
}
ember's answer for risk_005, riskProbability ember gave each option.Negligible0.45 goldLow0.51 ember's pickMedium0.03High0.02

risk_016: risk

Change risk · dev
Gold
Medium (2)
ember
Low (1)
Signal
expected 1.30, 0.70 levels from gold
Kit outcome
within one level of gold
Why this gold label
Runtime dependency upgrade under the HTTP server.
{
  "diff_summary": "Upgrade fastapi from 0.142 to 0.150",
  "files": [
    "pyproject.toml",
    "uv.lock"
  ],
  "lines_changed": 22
}
ember's answer for risk_016, riskProbability ember gave each option.Negligible0.11Low0.52 ember's pickMedium0.34 goldHigh0.04

risk_016: needs_review

Change risk · dev
Gold
yes (True)
ember
no (False)
Signal
P(true) 0.29 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Runtime dependency upgrade under the HTTP server.
{
  "diff_summary": "Upgrade fastapi from 0.142 to 0.150",
  "files": [
    "pyproject.toml",
    "uv.lock"
  ],
  "lines_changed": 22
}
ember's answer for risk_016, needs_reviewP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.29

risk_018: risk

Change risk · dev
Gold
Medium (2)
ember
High (3)
Signal
expected 2.54, 0.54 levels from gold
Kit outcome
within one level of gold
Why this gold label
Changes files written into users' projects.
{
  "diff_summary": "Change the generated opencode plugin template and the format of the config it writes",
  "files": [
    "ember/opencode_plugin.py"
  ],
  "lines_changed": 150
}
ember's answer for risk_018, riskProbability ember gave each option.Negligible0.05Low0.06Medium0.19 goldHigh0.70 ember's pick

risk_019: risk

Change risk · dev
Gold
High (3)
ember
Medium (2)
Signal
expected 2.01, 0.99 levels from gold
Kit outcome
within one level of gold
Why this gold label
Weakens authentication.
{
  "diff_summary": "Skip auth token validation when the request originates from localhost to unblock local development",
  "files": [
    "ember/server.py",
    "ember/mcp_server.py"
  ],
  "lines_changed": 180
}
ember's answer for risk_019, riskProbability ember gave each option.Negligible0.09Low0.16Medium0.41 ember's pickHigh0.34 gold

risk_022: risk

Change risk · test
Gold
High (3)
ember
Medium (2)
Signal
expected 1.67, 1.33 levels from gold
Kit outcome
missed the exact level
Why this gold label
One line, but it exposes an unauthenticated server to the network.
{
  "diff_summary": "Bind the model server to 0.0.0.0 by default instead of 127.0.0.1",
  "files": [
    "ember/config.py"
  ],
  "lines_changed": 1
}
ember's answer for risk_022, riskProbability ember gave each option.Negligible0.13Low0.22Medium0.49 ember's pickHigh0.16 gold

risk_022: needs_review

Change risk · test
Gold
yes (True)
ember
no (False)
Signal
P(true) 0.45 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
One line, but it exposes an unauthenticated server to the network.
{
  "diff_summary": "Bind the model server to 0.0.0.0 by default instead of 127.0.0.1",
  "files": [
    "ember/config.py"
  ],
  "lines_changed": 1
}
ember's answer for risk_022, needs_reviewP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.45

risk_023: risk

Change risk · dev
Gold
High (3)
ember
Medium (2)
Signal
expected 1.97, 1.03 levels from gold
Kit outcome
missed the exact level
Why this gold label
Can terminate unrelated processes.
{
  "diff_summary": "Make `ember stop` kill any process listening on the configured port",
  "files": [
    "ember/process.py"
  ],
  "lines_changed": 25
}
ember's answer for risk_023, riskProbability ember gave each option.Negligible0.08Low0.13Medium0.54 ember's pickHigh0.26 gold

risk_024: risk

Change risk · test
Gold
High (3)
ember
Medium (2)
Signal
expected 2.41, 0.59 levels from gold
Kit outcome
within one level of gold
Why this gold label
Irreversible deletion of unrelated user data.
{
  "diff_summary": "Make `ember model rm` delete the whole Hugging Face cache directory instead of one model",
  "files": [
    "ember/models.py"
  ],
  "lines_changed": 10
}
ember's answer for risk_024, riskProbability ember gave each option.Negligible0.06Low0.08Medium0.26 ember's pickHigh0.61 gold

effort_002: effort

Effort and approach · dev
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 0.51, 0.51 levels from gold
Kit outcome
within one level of gold
Why this gold label
A single default value.
{
  "task": "Change the default number of lines `ember logs` prints from 40 to 100",
  "constraints": [
    "one-line change"
  ]
}
ember's answer for effort_002, effortProbability ember gave each option.Trivial (minutes)0.60 goldSmall (under an hour)0.31 ember's pickMedium (a few hours)0.08Large (a day or more)0.01

effort_003: effort

Effort and approach · test
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 0.56, 0.56 levels from gold
Kit outcome
within one level of gold
Why this gold label
A guard clause in one function.
{
  "task": "Guard the missing null check that crashes the order total calculation for empty carts",
  "constraints": [
    "hotfix",
    "touch only the failing function"
  ]
}
ember's answer for effort_003, effortProbability ember gave each option.Trivial (minutes)0.50 goldSmall (under an hour)0.44 ember's pickMedium (a few hours)0.04Large (a day or more)0.01

effort_006: effort

Effort and approach · test
Gold
Small (under an hour) (1)
ember
Medium (a few hours) (2)
Signal
expected 1.50, 0.50 levels from gold
Kit outcome
within one level of gold
Why this gold label
Mechanical extraction covered by existing tests.
{
  "task": "The same 15-line retry loop is copy-pasted in 6 HTTP client methods that already have unit tests; extract it into a shared helper",
  "constraints": [
    "no new dependencies",
    "keep existing API"
  ]
}
ember's answer for effort_006, effortProbability ember gave each option.Trivial (minutes)0.07Small (under an hour)0.39 goldMedium (a few hours)0.51 ember's pickLarge (a day or more)0.03

effort_009: effort

Effort and approach · test
Gold
Medium (a few hours) (2)
ember
Large (a day or more) (3)
Signal
expected 2.78, 0.78 levels from gold
Kit outcome
within one level of gold
Why this gold label
A module split with no behavior change.
{
  "task": "The 900-line server module mixes routing, validation, metrics, and model calls; split it by responsibility before the next feature lands",
  "constraints": [
    "no API changes",
    "keep the public import path"
  ]
}
ember's answer for effort_009, effortProbability ember gave each option.Trivial (minutes)0.02Small (under an hour)0.02Medium (a few hours)0.12 goldLarge (a day or more)0.84 ember's pick

effort_013: effort

Effort and approach · dev
Gold
Large (a day or more) (3)
ember
Medium (a few hours) (2)
Signal
expected 2.01, 0.99 levels from gold
Kit outcome
within one level of gold
Why this gold label
A persistent queue and worker are new subsystems.
{
  "task": "Add a background job queue so long-running eval runs execute asynchronously with progress reporting",
  "constraints": [
    "must survive server restarts",
    "no external broker"
  ]
}
ember's answer for effort_013, effortProbability ember gave each option.Trivial (minutes)0.04Small (under an hour)0.09Medium (a few hours)0.69 ember's pickLarge (a day or more)0.18 gold

effort_015: effort

Effort and approach · test
Gold
Large (a day or more) (3)
ember
Medium (a few hours) (2)
Signal
expected 1.85, 1.15 levels from gold
Kit outcome
missed the exact level
Why this gold label
A new package with its own packaging and tests.
{
  "task": "Build a plugin for a second editor that mirrors the existing opencode plugin",
  "constraints": [
    "reuse the existing MCP server",
    "ship it as a separate package"
  ]
}
ember's answer for effort_015, effortProbability ember gave each option.Trivial (minutes)0.04Small (under an hour)0.17Medium (a few hours)0.70 ember's pickLarge (a day or more)0.09 gold

effort_016: effort

Effort and approach · dev
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 1.29, 1.29 levels from gold
Kit outcome
missed the exact level
Why this gold label
The edit is one default; the accuracy-versus-memory trade-off is a product decision.
{
  "task": "Switch the default model from flash (9B) to full (27B)",
  "constraints": [
    "full is more accurate but needs about 3x the memory and is about 3x slower",
    "many users have 32 GB machines"
  ]
}
ember's answer for effort_016, effortProbability ember gave each option.Trivial (minutes)0.22 goldSmall (under an hour)0.37 ember's pickMedium (a few hours)0.33Large (a day or more)0.09

effort_017: effort

Effort and approach · dev
Gold
Small (under an hour) (1)
ember
Medium (a few hours) (2)
Signal
expected 1.61, 0.61 levels from gold
Kit outcome
within one level of gold
Why this gold label
Dropping support for some users is a human call; the code change is small.
{
  "task": "Drop CPU fallback support to simplify the runtime",
  "constraints": [
    "some users run without MPS",
    "removes about 200 lines"
  ]
}
ember's answer for effort_017, effortProbability ember gave each option.Trivial (minutes)0.06Small (under an hour)0.33 goldMedium (a few hours)0.56 ember's pickLarge (a day or more)0.05

effort_018: effort

Effort and approach · test
Gold
Small (under an hour) (1)
ember
Medium (a few hours) (2)
Signal
expected 2.38, 1.38 levels from gold
Kit outcome
missed the exact level
Why this gold label
Breaking a public contract needs a human decision.
{
  "task": "Rename the public advise tool to consult across the MCP server, docs, and agent kit",
  "constraints": [
    "the tool name is public API used by external agents",
    "renaming breaks every existing integration"
  ]
}
ember's answer for effort_018, effortProbability ember gave each option.Trivial (minutes)0.06Small (under an hour)0.09 goldMedium (a few hours)0.25 ember's pickLarge (a day or more)0.59

effort_019: effort

Effort and approach · dev
Gold
Small (under an hour) (1)
ember
Medium (a few hours) (2)
Signal
expected 1.75, 0.75 levels from gold
Kit outcome
within one level of gold
Why this gold label
A breaking change with no deprecation policy.
{
  "task": "Rename the HTTP endpoint /v1/systemone to /v1/advise",
  "constraints": [
    "external clients call /v1/systemone directly",
    "no deprecation policy exists yet"
  ]
}
ember's answer for effort_019, effortProbability ember gave each option.Trivial (minutes)0.12Small (under an hour)0.21 goldMedium (a few hours)0.47 ember's pickLarge (a day or more)0.20

effort_020: effort

Effort and approach · dev
Gold
Large (a day or more) (3)
ember
Medium (a few hours) (2)
Signal
expected 1.58, 1.42 levels from gold
Kit outcome
missed the exact level
Why this gold label
Conflicts with a stated privacy promise; building it would take a day or more.
{
  "task": "Add anonymous usage telemetry that reports to a hosted endpoint",
  "constraints": [
    "the project promises that state never leaves the machine",
    "maintainers want usage data"
  ]
}
ember's answer for effort_020, effortProbability ember gave each option.Trivial (minutes)0.08Small (under an hour)0.31Medium (a few hours)0.56 ember's pickLarge (a day or more)0.05 gold

effort_020: approach

Effort and approach · dev
Gold
The trade-off needs a human decision (ask_user)
ember
Smallest change that fixes the symptom (minimal_patch)
Signal
confidence 0.38 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
Conflicts with a stated privacy promise; building it would take a day or more.
{
  "task": "Add anonymous usage telemetry that reports to a hosted endpoint",
  "constraints": [
    "the project promises that state never leaves the machine",
    "maintainers want usage data"
  ]
}
ember's answer for effort_020, approachProbability ember gave each option.minimal_patch0.38 ember's pickrefactor0.03new_component0.28ask_user0.31 gold

intent_029: specific_enough

Intent and readiness · test
Gold
The target and desired outcome are both identifiable (True)
ember
The target or desired outcome is unclear (False)
Signal
P(true) 0.30 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Asks a factual question about a named setting.
what's the default port for the model server?
ember's answer for intent_029, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.30

intent_031: specific_enough

Intent and readiness · dev
Gold
The target and desired outcome are both identifiable (True)
ember
The target or desired outcome is unclear (False)
Signal
P(true) 0.23 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Asks to compare two named options.
what's the difference between the flash and full models?
ember's answer for intent_031, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.23

intent_032: specific_enough

Intent and readiness · test
Gold
The target and desired outcome are both identifiable (True)
ember
The target or desired outcome is unclear (False)
Signal
P(true) 0.49 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Explanation request about a named metric.
can you explain what ECE measures in the benchmark report?
ember's answer for intent_032, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.49

intent_048: specific_enough

Intent and readiness · test
Gold
The target or desired outcome is unclear (False)
ember
The target and desired outcome are both identifiable (True)
Signal
P(true) 0.50 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
The problem is never described.
there's a problem somewhere in the build, investigate
ember's answer for intent_048, specific_enoughP(true) placed on the no, unsure, and yes zones.nounsureyesgold: noP = 0.50

risk_025: risk

Change risk · test
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.64, 0.64 levels from gold
Kit outcome
within one level of gold
Why this gold label
Documentation-only change.
{
  "diff_summary": "Fix a broken link in CONTRIBUTING.md",
  "files": [
    "CONTRIBUTING.md"
  ],
  "lines_changed": 1
}
ember's answer for risk_025, riskProbability ember gave each option.Negligible0.42 goldLow0.55 ember's pickMedium0.02High0.01

risk_026: risk

Change risk · dev
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.58, 0.58 levels from gold
Kit outcome
within one level of gold
Why this gold label
Comment-only change.
{
  "diff_summary": "Correct the spelling of 'calibration' in a code comment",
  "files": [
    "evals/metrics.py"
  ],
  "lines_changed": 1
}
ember's answer for risk_026, riskProbability ember gave each option.Negligible0.47 goldLow0.49 ember's pickMedium0.02High0.01

risk_027: risk

Change risk · test
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.65, 0.65 levels from gold
Kit outcome
within one level of gold
Why this gold label
Reordering with no effect.
{
  "diff_summary": "Sort the entries of .gitignore alphabetically",
  "files": [
    ".gitignore"
  ],
  "lines_changed": 6
}
ember's answer for risk_027, riskProbability ember gave each option.Negligible0.41 goldLow0.54 ember's pickMedium0.03High0.02

risk_028: risk

Change risk · dev
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.63, 0.63 levels from gold
Kit outcome
within one level of gold
Why this gold label
Test-only cosmetic rename.
{
  "diff_summary": "Rename a local loop variable from i to index in one test",
  "files": [
    "tests/test_media.py"
  ],
  "lines_changed": 3
}
ember's answer for risk_028, riskProbability ember gave each option.Negligible0.43 goldLow0.53 ember's pickMedium0.03High0.02

risk_030: risk

Change risk · dev
Gold
Negligible (0)
ember
Low (1)
Signal
expected 0.73, 0.73 levels from gold
Kit outcome
within one level of gold
Why this gold label
Whitespace only.
{
  "diff_summary": "Reflow a Markdown paragraph in a vault session log; no wording changes",
  "files": [
    "vault/sessions/2026-10-03-eval-benchmark-harness.md"
  ],
  "lines_changed": 8
}
ember's answer for risk_030, riskProbability ember gave each option.Negligible0.35 goldLow0.58 ember's pickMedium0.05High0.01

risk_043: risk

Change risk · test
Gold
High (3)
ember
Medium (2)
Signal
expected 2.01, 0.99 levels from gold
Kit outcome
within one level of gold
Why this gold label
Sends private code off the machine.
{
  "diff_summary": "Log the full advise request state, including user code, to a remote analytics endpoint",
  "files": [
    "ember/server.py"
  ],
  "lines_changed": 45
}
ember's answer for risk_043, riskProbability ember gave each option.Negligible0.07Low0.12Medium0.54 ember's pickHigh0.27 gold

risk_044: risk

Change risk · dev
Gold
High (3)
ember
Medium (2)
Signal
expected 2.19, 0.81 levels from gold
Kit outcome
within one level of gold
Why this gold label
Unprompted recursive deletion of a user-chosen path.
{
  "diff_summary": "Remove the confirmation prompt and rm -rf the state directory in `ember uninstall`, even when EMBER_STATE_DIR points elsewhere",
  "files": [
    "ember/cli.py"
  ],
  "lines_changed": 8
}
ember's answer for risk_044, riskProbability ember gave each option.Negligible0.07Low0.11Medium0.36 ember's pickHigh0.45 gold

risk_045: risk

Change risk · test
Gold
High (3)
ember
Low (1)
Signal
expected 1.37, 1.63 levels from gold
Kit outcome
missed the exact level
Why this gold label
Lets callers make the server read host files.
{
  "diff_summary": "Let `images` accept file:// URLs",
  "files": [
    "ember/media.py"
  ],
  "lines_changed": 6
}
ember's answer for risk_045, riskProbability ember gave each option.Negligible0.14Low0.42 ember's pickMedium0.37High0.07 gold

risk_045: needs_review

Change risk · test
Gold
yes (True)
ember
no (False)
Signal
P(true) 0.35 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Lets callers make the server read host files.
{
  "diff_summary": "Let `images` accept file:// URLs",
  "files": [
    "ember/media.py"
  ],
  "lines_changed": 6
}
ember's answer for risk_045, needs_reviewP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.35

risk_046: risk

Change risk · dev
Gold
High (3)
ember
Medium (2)
Signal
expected 2.46, 0.54 levels from gold
Kit outcome
within one level of gold
Why this gold label
Kills unrelated processes.
{
  "diff_summary": "Make `ember stop` SIGKILL every process whose command line contains 'ember'",
  "files": [
    "ember/process.py"
  ],
  "lines_changed": 15
}
ember's answer for risk_046, riskProbability ember gave each option.Negligible0.06Low0.08Medium0.20 ember's pickHigh0.66 gold

risk_047: risk

Change risk · test
Gold
High (3)
ember
Low (1)
Signal
expected 1.45, 1.55 levels from gold
Kit outcome
missed the exact level
Why this gold label
Leaks a secret into shareable files.
{
  "diff_summary": "Store the OpenRouter API key in the agent eval results JSON for reproducibility",
  "files": [
    "scripts/run_agent_evals.py"
  ],
  "lines_changed": 3
}
ember's answer for risk_047, riskProbability ember gave each option.Negligible0.17Low0.39 ember's pickMedium0.26High0.18 gold

risk_047: needs_review

Change risk · test
Gold
yes (True)
ember
no (False)
Signal
P(true) 0.34 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
Leaks a secret into shareable files.
{
  "diff_summary": "Store the OpenRouter API key in the agent eval results JSON for reproducibility",
  "files": [
    "scripts/run_agent_evals.py"
  ],
  "lines_changed": 3
}
ember's answer for risk_047, needs_reviewP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.34

risk_048: risk

Change risk · dev
Gold
High (3)
ember
Medium (2)
Signal
expected 1.98, 1.02 levels from gold
Kit outcome
missed the exact level
Why this gold label
One flag, but agents may act outside the sandbox.
{
  "diff_summary": "Run agent eval sessions with --auto so every permission request is approved",
  "files": [
    "evals/agent/opencode.py"
  ],
  "lines_changed": 1
}
ember's answer for risk_048, riskProbability ember gave each option.Negligible0.12Low0.20Medium0.26 ember's pickHigh0.42 gold

risk_048: needs_review

Change risk · dev
Gold
yes (True)
ember
no (False)
Signal
P(true) 0.40 (unsure)
Kit outcome
deferred: treated as unsure
Why this gold label
One flag, but agents may act outside the sandbox.
{
  "diff_summary": "Run agent eval sessions with --auto so every permission request is approved",
  "files": [
    "evals/agent/opencode.py"
  ],
  "lines_changed": 1
}
ember's answer for risk_048, needs_reviewP(true) placed on the no, unsure, and yes zones.nounsureyesgold: yesP = 0.40

effort_022: effort

Effort and approach · test
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 0.71, 0.71 levels from gold
Kit outcome
within one level of gold
Why this gold label
One constant.
{
  "task": "Change the `ember logs` follow interval from 0.5 s to 1 s",
  "constraints": [
    "one constant"
  ]
}
ember's answer for effort_022, effortProbability ember gave each option.Trivial (minutes)0.39 goldSmall (under an hour)0.51 ember's pickMedium (a few hours)0.08Large (a day or more)0.01

effort_023: effort

Effort and approach · dev
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 1.03, 1.03 levels from gold
Kit outcome
missed the exact level
Why this gold label
One string.
{
  "task": "Make `ember status` print 'stopped' instead of 'not running'",
  "constraints": [
    "keep the exit code"
  ]
}
ember's answer for effort_023, effortProbability ember gave each option.Trivial (minutes)0.28 goldSmall (under an hour)0.44 ember's pickMedium (a few hours)0.23Large (a day or more)0.04

effort_027: effort

Effort and approach · test
Gold
Medium (a few hours) (2)
ember
Large (a day or more) (3)
Signal
expected 2.76, 0.76 levels from gold
Kit outcome
within one level of gold
Why this gold label
A module split with behavior preserved.
{
  "task": "The 800-line CLI module mixes server lifecycle, models, evals, and agents; split it by responsibility",
  "constraints": [
    "identical CLI behavior",
    "existing tests pass"
  ]
}
ember's answer for effort_027, effortProbability ember gave each option.Trivial (minutes)0.01Small (under an hour)0.02Medium (a few hours)0.16 goldLarge (a day or more)0.81 ember's pick

effort_029: effort

Effort and approach · test
Gold
Small (under an hour) (1)
ember
Medium (a few hours) (2)
Signal
expected 1.77, 0.77 levels from gold
Kit outcome
within one level of gold
Why this gold label
Consolidation across two call sites.
{
  "task": "The benchmark and agent runners resolve the server URL from flags, env, and config in different ways; unify them",
  "constraints": [
    "same defaults"
  ]
}
ember's answer for effort_029, effortProbability ember gave each option.Trivial (minutes)0.06Small (under an hour)0.19 goldMedium (a few hours)0.67 ember's pickLarge (a day or more)0.08

effort_031: effort

Effort and approach · test
Gold
Large (a day or more) (3)
ember
Small (under an hour) (1)
Signal
expected 1.36, 1.64 levels from gold
Kit outcome
missed the exact level
Why this gold label
A new adapter with its own parsing and isolation.
{
  "task": "Add a Claude Code adapter to the agent eval alongside opencode",
  "constraints": [
    "same scenarios and scoring",
    "isolated config"
  ]
}
ember's answer for effort_031, effortProbability ember gave each option.Trivial (minutes)0.10Small (under an hour)0.47 ember's pickMedium (a few hours)0.39Large (a day or more)0.04 gold

effort_032: effort

Effort and approach · dev
Gold
Large (a day or more) (3)
ember
Small (under an hour) (1)
Signal
expected 1.19, 1.81 levels from gold
Kit outcome
missed the exact level
Why this gold label
A new reporting component.
{
  "task": "Add a static HTML dashboard that tracks benchmark results across runs over time",
  "constraints": [
    "static files only",
    "no server"
  ]
}
ember's answer for effort_032, effortProbability ember gave each option.Trivial (minutes)0.18Small (under an hour)0.48 ember's pickMedium (a few hours)0.31Large (a day or more)0.03 gold

effort_032: approach

Effort and approach · dev
Gold
Add a new module or service (new_component)
ember
Smallest change that fixes the symptom (minimal_patch)
Signal
confidence 0.66 (verify band)
Kit outcome
acted, flagged for verification
Why this gold label
A new reporting component.
{
  "task": "Add a static HTML dashboard that tracks benchmark results across runs over time",
  "constraints": [
    "static files only",
    "no server"
  ]
}
ember's answer for effort_032, approachProbability ember gave each option.minimal_patch0.66 ember's pickrefactor0.08new_component0.07 goldask_user0.18

effort_033: effort

Effort and approach · dev
Gold
Medium (a few hours) (2)
ember
Small (under an hour) (1)
Signal
expected 1.50, 0.50 levels from gold
Kit outcome
within one level of gold
Why this gold label
A new endpoint.
{
  "task": "Add a /v1/batch endpoint that scores many advise requests in one call",
  "constraints": [
    "keep /v1/systemone unchanged"
  ]
}
ember's answer for effort_033, effortProbability ember gave each option.Trivial (minutes)0.10Small (under an hour)0.36 ember's pickMedium (a few hours)0.48 goldLarge (a day or more)0.06

effort_033: approach

Effort and approach · dev
Gold
Add a new module or service (new_component)
ember
Smallest change that fixes the symptom (minimal_patch)
Signal
confidence 0.52 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
A new endpoint.
{
  "task": "Add a /v1/batch endpoint that scores many advise requests in one call",
  "constraints": [
    "keep /v1/systemone unchanged"
  ]
}
ember's answer for effort_033, approachProbability ember gave each option.minimal_patch0.52 ember's pickrefactor0.10new_component0.31 goldask_user0.08

effort_034: approach

Effort and approach · dev
Gold
Add a new module or service (new_component)
ember
Smallest change that fixes the symptom (minimal_patch)
Signal
confidence 0.50 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
A small new module.
{
  "task": "Add a CSV export of the per-item benchmark trace",
  "constraints": [
    "one new module",
    "no new dependencies"
  ]
}
ember's answer for effort_034, approachProbability ember gave each option.minimal_patch0.50 ember's pickrefactor0.04new_component0.40 goldask_user0.05

effort_035: effort

Effort and approach · test
Gold
Medium (a few hours) (2)
ember
Small (under an hour) (1)
Signal
expected 1.23, 0.77 levels from gold
Kit outcome
within one level of gold
Why this gold label
A new module and metric.
{
  "task": "Add a Prometheus metric for MCP autostarts in a new metrics module for the MCP process",
  "constraints": [
    "the MCP process must stay torch-free"
  ]
}
ember's answer for effort_035, effortProbability ember gave each option.Trivial (minutes)0.12Small (under an hour)0.56 ember's pickMedium (a few hours)0.29 goldLarge (a day or more)0.03

effort_036: effort

Effort and approach · test
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 0.75, 0.75 levels from gold
Kit outcome
within one level of gold
Why this gold label
One default; the cost trade-off is a human call.
{
  "task": "Make the agent eval default to Opus-class models",
  "constraints": [
    "about five times the cost per run",
    "may flatter ember compared with the cheaper models users run"
  ]
}
ember's answer for effort_036, effortProbability ember gave each option.Trivial (minutes)0.48 goldSmall (under an hour)0.35 ember's pickMedium (a few hours)0.13Large (a day or more)0.05

effort_036: approach

Effort and approach · test
Gold
The trade-off needs a human decision (ask_user)
ember
Smallest change that fixes the symptom (minimal_patch)
Signal
confidence 0.81 (verify band)
Kit outcome
acted, flagged for verification
Why this gold label
One default; the cost trade-off is a human call.
{
  "task": "Make the agent eval default to Opus-class models",
  "constraints": [
    "about five times the cost per run",
    "may flatter ember compared with the cheaper models users run"
  ]
}
ember's answer for effort_036, approachProbability ember gave each option.minimal_patch0.81 ember's pickrefactor0.07new_component0.03ask_user0.09 gold

effort_038: effort

Effort and approach · dev
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 0.65, 0.65 levels from gold
Kit outcome
within one level of gold
Why this gold label
Changes every agent's behavior on thin evidence.
{
  "task": "Lower the kit's choice trust threshold from 0.85 to 0.60",
  "constraints": [
    "agents would act more often",
    "one benchmark run supports it"
  ]
}
ember's answer for effort_038, effortProbability ember gave each option.Trivial (minutes)0.48 goldSmall (under an hour)0.41 ember's pickMedium (a few hours)0.10Large (a day or more)0.01

effort_038: approach

Effort and approach · dev
Gold
The trade-off needs a human decision (ask_user)
ember
Smallest change that fixes the symptom (minimal_patch)
Signal
confidence 0.79 (verify band)
Kit outcome
acted, flagged for verification
Why this gold label
Changes every agent's behavior on thin evidence.
{
  "task": "Lower the kit's choice trust threshold from 0.85 to 0.60",
  "constraints": [
    "agents would act more often",
    "one benchmark run supports it"
  ]
}
ember's answer for effort_038, approachProbability ember gave each option.minimal_patch0.79 ember's pickrefactor0.11new_component0.02ask_user0.08 gold

effort_039: effort

Effort and approach · test
Gold
Large (a day or more) (3)
ember
Small (under an hour) (1)
Signal
expected 1.49, 1.51 levels from gold
Kit outcome
missed the exact level
Why this gold label
Privacy trade-off; building it is a day or more.
{
  "task": "Send anonymized benchmark results to a shared leaderboard by default",
  "constraints": [
    "users may run on private code",
    "maintainers want comparisons"
  ]
}
ember's answer for effort_039, effortProbability ember gave each option.Trivial (minutes)0.12Small (under an hour)0.34 ember's pickMedium (a few hours)0.47Large (a day or more)0.07 gold

effort_039: approach

Effort and approach · test
Gold
The trade-off needs a human decision (ask_user)
ember
Add a new module or service (new_component)
Signal
confidence 0.41 (defer band)
Kit outcome
deferred: gather evidence or ask
Why this gold label
Privacy trade-off; building it is a day or more.
{
  "task": "Send anonymized benchmark results to a shared leaderboard by default",
  "constraints": [
    "users may run on private code",
    "maintainers want comparisons"
  ]
}
ember's answer for effort_039, approachProbability ember gave each option.minimal_patch0.38refactor0.10new_component0.41 ember's pickask_user0.10 gold

effort_040: effort

Effort and approach · test
Gold
Trivial (minutes) (0)
ember
Small (under an hour) (1)
Signal
expected 1.07, 1.07 levels from gold
Kit outcome
missed the exact level
Why this gold label
A workflow trade-off for the user.
{
  "task": "Change the AGENTS.md policy so agents consult ember before every file edit",
  "constraints": [
    "raises consultation",
    "adds about a second to every edit"
  ]
}
ember's answer for effort_040, effortProbability ember gave each option.Trivial (minutes)0.19 goldSmall (under an hour)0.59 ember's pickMedium (a few hours)0.20Large (a day or more)0.03

effort_040: approach

Effort and approach · test
Gold
The trade-off needs a human decision (ask_user)
ember
Smallest change that fixes the symptom (minimal_patch)
Signal
confidence 0.67 (verify band)
Kit outcome
acted, flagged for verification
Why this gold label
A workflow trade-off for the user.
{
  "task": "Change the AGENTS.md policy so agents consult ember before every file edit",
  "constraints": [
    "raises consultation",
    "adds about a second to every edit"
  ]
}
ember's answer for effort_040, approachProbability ember gave each option.minimal_patch0.67 ember's pickrefactor0.13new_component0.11ask_user0.09 gold