ember

4Benchmark design

The dataset has 264 items and 456 scored questions: 192 choice, 168 noul, and 96 score. Each line is self-contained: the state, the recipe's fixed question set, a gold label for every question, its source, and a one-line rationale. Other implementations can score it without this harness.

9 items reproduce the worked examples in the agent kit's playbook and sit in dev; the other 255 were written for this benchmark. Items are split into dev (130 items, for tuning) and test (134 items, for reporting), stratified so that every label of every question appears in both. The dataset card below answers the datasheet questions of [12].

Items per recipe and splitStacked bars of dev and test items for each recipe.dev (tuning)test (reporting)Intent and readiness282856Failure triage242448Change risk242448Routing and ownership202040Effort and approach182240vision_noul448vision_choice448vision_score448vision_video448
Figure 2 Items per recipe, split into dev and test.

Label balance

QuestionGold labelMeaningdevtest
intentquestionExplain or answer something; no code changes requested77
implementMake a change: add, modify, or configure something77
investigateLook into a problem and report findings77
fixRepair a specific broken behavior77
specific_enoughTrueThe target and desired outcome are both identifiable1619
FalseThe target or desired outcome is unclear129
failure_kindlogic_bugThe code under test returns a wrong result66
environmentMissing service, dependency, port, file, or configuration66
flakyTiming, ordering, or nondeterminism; may pass on retry66
test_bugThe test itself is wrong or outdated66
retryTrueyes66
Falseno1818
risk0Negligible66
1Low66
2Medium66
3High66
needs_reviewTrueyes1212
Falseno1212
ownerapiHTTP handlers and request validation44
storageDatabase, migrations, and persistence44
uiFrontend rendering and client state44
infraBuild, CI, deployment, and configuration44
unclearNot enough information to tell44
effort0Trivial (minutes)45
1Small (under an hour)67
2Medium (a few hours)56
3Large (a day or more)34
approachminimal_patchSmallest change that fixes the symptom46
refactorRestructure the code so the fix is natural46
new_componentAdd a new module or service55
ask_userThe trade-off needs a human decision55
dominant_redTrueRed is the dominant colour22
FalseRed is not dominant22
dominant_colourredThe image is mostly red or warm-red tones11
greenThe image is mostly green tones11
blueThe image is mostly blue tones11
mixedNo single colour clearly dominates11
alert_level0No alert (calm green)11
1Low (yellow)11
2Medium (orange)11
3High (red)11
colour_changedTrueThe dominant colour is different in the two frames22
FalseThe dominant colour is the same in both frames22

Recipes

Each recipe asks the same fixed question set for every item, because ember weighs the questions in one call jointly.

Intent and readiness

intent_readiness
When
Before changing files in response to a conversational request.
State
The user's latest message, verbatim.
Decides
What the user is asking for, and whether it is specific enough to act on without a clarifying question.
QuestionTypeAsksOptions
intentchoiceWhat is the user asking the coding agent to do?question: Explain or answer something; no code changes requested
implement: Make a change: add, modify, or configure something
investigate: Look into a problem and report findings
fix: Repair a specific broken behavior
specific_enoughnoulIs the request specific enough to act on without asking a clarifying question?true: The target and desired outcome are both identifiable
false: The target or desired outcome is unclear

Failure triage

failure_triage
When
Before retrying or fixing a failing test or command.
State
The test or command, and the relevant error lines.
Decides
The most likely cause, and whether a plain re-run would pass.
QuestionTypeAsksOptions
failure_kindchoiceWhat most likely caused this failure?logic_bug: The code under test returns a wrong result
environment: Missing service, dependency, port, file, or configuration
flaky: Timing, ordering, or nondeterminism; may pass on retry
test_bug: The test itself is wrong or outdated
retrynoulIs simply re-running likely to make it pass?yes or no

Change risk

change_risk
When
Before committing, pushing, or merging.
State
A diff summary, the files touched, and the lines changed.
Decides
How risky the change is to ship unreviewed, and whether a person should review it.
QuestionTypeAsksOptions
riskscoreHow risky is shipping this change without human review?0: Negligible
1: Low
2: Medium
3: High
needs_reviewnoulShould a human review this change before it is merged?yes or no

Routing and ownership

routing
When
When an error needs an owner.
State
The error message and the file or module it came from.
Decides
Which part of the system should handle it. The option set is a template that projects replace with their own.
QuestionTypeAsksOptions
ownerchoiceWhich part of the system should handle this?api: HTTP handlers and request validation
storage: Database, migrations, and persistence
ui: Frontend rendering and client state
infra: Build, CI, deployment, and configuration
unclear: Not enough information to tell

Effort and approach

effort_approach
When
When planning a task.
State
The task and its constraints.
Decides
How much work the task is, and which approach fits. The option sets are templates.
QuestionTypeAsksOptions
effortscoreHow much work is this task?0: Trivial (minutes)
1: Small (under an hour)
2: Medium (a few hours)
3: Large (a day or more)
approachchoiceWhich approach best fits the constraints?minimal_patch: Smallest change that fixes the symptom
refactor: Restructure the code so the fix is natural
new_component: Add a new module or service
ask_user: The trade-off needs a human decision

vision_noul

vision_noul
When
State
Decides
QuestionTypeAsksOptions
dominant_rednoulIs the image predominantly red?true: Red is the dominant colour
false: Red is not dominant

vision_choice

vision_choice
When
State
Decides
QuestionTypeAsksOptions
dominant_colourchoiceWhat is the dominant colour in the image?red: The image is mostly red or warm-red tones
green: The image is mostly green tones
blue: The image is mostly blue tones
mixed: No single colour clearly dominates

vision_score

vision_score
When
State
Decides
QuestionTypeAsksOptions
alert_levelscoreThe image shows a status indicator. How severe is the alert it is signalling?0: No alert (calm green)
1: Low (yellow)
2: Medium (orange)
3: High (red)

vision_video

vision_video
When
State
Decides
QuestionTypeAsksOptions
colour_changednoulThe frames are from a two-frame video clip. Did the dominant colour change between frame 1 and frame 2?true: The dominant colour is different in the two frames
false: The dominant colour is the same in both frames

Labelling protocol

  • Labels were judged from state and the option descriptions alone, and were fixed before any model run.
  • needs_review is true exactly when risk is Medium or High, and retry is true only for flaky failures.
  • Drafts whose evidence fit two options were rewritten or dropped while authoring, before any run.
  • A label changes only when a reviewer finds it wrong on its merits, never because the model disagreed. Every change alters the dataset's SHA-256, which each run records, so runs on different label sets cannot be compared by accident.

Dataset card

Why was it built?
To measure whether ember's probabilities are accurate and calibrated enough for agents to act on the agent kit's thresholds, and to compare model revisions and setups.
What does it contain?
264 items across five recipes, 456 scored questions, a gold label per question, a rationale per item.
How was it collected?
9 items reproduce the agent kit's worked examples; 255 were written by the ember authors. No user data.
How was it labelled?
From state and the option descriptions alone, fixed before any run, under the rules in the labelling protocol.
What is it for?
Tune on dev, report on test. It is not a measure of agent behaviour, and not a training set.
How is it maintained?
Relabelling happens only on review, changes the dataset's SHA-256, and invalidates comparisons with earlier runs.