4Benchmark design
The dataset has 264 items and 456 scored questions: 192 choice, 168 noul, and 96 score. Each line is self-contained: the state, the recipe's fixed question set, a gold label for every question, its source, and a one-line rationale. Other implementations can score it without this harness.
9 items reproduce the worked examples in the agent kit's playbook and sit in dev; the other 255 were written for this benchmark. Items are split into dev (130 items, for tuning) and test (134 items, for reporting), stratified so that every label of every question appears in both. The dataset card below answers the datasheet questions of [12].
dev and test.Label balance
| Question | Gold label | Meaning | dev | test |
|---|---|---|---|---|
intent | question | Explain or answer something; no code changes requested | 7 | 7 |
implement | Make a change: add, modify, or configure something | 7 | 7 | |
investigate | Look into a problem and report findings | 7 | 7 | |
fix | Repair a specific broken behavior | 7 | 7 | |
specific_enough | True | The target and desired outcome are both identifiable | 16 | 19 |
False | The target or desired outcome is unclear | 12 | 9 | |
failure_kind | logic_bug | The code under test returns a wrong result | 6 | 6 |
environment | Missing service, dependency, port, file, or configuration | 6 | 6 | |
flaky | Timing, ordering, or nondeterminism; may pass on retry | 6 | 6 | |
test_bug | The test itself is wrong or outdated | 6 | 6 | |
retry | True | yes | 6 | 6 |
False | no | 18 | 18 | |
risk | 0 | Negligible | 6 | 6 |
1 | Low | 6 | 6 | |
2 | Medium | 6 | 6 | |
3 | High | 6 | 6 | |
needs_review | True | yes | 12 | 12 |
False | no | 12 | 12 | |
owner | api | HTTP handlers and request validation | 4 | 4 |
storage | Database, migrations, and persistence | 4 | 4 | |
ui | Frontend rendering and client state | 4 | 4 | |
infra | Build, CI, deployment, and configuration | 4 | 4 | |
unclear | Not enough information to tell | 4 | 4 | |
effort | 0 | Trivial (minutes) | 4 | 5 |
1 | Small (under an hour) | 6 | 7 | |
2 | Medium (a few hours) | 5 | 6 | |
3 | Large (a day or more) | 3 | 4 | |
approach | minimal_patch | Smallest change that fixes the symptom | 4 | 6 |
refactor | Restructure the code so the fix is natural | 4 | 6 | |
new_component | Add a new module or service | 5 | 5 | |
ask_user | The trade-off needs a human decision | 5 | 5 | |
dominant_red | True | Red is the dominant colour | 2 | 2 |
False | Red is not dominant | 2 | 2 | |
dominant_colour | red | The image is mostly red or warm-red tones | 1 | 1 |
green | The image is mostly green tones | 1 | 1 | |
blue | The image is mostly blue tones | 1 | 1 | |
mixed | No single colour clearly dominates | 1 | 1 | |
alert_level | 0 | No alert (calm green) | 1 | 1 |
1 | Low (yellow) | 1 | 1 | |
2 | Medium (orange) | 1 | 1 | |
3 | High (red) | 1 | 1 | |
colour_changed | True | The dominant colour is different in the two frames | 2 | 2 |
False | The dominant colour is the same in both frames | 2 | 2 |
Recipes
Each recipe asks the same fixed question set for every item, because ember weighs the questions in one call jointly.
Intent and readiness
intent_readiness- When
- Before changing files in response to a conversational request.
- State
- The user's latest message, verbatim.
- Decides
- What the user is asking for, and whether it is specific enough to act on without a clarifying question.
| Question | Type | Asks | Options |
|---|---|---|---|
intent | choice | What is the user asking the coding agent to do? | question: Explain or answer something; no code changes requested implement: Make a change: add, modify, or configure something investigate: Look into a problem and report findings fix: Repair a specific broken behavior |
specific_enough | noul | Is the request specific enough to act on without asking a clarifying question? | true: The target and desired outcome are both identifiable false: The target or desired outcome is unclear |
Failure triage
failure_triage- When
- Before retrying or fixing a failing test or command.
- State
- The test or command, and the relevant error lines.
- Decides
- The most likely cause, and whether a plain re-run would pass.
| Question | Type | Asks | Options |
|---|---|---|---|
failure_kind | choice | What most likely caused this failure? | logic_bug: The code under test returns a wrong result environment: Missing service, dependency, port, file, or configuration flaky: Timing, ordering, or nondeterminism; may pass on retry test_bug: The test itself is wrong or outdated |
retry | noul | Is simply re-running likely to make it pass? | yes or no |
Change risk
change_risk- When
- Before committing, pushing, or merging.
- State
- A diff summary, the files touched, and the lines changed.
- Decides
- How risky the change is to ship unreviewed, and whether a person should review it.
| Question | Type | Asks | Options |
|---|---|---|---|
risk | score | How risky is shipping this change without human review? | 0: Negligible 1: Low 2: Medium 3: High |
needs_review | noul | Should a human review this change before it is merged? | yes or no |
Routing and ownership
routing- When
- When an error needs an owner.
- State
- The error message and the file or module it came from.
- Decides
- Which part of the system should handle it. The option set is a template that projects replace with their own.
| Question | Type | Asks | Options |
|---|---|---|---|
owner | choice | Which part of the system should handle this? | api: HTTP handlers and request validation storage: Database, migrations, and persistence ui: Frontend rendering and client state infra: Build, CI, deployment, and configuration unclear: Not enough information to tell |
Effort and approach
effort_approach- When
- When planning a task.
- State
- The task and its constraints.
- Decides
- How much work the task is, and which approach fits. The option sets are templates.
| Question | Type | Asks | Options |
|---|---|---|---|
effort | score | How much work is this task? | 0: Trivial (minutes) 1: Small (under an hour) 2: Medium (a few hours) 3: Large (a day or more) |
approach | choice | Which approach best fits the constraints? | minimal_patch: Smallest change that fixes the symptom refactor: Restructure the code so the fix is natural new_component: Add a new module or service ask_user: The trade-off needs a human decision |
vision_noul
vision_noul- When
- State
- Decides
| Question | Type | Asks | Options |
|---|---|---|---|
dominant_red | noul | Is the image predominantly red? | true: Red is the dominant colour false: Red is not dominant |
vision_choice
vision_choice- When
- State
- Decides
| Question | Type | Asks | Options |
|---|---|---|---|
dominant_colour | choice | What is the dominant colour in the image? | red: The image is mostly red or warm-red tones green: The image is mostly green tones blue: The image is mostly blue tones mixed: No single colour clearly dominates |
vision_score
vision_score- When
- State
- Decides
| Question | Type | Asks | Options |
|---|---|---|---|
alert_level | score | The image shows a status indicator. How severe is the alert it is signalling? | 0: No alert (calm green) 1: Low (yellow) 2: Medium (orange) 3: High (red) |
vision_video
vision_video- When
- State
- Decides
| Question | Type | Asks | Options |
|---|---|---|---|
colour_changed | noul | The frames are from a two-frame video clip. Did the dominant colour change between frame 1 and frame 2? | true: The dominant colour is different in the two frames false: The dominant colour is the same in both frames |
Labelling protocol
- Labels were judged from
stateand the option descriptions alone, and were fixed before any model run. needs_reviewis true exactly whenriskis Medium or High, andretryis true only forflakyfailures.- Drafts whose evidence fit two options were rewritten or dropped while authoring, before any run.
- A label changes only when a reviewer finds it wrong on its merits, never because the model disagreed. Every change alters the dataset's SHA-256, which each run records, so runs on different label sets cannot be compared by accident.
Dataset card
- Why was it built?
- To measure whether ember's probabilities are accurate and calibrated enough for agents to act on the agent kit's thresholds, and to compare model revisions and setups.
- What does it contain?
- 264 items across five recipes, 456 scored questions, a gold label per question, a rationale per item.
- How was it collected?
- 9 items reproduce the agent kit's worked examples; 255 were written by the ember authors. No user data.
- How was it labelled?
- From
stateand the option descriptions alone, fixed before any run, under the rules in the labelling protocol. - What is it for?
- Tune on
dev, report ontest. It is not a measure of agent behaviour, and not a training set. - How is it maintained?
- Relabelling happens only on review, changes the dataset's SHA-256, and invalidates comparisons with earlier runs.