ember

5Metrics

evals/metrics.py computes every metric from the run's trace; the definitions follow the cited sources.

Accuracy and 95% interval

The share of scored questions answered correctly. The interval is a percentile bootstrap over questions: 1,000 resamples with replacement (seed 0), reporting the 2.5th and 97.5th percentiles [9].

acc = (1/N) * sum_i [prediction_i = gold_i]

Applies to: every question.

Macro-F1

The unweighted mean of per-class F1, so rare options count as much as common ones; computed as the arithmetic mean of per-class scores [8].

F1_c = 2TP_c / (2TP_c + FP_c + FN_c); Macro-F1 = mean over classes

Applies to: choice.

Expected and maximum calibration error

Answers are binned by top-label confidence into B = 10 equal-width bins; ECE is the count-weighted gap between each bin's accuracy and its mean confidence [3] [4]. A noul answer's top-label confidence is max(P, 1 - P). With bins of a few items the estimate is noisy and depends on binning [5], so read it with the reliability diagrams.

ECE = sum_b (n_b / N) * |acc(b) - conf(b)|; MCE = max_b |acc(b) - conf(b)|

Applies to: choice, noul.

Brier score

The squared error between the predicted distribution and the one-hot gold label, a proper scoring rule that rewards both calibration and sharpness [6]. 0 is perfect; the worst is 2 for choice and 1 for noul (where K = 2 is folded into P(true)).

BS = (1/N) * sum_i sum_k (p_ik - y_ik)^2

Applies to: choice, noul.

Ranked probability score

The proper scoring rule for ordered levels: it compares cumulative distributions, so a near miss costs less than a far one [7]. 0 is perfect and 1 the worst.

RPS = (1/(K-1)) * sum_k (cumulative p_k - cumulative o_k)^2

Applies to: score.

Mean absolute error

How far the expected score lands from the gold level. Exact-level accuracy is strict for ordinal judgments, so read it with MAE, RPS, and the share within one level.

MAE = (1/N) * sum_i |expected_i - gold_i|, in levels

Applies to: score.

Coverage and accuracy at a threshold

Selective prediction [10]: how often an answer clears the threshold at which the kit tells an agent to act, and how accurate those answers are. The curve sweeps the threshold; the markers are the kit's published values.

coverage(t) = share of answers with confidence >= t; accuracy(t) = accuracy of those answers

Applies to: choice, noul.