5Metrics
evals/metrics.py computes every metric from the run's trace; the definitions follow the cited sources.
Accuracy and 95% interval
The share of scored questions answered correctly. The interval is a percentile bootstrap over questions: 1,000 resamples with replacement (seed 0), reporting the 2.5th and 97.5th percentiles [9].
acc = (1/N) * sum_i [prediction_i = gold_i]Applies to: every question.
Macro-F1
The unweighted mean of per-class F1, so rare options count as much as common ones; computed as the arithmetic mean of per-class scores [8].
F1_c = 2TP_c / (2TP_c + FP_c + FN_c); Macro-F1 = mean over classesApplies to: choice.
Expected and maximum calibration error
Answers are binned by top-label confidence into B = 10 equal-width bins; ECE is the count-weighted gap between each bin's accuracy and its mean confidence [3] [4]. A noul answer's top-label confidence is max(P, 1 - P). With bins of a few items the estimate is noisy and depends on binning [5], so read it with the reliability diagrams.
ECE = sum_b (n_b / N) * |acc(b) - conf(b)|; MCE = max_b |acc(b) - conf(b)|Applies to: choice, noul.
Brier score
The squared error between the predicted distribution and the one-hot gold label, a proper scoring rule that rewards both calibration and sharpness [6]. 0 is perfect; the worst is 2 for choice and 1 for noul (where K = 2 is folded into P(true)).
BS = (1/N) * sum_i sum_k (p_ik - y_ik)^2Applies to: choice, noul.
Ranked probability score
The proper scoring rule for ordered levels: it compares cumulative distributions, so a near miss costs less than a far one [7]. 0 is perfect and 1 the worst.
RPS = (1/(K-1)) * sum_k (cumulative p_k - cumulative o_k)^2Applies to: score.
Mean absolute error
How far the expected score lands from the gold level. Exact-level accuracy is strict for ordinal judgments, so read it with MAE, RPS, and the share within one level.
MAE = (1/N) * sum_i |expected_i - gold_i|, in levelsApplies to: score.
Coverage and accuracy at a threshold
Selective prediction [10]: how often an answer clears the threshold at which the kit tells an agent to act, and how accurate those answers are. The curve sweeps the threshold; the markers are the kit's published values.
coverage(t) = share of answers with confidence >= t; accuracy(t) = accuracy of those answersApplies to: choice, noul.