Instrument
instrument.delong_se()
Return the DeLong standard error of a single AUC.
Usage
instrument.delong_se(
labels,
probs,
)Parameters
labels: npt.ArrayLike-
Binary labels (higher score = positive class).
probs: npt.ArrayLike-
Per-row scores aligned with
labels.
Returns
float-
The DeLong standard error of the AUC of
probsagainstlabels.
instrument.paired_delong()
Compare two score vectors on identical labels with a paired fast-DeLong test.
Usage
instrument.paired_delong(labels, probs_a, probs_b)Parameters
labels: npt.ArrayLike-
Binary labels shared by both score vectors (higher score = positive).
probs_a: npt.ArrayLike-
Per-row scores of the first classifier.
probs_b: npt.ArrayLike- Per-row scores of the second classifier.
Returns
A: PairedDeLong-
class:PairedDeLong with both AUCs, their delta, the per-AUC and
PairedDeLong-
covariance-aware standard errors, the pairing correlation, the z-statistic,
and the 95% confidence interval on the delta.
instrument.PairedDeLong
Result of a paired fast-DeLong comparison of two score vectors on shared labels.
Usage
instrument.PairedDeLong(
auc_a, auc_b, delta, se_a, se_b, cov, se_delta, rho, z, ci95
)delta is auc_b - auc_a and se_delta is its covariance-aware DeLong standard error, sqrt(var_a + var_b - 2 * cov); ci95 is the two-sided 95% confidence interval on delta.
Attributes
auc_a: float-
DeLong AUC of the first score vector.
auc_b: float-
DeLong AUC of the second score vector.
delta: float-
auc_b - auc_a. se_a: float-
DeLong standard error of
auc_a. se_b: float-
DeLong standard error of
auc_b. cov: float-
Covariance of the two AUC estimates.
se_delta: float-
Covariance-aware standard error of
delta. rho: float-
Correlation of the two AUC estimates (NaN when either variance is zero).
z: float-
delta / se_delta(NaN whense_deltais zero). ci95: tuple[float, float]-
Two-sided 95% confidence interval on
delta.
instrument.bootstrap_ci()
Nonparametric bootstrap 95% CI of the AUC (E27 harness resampling).
Usage
instrument.bootstrap_ci(scores, labels, *, iters=2000, seed=1729)Resamples row indices with replacement iters times, keeps resamples that retain both classes, and returns the 2.5th and 97.5th percentiles of the bootstrap AUC distribution.
Parameters
scores: list[float]-
Per-row scores (higher = positive class).
labels: list[int]-
Binary labels aligned with
scores. iters: int = 2000-
Number of bootstrap resamples.
seed: int = 1729- Seed for the resampling RNG.
Returns
tuple[float, float]-
The
(lo, hi)bounds of the bootstrap 95% confidence interval.
instrument.mde()
Minimum detectable effect for a two-sided z-test at the given standard error.
Usage
instrument.mde(se, *, alpha=0.05, power=0.8)Computes (z_{1-alpha/2} + z_{power}) * se with the normal quantiles from the standard library; the defaults (alpha=0.05, power=0.8) reproduce the card’s 2.8016 * se constant to four decimals.
Parameters
se: float-
Standard error of the estimate under test.
alpha: float = 0.05-
Two-sided significance level.
power: float = 0.8- Desired statistical power.
Returns
float-
The smallest effect the test can resolve at
alphaandpower.
instrument.actionable()
Apply the instrument card’s two-part actionability rule to an AUC delta.
Usage
instrument.actionable(delta, se_delta, frame_mde)The delta is actionable when its two-sided 95% confidence interval excludes zero and its magnitude clears the frame’s minimum detectable effect.
Parameters
delta: float-
Observed AUC delta.
se_delta: float-
Covariance-aware standard error of
delta. frame_mde: float- The frame’s minimum detectable effect.
Returns
bool-
Truewhen the delta is both significant and large enough to act on.
instrument.InstrumentCard
The measurement-instrument card: the frame noise floors comparative verdicts consult.
Usage
instrument.InstrumentCard(mde, stopping, mde_paired)Read from the sidecar JSON at INSTRUMENT_CARD (~/.cc-steer/experiments/instrument-card-v1.json) via load(). The paired path measures its own MDE from the DeLong standard error, so the card matters exactly where pairing is impossible: mde is the unpaired floor of the n=628 watcher frame — the fallback when a counterpart’s per-row probs were never persisted — with the other frames’ floors under stopping.
Attributes
mde: float-
The unpaired minimum detectable effect of the watcher frame.
stopping: Mapping[str, float]-
The per-frame unpaired floors (mde,
mde_gate_frame,mde_golden). mde_paired: Mapping[str, object]- The card’s paired-MDE reference table with measured-rho examples.
Methods
| Name | Description |
|---|---|
| load() |
Read the card sidecar; path overrides INSTRUMENT_CARD.
|
load()
Read the card sidecar; path overrides INSTRUMENT_CARD.
Usage
load(path=None)instrument.Comparison
A card-governed comparative verdict between two checkpoints’ AUCs on one frame.
Usage
instrument.Comparison(auc_a, auc_b, delta, mde, actionable, verdict, paired)verdict speaks the card’s language: an actionable gain or regression names the delta, confidence interval, and MDE it cleared; a sub-threshold delta reads within noise floor (MDE <value>) and is never a win, a loss, or a rejection.
Attributes
auc_a: float-
AUC of the first arm (the incumbent or baseline).
auc_b: float-
AUC of the second arm (the candidate).
delta: float-
auc_b - auc_a. mde: float-
The minimum detectable effect the verdict applied —
2.8016 * se_deltaon the paired path, the card’s unpaired frame floor on the fallback. actionable: bool-
Whether the 95% CI excludes zero and
|delta| >= mde. verdict: str-
The card-rule verdict line.
paired: PairedDeLong | None-
The full paired DeLong record, or
Noneon the unpaired fallback.
instrument.paired_verdict()
The card rule over a paired DeLong record — the production comparative verdict.
Usage
instrument.paired_verdict(paired)The MDE is measured, not assumed: 2.8016 times the record’s covariance-aware se_delta, which already carries the observed pairing correlation. The verdict line names the delta, CI, rho, and MDE when actionable, and reports a sub-threshold delta as within the noise floor — never a win or a loss.
Parameters
paired: PairedDeLong- The paired comparison from paired_delong() over both arms’ persisted per-row probs on the identical frame.
Returns
The: Comparison- class:Comparison carrying the applied MDE and the verdict line.
instrument.unpaired_verdict()
The card rule when a counterpart’s per-row probs were never persisted.
Usage
instrument.unpaired_verdict(auc_a, auc_b, *, frame_mde)With no pairing there is no measured rho, so the card’s unpaired frame floor is the MDE. |delta| >= frame_mde already implies the unpaired 95% CI excludes zero (the MDE sits at 2.8016 standard errors, the CI bound at 1.96), so the card’s two-part rule reduces to the MDE test.
Parameters
auc_a: float-
Point AUC of the first arm (the incumbent or baseline).
auc_b: float-
Point AUC of the second arm (the candidate).
frame_mde: float- The frame’s unpaired MDE from InstrumentCard.
Returns
The: Comparison-
class:Comparison with
paired=None.
retrain.sweep.ArmScore
One arm’s frozen-frame measurement: the scalar metric and the per-row probs behind it.
Usage
retrain.sweep.ArmScore()Attributes
metric: float-
The arm’s sentinel AUC over the frozen frame.
probs: dict[str, float]-
Per-row
P(NO_STEER)keyed by frame row id — the vector any later paired comparison against another arm reads. frame_digest: str- The frozen frame’s order-invariant content digest.
retrain.sweep.compare_score_reports()
Card-governed verdict between two arms’ score reports — paired DeLong as the production default.
Usage
retrain.sweep.compare_score_reports(
report_a, report_b, *, eval_root=None, card_path=None
)When both reports carry per-row probs, the verdict is a paired DeLong test on the frozen watcher frame with measured rho (cc_steer.instrument.paired_verdict()). A report predating per-row persistence (no probs key) falls back to the point-AUC delta against the unpaired watcher-frame MDE read from the instrument-card sidecar (cc_steer.instrument.INSTRUMENT_CARD; card_path overrides). A sub-threshold delta is reported as within the noise floor, never as a win or a loss.
Parameters
report_a: Path-
Path to the baseline arm’s
SCORE_REPORT_FILE. report_b: Path-
Path to the candidate arm’s
SCORE_REPORT_FILE. eval_root: Path | None = None-
Frozen-eval root override for loading the frame the probs pair on.
card_path: Path | None = None- Instrument-card sidecar override for the unpaired fallback floor.
Returns
The: instrument.Comparison-
class:
~cc_steer.instrument.Comparisonwithdelta = b - a.
Raises
FrameMismatch- A report’s probs were scored against a different frame digest.