acsi
Certification Report
oss-issue-summary · issued 2026-07-20T05:05:45Z
BLOCK candidate claude-sonnet-5 vs. baseline claude-opus-4-1

BLOCK at n=279, covering 100.0% of production template distribution, 95% CI [40.1%, 52.0%]. This certifies the sampled workload against the stated assertions; it does not certify unsampled inputs.

In plain terms: we tested 279 real inputs; between 40.1% and 52.0% of the new model's answers differed beyond normal randomness.

149 of 279 sampled pairs (53.4%) regressed — 149 by a failing assertion.

128 of 279 pairs (45.9%) were unresolved — the panel could not decide: 66 unresolved-only, 62 also assertion-flagged. Counted conservatively toward the verdict, never as a judge-flagged regression.

How to read this certificate

A pair is one real input answered twice — once by the current model (baseline), once by the candidate you are considering.

A pair regressed when the candidate's answer broke a required rule (an assertion) or a panel of independent judges graded it worse than the baseline's.

A pair is unresolved when the judges could not reach enough valid, agreeing votes to decide it; these are counted against the candidate so the verdict stays on the safe side.

BLOCK means do not switch to the candidate until the named failures below are fixed. PASS means the candidate is safe to switch to for the inputs tested here.

Pass criteria

All three must pass to certify. This run is a BLOCK.

BLOCK

Critical assertion failures

No sampled response may fail a critical check (schema, safety). Even one failure fails the run.

Observed
149 critical failure(s) observed
Allowed
0 allowed
BLOCK

Regression vs. baseline noise

The candidate must not disagree with the baseline more than the baseline already disagrees with itself, plus a small tolerance.

Observed
The new model differed on up to 52.0% of inputs. The old model's own randomness allows 8.5%. 128 undecided pairs are counted against the new model to be safe.
Allowed
8.5% allowed
BLOCK

Critical failure concentration

No single critical failure pattern may cover more than a small share of the sampled traffic.

Observed
one pattern at 2.2% of traffic; one pattern at 3.9% of traffic; one pattern at 29.0% of traffic; one pattern at 18.3% of traffic
Allowed
1.0% allowed
Failure clusters

Regressions grouped by mechanism. Expand any cluster for the exact input, the candidate's failing response, and why it failed.

The 66 unresolved-only pairs split into 2 clusters below (65 + 1 = 66).

Unresolved — panel could not decide unresolved 65 pairs · 23.3%
Input prompt
Title: [DevTools Bug]: Profiler's "What changed" context becomes unusable when inspecting commits with long render histories

Body: ### Website or app

React App with a component that re-renders frequently

### Repro steps

**Note:** This is more like an accessibility and usability issue rather t...
Input prompt
Title: [Flake][sig-scheduling] k8s.io/kubernetes/test/integration/scheduler.preemption

Body: ### Which jobs are flaking?

* sig-release-master-blocking
* integration-arm64-master

### Which tests are flaking?

* [[sig-scheduling] k8s.io/kubernetes/test/integration/scheduler.preemption](https://p...
Input prompt
Title: [ICE] while building proc-macro2's build script

Body: