Granite 3.3 2B Instruct
15 wins · 0 lossesTwo-sided exact test: p=0.00006103515625. Status: AUDITED.
PUBLIC EVIDENCE DASHBOARD
Paired comparisons on mechanically verifiable tasks. Diagnostics, pretests and invalidated campaigns are never pooled with primary evidence.
PRIMARY EVIDENCE
Accuracy is directly labelled. No hover is required.
PAIRED DECISION
A higher score is not enough: the decision accounts for discordant pairs and sample size.
Two-sided exact test: p=0.00006103515625. Status: AUDITED.
Two-sided exact test: p=0.00048828125. Status: AUDITED.
Two-sided exact test: p=0.0625. Status: AUDITED.
Two-sided exact test: p=1. Status: AUDITED_SMALL_SAMPLE.
| Model | Domain | n | Alone | Nexus | Δ | Decision | exact p |
|---|---|---|---|---|---|---|---|
| Granite 3.3 2B Instruct | mixed verifiable | 18 | 3/18 (16.7 %) | 18/18 (100 %) | +83.3 pp | GAIN | 0.00006103515625 |
| Gemma 3 12B IT QAT | mixed verifiable | 18 | 3/18 (16.7 %) | 15/18 (83.3 %) | +66.7 pp | GAIN | 0.00048828125 |
| Llama 3.1 8B | mixed verifiable | 6 | 1/6 (16.7 %) | 6/6 (100 %) | +83.3 pp | INCONCLUSIVE | 0.0625 |
| Granite 3.3 2B Instruct | code | 5 | 1/5 (20 %) | 1/5 (20 %) | +0.0 pp | INCONCLUSIVE | 1 |
AUDIT LEDGER
These campaigns remain visible, but are excluded from primary averages.
HOW TO READ THIS
They show a reproducible gain on some verifiable problems when Nexus has an appropriate executor or verifier.
They do not show universal improvement in reasoning, creativity, code or open-domain knowledge.
Publication rule: previously observed answers are never reused to score a new campaign; sealed sets are not used for training.