PUBLIC EVIDENCE DASHBOARD

A verified layer can help a small LLM — within a measurable domain.

Paired comparisons on mechanically verifiable tasks. Diagnostics, pretests and invalidated campaigns are never pooled with primary evidence.

Static audited snapshot loaded · 4 primary comparisons · data 2026-08-02

PRIMARY EVIDENCE

LLM alone vs LLM + Nexus Safe

Accuracy is directly labelled. No hover is required.

47audited questions / questions auditées
2/4comparisons classified as GAIN
+68.1 ppdescriptive aggregate gap
LLM aloneLLM + Nexus Safe

PAIRED DECISION

Gain, loss or inconclusive result

A higher score is not enough: the decision accounts for discordant pairs and sample size.

GAIN

Granite 3.3 2B Instruct

15 wins · 0 losses

Two-sided exact test: p=0.00006103515625. Status: AUDITED.

GAIN

Gemma 3 12B IT QAT

12 wins · 0 losses

Two-sided exact test: p=0.00048828125. Status: AUDITED.

INCONCLUSIVE

Llama 3.1 8B

5 wins · 0 losses

Two-sided exact test: p=0.0625. Status: AUDITED.

INCONCLUSIVE

Granite 3.3 2B Instruct · code

0 wins · 0 losses

Two-sided exact test: p=1. Status: AUDITED_SMALL_SAMPLE.

PORTABLE DATA

Complete primary comparison table

JSON · CSV

Audited results published without questions, answers or targets.
ModelDomainnAloneNexusΔDecisionexact p
Granite 3.3 2B Instructmixed verifiable183/18 (16.7 %)18/18 (100 %)+83.3 ppGAIN0.00006103515625
Gemma 3 12B IT QATmixed verifiable183/18 (16.7 %)15/18 (83.3 %)+66.7 ppGAIN0.00048828125
Llama 3.1 8Bmixed verifiable61/6 (16.7 %)6/6 (100 %)+83.3 ppINCONCLUSIVE0.0625
Granite 3.3 2B Instructcode51/5 (20 %)1/5 (20 %)+0.0 ppINCONCLUSIVE1

AUDIT LEDGER

Informative evidence that is not a product claim

These campaigns remain visible, but are excluded from primary averages.

Excluded or running campaigns

HOW TO READ THIS

What these numbers do — and do not — show

They show a reproducible gain on some verifiable problems when Nexus has an appropriate executor or verifier.

They do not show universal improvement in reasoning, creativity, code or open-domain knowledge.

Publication rule: previously observed answers are never reused to score a new campaign; sealed sets are not used for training.