Primus Decision 0.1 · Experiments

Every measurement, with its class

Post-release experiments on Decision 0.1, with their methods and evidence classes. These diagnostics preserve the original weights and sealed record; new versions and task studies use separately reported evaluations.

M measured by usR reproduced by usX reported externally

E01M, measured by us

The sealed result checked with a separate metric implementation

A separately written metric implementation used by AAME, applied to a fresh run of the verified release bytes.

Every recorded cell of the sealed result is reproduced; the largest difference over 188 cells is 1.1e-16.

workflowaccuracyBrierNLLECEscore MAE
Agent-trace observability0.7320.0560.7010.1600.210
Customer service0.7280.0920.6920.1120.297
Invoice processing0.8080.0500.4740.0970.280
Security incidents0.7360.0380.6800.1450.311

Evaluation context: One frozen model and one sealed evaluation record.

E02M, measured by us

Case-level bootstrap intervals

2,000 resamples of the 400 test cases, whole cases kept together, seed 0, 95 % percentile intervals.

metricpoint95 % intervalLaya, reproduced
accuracy0.7510.729 – 0.7740.766
NLL0.6370.612 – 0.6620.707
Brier0.0590.052 – 0.0660.066
ECE0.1270.108 – 0.1460.213
score MAE0.2750.256 – 0.2930.242

Evaluation context: Case-sampling intervals for one ensemble, with one seed per member.

E03M, measured by us

Calibration and abstention

Reliability over the 2,000 decisions (15 bins, confidence = largest probability) and selective prediction with the raw confidence.

outputmean confidenceaccuracyECE
raw (the default)0.6250.7510.127
temperature profile applied0.7800.7510.040

The temperature profile lowers ECE and NLL at the cost of Brier (0.059 → 0.114) and score MAE (0.275 → 0.320); it is off by default.

Accuracy on the answered decisions vs coverage
0.700.800.901.0010%30%50%70%90%coverage (most confident first)accuracy10 %: 1.00020 %: 0.98830 %: 0.95540 %: 0.93550 %: 0.90960 %: 0.87870 %: 0.84180 %: 0.81690 %: 0.783100 %: 0.7510.909 on the most confident half

90.9% accuracy on the most-confident 50% of decisions. AURC 0.102.

Evaluation context: Thresholds evaluated on this test split; set application thresholds from recorded outcomes.

E04M, measured by us

Latency, throughput and memory: six runs

Public runtime on the release bytes, 4-vCPU Intel Xeon 2.10 GHz, 2 threads, model loaded, one five-decision case at a time; two methods.

runmethodp50p95batch 32 / casepeak RSS
Ascript, idle machine115 ms134 ms36 ms1,563 MB
Bscript, packaged copy139 ms182 ms60 ms1,499 MB
Cprobe130 ms287 ms49 ms1,403 MB
Dprobe128 ms280 ms46 ms1,403 MB
Eprobe137 ms285 ms39 ms1,404 MB
Fprobe, final tree162 ms316 ms65 ms1,403 MB

Run A (115 ms) is the fastest of six runs, on an idle machine; it is never quoted alone. 115–162 ms p50 across six recorded runs; read any single figure as ±25 %. Same-machine comparison with Laya on the benchmark page.

Evaluation context: CPU measurements on the stated machine and thread settings.

E05M, measured by us

Robustness and consistency

Synthetic, seed-deterministic rewrites of the test inputs; accuracy before and after on the same subset.

Accuracy change per transform (points)
option order of every `choice` question shuffledoption order of every `choice` question shuffled: 0.0 points; 100.0 % same answer0.0instructions paraphrasedinstructions paraphrased: -0.1 points; 96.7 % same answer-0.1synonyms in state valuessynonyms in state values: -0.1 points; 99.6 % same answer-0.1dates in long formatdates in long format: -0.5 points; 96.9 % same answer-0.5dates in US formatdates in US format: -0.4 points; 97.0 % same answer-0.4three noise fields addedthree noise fields added: -0.8 points; 96.6 % same answer-0.8five noise fields addedfive noise fields added: -1.1 points; 95.8 % same answer-1.1nesting layout changednesting layout changed: -1.3 points; 95.6 % same answer-1.3keys renamed to kebab-casekeys renamed to kebab-case: -1.4 points; 96.1 % same answer-1.4relation statements rewrittenrelation statements rewritten: -1.6 points; 94.8 % same answer-1.6keys replaced by synonymskeys replaced by synonyms: -4.1 points; 90.0 % same answer-4.1keys renamed to camelCasekeys renamed to camelCase: -5.4 points; 89.8 % same answer-5.4

Answers are invariant under the measured option-order shuffle. Identical runs agree bit for bit; batched and one-at-a-time inference never change a decision.

Evaluation context: Named, controlled input rewrites on the original test cases.

E06R, reproduced by us

Laya's checkpoint, reproduced on the same harness

The released checkpoint (Hub revision dd079950, 843 MB weights) run through its author's own predict call on the pinned test split, CPU, 4 threads, scored two ways.

metricpublished cardauthor's harnessour protocol
accuracy0.7660.7660.766
soft accuracy0.4710.5090.509
Brier0.0620.0660.066
ECE0.2130.2130.213
score MAE0.2420.2420.242
NLL——0.707
within-1—0.9950.995

Accuracy, ECE, score MAE and every per-workflow accuracy reproduce the card exactly; the card's soft accuracy and Brier do not, so the reproduced values are the ones compared.

Evaluation context: The identified Laya checkpoint and CPU configuration.

E07M, measured by us

What the experiments establish

Recorded task performance, operating profile and robustness of the frozen component.

Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison.

The public runtime also records deterministic predictions, answer invariance under choice-option shuffling, and confidence-based selection. The full benchmark page reports methods and task detail.

Full results and evaluation protocol

E08M, measured by us

Primus for invoice decisions

26 September 2026 · three controlled synthetic experiments: pilot, answer-format audit and broader invoice test.

The released weights answered 26/32 structured reconciliation questions correctly (81.25%) in the broader test without retraining. The study publishes structured and narrative conditions, duplicate-ID questions, the TF-IDF classifier, Qwen3-1.7B and explicit rules. Its preprocessing and prediction records accompany the result.

Read the case study and evidence