Primus Decision 0.1 is open source.Apache-2.0, weights and protocol on GitHub · Hugging Face
Primus Decision 0.1 · Experiments
Every measurement, with its class
Post-release experiments on Decision 0.1, with their methods and evidence classes. These diagnostics preserve the original weights and sealed record; new versions and task studies use separately reported evaluations.
● MM measured by us◐ RR reproduced by us○ XX reported externally
E01● MM, measured by us
The sealed result checked with a separate metric implementation
A separately written metric implementation used by AAME, applied to a fresh run of the verified release bytes.
Every recorded cell of the sealed result is reproduced; the largest difference over 188 cells is 1.1e-16.
workflow
accuracy
Brier
NLL
ECE
score MAE
Agent-trace observability
0.732
0.056
0.701
0.160
0.210
Customer service
0.728
0.092
0.692
0.112
0.297
Invoice processing
0.808
0.050
0.474
0.097
0.280
Security incidents
0.736
0.038
0.680
0.145
0.311
Evaluation context: One frozen model and one sealed evaluation record.
E02● MM, measured by us
Case-level bootstrap intervals
2,000 resamples of the 400 test cases, whole cases kept together, seed 0, 95 % percentile intervals.
metric
point
95 % interval
Laya, reproduced
accuracy
0.751
0.729 – 0.774
0.766
NLL
0.637
0.612 – 0.662
0.707
Brier
0.059
0.052 – 0.066
0.066
ECE
0.127
0.108 – 0.146
0.213
score MAE
0.275
0.256 – 0.293
0.242
Evaluation context: Case-sampling intervals for one ensemble, with one seed per member.
E03● MM, measured by us
Calibration and abstention
Reliability over the 2,000 decisions (15 bins, confidence = largest probability) and selective prediction with the raw confidence.
output
mean confidence
accuracy
ECE
raw (the default)
0.625
0.751
0.127
temperature profile applied
0.780
0.751
0.040
The temperature profile lowers ECE and NLL at the cost of Brier (0.059 → 0.114) and score MAE (0.275 → 0.320); it is off by default.
Accuracy on the answered decisions vs coverage
90.9% accuracy on the most-confident 50% of decisions. AURC 0.102.
Evaluation context: Thresholds evaluated on this test split; set application thresholds from recorded outcomes.
E04● MM, measured by us
Latency, throughput and memory: six runs
Public runtime on the release bytes, 4-vCPU Intel Xeon 2.10 GHz, 2 threads, model loaded, one five-decision case at a time; two methods.
run
method
p50
p95
batch 32 / case
peak RSS
A
script, idle machine
115 ms
134 ms
36 ms
1,563 MB
B
script, packaged copy
139 ms
182 ms
60 ms
1,499 MB
C
probe
130 ms
287 ms
49 ms
1,403 MB
D
probe
128 ms
280 ms
46 ms
1,403 MB
E
probe
137 ms
285 ms
39 ms
1,404 MB
F
probe, final tree
162 ms
316 ms
65 ms
1,403 MB
Run A (115 ms) is the fastest of six runs, on an idle machine; it is never quoted alone. 115–162 ms p50 across six recorded runs; read any single figure as ±25 %. Same-machine comparison with Laya on the benchmark page.
Evaluation context: CPU measurements on the stated machine and thread settings.
E05● MM, measured by us
Robustness and consistency
Synthetic, seed-deterministic rewrites of the test inputs; accuracy before and after on the same subset.
Accuracy change per transform (points)
Answers are invariant under the measured option-order shuffle. Identical runs agree bit for bit; batched and one-at-a-time inference never change a decision.
Evaluation context: Named, controlled input rewrites on the original test cases.
E06◐ RR, reproduced by us
Laya's checkpoint, reproduced on the same harness
The released checkpoint (Hub revision dd079950, 843 MB weights) run through its author's own predict call on the pinned test split, CPU, 4 threads, scored two ways.
metric
published card
author's harness
our protocol
accuracy
0.766
0.766
0.766
soft accuracy
0.471
0.509
0.509
Brier
0.062
0.066
0.066
ECE
0.213
0.213
0.213
score MAE
0.242
0.242
0.242
NLL
—
—
0.707
within-1
—
0.995
0.995
Accuracy, ECE, score MAE and every per-workflow accuracy reproduce the card exactly; the card's soft accuracy and Brier do not, so the reproduced values are the ones compared.
Evaluation context: The identified Laya checkpoint and CPU configuration.
E07● MM, measured by us
What the experiments establish
Recorded task performance, operating profile and robustness of the frozen component.
Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison.
The public runtime also records deterministic predictions, answer invariance under choice-option shuffling, and confidence-based selection. The full benchmark page reports methods and task detail.
26 September 2026 · three controlled synthetic experiments: pilot, answer-format audit and broader invoice test.
The released weights answered 26/32 structured reconciliation questions correctly (81.25%) in the broader test without retraining. The study publishes structured and narrative conditions, duplicate-ID questions, the TF-IDF classifier, Qwen3-1.7B and explicit rules. Its preprocessing and prediction records accompany the result.