Primus Decision 0.1 · Benchmarks

Benchmarks

One sealed run on the official test split of LocalLLaMA/typed-decisions: 400 cases, 2,000 decisions, recorded in one sealed release evaluation after the model was frozen. Three evidence classes are kept apart on this page.

M measured by usR reproduced by usX reported externally

  • Open source · Apache-2.0
  • Local CPU inference
  • Non-transformer
What these scores measure
The benchmark is synthetic and teacher-labelled: its gold labels are the mean of three samples from a 4B-class teacher. These scores measure agreement with those reference labels. They describe this evaluation of the frozen component; new tasks and versions establish their own results.

The result against Laya, reproduced

Same test split, same harness, raw probabilities. Metrics and measurement conditions are reported below. M, measured by us R, reproduced by us

Accuracy
0.751 vs 0.766
Recorded accuracy
Primus vs Laya. Official test split; one frozen Primus ensemble.
NLL
0.637 vs 0.707
Primus lower
Primus vs Laya. Lower is better.
Raw ECE
0.127 vs 0.213
Primus lower
Primus vs Laya. Lower is better.
Score MAE
0.275 vs 0.242
Laya lower
Primus vs Laya. Mean absolute error for ordinal questions; lower is better.

Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison M, measured by us R, reproduced by us Sources: SEALED_RESULT.json, EXPERIMENTS.md §2, §4, §6

Reproduce the sealed result. The frozen model, the metric definitions (primus_decision/metrics.py) and the protocol are in the release; the dataset loader and evaluation driver are not. Re-scoring means reading the official test split at the pinned revision and applying PROTOCOL.md; the numbers to expect are in REPRODUCIBILITY.md §6.

BENCHMARKS.md

Direct comparison

Measured or reproduced on the same harness. Definitions in the release's PROTOCOL.md. No row is marked as a winner; the Primus raw row is the reference row.

Systems measured on the same harness
systemclassaccuracy ↑soft accuracy ↑Brier ↓ECE ↓NLL ↓score MAE ↓within-1 ↑
Primus Decision 0.1, rawM, measured by us0.7510.5630.0590.1270.6370.2750.984
Primus Decision 0.1, temperature profile (optional, off by default)M, measured by us0.7510.6180.1140.0400.5680.3200.970
Laya-typed-decisions, reproduced on the same harnessR, reproduced by us0.7660.5090.0660.2130.7070.2420.995

Externally reported systems and evaluation settings

Figures reported by the named publishers. Generalists answer in a zero-shot setting; specialists are fitted to the benchmark workflows. Their training histories and deployment configurations provide context for reading these results. Source: BENCHMARKS.md

Externally reported figures
systemclassaccuracy ↑soft accuracy ↑Brier ↓ECE ↓NLL ↓score MAE ↓within-1 ↑
Laya-typed-decisions, published model cardX, reported externally0.7660.471 †0.062 †0.213—0.242—
TypeSafe Jev 1.13.0, generalist, publishedX, reported externally0.7270.5800.1480.144—0.3910.952
meraGPT Decider 1, generalist, zero-shot, proprietary, publishedX, reported externally0.7680.6080.0520.180—0.2190.984

† Laya's published soft accuracy and Brier do not reproduce (0.509 and 0.066 under both its author's harness and ours); read the reproduced row above. X, reported externally published figures are quoted from their publishers and not re-measured.

Case-sampling uncertainty

95 % case-level bootstrap intervals (2,000 resamples of the 400 cases, seed 0) for Primus, with the reproduced Laya figure as a dot. M, measured by us R, reproduced by us Source: EXPERIMENTS.md §2, §6

Primus bar and dot: series 1; Laya dot: series 2. The intervals describe case sampling for one frozen ensemble, with one seed per member. They are not a paired comparison or an equivalence test.

By workflow and by question type

Accuracy by workflow (500 decisions each)
Primus Decision 0.1Laya, reproduced
agent-trace observability0.7300.732+1 of 500customer service0.7280.764−18 of 500invoice processing0.8040.808+2 of 500security incidents0.7360.766−15 of 500

Recorded per-workflow accuracies and decision-count differences, grouped into test cases.

Accuracy by question type
Primus Decision 0.1Laya, reproduced
choice0.7330.738+3 of 600noul0.8370.857−12 of 600score0.6960.723−21 of 800

Recorded results for categorical choice, yes/no probability and ordinal scoring.

Calibration and abstention

Raw probabilities are the default. The optional temperature profile changes calibration and probability-error metrics as recorded below.

0.127
Raw ECE 0.127, the default output

Mean confidence 0.625 against accuracy 0.751: under-confident. Maximum calibration error 0.328. M, measured by us Source: EXPERIMENTS.md §3

0.040
Calibrated ECE 0.040, optional temperature profile

Cost: Brier 0.059 → 0.114, score MAE 0.275 → 0.320; NLL 0.637 → 0.568. Fitted on the validation part only; off by default. M, measured by us Source: EXPERIMENTS.md §3

Accuracy on the answered decisions vs coverage (raw confidence, most confident first)
0.700.800.901.0010%30%50%70%90%coverage (most confident first)accuracy10 %: 1.00020 %: 0.98830 %: 0.95540 %: 0.93550 %: 0.90960 %: 0.87870 %: 0.84180 %: 0.81690 %: 0.783100 %: 0.7510.909 on the most confident half

Area under the risk–coverage curve 0.102 (raw). 90.9% accuracy on the most-confident 50% of decisions.

confidence thresholddecisions answeredaccuracy on them
0.573 %0.837
0.647 %0.915
0.728 %0.960
0.818 %0.992
0.911 %0.996

These figures describe this test split; a deployment sets its own threshold on its own Decision History. Source: EXPERIMENTS.md §3

Latency, throughput, memory

All runs use the public runtime on the release bytes, a 4-vCPU Intel Xeon at 2.10 GHz, 2 threads and the model already loaded; one five-decision case at a time; page cache dropped before loading. Two methods: the packaged script (100 distinct test cases after a warm-up call) and the probe (60 timings over a 120-case fixture). M, measured by us Source: EXPERIMENTS.md §4

runmethoddate (UTC)loadfirst inferenceone case p50p95batch of 32, per casewhole splitRSS after loadpeak RSS
Ascript, idle machine2026-09-23 08:370.28 s370 ms115 ms134 ms36 ms11.7 s (34 cases/s)602 MB1,563 MB
Bscript, packaged copy2026-09-23 12:291.53 s663 ms139 ms182 ms60 ms20.1 s (20 cases/s)578 MB1,499 MB
Cprobe2026-09-22 19:392.07 s836 ms130 ms287 ms49 ms–539 MB1,403 MB
Dprobe2026-09-22 20:200.18 s246 ms128 ms280 ms46 ms–539 MB1,403 MB
Eprobe2026-09-22 22:540.17 s314 ms137 ms285 ms39 ms–539 MB1,404 MB
Fprobe, final tree2026-09-23 11:280.84 s1005 ms162 ms316 ms65 ms–539 MB1,403 MB

Run A (115 ms) is the fastest of six runs on an idle machine; run F (162 ms) the slowest, on the final release build. Neither is quoted alone.

One-case latency: 115–162 ms p50 across six recorded runs, about 140 ms; read any single figure as ±25 %. The spread comes from the shared cloud machine, not from the model: the answers are deterministic.

One five-decision case, p50, same machine (ms)
Primus Decision 0.1, 2 threads (range marker: 115–162 ms across six runs)Laya-typed-decisions, reproduced, 4 threads
Primus Decision 0.1Primus Decision 0.1: 140 ms~140 ms (115–162)Laya, reproducedLaya, reproduced: 3,291 ms3,291 ms

Laya's released checkpoint on the same machine, CPU, 4 threads (its runtime's CPU default): p50 3,291 ms, p95 4,656 ms, peak RSS 3,029 MB; 20–29× the Primus per-case figures depending on the Primus run. Its author reports GPU figures; none were measured here. Reproduced by us. Source: EXPERIMENTS.md §4

The sealed runner's own 134.04 ms per case is a different quantity: loading plus the whole split in batches of 32, divided by 400. It is not a per-case p50. Source: BENCHMARKS.md

36–65 ms
per case in batches of 32 M, measured by us
20–34
cases per second over the whole split (100–171 decisions) M, measured by us
0.5–0.6 GB
resident after loading M, measured by us
1.4–1.6 GB
peak during inference, almost all of it the featurizers M, measured by us

Size

3,715,074
neural parameters: 1,932,609 S4D + 1,782,465 GRU, all active on every prediction M, measured by us
15 MB
float32 weights of the two networks M, measured by us
2 × 71.4 MB
LSA featurizers (word and character TF-IDF, 256-d SVD), fitted tables M, measured by us
157.9 MB
installed; 114.5 MB tar.gz M, measured by us

421M → 3.7M neural parameters (~113× fewer) by Laya's model card; Laya's weights file is 843 MB; Primus's original complete installed release is 157.9 MB, including its fitted feature tables. These file measurements have different contents. X, reported externally M, measured by us Sources: MODEL_SIZE_AND_PARAMETERS.md, EXPERIMENTS.md §6

Robustness

Synthetic, seed-deterministic rewrites of the test inputs; accuracy before and after on the same subset. M, measured by us Source: EXPERIMENTS.md §5

Accuracy change per transform (points)
option order of every `choice` question shuffledoption order of every `choice` question shuffled: 0.0 points; 100.0 % same answer0.0instructions paraphrasedinstructions paraphrased: -0.1 points; 96.7 % same answer-0.1synonyms in state valuessynonyms in state values: -0.1 points; 99.6 % same answer-0.1dates in long formatdates in long format: -0.5 points; 96.9 % same answer-0.5dates in US formatdates in US format: -0.4 points; 97.0 % same answer-0.4three noise fields addedthree noise fields added: -0.8 points; 96.6 % same answer-0.8five noise fields addedfive noise fields added: -1.1 points; 95.8 % same answer-1.1nesting layout changednesting layout changed: -1.3 points; 95.6 % same answer-1.3keys renamed to kebab-casekeys renamed to kebab-case: -1.4 points; 96.1 % same answer-1.4relation statements rewrittenrelation statements rewritten: -1.6 points; 94.8 % same answer-1.6keys replaced by synonymskeys replaced by synonyms: -4.1 points; 90.0 % same answer-4.1keys renamed to camelCasekeys renamed to camelCase: -5.4 points; 89.8 % same answer-5.4

Answers are invariant under the measured option-order shuffle (largest probability shift 3.7e-8). Wording, dates, noise fields and layout cost at most 1.3 points; kebab-case keys and rewritten relation statements 1.4 and 1.6. Deterministic: identical runs agree bit for bit; batched and one-at-a-time inference differ by at most 1.1e-7 and never change a decision.

Field names affect the lexical features. Key-synonym and camelCase rewrites changed accuracy by −4.1 and −5.4 points; evaluate the application's input mapping and adapt it or the model as needed.

Invoice decisions: a separate task study

The released weights answered 26/32 structured reconciliation questions correctly (81.25%) in the broader synthetic invoice test without retraining. See every task, presentation, baseline and evaluation condition.

Original evaluation protocol

The original release targets were fixed before the test split was opened. The record below preserves their outcome for Decision 0.1 and remains separate from subsequent experiments.

1
The rule was written first

Primary gate: accuracy above the reference. Secondary gate: Brier or raw ECE below it. Committed before the test split was read.

2
The model was sealed

Architecture, hyper-parameters and calibration frozen with content hashes. The original result was recorded in one sealed release evaluation.

3
The result was published as it came

The recorded accuracy target was not reached; the Brier/ECE target was reached. These are outcomes of this frozen experiment.

Accuracy above the reference · primary gate0.751 against 0.766 · not met
Brier below the reference · secondary gate0.059 against 0.062 · met
Raw ECE below the reference · secondary gate0.127 against 0.213 · met
Score MAE below the reference · informational, not part of the rule0.275 against 0.242 · not met
The sealed runner did not authorize its composite superiority claim over Laya. This historical rule describes the 0.1 comparison; new experiments and versions establish their own results.

Evaluation record

  • One frozen ensemble, with one seed per member and case-level sampling intervals.
  • Externally reported generalists listed with their evaluation settings.
  • Named input transformations, with measured accuracy changes.
  • Each future Primus release will carry its own evaluation.

Reproduce it

latency method on your machine
$ python -m pip install pandas pyarrow huggingface_hub
$ hf download LocalLLaMA/typed-decisions all/test-00000-of-00001.parquet --repo-type dataset --revision ea9306458d6e9563628369a3d1e72e362fb381d2 --local-dir typed-decisions
$ python examples/measure_latency.py --parquet typed-decisions/all/test-00000-of-00001.parquet

Run it in the release folder with its virtual environment active (see Download). The script also needs pandas and pyarrow, which requirements.txt does not list; hf fetches the official test split at the pinned revision.

dataset
LocalLLaMA/typed-decisions @ ea9306458d6e9563628369a3d1e72e362fb381d2
sealed test file
4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c
release tree
fa0406579f5b5338c2f44de0c42a0626d1a1e1a8