Benchmarks
One sealed run on the official test split of LocalLLaMA/typed-decisions: 400 cases, 2,000 decisions, recorded in one sealed release evaluation after the model was frozen. Three evidence classes are kept apart on this page.
M measured by usR reproduced by usX reported externally
- Open source · Apache-2.0
- Local CPU inference
- Non-transformer
The result against Laya, reproduced
Same test split, same harness, raw probabilities. Metrics and measurement conditions are reported below. M, measured by us R, reproduced by us
Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison M, measured by us R, reproduced by us Sources: SEALED_RESULT.json, EXPERIMENTS.md §2, §4, §6
Reproduce the sealed result. The frozen model, the metric definitions (primus_decision/metrics.py) and the protocol are in the release; the dataset loader and evaluation driver are not. Re-scoring means reading the official test split at the pinned revision and applying PROTOCOL.md; the numbers to expect are in REPRODUCIBILITY.md §6.
BENCHMARKS.mdDirect comparison
Measured or reproduced on the same harness. Definitions in the release's PROTOCOL.md. No row is marked as a winner; the Primus raw row is the reference row.
| system | class | accuracy ↑ | soft accuracy ↑ | Brier ↓ | ECE ↓ | NLL ↓ | score MAE ↓ | within-1 ↑ |
|---|---|---|---|---|---|---|---|---|
| Primus Decision 0.1, raw | M, measured by us | 0.751 | 0.563 | 0.059 | 0.127 | 0.637 | 0.275 | 0.984 |
| Primus Decision 0.1, temperature profile (optional, off by default) | M, measured by us | 0.751 | 0.618 | 0.114 | 0.040 | 0.568 | 0.320 | 0.970 |
| Laya-typed-decisions, reproduced on the same harness | R, reproduced by us | 0.766 | 0.509 | 0.066 | 0.213 | 0.707 | 0.242 | 0.995 |
Externally reported systems and evaluation settings
Figures reported by the named publishers. Generalists answer in a zero-shot setting; specialists are fitted to the benchmark workflows. Their training histories and deployment configurations provide context for reading these results. Source: BENCHMARKS.md
| system | class | accuracy ↑ | soft accuracy ↑ | Brier ↓ | ECE ↓ | NLL ↓ | score MAE ↓ | within-1 ↑ |
|---|---|---|---|---|---|---|---|---|
| Laya-typed-decisions, published model card | X, reported externally | 0.766 | 0.471 † | 0.062 † | 0.213 | — | 0.242 | — |
| TypeSafe Jev 1.13.0, generalist, published | X, reported externally | 0.727 | 0.580 | 0.148 | 0.144 | — | 0.391 | 0.952 |
| meraGPT Decider 1, generalist, zero-shot, proprietary, published | X, reported externally | 0.768 | 0.608 | 0.052 | 0.180 | — | 0.219 | 0.984 |
† Laya's published soft accuracy and Brier do not reproduce (0.509 and 0.066 under both its author's harness and ours); read the reproduced row above. X, reported externally published figures are quoted from their publishers and not re-measured.
Case-sampling uncertainty
95 % case-level bootstrap intervals (2,000 resamples of the 400 cases, seed 0) for Primus, with the reproduced Laya figure as a dot. M, measured by us R, reproduced by us Source: EXPERIMENTS.md §2, §6
Primus bar and dot: series 1; Laya dot: series 2. The intervals describe case sampling for one frozen ensemble, with one seed per member. They are not a paired comparison or an equivalence test.
By workflow and by question type
Recorded per-workflow accuracies and decision-count differences, grouped into test cases.
Recorded results for categorical choice, yes/no probability and ordinal scoring.
Calibration and abstention
Raw probabilities are the default. The optional temperature profile changes calibration and probability-error metrics as recorded below.
Mean confidence 0.625 against accuracy 0.751: under-confident. Maximum calibration error 0.328. M, measured by us Source: EXPERIMENTS.md §3
Cost: Brier 0.059 → 0.114, score MAE 0.275 → 0.320; NLL 0.637 → 0.568. Fitted on the validation part only; off by default. M, measured by us Source: EXPERIMENTS.md §3
Area under the risk–coverage curve 0.102 (raw). 90.9% accuracy on the most-confident 50% of decisions.
| confidence threshold | decisions answered | accuracy on them |
|---|---|---|
| 0.5 | 73 % | 0.837 |
| 0.6 | 47 % | 0.915 |
| 0.7 | 28 % | 0.960 |
| 0.8 | 18 % | 0.992 |
| 0.9 | 11 % | 0.996 |
These figures describe this test split; a deployment sets its own threshold on its own Decision History. Source: EXPERIMENTS.md §3
Latency, throughput, memory
All runs use the public runtime on the release bytes, a 4-vCPU Intel Xeon at 2.10 GHz, 2 threads and the model already loaded; one five-decision case at a time; page cache dropped before loading. Two methods: the packaged script (100 distinct test cases after a warm-up call) and the probe (60 timings over a 120-case fixture). M, measured by us Source: EXPERIMENTS.md §4
| run | method | date (UTC) | load | first inference | one case p50 | p95 | batch of 32, per case | whole split | RSS after load | peak RSS |
|---|---|---|---|---|---|---|---|---|---|---|
| A | script, idle machine | 2026-09-23 08:37 | 0.28 s | 370 ms | 115 ms | 134 ms | 36 ms | 11.7 s (34 cases/s) | 602 MB | 1,563 MB |
| B | script, packaged copy | 2026-09-23 12:29 | 1.53 s | 663 ms | 139 ms | 182 ms | 60 ms | 20.1 s (20 cases/s) | 578 MB | 1,499 MB |
| C | probe | 2026-09-22 19:39 | 2.07 s | 836 ms | 130 ms | 287 ms | 49 ms | – | 539 MB | 1,403 MB |
| D | probe | 2026-09-22 20:20 | 0.18 s | 246 ms | 128 ms | 280 ms | 46 ms | – | 539 MB | 1,403 MB |
| E | probe | 2026-09-22 22:54 | 0.17 s | 314 ms | 137 ms | 285 ms | 39 ms | – | 539 MB | 1,404 MB |
| F | probe, final tree | 2026-09-23 11:28 | 0.84 s | 1005 ms | 162 ms | 316 ms | 65 ms | – | 539 MB | 1,403 MB |
Run A (115 ms) is the fastest of six runs on an idle machine; run F (162 ms) the slowest, on the final release build. Neither is quoted alone.
One-case latency: 115–162 ms p50 across six recorded runs, about 140 ms; read any single figure as ±25 %. The spread comes from the shared cloud machine, not from the model: the answers are deterministic.
Laya's released checkpoint on the same machine, CPU, 4 threads (its runtime's CPU default): p50 3,291 ms, p95 4,656 ms, peak RSS 3,029 MB; 20–29× the Primus per-case figures depending on the Primus run. Its author reports GPU figures; none were measured here. Reproduced by us. Source: EXPERIMENTS.md §4
The sealed runner's own 134.04 ms per case is a different quantity: loading plus the whole split in batches of 32, divided by 400. It is not a per-case p50. Source: BENCHMARKS.md
Size
421M → 3.7M neural parameters (~113× fewer) by Laya's model card; Laya's weights file is 843 MB; Primus's original complete installed release is 157.9 MB, including its fitted feature tables. These file measurements have different contents. X, reported externally M, measured by us Sources: MODEL_SIZE_AND_PARAMETERS.md, EXPERIMENTS.md §6
Robustness
Synthetic, seed-deterministic rewrites of the test inputs; accuracy before and after on the same subset. M, measured by us Source: EXPERIMENTS.md §5
Answers are invariant under the measured option-order shuffle (largest probability shift 3.7e-8). Wording, dates, noise fields and layout cost at most 1.3 points; kebab-case keys and rewritten relation statements 1.4 and 1.6. Deterministic: identical runs agree bit for bit; batched and one-at-a-time inference differ by at most 1.1e-7 and never change a decision.
Field names affect the lexical features. Key-synonym and camelCase rewrites changed accuracy by −4.1 and −5.4 points; evaluate the application's input mapping and adapt it or the model as needed.
The released weights answered 26/32 structured reconciliation questions correctly (81.25%) in the broader synthetic invoice test without retraining. See every task, presentation, baseline and evaluation condition.
Original evaluation protocol
The original release targets were fixed before the test split was opened. The record below preserves their outcome for Decision 0.1 and remains separate from subsequent experiments.
Primary gate: accuracy above the reference. Secondary gate: Brier or raw ECE below it. Committed before the test split was read.
Architecture, hyper-parameters and calibration frozen with content hashes. The original result was recorded in one sealed release evaluation.
The recorded accuracy target was not reached; the Brier/ECE target was reached. These are outcomes of this frozen experiment.
Evaluation record
- One frozen ensemble, with one seed per member and case-level sampling intervals.
- Externally reported generalists listed with their evaluation settings.
- Named input transformations, with measured accuracy changes.
- Each future Primus release will carry its own evaluation.
Reproduce it
$ python -m pip install pandas pyarrow huggingface_hub
$ hf download LocalLLaMA/typed-decisions all/test-00000-of-00001.parquet --repo-type dataset --revision ea9306458d6e9563628369a3d1e72e362fb381d2 --local-dir typed-decisions
$ python examples/measure_latency.py --parquet typed-decisions/all/test-00000-of-00001.parquetRun it in the release folder with its virtual environment active (see Download). The script also needs pandas and pyarrow, which requirements.txt does not list; hf fetches the official test split at the pinned revision.
- dataset
- LocalLLaMA/typed-decisions @ ea9306458d6e9563628369a3d1e72e362fb381d2
- sealed test file
- 4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c
- release tree
- fa0406579f5b5338c2f44de0c42a0626d1a1e1a8