Introducing Primus Decision 0.1
An open-source 3.7M-parameter model for probabilistic decisions inside software.
- Open source · Apache-2.0
- Local CPU inference
- Non-transformer
Much of what software asks an AI system is not a request for prose. Is this invoice a duplicate? Should this support ticket be escalated? Should this agent be stopped? Each of these has a small, known set of valid answers, and what the caller needs is how likely each one is, not a paragraph that sounds confident.
Primus Decision 0.1 is our first open-source model, built for exactly that. Give it a JSON state and a typed question, and it returns a probability distribution over the valid answers. It is small (3.7M neural parameters), has no transformer, runs on a CPU, and is released under Apache-2.0 with its protocol, interfaces and measured results.
Typed probabilistic decisions
There are three question types. A noul question asks whether a statement is true and returns P(false) and P(true). A choice question returns one probability per offered option. A score question returns one probability per level of an ordinal scale, plus the expected level. For the invoice in the release's README, which exceeds its purchase order by 20 %, the model answers:
- duplicate: false 0.986, true 0.014
- disposition: manual review 0.708, hold 0.206, reject 0.084, approve 0.002
- discrepancy severity: severe 0.798, expected level 2.672
The bars are the model's actual raw probabilities for this invoice, from the released model. Highlighted fields are the ones each question is about; the model reads the whole state. Pick a question, or let it cycle.
Why a dedicated decision model?
A language model can also be constrained to choose among options, and a conventional classifier can be trained per question. Primus explores a different trade-off: one purpose-built model whose output is a probability distribution over the answers a caller offers, for any of the benchmark's question schemas, with CPU-friendly inference and a public runtime small enough to inspect and reproduce. Decision 0.1 demonstrates a compact implementation of this approach on its target benchmark, with task results and runtime trade-offs published for inspection.
The model
Two small encoders, trained from scratch on the benchmark's training part: a diagonal state-space (S4D) model and a bidirectional GRU, each paired with an LSA case vector. Their probabilities are averaged. There is no pretrained encoder and no token-to-token attention; the only attention pools one question's token states.
How it was evaluated
The evaluation protocol was fixed before the official test split of LocalLLaMA/typed-decisions was opened. Architecture, hyper-parameters and calibration were frozen with content hashes. The benchmark page preserves the original protocol and its target outcomes.
The result, from one sealed run on the official test split: 75.1% accuracy (2,000 decisions), Brier 0.059, Raw ECE 0.127, NLL 0.637. M, measured by us Original protocol and evaluation record.
Performance and operating trade-offs
Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison.
We reproduced Laya's released checkpoint on our own harness before comparing, because two cells of its published card (soft accuracy and Brier) do not reproduce. The comparison row is the reproduced one. R, reproduced by us
Calibration and selective automation
The raw model is under-confident: mean confidence 0.625 against accuracy 0.751. That caution is usable. Answering only the most confident half of the decisions gives 0.909 accuracy on them; a deployment would set its own threshold on its own Decision History. An optional temperature profile brings ECE from 0.127 to 0.040 but worsens Brier (0.059 → 0.114) and score MAE (0.275 → 0.320), so it is off by default.
Speed and size
~140 ms p50 CPU inference for one five-decision case with the model loaded, on a 4-vCPU Intel Xeon with 2 threads; 115–162 ms p50 across six recorded runs. The networks are 15 MB of float32 weights; with the two fitted LSA tables the package is 157.9 MB installed.
Verify it
Every file is checksummed. The public package reproduces our internal frozen model bit for bit on 600 decisions, and a separately written metric implementation used by AAME reproduces the sealed record within approximately 1e-16. A fresh clone of the hosted repository was hash-checked on 2026-09-23: release tree fa0406579f5b…, tag primus-decision-0.1.
What comes next
Update, 26 September 2026. The invoice case study evaluates the released weights without retraining against Qwen3-1.7B, a classifier and explicit rules, with all conditions and evidence published.
Our next experiments investigate input representations, ordinal decisions and consistency across training runs. New model versions will be evaluated separately while the Decision 0.1 artifacts remain frozen. The broader Primus programme connects this work with research on memory, reasoning, learning and action.
Get the model
Download, verify and run
CPU only, Python 3.11+. No GPU, no API key, no sign-up.
$ python3 -m venv .venv && . .venv/bin/activate && pip install -U huggingface_hub
$ hf download The-Aame/primus-decision-0.1 --local-dir primus-decision-0.1 && cd primus-decision-0.1
$ grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c - && \
python -m pip install -r requirements.txt --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple && python examples/predict_example.pyYou get three probability distributions for one invoice, 20 % over its purchase order. The manifest covers the published documentation and evidence. The check skips the Hub-specific README.md and .gitattributes; frozen model and runtime artifacts match tag primus-decision-0.1. No Git LFS is needed.
On Linux, sha256sum works in place of shasum -a 256. On Debian and Ubuntu, python3 -m venv needs the python3-venv package.
~140 ms p50 CPU inference per five-decision case, measured by us (M) on a 4-vCPU Intel Xeon 2.10 GHz, 2 threads, model loaded; 115–162 ms p50 across six recorded runs. The CPU build of torch is about 200 MB.