Article · 2026-09-23

Introducing Primus Decision 0.1

An open-source 3.7M-parameter model for probabilistic decisions inside software.

  • Open source · Apache-2.0
  • Local CPU inference
  • Non-transformer

Much of what software asks an AI system is not a request for prose. Is this invoice a duplicate? Should this support ticket be escalated? Should this agent be stopped? Each of these has a small, known set of valid answers, and what the caller needs is how likely each one is, not a paragraph that sounds confident.

Primus Decision 0.1 is our first open-source model, built for exactly that. Give it a JSON state and a typed question, and it returns a probability distribution over the valid answers. It is small (3.7M neural parameters), has no transformer, runs on a CPU, and is released under Apache-2.0 with its protocol, interfaces and measured results.

Typed probabilistic decisions

There are three question types. A noul question asks whether a statement is true and returns P(false) and P(true). A choice question returns one probability per offered option. A score question returns one probability per level of an ordinal scale, plus the expected level. For the invoice in the release's README, which exceeds its purchase order by 20 %, the model answers:

  • duplicate: false 0.986, true 0.014
  • disposition: manual review 0.708, hold 0.206, reject 0.084, approve 0.002
  • discrepancy severity: severe 0.798, expected level 2.672
1
The state
invoice_processing · the README example
Invoice
INV-2041
received
Acme Supplies
Invoice numberINV-2041
Invoice amount$1,200.00
PO amount on invoice$1,000.00
Purchase order PO-7781$1,000.00
Over the purchase order+$200.00 · +20%
2
Two small encoders
3.7M parameters · no transformer
S4D member1.93M
diagonal state-space · 2 blocks per direction
GRU member1.78M
bidirectional · 2 layers
LSA case vector256-d
TF-IDF → SVD, fitted on the train split
3
A typed answer
probabilities, not text
disposition · choice
What should happen to this invoice?
  • approve0.002
  • hold0.206
  • manual review0.708
  • reject0.084
confidence
0.708
answer manual review

The bars are the model's actual raw probabilities for this invoice, from the released model. Highlighted fields are the ones each question is about; the model reads the whole state. Pick a question, or let it cycle.

Why a dedicated decision model?

A language model can also be constrained to choose among options, and a conventional classifier can be trained per question. Primus explores a different trade-off: one purpose-built model whose output is a probability distribution over the answers a caller offers, for any of the benchmark's question schemas, with CPU-friendly inference and a public runtime small enough to inspect and reproduce. Decision 0.1 demonstrates a compact implementation of this approach on its target benchmark, with task results and runtime trade-offs published for inspection.

The model

Two small encoders, trained from scratch on the benchmark's training part: a diagonal state-space (S4D) model and a bidirectional GRU, each paired with an LSA case vector. Their probabilities are averaged. There is no pretrained encoder and no token-to-token attention; the only attention pools one question's token states.

JSON stateone workflowflattened textpath: value lines + relationsS4D member2 diagonal state-space blocksGRU member2 bidirectional GRU layersLSA case vectorTF-IDF → 256-d SVD, per memberoption scorer + softmaxper member, over offered optionsaverage 0.5 / 0.5raw distribution; optional profilequestion-conditioned attention pooling in both members · no token-to-token attention · no transformer · no pretrained encoder

How it was evaluated

The evaluation protocol was fixed before the official test split of LocalLLaMA/typed-decisions was opened. Architecture, hyper-parameters and calibration were frozen with content hashes. The benchmark page preserves the original protocol and its target outcomes.

The result, from one sealed run on the official test split: 75.1% accuracy (2,000 decisions), Brier 0.059, Raw ECE 0.127, NLL 0.637. M, measured by us Original protocol and evaluation record.

Performance and operating trade-offs

Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison.

We reproduced Laya's released checkpoint on our own harness before comparing, because two cells of its published card (soft accuracy and Brier) do not reproduce. The comparison row is the reproduced one. R, reproduced by us

Calibration and selective automation

The raw model is under-confident: mean confidence 0.625 against accuracy 0.751. That caution is usable. Answering only the most confident half of the decisions gives 0.909 accuracy on them; a deployment would set its own threshold on its own Decision History. An optional temperature profile brings ECE from 0.127 to 0.040 but worsens Brier (0.059 → 0.114) and score MAE (0.275 → 0.320), so it is off by default.

Speed and size

~140 ms p50 CPU inference for one five-decision case with the model loaded, on a 4-vCPU Intel Xeon with 2 threads; 115–162 ms p50 across six recorded runs. The networks are 15 MB of float32 weights; with the two fitted LSA tables the package is 157.9 MB installed.

Verify it

Every file is checksummed. The public package reproduces our internal frozen model bit for bit on 600 decisions, and a separately written metric implementation used by AAME reproduces the sealed record within approximately 1e-16. A fresh clone of the hosted repository was hash-checked on 2026-09-23: release tree fa0406579f5b…, tag primus-decision-0.1.

What comes next

Update, 26 September 2026. The invoice case study evaluates the released weights without retraining against Qwen3-1.7B, a classifier and explicit rules, with all conditions and evidence published.

Our next experiments investigate input representations, ordinal decisions and consistency across training runs. New model versions will be evaluated separately while the Decision 0.1 artifacts remain frozen. The broader Primus programme connects this work with research on memory, reasoning, learning and action.

Get the model

Download, verify and run

CPU only, Python 3.11+. No GPU, no API key, no sign-up.

Hugging Face · shell
$ python3 -m venv .venv && . .venv/bin/activate && pip install -U huggingface_hub
$ hf download The-Aame/primus-decision-0.1 --local-dir primus-decision-0.1 && cd primus-decision-0.1
$ grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c - && \
  python -m pip install -r requirements.txt --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple && python examples/predict_example.py

You get three probability distributions for one invoice, 20 % over its purchase order. The manifest covers the published documentation and evidence. The check skips the Hub-specific README.md and .gitattributes; frozen model and runtime artifacts match tag primus-decision-0.1. No Git LFS is needed.

On Linux, sha256sum works in place of shasum -a 256. On Debian and Ubuntu, python3 -m venv needs the python3-venv package.

~140 ms p50 CPU inference per five-decision case, measured by us (M) on a 4-vCPU Intel Xeon 2.10 GHz, 2 threads, model loaded; 115–162 ms p50 across six recorded runs. The CPU build of torch is about 200 MB.

Benchmarks · Experiments · Download and cite