Primus Decision 0.1 — Research Alpha
Our first public Primus research component. A 3.7M-neural-parameter non-transformer model for typed probabilistic decisions: give it a structured JSON state and a supported question, and it returns a probability for each valid answer. It runs locally on a CPU with no LLM call at inference. Model artifacts, runtime, interfaces and evaluation evidence are public under Apache-2.0.
- Open source · Apache-2.0
- Local CPU inference
- Non-transformer
Trained on the official train split only [1], recorded in one sealed release evaluation, released under Apache-2.0 with checksums, a public protocol, interfaces and measured results.
These results describe the frozen Decision 0.1 component on its named benchmark. Each new experiment and model version establishes its own results. M, measured by us Source: SEALED_RESULT.json
Explore accuracy, probability quality, model size and CPU inference by task. The tables identify results measured by AAME, checkpoints reproduced by AAME, and figures reported by other publishers.
M measured by usR reproduced by usX reported externally
How these numbers were measured
- [1] Training split: 1,005 fitting cases and 195 validation cases; no additional training data or pretrained encoder. Source: PROTOCOL.md
- [2] 3.7M neural parameters: 3,715,074 neural parameters; the two LSA featurizers (71.4 MB each) are fitted tables, not counted. Source: MODEL_SIZE_AND_PARAMETERS.md §1–2
- [3] 75.1% accuracy, Brier 0.059, Raw ECE 0.127, NLL 0.637 on the official test split of LocalLLaMA/typed-decisions (400 cases, 2,000 decisions), one sealed run, raw probabilities. Source: SEALED_RESULT.json
- [4] ~140 ms p50 CPU inference: one five-decision case, model loaded, 4-vCPU Intel Xeon 2.10 GHz, 2 threads; 115–162 ms p50 across six recorded runs with two methods. Source: EXPERIMENTS.md §4
What it answers
| question type | asks | returns |
|---|---|---|
| noul | is this statement true? | P(false), P(true) |
| choice | which of these options applies? | one probability per option |
| score | which level on an ordinal scale? | one probability per level, plus the expected level |
Questions use one of the 20 benchmark schemas: four workflows (agent-trace observability, customer service, invoice processing, security incidents), five questions each. For a supported choice question, supply any subset of its trained option keys in any order. Question and option descriptions can be supplied with the request; the runtime validates the schema.
Invoice case study
The released weights answered 26/32 structured reconciliation questions correctly (81.25%) without retraining in a broader synthetic invoice experiment, following 32/32 in the initial pilot. The study reports all structured and narrative results, duplicate-ID checks, Qwen3-1.7B, a TF-IDF classifier and explicit rules together.
Architecture
No transformer and no pretrained encoder anywhere: both encoders are trained from scratch on the benchmark's training part, and the only attention is the pooling of one question's token states. M, measured by us Source: ARCHITECTURE.md
Evidence
Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison M, measured by us R, reproduced by us Sources: SEALED_RESULT.json, EXPERIMENTS.md §2, §4, §6
- Accuracy. 0.751 against 0.766 for the reproduced Laya checkpoint. The recorded difference is 30 decisions out of 2,000. Primus's interval, 0.729–0.774, describes case-sampling uncertainty for this frozen run. M, measured by us R, reproduced by us Source: EXPERIMENTS.md §2, §6
- Probability quality. Primus records lower NLL (0.637 vs 0.707), raw ECE (0.127 vs 0.213) and Brier (0.059 vs 0.066) in this comparison. M, measured by us R, reproduced by us
- Task detail. Laya records the lower ordinal score error (MAE 0.242 against 0.275) and higher accuracy in the customer-service and security-incident workflows. Agent-trace and invoice differences are one and two decisions out of 500. M, measured by us R, reproduced by us Sources: BENCHMARKS.md, EXPERIMENTS.md §6
- Size. 421M → 3.7M neural parameters (~113× fewer) by Laya's model card; Laya's weights file is 843 MB; Primus's original complete installed release is 157.9 MB, including two 71.4 MB feature tables. These file measurements have different contents. X, reported externally M, measured by us Source: EXPERIMENTS.md §6
- Abstention. Confidence is usable for abstention: 90.9% accuracy on the most-confident 50% of decisions. M, measured by us Source: EXPERIMENTS.md §3
Recorded accuracy by workflow. Results are grouped into cases; comparative statistical claims require a suitable paired analysis.
Calibration
Mean confidence 0.625 against accuracy 0.751: the raw model is under-confident. M, measured by us
Cost: Brier 0.059 → 0.114, score MAE 0.275 → 0.320; NLL 0.637 → 0.568. Fitted on the validation part only; off by default; raw probabilities are the default evaluation. M, measured by us
Area under the risk–coverage curve 0.102. 90.9% accuracy on the most-confident 50% of decisions. These figures describe this test split; a deployment sets its own threshold on its own Decision History.
Evaluation and integration
The original benchmark uses four synthetic workflows with teacher-generated reference labels. The invoice study adds generated records with exact arithmetic and membership labels. These are different evaluations of the released component.
The model's lexical features learn field-name conventions. Key-synonym and camelCase rewrites changed accuracy by −4.1 and −5.4 points. Evaluate the application's records, field conventions and decision criteria, then adapt the input mapping or model as needed.
The installed release is 157.9 MB, with recorded inference peaks of 1.4–1.6 GB. Use compact structured states for the tested input path. The results cover one frozen ensemble, one seed per member, with case-level sampling intervals.
Use application outcomes to set confidence thresholds and retain human review for consequential decisions. The public interface and Decision History format support integration and evaluation.
Download and verify
Download, verify and run
CPU only, Python 3.11+. No GPU, no API key, no sign-up.
$ python3 -m venv .venv && . .venv/bin/activate && pip install -U huggingface_hub
$ hf download The-Aame/primus-decision-0.1 --local-dir primus-decision-0.1 && cd primus-decision-0.1
$ grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c - && \
python -m pip install -r requirements.txt --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple && python examples/predict_example.pyYou get three probability distributions for one invoice, 20 % over its purchase order. The manifest covers the published documentation and evidence. The check skips the Hub-specific README.md and .gitattributes; frozen model and runtime artifacts match tag primus-decision-0.1. No Git LFS is needed.
On Linux, sha256sum works in place of shasum -a 256. On Debian and Ubuntu, python3 -m venv needs the python3-venv package.
~140 ms p50 CPU inference per five-decision case, measured by us (M) on a 4-vCPU Intel Xeon 2.10 GHz, 2 threads, model loaded; 115–162 ms p50 across six recorded runs. The CPU build of torch is about 200 MB.
Take the model from Hugging Face or from the release archive. GitHub's auto-generated "Source code" archives, like a clone made without Git LFS, hold LFS pointer files instead of the two featurizers; loading one fails with a message that says so.
Provenance
| frozen | 2026-09-22T06:22:58Z |
|---|---|
| sealed evaluation | 2026-09-22T06:32:28Z |
| dataset | LocalLLaMA/typed-decisions @ ea930645… |
| sealed test file sha256 | 4f294f218ea1da27… |
| release tree | fa0406579f5b5338c2f44de0c42a0626d1a1e1a8 |
| commit | f0bf642cae6c2399da64cc86f1585d468a9e5b40 |
| tag | primus-decision-0.1 |
| equivalence | 600 / 600 decisions bit-identical to the internal frozen package |
| hash-checked | 2026-09-23, a fresh clone of the hosted repository |
How to check us
- Sealed protocolThe protocol was fixed before the test split was opened. One sealed release evaluation records the original result; later verification and diagnostics are identified separately.
- Reproduced to 1e-16A separately written metric implementation used by AAME reproduces the sealed result within approximately 1e-16 on a fresh run of the verified model.
- Bit-identical packageThe public package produces bit-identical outputs to the internal frozen package on 600 decisions.
- Every number labelledMeasured by us, reproduced by us, or reported externally, on every page.
- Check it yourself
grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c -In the release folder; on Linux, sha256sum works too.
The Road to Primus
Primus Decision 0.1 is our first public research alpha: one trained decision component in a longer research programme led by Sai Tilak Pally at AAME.
Our goal is a broader architecture connecting persistent memory, recurrent graph reasoning, temporal state and history, salience and priority mechanisms, imagination and simulation, consolidation and learning cycles, language and grounding, planning and decision layers, and agent and tool interfaces. These are intended research directions; each public release will document the capabilities it demonstrates.
Public vs Protected
We publish architecture descriptions, interfaces, benchmarks and reproducibility evidence for released components so their claims can be examined. Unreleased components, training methods, system integration details, internal datasets and implementation specifics remain private until AAME chooses to publish them.
Researchers can inspect and extend the Apache-2.0 release, run new evaluations, and share reproducible results. Give changed models or preprocessing a distinct version and report the data, conditions and relevant baselines. Source: PUBLIC_BOUNDARY.md
Frequently asked questions
What is Primus Decision 0.1?
AAME's first public Primus component: a compact non-transformer model that turns structured states and supported questions into typed probability distributions.
Does it need an LLM or cloud API to run?
Inference runs locally on a CPU with the released S4D/GRU ensemble and feature tables. No LLM or hosted API is called. The original benchmark training labels were teacher-generated.
Is 75.1% the maximum Primus can achieve?
It is the measured accuracy of the frozen 0.1 model on one named benchmark. The invoice study reports different task-specific results; future experiments and versions establish their own scores.
Can I test and extend it?
Yes. Run new records within the supported schemas, inspect the outputs, or modify the published artifacts under Apache-2.0. The case study and evidence provide a reproducible starting point.
Is this the complete Primus architecture?
Decision 0.1 is the first public decision component. The broader architecture is the research programme described above.
Cite
@software{primus_decision_0_1,
author = {Pally, Sai Tilak and {AAME}},
title = {Primus Decision 0.1 (Research Alpha): an open 3.7M-parameter non-transformer model for typed probabilistic decisions},
year = {2026},
version = {0.1},
license = {Apache-2.0},
url = {https://github.com/pally-sai-tilak/primus-decision}
}