Primus · open-source release · Apache-2.0

Primus Decision 0.1 — Research Alpha

Our first public Primus research component. A 3.7M-neural-parameter non-transformer model for typed probabilistic decisions: give it a structured JSON state and a supported question, and it returns a probability for each valid answer. It runs locally on a CPU with no LLM call at inference. Model artifacts, runtime, interfaces and evaluation evidence are public under Apache-2.0.

  • Open source · Apache-2.0
  • Local CPU inference
  • Non-transformer

Trained on the official train split only [1], recorded in one sealed release evaluation, released under Apache-2.0 with checksums, a public protocol, interfaces and measured results.

3.7M
neural parameters M, measured by us 2
75.1%
accuracy, official test split M, measured by us 3
0.059
Brier M, measured by us 3
0.127
raw ECE M, measured by us 3
~140 ms
p50 CPU inference M, measured by us 4

These results describe the frozen Decision 0.1 component on its named benchmark. Each new experiment and model version establishes its own results. M, measured by us Source: SEALED_RESULT.json

Compare the measured results

Explore accuracy, probability quality, model size and CPU inference by task. The tables identify results measured by AAME, checkpoints reproduced by AAME, and figures reported by other publishers.

Full comparison and methods · Invoice case study

M measured by usR reproduced by usX reported externally

How these numbers were measured
  1. [1] Training split: 1,005 fitting cases and 195 validation cases; no additional training data or pretrained encoder. Source: PROTOCOL.md
  2. [2] 3.7M neural parameters: 3,715,074 neural parameters; the two LSA featurizers (71.4 MB each) are fitted tables, not counted. Source: MODEL_SIZE_AND_PARAMETERS.md §1–2
  3. [3] 75.1% accuracy, Brier 0.059, Raw ECE 0.127, NLL 0.637 on the official test split of LocalLLaMA/typed-decisions (400 cases, 2,000 decisions), one sealed run, raw probabilities. Source: SEALED_RESULT.json
  4. [4] ~140 ms p50 CPU inference: one five-decision case, model loaded, 4-vCPU Intel Xeon 2.10 GHz, 2 threads; 115–162 ms p50 across six recorded runs with two methods. Source: EXPERIMENTS.md §4

What it answers

question typeasksreturns
noulis this statement true?P(false), P(true)
choicewhich of these options applies?one probability per option
scorewhich level on an ordinal scale?one probability per level, plus the expected level

Questions use one of the 20 benchmark schemas: four workflows (agent-trace observability, customer service, invoice processing, security incidents), five questions each. For a supported choice question, supply any subset of its trained option keys in any order. Question and option descriptions can be supplied with the request; the runtime validates the schema.

Invoice case study

The released weights answered 26/32 structured reconciliation questions correctly (81.25%) without retraining in a broader synthetic invoice experiment, following 32/32 in the initial pilot. The study reports all structured and narrative results, duplicate-ID checks, Qwen3-1.7B, a TF-IDF classifier and explicit rules together.

Read the case study and examine its evidence

Architecture

JSON stateone workflowflattened textpath: value lines + relationsS4D member2 diagonal state-space blocksGRU member2 bidirectional GRU layersLSA case vectorTF-IDF → 256-d SVD, per memberoption scorer + softmaxper member, over offered optionsaverage 0.5 / 0.5raw distribution; optional profilequestion-conditioned attention pooling in both members · no token-to-token attention · no transformer · no pretrained encoder

No transformer and no pretrained encoder anywhere: both encoders are trained from scratch on the benchmark's training part, and the only attention is the pooling of one question's token states. M, measured by us Source: ARCHITECTURE.md

Architecture page

Evidence

Primus records 75.1% accuracy, NLL 0.637 and raw ECE 0.127 with 3,715,074 neural parameters. AAME's reproduction of the 421M-parameter Laya checkpoint on the same harness records 76.6% accuracy, NLL 0.707 and ECE 0.213. Primus uses about 113× fewer neural parameters; its separate feature tables bring the installed release to 157.9 MB. Full task results and measurement conditions accompany the comparison M, measured by us R, reproduced by us Sources: SEALED_RESULT.json, EXPERIMENTS.md §2, §4, §6

  • Accuracy. 0.751 against 0.766 for the reproduced Laya checkpoint. The recorded difference is 30 decisions out of 2,000. Primus's interval, 0.729–0.774, describes case-sampling uncertainty for this frozen run. M, measured by us R, reproduced by us Source: EXPERIMENTS.md §2, §6
  • Probability quality. Primus records lower NLL (0.637 vs 0.707), raw ECE (0.127 vs 0.213) and Brier (0.059 vs 0.066) in this comparison. M, measured by us R, reproduced by us
  • Task detail. Laya records the lower ordinal score error (MAE 0.242 against 0.275) and higher accuracy in the customer-service and security-incident workflows. Agent-trace and invoice differences are one and two decisions out of 500. M, measured by us R, reproduced by us Sources: BENCHMARKS.md, EXPERIMENTS.md §6
  • Size. 421M → 3.7M neural parameters (~113× fewer) by Laya's model card; Laya's weights file is 843 MB; Primus's original complete installed release is 157.9 MB, including two 71.4 MB feature tables. These file measurements have different contents. X, reported externally M, measured by us Source: EXPERIMENTS.md §6
  • Abstention. Confidence is usable for abstention: 90.9% accuracy on the most-confident 50% of decisions. M, measured by us Source: EXPERIMENTS.md §3
Accuracy by workflow, official test split (500 decisions each)
Primus Decision 0.1Laya, reproduced
agent-trace observability0.7300.732+1 of 500customer service0.7280.764−18 of 500invoice processing0.8040.808+2 of 500security incidents0.7360.766−15 of 500

Recorded accuracy by workflow. Results are grouped into cases; comparative statistical claims require a suitable paired analysis.

Full tables, intervals and the same-machine latency runs

Calibration

0.127
Raw ECE 0.127 (the default output)

Mean confidence 0.625 against accuracy 0.751: the raw model is under-confident. M, measured by us

0.040
Calibrated ECE 0.040, optional temperature profile

Cost: Brier 0.059 → 0.114, score MAE 0.275 → 0.320; NLL 0.637 → 0.568. Fitted on the validation part only; off by default; raw probabilities are the default evaluation. M, measured by us

Accuracy on the answered decisions vs coverage (raw confidence, most confident first)
0.700.800.901.0010%30%50%70%90%coverage (most confident first)accuracy10 %: 1.00020 %: 0.98830 %: 0.95540 %: 0.93550 %: 0.90960 %: 0.87870 %: 0.84180 %: 0.81690 %: 0.783100 %: 0.7510.909 on the most confident half

Area under the risk–coverage curve 0.102. 90.9% accuracy on the most-confident 50% of decisions. These figures describe this test split; a deployment sets its own threshold on its own Decision History.

Evaluation and integration

The original benchmark uses four synthetic workflows with teacher-generated reference labels. The invoice study adds generated records with exact arithmetic and membership labels. These are different evaluations of the released component.

The model's lexical features learn field-name conventions. Key-synonym and camelCase rewrites changed accuracy by −4.1 and −5.4 points. Evaluate the application's records, field conventions and decision criteria, then adapt the input mapping or model as needed.

The installed release is 157.9 MB, with recorded inference peaks of 1.4–1.6 GB. Use compact structured states for the tested input path. The results cover one frozen ensemble, one seed per member, with case-level sampling intervals.

Use application outcomes to set confidence thresholds and retain human review for consequential decisions. The public interface and Decision History format support integration and evaluation.

Download and verify

Download, verify and run

CPU only, Python 3.11+. No GPU, no API key, no sign-up.

Hugging Face · shell
$ python3 -m venv .venv && . .venv/bin/activate && pip install -U huggingface_hub
$ hf download The-Aame/primus-decision-0.1 --local-dir primus-decision-0.1 && cd primus-decision-0.1
$ grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c - && \
  python -m pip install -r requirements.txt --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple && python examples/predict_example.py

You get three probability distributions for one invoice, 20 % over its purchase order. The manifest covers the published documentation and evidence. The check skips the Hub-specific README.md and .gitattributes; frozen model and runtime artifacts match tag primus-decision-0.1. No Git LFS is needed.

On Linux, sha256sum works in place of shasum -a 256. On Debian and Ubuntu, python3 -m venv needs the python3-venv package.

~140 ms p50 CPU inference per five-decision case, measured by us (M) on a 4-vCPU Intel Xeon 2.10 GHz, 2 threads, model loaded; 115–162 ms p50 across six recorded runs. The CPU build of torch is about 200 MB.

Take the model from Hugging Face or from the release archive. GitHub's auto-generated "Source code" archives, like a clone made without Git LFS, hold LFS pointer files instead of the two featurizers; loading one fails with a message that says so.

Provenance

frozen2026-09-22T06:22:58Z
sealed evaluation2026-09-22T06:32:28Z
datasetLocalLLaMA/typed-decisions @ ea930645…
sealed test file sha2564f294f218ea1da27…
release treefa0406579f5b5338c2f44de0c42a0626d1a1e1a8
commitf0bf642cae6c2399da64cc86f1585d468a9e5b40
tagprimus-decision-0.1
equivalence600 / 600 decisions bit-identical to the internal frozen package
hash-checked2026-09-23, a fresh clone of the hosted repository

How to check us

  • Sealed protocol
    The protocol was fixed before the test split was opened. One sealed release evaluation records the original result; later verification and diagnostics are identified separately.
  • Reproduced to 1e-16
    A separately written metric implementation used by AAME reproduces the sealed result within approximately 1e-16 on a fresh run of the verified model.
  • Bit-identical package
    The public package produces bit-identical outputs to the internal frozen package on 600 decisions.
  • Every number labelled
    Measured by us, reproduced by us, or reported externally, on every page.
  • Check it yourself
    grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c -
    In the release folder; on Linux, sha256sum works too.

The Road to Primus

Primus Decision 0.1 is our first public research alpha: one trained decision component in a longer research programme led by Sai Tilak Pally at AAME.

Our goal is a broader architecture connecting persistent memory, recurrent graph reasoning, temporal state and history, salience and priority mechanisms, imagination and simulation, consolidation and learning cycles, language and grounding, planning and decision layers, and agent and tool interfaces. These are intended research directions; each public release will document the capabilities it demonstrates.

Public vs Protected

We publish architecture descriptions, interfaces, benchmarks and reproducibility evidence for released components so their claims can be examined. Unreleased components, training methods, system integration details, internal datasets and implementation specifics remain private until AAME chooses to publish them.

Researchers can inspect and extend the Apache-2.0 release, run new evaluations, and share reproducible results. Give changed models or preprocessing a distinct version and report the data, conditions and relevant baselines. Source: PUBLIC_BOUNDARY.md

Frequently asked questions

What is Primus Decision 0.1?

AAME's first public Primus component: a compact non-transformer model that turns structured states and supported questions into typed probability distributions.

Does it need an LLM or cloud API to run?

Inference runs locally on a CPU with the released S4D/GRU ensemble and feature tables. No LLM or hosted API is called. The original benchmark training labels were teacher-generated.

Is 75.1% the maximum Primus can achieve?

It is the measured accuracy of the frozen 0.1 model on one named benchmark. The invoice study reports different task-specific results; future experiments and versions establish their own scores.

Can I test and extend it?

Yes. Run new records within the supported schemas, inspect the outputs, or modify the published artifacts under Apache-2.0. The case study and evidence provide a reproducible starting point.

Is this the complete Primus architecture?

Decision 0.1 is the first public decision component. The broader architecture is the research programme described above.

Cite

BibTeX · CITATION.cff in the repository root
@software{primus_decision_0_1,
  author  = {Pally, Sai Tilak and {AAME}},
  title   = {Primus Decision 0.1 (Research Alpha): an open 3.7M-parameter non-transformer model for typed probabilistic decisions},
  year    = {2026},
  version = {0.1},
  license = {Apache-2.0},
  url     = {https://github.com/pally-sai-tilak/primus-decision}
}