Case study · 26 September 2026

Primus for invoice decisions

Local invoice decisions from the released Primus weights, without retraining. Three controlled synthetic experiments, with task results, baselines and public evidence.

M, measured by us

Primus Decision 0.1 performs structured invoice checks locally with 3,715,074 neural parameters and no LLM call at inference. In our broader invoice experiment, it answered 26 of 32 structured reconciliation questions correctly (81.25%). The initial, simpler pilot recorded 32/32, including after question rewording.

We built the study to make that capability inspectable: generated records, questions, model answers, comparison systems and evaluation conditions accompany the results. Qwen3-1.7B runs as a standalone competitor; its outputs never enter Primus.

Complete report on GitHub · Download and reproduce the evidence · Public model

The decision task

Given an invoice, purchase order and delivery record, the model returns a probability that the items, quantities and amounts reconcile. A second question checks whether the invoice ID appears in the supplied history. Labels come from exact arithmetic and membership checks on synthetic records, computed before inference.

Primus and the TF-IDF/logistic-regression classifier retain their original task supervision; neither was fitted on these evaluation invoices. Qwen uses its official non-thinking chat template without demonstrations, additional training or external tools. These results compare the stated deployment configurations and their different training histories.

Broader test: 32 new invoices

The broader test covers delivery shortages, price discrepancies, offsetting item-level price discrepancies with an unchanged grand total, and distracting historical amounts. Every invoice appears as both structured records and a narrative containing the same facts. Each row below has 16 true and 16 false cases.

Broader invoice results: correct answers out of 32
PresentationQuestionPrimus 0.1TF-IDF classifierQwen token scoringQwen free answerExplicit rules
StructuredReconciliation26/3227/3216/3216/3232/32
StructuredDuplicate ID16/3216/3218/3218/3232/32
NarrativeReconciliation17/3224/3218/3218/3232/32
NarrativeDuplicate ID16/3216/3216/3214/3232/32

Structured reconciliation is Primus's strongest result in this experiment: 8/8 delivery-quantity cases and 6/8 in each other family. The same facts in narrative form score 17/32, identifying input representation as a concrete direction for further work. The classifier records 27/32 structured and 24/32 narrative answers; explicit rules provide the exact reference solution for these arithmetic and membership checks.

The invoices form 16 matching/mismatching pairs. A pair counts as solved only when both answers are correct: Primus solves 10/16 structured pairs and 1/16 narrative pairs; the classifier solves 11/16 and 8/16; both Qwen answer methods solve 0/16 and 2/16. This checks whether each system distinguishes the changed facts within a pair.

Initial pilot and reworded questions

The pilot used 32 synthetic invoices varying duplicate ID, price discrepancy, delivery shortage, payment timing and vendor. Six questions per invoice produced 192 related answer instances across supported, reworded and new-question tracks.

Pilot invoice results: correct answers out of 32
TrackQuestionPrimus 0.1Qwen3-1.7BTF-IDF classifierExplicit rules
SupportedReconciliation32/328/3232/3232/32
RewordedReconciliation32/328/3232/3232/32
SupportedDuplicate ID18/3219/3216/3232/32
RewordedDuplicate ID18/3216/3216/3232/32

Reconciliation has 8 true and 24 false labels in the pilot. Its 32/32 result describes that template family; the balanced broader experiment extends the evidence. Rewording preserves the supported question IDs and answer meanings, testing stability within the released interface.

The original pilot also records a separate development candidate on payment-timing and delivery-shortage questions (17/32 and 14/32). Those question IDs are outside the public 0.1 schema contract. The evidence identifies all candidate predictions separately from the released model; candidate implementation and checkpoints remain private.

Checking Qwen's answer format

A follow-up audit tested six answer-format conditions on the same 64 supported-question instances. The original replay reproduced every probability and label exactly.

Qwen answer-format audit
Answer methodDuplicate IDReconciliation
Original lowercase-token replay19/328/32
Lowercase and capitalized aliases17/328/32
Free greedy answer14/328/32
Yes/no token aliases20/328/32
A=false, B=true16/328/32
A=true, B=false16/328/32

This was a retrospective diagnostic on already inspected cases. The broader experiment subsequently reports both capitalization-aware token scoring and free answers. The exact checkpoint, non-thinking template and answer methods are documented so readers can reproduce the result and extend the comparison to other configurations.

Evaluation conditions

  • Primus: released Decision 0.1 weights, local CPU float32 inference, two threads. The S4D/GRU ensemble and fitted TF-IDF/LSA features make no LLM call. Original benchmark training uses teacher-generated labels.
  • Qwen: Qwen/Qwen3-1.7B at revision 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e, Apple GPU float16, official non-thinking template. Free answers use greedy generation capped at 12 tokens; invalid answers count as incorrect.
  • Preprocessing: weights were used without retraining. The pilot used the default cap of 40 generic derived-relation lines. The broader test used a cap of 6 for both presentations and fit complete states within the 640-token input. Changes between experiments therefore reflect both cases and preprocessing.
  • Evidence units: 32 pilot invoices and 32 new broader-test invoices. The answer-format audit reuses the pilot. Multiple questions, pairs and presentations are related observations, so each task and condition is reported separately.
  • Timing: pilot devices and timing boundaries differ. Recorded timings are descriptive; the case study compares task predictions.

These are AAME's controlled, assistant-generated synthetic experiments. Cases and procedures were frozen locally before their respective predictions; the records document that sequence. The pilot evidence discloses an implementation correction to the rule baseline that left predictions unchanged.

What this gives us

A compact, locally runnable Primus component can make supported structured invoice decisions, with results inspectable per case. The next step is to extend performance across representations and richer records, then evaluate externally sourced invoices with human-checked labels.

The original 75.1% Typed Decisions result remains separate: 400 cases and 2,000 decisions. This study adds task-specific evidence for the released component; further Primus releases will be evaluated against the capabilities they add.

Public evidence and protected work

The evidence includes generated cases, recorded probabilities and answers, every comparison condition, metrics, model revisions, checksums and public-model evaluation code. Unreleased candidate implementation, private training code, checkpoints, internal datasets and broader system integration remain protected. The Road to Primus describes the broader research direction and public boundary.

Read the full GitHub report · Download and verify the evidence · Run a reproduction