Verify and reproduce
Download the frozen model, verify its hashes and run the public example on a CPU. Compare your predictions with the sealed record, or reproduce the invoice study using its published evaluation scripts.
Get the model, check every file, run the example
Hugging Face · shell$ python3 -m venv .venv && . .venv/bin/activate && pip install -U huggingface_hub $ hf download The-Aame/primus-decision-0.1 --local-dir primus-decision-0.1 && cd primus-decision-0.1 $ grep -vE ' (README\.md|\.gitattributes)$' SHA256SUMS | shasum -a 256 -c - && \ python -m pip install -r requirements.txt --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple && python examples/predict_example.pyYou get three probability distributions for one invoice, 20 % over its purchase order. The manifest covers the published documentation and evidence. The check skips the Hub-specific README.md and .gitattributes; frozen model and runtime artifacts match tag primus-decision-0.1. No Git LFS is needed.
On Linux, sha256sum works in place of shasum -a 256. On Debian and Ubuntu, python3 -m venv needs the python3-venv package.
~140 ms p50 CPU inference per five-decision case, measured by us (M) on a 4-vCPU Intel Xeon 2.10 GHz, 2 threads, model loaded; 115–162 ms p50 across six recorded runs. The CPU build of torch is about 200 MB.
Compare
The example prints the model's answer for one invoice: three probability distributions. Compare them with the documented answer, the response on the Docs page, which gives them to three decimals.
The sealed record, SEALED_RESULT.json, holds the one sealed run on the official test split, raw probabilities: accuracy 0.751, Brier 0.059, ECE 0.127, NLL 0.637, score MAE 0.275. A separately written metric implementation used by AAME reproduced the sealed record within approximately 1e-16 on a fresh run of the verified model. M, measured by us
For the original benchmark, combine a dataset loader with the official test split at its pinned revision (step 3 fetches it), each decision's gold distribution in option order, and the public metric definitions in primus_decision/metrics.py and PROTOCOL.md.
Time it on your machine
The latency method of the release, examples/measure_latency.py, in the release folder with the environment from step 1 active. It needs pandas and pyarrow, which requirements.txt does not list.
latency method on your machine$ python -m pip install pandas pyarrow huggingface_hub $ hf download LocalLLaMA/typed-decisions all/test-00000-of-00001.parquet --repo-type dataset --revision ea9306458d6e9563628369a3d1e72e362fb381d2 --local-dir typed-decisions $ python examples/measure_latency.py --parquet typed-decisions/all/test-00000-of-00001.parquetThe recorded runs, on a 4-vCPU Intel Xeon 2.10 GHz with 2 threads: about 140 ms per five-decision case, 115–162 ms p50 across six recorded runs. Your machine gives its own figure. M, measured by us
Tell us what you got
Open an issue titled "Reproduction report" with your machine, OS, Python and torch versions, what you ran and what you got. The button opens GitHub's issue form with those fields filled in; GitHub asks you to sign in.