Primus Decision 0.1 · Demo

Replay Primus Decision 0.1

Explore 25 recorded outputs from the released model: the README example and 24 official test cases, six per workflow. This browser replay presents precomputed predictions; run the public model locally to try new records. Every answer was produced by tag primus-decision-0.1 (tree fa040657…) on 2026-09-23 with the public runtime. Nothing is simulated and nothing you select leaves your browser.

Output

Profile: Raw ECE 0.127 → Calibrated ECE 0.040 on the test split, but Brier 0.059 → 0.114 and score MAE 0.275 → 0.320; off by default. M, measured by us Source: EXPERIMENTS.md §3

request · invoice processing · README example
state
{
 "invoice": {
  "invoice_number": "INV-2041",
  "amount_usd": 1200,
  "po_amount_usd": 1000,
  "status": "received",
  "vendor": "Acme Supplies"
 },
 "purchase_order": {
  "po_number": "PO-7781",
  "amount_usd": 1000
 }
}
questions
{
 "duplicate": {
  "type": "noul",
  "instructions": "Is this invoice a duplicate of one already paid?"
 },
 "disposition": {
  "type": "choice",
  "instructions": "What should happen to this invoice?",
  "criteria": {
   "approve": "Approve and pay.",
   "hold": "Hold until the discrepancy is resolved.",
   "manual_review": "Send to a person for review.",
   "reject": "Reject the invoice."
  }
 },
 "discrepancy_severity": {
  "type": "score",
  "instructions": "How severe is the discrepancy between invoice and purchase order?",
  "criteria": [
   "No discrepancy.",
   "Minor.",
   "Material.",
   "Severe."
  ]
 }
}
answers · raw probabilities · 3 decimals
duplicate noul
confidence 0.986

Is this invoice a duplicate of one already paid?

  • false0.986
  • true0.014
disposition choice
confidence 0.708

What should happen to this invoice?

  • approve0.002
  • hold0.206
  • manual_review0.708
  • reject0.084
discrepancy_severity score
confidence 0.798 · expected level 2.672

How severe is the discrepancy between invoice and purchase order?

  • 00.043
  • 10.041
  • 20.118
  • 30.798

Provenance: tag primus-decision-0.1 · tree fa0406579f5b5338c2f44de0c42a0626d1a1e1a8 · generated 2026-09-23T15:44:05Z · 25 cases (test-00000-of-00001.parquet).

Gold labels come from the public benchmark's test split (Apache-2.0) and are the mean of three samples from a 4B-class teacher; these examples illustrate the interface. The measured accuracy over all 2,000 decisions is on the benchmark page.