Real ledgers are the obvious training ground for fraud detection and the worst available one: confirmed fraud is rare, labels are incomplete, and privacy law stands between you and the data. Synthetic data solves access - but only measures anything if it is built to let a model be wrong.
Nobody uninstalls fraud detection for missing something; they uninstall it after the third good supplier gets frozen for nothing. Measuring precision requires innocent look-alikes of every fraud pattern in the training data - open receivables that are simply slow, voided cheques that were reissued, non-PO invoices that are just the rent - so false positives exist to be counted.
Every synthetic cfo package regenerates as a fraud-free twin: same company, same seed, fraud set to none, answer key empty. On an innocent company the correct number of flags is zero, so the twin measures the false-positive rate directly - the number an investigation team actually feels. No real company can hand you a ledger certified clean; a generated one can. On the published run, generic anomaly detection flagged hundreds of records on the twin, which is exactly why that baseline earns a failing grade.
More than 25 schemes are planted per world - across procurement, revenue, payroll, inventory, treasury and the intercompany layer - each separable only through behaviour and event trails, and each recorded in a sealed answer key at plant time. Evaluation mode ships the world with annotations structurally withheld, so a model is graded on detection, never on reading a marker column. Scores come back as precision, recall, F1 and a per-vector breakdown against the exact ground truth.
Before trusting any of this, read how real detectors score on the data - the proof page publishes the full experiment, including the failures. Then download the free sample and run your own model against the shipped answer key.