synthetic cfo Back to synthetic cfo
Synthetic ERP data

Synthetic ERP data, generated from accounting rules

Synthetic ERP data is enterprise resource planning data - purchase orders, invoices, payments, journal entries, bank statements - that never belonged to a real company. There are two ways to make it, and they are not interchangeable.

Masking real data vs generating from rules

Most synthetic data tools start from a real database: they learn its statistical patterns and emit a privacy-safe copy. That works when you have a source system and permission to touch it. It fails when you have neither, and it inherits whatever the source contained.

synthetic cfo takes the other road: a cold start engine that generates complete company worlds from accounting rules alone - IFRS and GAAP treatments, real ERP table architecture, and double-entry logic. No seed dataset exists at any stage, so there is nothing to anonymise, nothing to re-identify, and nothing to clear with a privacy team.

What a generated world contains

Seven business cycles on three platforms (SAP ECC, SAP S/4HANA, Oracle Cloud): procure-to-pay, order-to-cash, cash and treasury, record-to-report, payroll, inventory and fixed assets. Every foreign key resolves, every debit matches a credit, the trial balance nets to zero, and the financial statements articulate from the operational subledgers with a tie-out workpaper. The world carries two accounting layers and says so: the record-to-report overlay (manual journals, trial balance, close) is not re-consolidated with the subledgers, so the trial balance ties to the general ledger journal layer and never to the payables or receivables documents. Documents carry a real approval lifecycle - release events, invoice holds - and every bank statement line mirrors an operational payment, receipt, payroll run or tax remittance, with the per-bank reconciliation proven to the cent at generation time.

Multi-year fraud arcs run up to five fiscal years of the same company, and a two-ERP intercompany group runs as an arc too, so multi-entity and multi-year come in one archive with an elimination schedule that carries innocent open items beside the planted fraud.

Why the answer key is the product

More than 25 fraud schemes are planted into the data - ghost vendors, duplicate invoices, kiting, lapping, channel stuffing, self-approval above authority - and every instance is recorded in a ground-truth answer key at the moment it is planted. Innocent look-alikes of every fraud condition exist beside the real thing, so a detector or a model can be wrong, which is what makes a score mean something. The same seed and configuration regenerate the identical dataset byte for byte, with a certificate of file hashes to prove it.

How to get it

A complete sample package is free to download with no account, and the free developer plan generates on all three platforms. Programmatic access ships on every plan: an API key plus the Python client (one pip command on the evaluate page) generates, polls and downloads from a few lines of Python, or from one terminal command with no code at all.

Keep reading