synthetic cfo Back to synthetic cfo
How to generate it

How to generate synthetic ERP data: four approaches, compared honestly

Most guides to synthetic data list tools. This one lists methods, because the method decides what the data can and cannot do long before a vendor does. There are four ways to make ERP data that never belonged to a real company, and each is the right answer to a different question.

1. Statistical synthesis from a source database

A model learns the patterns and relationships in a real database and emits a privacy-safe copy. This is what most synthetic data platforms do, and it is the right choice when you have a production system, permission to touch it, and a need for a copy that behaves like it. What it cannot do: run without the source, or know anything the source did not contain. A real ledger carries no fraud labels, so the copy carries none either, and whatever the source lacked, the copy lacks too.

2. Mock data generators

Field-by-field generators produce plausible names, dates, amounts and identifiers in seconds and are free or nearly so. They are the right choice for user-interface fixtures and unit tests. What they cannot do: keep a debit equal to a credit, make a goods receipt follow a purchase order, or hold a foreign key across tables unless you script every relationship yourself, at which point you are writing the generator.

3. Prompting a language model

A large model will write you invoices, journal descriptions and support tickets on request, in any format, and it is genuinely good at unstructured text. What it cannot do: keep a ledger balanced across thousands of documents, produce the same output twice, or be trusted with a number. Every figure is a plausible-looking token, and at volume the cost and the retries climb. This method now also ships inside the databases themselves, as a procedure that fills the empty tables of a schema with rows a language model writes from the table definitions, a record count and a prompt, keeping the foreign keys you declared. That is the right tool for populating a clone or a demo, and its own documentation says the model can hallucinate and a prompt does not guarantee the result. It is still this method: the rows fit the tables, the books do not have to balance, nothing is labelled, and nothing regenerates. synthetic cfo's first rule is that no language model ever generates financial data; a model may only translate a plain-language request into a validated configuration that deterministic code then runs.

4. Generation from accounting rules

No source database exists at any stage. A company is built from double-entry logic, an accounting framework, a statutory tax treatment and a real ERP table architecture, so the books balance by construction, every foreign key resolves, and every fraud scheme is recorded in an answer key at the moment it is planted, with innocent look-alikes beside it. The same seed regenerates the package byte for byte. This is the right choice when there is no source, when labels are the point, or when the data has to be defensible to an auditor. What it cannot do: be your company. It is a realistic company you configured, calibrated to plausible ratios rather than to your economics, and its record-to-report overlay is a second accounting layer that is not re-consolidated with the operational subledgers, which it says on its own integrity screens.

Which one, in one table
You needUse
A privacy-safe copy of a system you already haveStatistical synthesis from the source
Fixtures for a screen or a unit testA mock data generator
Narrative text, documents, conversationsA language model
Balanced books with a labelled fraud answer key, and no sourceGeneration from accounting rules
A migration test across two ERP architectures from one seedGeneration from accounting rules
Training data where a detector can be measurably wrongGeneration from accounting rules, with innocents planted
Four tests to run on any of them

Whichever route you take, check the output the way an auditor would. Do the books balance to the cent, with every foreign key resolving? Were the labels recorded when the scheme was planted, or derived afterwards by a query that defines them? Do innocent instances of each fraud condition exist, so a rule can be wrong? Does the same seed regenerate the identical bytes? The forensic-grade page walks each test, and the proof page publishes them run against our own data, including the detector that fails.

Keep reading