Most synthetic data tools make data that looks right. Forensic-grade is a higher bar: data that holds up when a professional, or an automated audit tool, actually works it. Four tests separate the two, and any dataset can be checked against them, whoever made it: a vendor, a script, a language model, or the generator built into the database itself.
Every debit has a matching credit, every foreign key resolves, the trial balance nets to zero and ties to the general ledger journal layer, and the financial statements articulate from the operational subledgers with a workpaper showing the tie-out. Where a dataset carries two accounting layers, as this one does, forensic-grade means saying so rather than faking a consolidation. If Table A shows an invoice and Table B shows its tax, Table C must show the cash clearing exactly. One violated constraint breaks the illusion for any audit tool.
A fraud label must be written at the moment the scheme is planted - never derived afterwards by querying the finished data. Derived labels are a tautology: if every open receivable is labelled phantom revenue, a detector that flags open receivables cannot be wrong, and a model trained on the data learns that an open receivable means fraud, which is false in every real ledger on earth.
Real books are full of innocent instances of nearly every fraud condition: receivables open because they are not yet due, cheques voided over a misprint and reissued, rent invoiced with no purchase order, year-end intercompany sales still in transit. Forensic-grade data plants these beside the fraud, so precision is measurable and a detector can be wrong. Two conditions are deliberate closed sets where the join is the analytic, self-approval and the ghost vendor, and the data should say so rather than hide it.
The same seed and configuration must regenerate the identical dataset byte for byte, with published file hashes. That is what makes a benchmark result defensible rather than a screenshot.
Take each label family. Write the simplest query that selects those rows. Count how many rows the query selects that are NOT labelled. If the answer is zero, the answer key is inside the data and every score measures the generator, not the model. synthetic cfo publishes this test run against its own data, along with real detectors' scores, on the proof page - including the round where the test caught our own engine and forced a rebuild.