synthetic cfo Back to synthetic cfo
Beat the textbook rules

Beat the textbook rules

Five frozen benchmark tracks (SCFO-BENCH v1.2, engine v1.8.1). Configuration and seed are immutable, so everyone is graded on byte-identical data. The house baselines are the three detectors from the proof, run here exactly as your model would be: against a database copy with no answer key in the directory. Beat the textbook rules and publish it.

B-01 Sprint · SAP ECC · 1,500 rows · medium fraud · seed 20260101 · 426 labelled fraud records

The entry track: a compact SAP ECC manufacturer with the full four-cycle fraud surface. Fast to generate, fast to evaluate - start here.

#DetectorPrecisionRecallF1GradeWhen
1Targeted forensic rules (ceiling) house baseline0.430.780.55C
2Textbook CAAT rules house baseline0.460.610.52C
3Generic statistical anomaly detection house baseline0.040.150.06F
B-02 P2P Core · SAP ECC · 5,000 rows · medium fraud · seed 20260202 · 1443 labelled fraud records

The standard reference track: 5,000 rows of SAP ECC manufacturing with medium-intensity fraud across procurement, revenue, treasury and the close. The headline comparison number.

#DetectorPrecisionRecallF1GradeWhen
1Textbook CAAT rules house baseline0.520.690.60C
2Targeted forensic rules (ceiling) house baseline0.440.840.58C
3Generic statistical anomaly detection house baseline0.050.160.07F
B-03 Insider Threat · SAP ECC · 5,000 rows · high fraud · seed 20260303 · 2544 labelled fraud records

All five modules including HR/payroll at high intensity: ghost employees, payroll diversion, collusion between the vendor master and the payroll file. The hardest SAP track.

#DetectorPrecisionRecallF1GradeWhen
1Targeted forensic rules (ceiling) house baseline0.560.710.63C
2Textbook CAAT rules house baseline0.600.530.56C
3Generic statistical anomaly detection house baseline0.040.090.06F
B-04 Oracle Cloud · ORACLE · 5,000 rows · medium fraud · seed 20260404 · 1012 labelled fraud records

The Oracle Fusion track: retail seasonality, module schemas (AP/AR/CE/GL), medium intensity. Tests whether a model's skills transfer across ERP architectures.

#DetectorPrecisionRecallF1GradeWhen
1Textbook CAAT rules house baseline0.480.790.59C
2Targeted forensic rules (ceiling) house baseline0.370.940.53C
3Generic statistical anomaly detection house baseline0.030.140.05F
B-05 Universal Journal · SAP S4HANA · 5,000 rows · medium fraud · seed 20260505 · 1068 labelled fraud records

S/4HANA with ACDOCA: the post-migration world. Pharma procurement discipline, medium intensity. Models trained on ECC structure meet the single-table journal here.

#DetectorPrecisionRecallF1GradeWhen
1Targeted forensic rules (ceiling) house baseline0.360.750.48D
2Textbook CAAT rules house baseline0.400.570.47D
3Generic statistical anomaly detection house baseline0.020.080.03F
How scores get here, and what they cannot prove
Computed, not claimedEvery score on this page was computed on our server from the record ids a signed-in submitter uploaded, graded against the withheld answer key by the same scorer that produced the house baselines. Nobody types their own numbers in.
Published by choiceScores appear here only when the submitter ticked publish and named their detector. The name is theirs; no email address is ever shown.
A perfect score proves nothingThe answer key ships inside every generated package, so anyone can submit the key itself and score 100 percent. That is why the bar that matters is honest daylight between your detector and the house baselines, not a flawless row.
Put your detector on the board

Create a free account, open the Benchmark tab, pick an official track and paste the record ids your detector flags. You get the full scorecard either way; ticking publish with a model name puts it here.

Track definitions are frozen; engine version changes create a new suite version and reset comparability, never silently. Baselines regenerate byte-identically from proof/bench_baselines.py.