Proof
Measured the way a sceptic would.
Three public benchmarks, no human in the build, scored on data the agent never saw.
The protocol
Three rules that keep the answer key locked.
- 01
Never seen
Each detector is scored on recent history the agent never saw.
- 02
Never tuned on
Scored when the detector registers, and never fed back into the build.
- 03
Scored strictly
Every alert must match one real transaction.
Benchmark
IBM AML, HI-Small
- Transactions
- 5.08 million
- Laundering
- 0.10%
- Split
- Temporal 60 / 20 / 20
- Scored on
- 1,015,564 held-out transactions
Minority-class F1, per transaction
0.557
About 5.5 hours, unattended · $4.63 in model calls
| Model | F1 |
|---|---|
| GIN | 0.287 |
| GIN + edge updates | 0.477 |
| Lucir, built by the agent | 0.557 |
| PNA | 0.568 |
| LightGBM + graph features | 0.629 |
| XGBoost + graph features | 0.632 |
About a point below PNA, a purpose-built graph network. The best published methods are further ahead.
Precision 0.897: about nine in ten alerts are laundering.
Scoring: 0.584 at registration under our earlier, looser matching; 0.557 is the same detector re-scored strictly. Third unattended attempt; the first two registered nothing.
Benchmark
Elliptic Bitcoin
- Transactions
- 203,769
- Time steps
- 49
- Split
- Train 1 to 34, test 35 to 49
- Features
- 166 per transaction
Source: Weber et al., Anti-Money Laundering in Bitcoin, KDD 2019 Workshop on Anomaly Detection in Finance
Illicit-class F1
0.625
8 minutes · $0.42 in model calls
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Logistic regression, all features | 0.404 | 0.593 | 0.481 |
| Lucir, built by the agent | 0.816 | 0.506 | 0.625 |
| GCN | 0.812 | 0.512 | 0.628 |
| MLP, all features | 0.694 | 0.617 | 0.653 |
| Skip-GCN | 0.812 | 0.623 | 0.705 |
| EvolveGCN | 0.850 | 0.624 | 0.720 |
| Random forest, all features | 0.956 | 0.670 | 0.788 |
Level with the paper's GCN, in eight minutes. The random forest is still well ahead.
A dark market closed at step 43; the paper reports no model catches what follows.
Benchmark
IEEE-CIS card fraud
- Transactions
- 590,540 labelled
- Fraud
- 3.5%
- Split
- Last 20% by time held out
- Scored on
- 118,107 held-out transactions
Source: Vesta Corporation and IEEE-CIS, Fraud Detection competition, Kaggle 2019
ROC-AUC, every held-out transaction
0.926
61 minutes, unattended
| Model | ROC-AUC | PR-AUC | F1 |
|---|---|---|---|
| Plain gradient boosting, raw columns | 0.920 | 0.582 | 0.553 |
| 1st-place recipe, past data only | 0.924 | 0.583 | n/a |
| Lucir, built by the agent | 0.926 | 0.576 | 0.550 |
| 1st-place recipe, whole period | 0.945 | 0.663 | n/a |
Level on ROC-AUC with the published winning recipe when that recipe may only use earlier data, as live monitoring must, and a little behind it on PR-AUC.
Most of the winners' lead came from aggregates computed over the whole period, future included, which Kaggle's rules allowed. What the leaderboard knew
A correction
We withdrew a number. Here is the one that holds up.
In August our internal score was 0.567. We then found three ways the holdout leaked into the build, closed them, and rebuilt: 0.584 on the same scoring.
Then we made the scoring strict, one flag to one transaction. That gives 0.557, the only figure we stand behind.
Read the full accountNext
A third benchmark, when it is honest.
Known typologies injected into real bank history. Published once injected cases leave no trace.