UK Finance data puts total payment fraud losses at £1.28bn, up 4% year on year, with APP fraud alone up 19%. Most fraud operations haven't grown anywhere near proportionally to that.
With ever-increasing traffic, it's not possible to cover every case in the time allocated. Quality drops, and in some instances cases get skipped entirely. All of which drives up the false positive rate, and creates blindness to mules, because there's no time left for complex investigation.

Key takeaways
|
When resourcing becomes an evidential problem
Experian's research found that of all fraud cases reported, only 37% were confirmed and 63% remained suspected, indicating operational teams may be prioritising efficiency over thorough investigation.
The moment analysts fast-track or skip investigative steps due to high queue volumes, without approval, that changes from a resourcing decision into something closer to systemic negligence under ECCTA's test for "reasonable procedures."
- Outdated software that isn't upgraded to automated tools, behavioural checks, or network mapping, despite known capacity strain, is a genuine exposure point
- Sacrificing thorough checks to clear volume faster works against the "reasonable procedures" standard, even under real resourcing pressure
- The pressure is legitimate. What a team does with it, and how that's governed, is what gets tested
And queue pressure, on its own, doesn't satisfy the reasonable procedures test.
When your fraud rules engine becomes the problem
We often find that rules-based engines tend to accumulate. New rules get added, old ones aren’t retired, and referral volumes drift upwards as a result, often going unnoticed until the numbers are already well past a manageable threshold.
Thresholds worth checking your own numbers against:
- 90%+ junk alert rate — your rules are generating noise, not signal
- ~30 good customers blocked per fraudster caught — the ratio has tipped too far toward false positives
- 15-20%+ referral rate — more of a capacity question than a rules-design one, but still worth watching
Where fraud models and rules collide
A high-performing fraud model and a well-constructed rules engine can both look defensible. But combined and without coordination, referral volume can climb well beyond what either would generate alone.
A model flags a transaction as low risk, say a score of 15 out of 100. The rules engine flags it anyway over something minor, a different shipping address, for example, and it gets referred regardless of what the model concluded.
That's also the diagnostic: if a large chunk of your referred transactions have excellent ML scores but were flagged solely by a static rule, your rules are overriding your model's intelligence rather than working alongside it. Part of what makes this hard to unpick is that most models can't tell you why a case was referred. We've found a way around this. Archetype, Jaywing's fraud modelling platform, surfaces the strongest impact factors alongside every score, so investigators can see exactly what drove the referral and whether the rules engine is adding value or adding noise.
Further reading: Smarter fraud and AML convergence: Escaping the silos
Where network and behavioural fraud prevention earns its place
This is where I think a lot of fraud teams are underinvested, largely because rebuilding a rules engine is a harder internal sell than adding one more rule to the existing pile.
The evidence for graph-based approaches is hard to ignore. HSBC's research presented at KDD 2025 showed that deploying a graph neural network for transaction monitoring reduced AML false positives by 30%, while increasing detection of previously unknown suspicious networks by 20%. The improvement comes from incorporating relationships between accounts, counterparties, and geographies, providing context that rule-based and tabular systems can't capture.
That's the combination most rules engines can't achieve on their own. Every added rule tends to catch more good customers along with the bad. A graph-based approach changes what lands on an investigator's desk in the first place, so the time available gets spent on cases that genuinely warrant it.
It's the same principle behind how Jaywing builds fraud models — starting from what's already in the application and bureau data, instead of requiring new systems or additional data sources to get there.
Further reading: Identifying hidden fraud networks: Why fraud detection needs a network-based approach
The two costs of false positives, and why one is harder to fix
False positives carry two costs that almost always get measured separately: investigator time, and customer abandonment at the point of friction. They're connected, but customer abandonment is the harder one to change.
- Investigator time responds well to better software, cleaner queues, and workflow automation
- Customer abandonment is driven by human psychology, and customers are sensitive to any form of friction, however small
- Reducing friction generally means reducing rules, which reopens the sensitivity-versus-workload trade-off
The alternative is changing how friction gets presented. I'd call this positive friction: designed to help and protect the customer rather than stop them. Most firms are hesitant to impede a customer journey at all, and that hesitation is understandable. But when messaging and information tell a customer that a check is there for their own safety, the response changes. Done well, it generates goodwill rather than abandonment. The customer feels protected rather than obstructed, and the firm gets the investigative step it needs without the drop-off.
That's difficult to execute well. The messaging has to be genuine, timely, and framed around the customer's interest rather than the firm's process. But it's a stronger answer than stripping rules out and accepting the added risk, and it's one more firms should be testing rather than assuming it won't land.
So why doesn't this move faster?
If a fraud model can genuinely outperform an existing score and cut false positives substantially, what's slowing adoption?
Legal risk plays a real part. Strict customer treatment obligations, "treat customers fairly" among them, combined with legacy IT systems, make sticking with the current rules engine feel like the safer option, even when everyone involved knows a model-led approach would perform better.
The IT piece is worth dwelling on. Legacy infrastructure doesn't just slow down implementation, it actively constrains what's possible. A model-led approach requires data pipelines, decision orchestration, and feedback loops that older systems often can't support without significant re-engineering work sitting upstream of the fraud problem itself. That's not a fraud team decision. It's a technology investment decision, and it competes with every other infrastructure priority on the roadmap.
Most organisations I speak to genuinely want to move on this. Where they get stuck is the how: moving from a rules engine that's grown organically over a decade to a model-led approach that regulators, auditors, and internal governance teams are all comfortable signing off on. That's where Jaywing's work tends to start, with the modelling and governance infrastructure that makes the move tractable, rather than the aspiration itself.
But the cost of staying still compounds too. Every month a rules engine runs past the thresholds we've described, the false positive burden accumulates, investigator capacity erodes further, and the gap between what the model could do and what the rules engine is actually doing widens. The case for starting doesn't get weaker the longer it's left.
You may also like: Tackling telecom-enabled fraud through smarter data collaboration

Where to go from here
The problems this piece describes, rules engines generating noise, referral volumes drifting upward, investigators overwhelmed, are fixed by building models that can do more of the heavy lifting, cleanly and explainably.
Jaywing's fraud models, built on our proprietary Archetype platform, exploit the non-linear patterns in application and bureau data that traditional scorecards miss, and deliver fully explainable outputs at both model and individual decision level. That’s key for investigator confidence, as well as for the regulatory explainability requirements that otherwise make a joint outcome model approach so difficult to deploy. Virgin Money is one example of what's possible when that explainability piece is genuinely solved:
- A predictive power of 93% for Virgin Money
- Offered a relative uplift of 31% over the incumbent approach
- Detected over 86.6% of fraud cases by reviewing just the top 10% of case
“The Archetype fraud model is the best of the lot, giving us an exceptionally strong weapon in the fight against financial crime in the application space.” - Nick Martin, Head of Analytics, Virgin Money Credit Card
We've delivered fraud models for over 30 financial services brands. Across those engagements, the consistent outcomes are fewer false positives, higher detection rates, and investigation resource focused where it's needed.
If your referral volumes, false positive ratios, or investigator capacity are telling you something needs to change, we'd be worth talking to.