- Fraud losses keep climbing despite record investment in detection: the FTC reported nearly $16 billion in consumer fraud losses in 2025, up about 25% year over year, so the industry has a detection surplus and a prevention deficit.
- The systems gap is the time and inconsistency created when state, decision logic and enforcement run in separate systems, leaving a window of tens to hundreds of milliseconds that fraud attacks are built to exploit.
- Irreversible instant rails such as FedNow, SEPA Instant and UPI, plus rules like the UK PSR’s APP fraud reimbursement and DORA, have turned batch-adjacent decisioning from a competitive disadvantage into a compliance gap.
- Faster infrastructure alone does not close the gap. Closing it requires a single transaction boundary for state, logic and enforcement, deterministic evaluation against current state, and outcomes recorded atomically as authoritative truth.
- For agentic AI, roughly 90% of production effectiveness comes from the data and decisioning layer: agents need real-time authoritative state, deterministic evaluation of their recommendations, and a complete interaction chain for audit and retraining.
TL;DR
Key takeaway
Fraud losses keep rising despite record investment in detection, because in most financial services stacks the systems that hold state, apply decision logic and enforce outcomes are separate. On instant payment rails, that gap has become a compliance problem. Closing it means putting the decision inside the authorization path, within a single transaction boundary.
Financial services organizations have never spent more on detecting bad outcomes, and they have never lost more money to them.
The FTC reported nearly $16 billion in consumer fraud losses in 2025, up about 25% year over year.1 FinCEN’s suspicious activity reports rose 72% between 2020 and 2025. According to the Nilson Report, global payment card fraud losses are projected to total $407.6 billion over the next ten years.2 None of this is happening because banks stopped investing in fraud detection. It’s happening despite record investment in fraud models, instant payment rails, and AI-driven risk scoring.
That contradiction is worth exploring. The industry doesn’t have a detection problem. It has a detection surplus and a prevention deficit. Models are catching more than ever, but what they catch rarely translates into a decision that’s made, recorded, and enforced before the money is gone.
This is what I mean by the headline. The model is rarely the problem. So, what do I mean by the “gap between systems”?
What is the systems gap?
The systems gap is the time and inconsistency introduced when state, decision logic and enforcement run in separate systems, so a transaction can be approved before the decision about it is current, complete and enforceable.
A useful decision counts only if three things happen at once: the system evaluates the current state, applies decision logic to it, and enforces the outcome before the transaction is authorized. In most financial services stacks, those three things don’t happen at once, and that’s because they happen in three different systems, stitched together by APIs and message queues. That means each has its own timing, its own consistency model, and its own idea of what “current” means.
If you simply walk one transaction through a typical stack, you can watch the milliseconds start adding up. An event lands in a streaming platform, and a fraud model scores it. Sounds great until you realize it’s happening asynchronously, and against whatever state was available when the query ran. Then the score is written somewhere else, and a rules engine or a downstream service reads that score, along with account state pulled from a separate database. Finally, the system decides whether to allow the transaction. By the time this enforcement happens, several hops and several dozen milliseconds have passed, and the state each system used to make its part of the decision was already slightly different from the state the next system saw.
This is where the race condition lives. Two concurrent transactions check the same balance or velocity counter within the same window, and neither check reflects the other because neither has committed yet. That window, which is typically tens to hundreds of milliseconds, is exactly what fraud attacks are built to exploit. It isn’t a bug or a staffing gap. It’s the result of an architecture where state, logic, and enforcement don’t share a transaction boundary.
What’s different about 2026?
This gap isn’t new. It’s existed for years. What’s changed is that money rails no longer give the industry time to close it after the fact.
FedNow has crossed 1,800 participating institutions,3 with 95% of them being small to medium banks and credit unions.4 RTP volume keeps climbing. SEPA Instant is now mandatory across the eurozone.5 UPI is processing more than 24.5 billion transactions a month,6 and RBI’s Digital Payment Security Controls Directions 2026 now mandate real-time screening on top of that volume.7 These rails share one property that batch-era payments didn’t: irreversibility at scale. No overnight batch window remains to catch what real-time processing missed.
Regulation is catching up to that reality. The UK’s PSR now requires reimbursement for authorized push payment fraud,8 which means a bank’s failure to stop a scam in real time becomes a direct balance-sheet cost, not a customer-service problem. DORA makes operational resilience (including the resilience of decisioning infrastructure) a supervisory requirement, not a best practice.9 ISO 20022 is now complete on Fedwire and CHIPS, carrying richer data that regulators and counterparties expect institutions to act on immediately, not reconcile later.
Put together, batch-adjacent decisioning has moved well beyond competitive disadvantage, and has become a compliance gap.
Why faster infrastructure doesn’t close the gap
The reality is that the industry has, until now, accepted a classic “two from three” trade-off. In this case, it’s pick two of scale, latency and data consistency.
Fast-but-eventually-consistent architectures produce duplicate postings, negative balances, and, in the worst cases, sanctioned parties who complete a transaction because the screening result arrived after the money moved. Correct-but-slow architectures get the right answer after the authorization window has already closed, which is functionally the same as getting it wrong. And, to be clear, neither of these failure modes is theoretical or alarmist. Both of them show up in the industry loss numbers stated above.
The instinct this produces is to keep buying faster infrastructure, which is why so many companies focus on their faster streaming, faster models, and faster networks. But speed alone doesn’t close the gap, because speed doesn’t address the underlying issue: nothing in the stack has the authority to make a decision that’s simultaneously current, deterministic, and enforceable. Every vendor in this space claims low latency. Very few can make the resulting decision authoritative, consistent under load, and auditable after the fact. It’s that combination, not just the raw speed, that is the actual differentiator.
What closing the gap requires
This isn’t a call to buy a specific product. It’s a call to evaluate what’s already in the stack against a short list of requirements, because most current architectures will fail at least one.
In practice, that comes down to three requirements:
- State, decision logic and enforcement operate within a single transaction boundary, not across a chain of asynchronous hops.
- Model and rule outputs are evaluated deterministically against the state that is current at the moment of decision, not the state that was current when a query was issued.
- Outcomes are recorded atomically as authoritative truth, not as an event that downstream systems interpret independently.
Institutions that have closed this gap are seeing it show up directly in their loss numbers. Among Volt customers, I’ve seen an 83% reduction in fraud inside the authorization window at one institution, and a tier-1 bank moved its P99 decision latency from 150ms to under 12ms while sustaining 120,000+ events per second. Neither result came from a better model, but from inserting the decision to allow or reject a transaction in the authorization path.
The agentic AI problem, and the 90/10 solution
Agentic AI is the newest layer being added to fraud and risk stacks, but it’s worth thinking about where it best fits. I’d argue that roughly 90% of what makes an agentic AI system effective in production is the data and decisioning infrastructure underneath it. The model is the last 10%. Now, I’m not trying to be prescriptive here. It could be 80/20 or 95/5. That’s not the point.
The point is that current models are genuinely capable of sophisticated fraud reasoning. What matters is what they’re reasoning over, and whether they can sit in the authorization path without degrading the end user experience. An agent that queries a stale data warehouse is reasoning from yesterday’s truth, and it will hallucinate in proportion to how stale that truth is. An agent that produces a confident, well-argued recommendation to block a transaction still hasn’t made a “decision”. It’s made a “suggestion”, and if nothing evaluates that suggestion deterministically against current state, it’s just a new, more articulate source of latency.
Closing this gap for agentic AI means three things happening together. Agents are supplied real-time, authoritative operational state (such as current balances, live transaction history, and active alerts) through structured tools during their reasoning, not from a batch-indexed store after the fact. Their recommendations are evaluated deterministically against business rules and current state before anything is enforced. And the complete interaction chain is captured by design, so it’s available for audit and for retraining the next generation of models and agents.
This is the difference between agentic AI as a pilot and agentic AI as a production system in a regulated institution. Sophistication in reasoning doesn’t buy trust. Real-time, authoritative grounding and deterministic evaluation do.
The constraint was never the model
None of this is a call for another point solution. Institutions can keep buying better detection models, and the losses above will keep climbing anyway, because the model was never the constraint. The constraint is the distance between the moment a system knows something and the moment it acts on it: reliably, consistently, and at the scale and speed these new rails demand.
Until that gap closes, more detection just means more accurate documentation of losses the system couldn’t stop in time.
See how the decision moves inside the authorization path
Sources
- Federal Trade Commission, Consumer Sentinel data release: “FTC data show people reported losing $3.5 billion to imposter scams in 2025”, June 2026.
- The Nilson Report, global card fraud losses and ten-year projection, January 2026.
- Federal Reserve Financial Services, FedNow Service participants and service providers.
- Federal Reserve FedNow Service, “FedNow Service achieves new participation milestone: 1,000-plus financial institutions”.
- Regulation (EU) 2024/886 on instant credit transfers in euro (Instant Payments Regulation), EUR-Lex.
- National Payments Corporation of India, UPI product statistics.
- Reserve Bank of India, Digital Payment Security Controls Directions.
- Payment Systems Regulator, APP scams reimbursement policy and publications.
- Regulation (EU) 2022/2554 on digital operational resilience for the financial sector (DORA), EUR-Lex.




