According to S&P Global Market Intelligence's 2025 survey of over 1,000 IT and business leaders across North America and Europe, 42% of companies abandoned most of their AI initiatives, more than double the 17% recorded in 2024. The average organisation scrapped 46% of its AI proofs of concept before they ever reached production.
MIT's 2025 research on generative AI in business found much the same pattern from a different angle: 95% of organisations using generative AI pilots saw no measurable return on their investment.
Banking sits squarely inside these numbers. Most banks we talk to have already run several AI pilots. Fraud detection, document processing, credit scoring, customer service, the use cases are rarely the problem. What's harder to find is a pilot that made it into live, governed, revenue-affecting operation.
The instinct when a pilot stalls is often to run another one, in a different department, with a different vendor, hoping the next attempt clears the same wall the last one hit. That rarely works, because the wall is the absence of a safe, repeatable path from a working demo to a production system that a bank's risk, compliance, and IT operations teams can all stand behind, paired with a clear definition of what success actually looks like in business terms.
Key takeaways
- Many AI pilots in banking prove technically successful and still never reach production, often because the path from a working demo to a governed live system was never mapped out as clearly as the pilot itself.
- A safe path to production means traceable decisions, drift monitoring, and human oversight built in from the pilot stage, not added once DORA or the EU AI Act make it unavoidable.
- The business metrics that justify scaling an AI system, cost-to-serve, exception rates, risk-adjusted return, are different from the ones that prove a pilot works, and need to be defined before the pilot starts.
- Banks that consistently reach production plan the transition before running the pilot, using a readiness assessment and phased rollout rather than negotiating the path forward after the fact.
The state of AI in banking: plenty of PoCs, very little production
More than 75% of UK financial firms already tried leveraging AI in 2024 in one form or another and 92% of EU banks are using AI agents in production. This is not surprising at all, given that fairly simple AI tools can automate routine banking work with manual processes and significantly reduce the burden on human employees, for example by providing 24/7 support in customer service.
It’s worth being precise about what these adoption figures actually measure. ‘Using AI’ in these surveys spans everything from a chatbot handling routine queries to a model embedded in core credit decisioning, and most of that footprint sits in lower-stakes, advisory use cases. The Bank of England and FCA’s own survey found that 62% of AI use cases are rated low materiality by the firms using them, with only 16% rated high materiality, which is a more precise way of putting it: broad AI adoption and a scarcity of AI running inside critical, governed production processes coexist quite comfortably.
However, MIT’s 2025 research on generative AI in business found a steep drop-off between evaluation and deployment for custom, enterprise-grade tools specifically, the category most banking use cases fall into. 60% of organisations evaluated these tools, 20% got as far as a pilot, and only 5% reached production. Generic tools like ChatGPT saw higher adoption, but they solve individual productivity problems, not the fraud detection, credit scoring or document processing challenges that actually move a bank’s numbers.
Financial services shows up in that same research as one of the sectors furthest from structural change. The report’s disruption index, built from indicators like market share shifts, new AI-driven business models, and changes in customer behaviour, places financial services near the bottom of nine major sectors, describing the pattern as backend automation with customer relationships left largely untouched. In other words: banks are running plenty of pilots behind the scenes, but very little of that activity is reaching the point where it changes how the bank actually operates or competes.
This matches what we see when we talk to banks directly: many banks face challenges in integrating AI with internal systems and meeting data governance requirements. It holds across most of the use cases banks are piloting right now:
- Fraud detection: models built for transaction monitoring get validated on their ability to flag suspicious activity and spot unusual behaviour or shifting spending patterns, with success measured almost entirely by how well the model reduces false positives against a known set of fraud tactics, without much thought for what happens once it has to run continuously against live accounts and feed into anti-money laundering processes that compliance teams are accountable for
- Document processing: a pilot for document verification, extracting data from bank statements using natural language processing, succeeds against a curated set of documents and stalls the moment it meets the inconsistent formats and edge cases a production system actually has to handle
- Customer service: a virtual assistant pilot handles routine inquiries well enough in a demo, but scaling it across real customer interactions and digital channels, in a way that holds up across the full customer journey rather than a scripted subset of it, turns out to be a much harder problem
- Corporate banking and wealth management: the same pattern shows up here too, where pilots for things like portfolio insights or client onboarding tend to hit the same wall once they move past a small group of test users
The pattern is consistent enough that it’s worth naming directly. Proving a use case works has stopped being the hard part of banking AI. What actually determines whether a project survives is everything that happens after that proof, in the stretch before the system ever touches a live customer or a real transaction.
Enhancing the architecture of an application that enables over 20 million invoices to be processed each day
Read the case studyWhy AI PoCs in banks stall before production
Most banking pilots are scoped to answer a narrow question in the shortest time possible: can this model do the thing it’s supposed to do. That scope works well for getting a green light from leadership, and works badly as an AI readiness exercise and a foundation for anything that has to hold up under audit, high transaction volumes, or a regulator asking how a decision was made. Shortcuts that are perfectly reasonable in a three-week pilot, a manually curated dataset without focusing on data quality and data protection, a workaround for a system integration, no formal model documentation, become liabilities the moment someone proposes putting the same system in front of real customers.
The caution that keeps AI pilots in banks stuck
Banks are, understandably, risk-averse with anything that touches customer money or customer data. The problem is that this caution usually shows up as an instinct to slow down, rather than as a defined set of conditions a project needs to meet to move forward safely. Without that structure, an artificial intelligence pilot that works doesn’t get a clear route to production. It gets parked, revisited in the next planning cycle, and quietly deprioritised in favour of the next pilot, because pausing tends to be the safer option in the absence of a clearly specified threshold for what ‘safe enough to launch’ actually looks like.
What it actually costs to scale AI PoCs in banks
A pilot’s budget covers a handful of weeks and a small team, so the cost of running the same system at the scale of an entire bank often isn’t scrutinised until much later in the process. Production changes that calculation completely: licensing, infrastructure, monitoring, and the people needed to keep a model behaving correctly against live data all show up as recurring costs rather than a one-off project expense. That’s often the first moment the business case gets recalculated in the terms production actually requires, rather than the terms that got the pilot approved.
What a safe path to production looks like
A pilot proves a model works, and production proves an organisation can be held accountable for it, which is a different and considerably higher bar, and meeting that bar means building proper governance frameworks around the machine learning models involved rather than treating governance as a formality to satisfy once the technology has already been signed off.
For a bank, that means the system needs traceable decisions, so a credit or fraud call can be reconstructed and explained after the fact. That accountability matters more than it might first appear: the Bank of England and FCA’s 2024 survey found that 55% of AI use cases in UK financial services already involve some degree of automated decision-making, which means the majority of AI in this sector isn’t advisory, it’s shaping or making calls that affect real customers. Meeting regulatory expectations at that level usually calls for some form of explainable AI, where the model’s output can be traced back to the factors that drove it rather than treated as a black box result. It also means monitoring that catches a model drifting away from the data it was trained on, rather than someone noticing months later because complaints started rising, and clear human oversight at the points where a regulator, or the bank’s own risk function, expects a person to be able to intervene.
None of this is optional under current EU rules. DORA requires financial institutions to treat AI systems supporting critical functions as part of their ICT risk management framework, with the same regulatory compliance rigour applied to resilience and incident response as any other critical system. The EU AI Act adds requirements specific to the use case: credit scoring, for instance, is classified as high-risk, which brings its own set of regulatory requirements around documentation, human oversight and conformity assessment that a pilot was never built to satisfy.
Treating these as compliance risk to be managed from the outset, rather than paperwork to complete once a system is nearly live, is what actually separates a governed production system from a pilot wearing a production label.
We’ve covered this in more depth in the article about AI adoption in banking and resilience, including why on-premise deployment is becoming the practical answer for banks that need to meet these obligations without giving up control of their data.
The banks that get this right build it into the pilot from day one, even though that makes the pilot itself slower and slightly more expensive to run. The alternative, adding it once a system is already earmarked for production, tends to mean a compliance review surfacing a long list of gaps at the exact moment the project has the most momentum and the highest cost of turning back.
A secure on-premise foundation for compliant AI development
Read the case studyMeasuring the effect: what to track instead of PoC success
A pilot is usually judged against a technical benchmark, accuracy, precision, speed against a test set, and those numbers are the wrong ones to carry into a production business case. They tell you the model works, but there is almost nothing about what it’s worth to the bank’s actual business strategy.
The metrics that actually justify scaling a system are business ones, and they tend to be less flattering than the pilot numbers once real volumes hit them. The ones worth tracking before a project ever reaches a funding decision:
- Cost-to-serve for the process the AI is touching, measured against real volumes rather than the pilot’s sample size
- Processing time from the customer’s perspective, not the model’s inference time, since a fast model still means a slow process if a human has to review most of its outputs
- Exception rate, the share of cases still requiring a person to step in, which is almost always higher in production than it was against a clean test set
- Internal efficiency gains, measured as the actual reduction in manual work across the internal processes the system was meant to change, not the theoretical gain a pilot demonstrated in isolation
- Risk-adjusted return, accounting for any new compliance exposure the system introduces, since a system that saves money on paper while creating new regulatory risk hasn’t necessarily made the bank better off
Defining these before a pilot starts, rather than after it succeeds, changes the nature of the project entirely. It gives the team building the pilot a target that actually matters to the business, and it gives everyone deciding whether to fund production a number they can hold the result against instead of a technical score that was never designed to answer that question.
Getting from pilot to production
The banks that do get AI into production successfully tend not to run a fundamentally different kind of pilot. They run the same pilot inside a different process, one where the transition to production was planned before the pilot ever started rather than negotiated after it succeeded, and where AI adoption was treated as an operational commitment rather than a one-off experiment from the beginning.
That starts with an AI readiness assessment before scaling, not after: a clear-eyed look at whether the data, infrastructure, and governance a production system needs are actually in place, rather than assuming they’ll sort themselves out once the model has proven its point. It continues with a phased rollout that has defined checkpoints, so the decision to expand isn’t a single high-stakes leap from pilot to full deployment but a series of smaller, reviewable steps. And it depends on infrastructure decisions, cloud versus on-premise among them, being made deliberately for the production system’s actual requirements, rather than inherited by default from whatever the pilot happened to be built on. Embedding AI into a bank’s actual operations is what separates this kind of rollout from a pilot that simply gets a longer leash.
None of this makes the pilot itself faster. It makes the six months after the pilot succeeds look completely different, because the questions that would otherwise stall the project for a year have already been answered.
If your organisation is trying to work out where it currently stands on that path, our [AI Readiness Assessment] gives a structured picture of the gaps most likely to stall a project before it reaches production, and our on-premise AI Advisory service works through the infrastructure decision directly, for the cases where data sovereignty or regulatory exposure make cloud-first the wrong default.
Conclusion
Banks don’t have a shortage of AI pilots, and running another one rarely fixes what stalled the last. What separates the institutions getting real value from AI is a deliberate path from proof of concept to governed production, built before the pilot starts rather than negotiated after it succeeds, and a definition of success that was written in business terms from day one.
That’s a different discipline from running experiments, and it’s the one most banks haven’t yet built. If you’re looking at a pilot that worked and still hasn’t gone anywhere, or planning the next one and want to avoid the same outcome, the AI Readiness Assessment is a practical place to start.
AI Readiness Assessment
Gain a clear view of how prepared your data is to support and scale AI initiatives in your organisation.