The FCA published the Mills Review on 6 July 2026 which draws on 140 written submissions, consumer research, and industry engagement to map how AI could reshape retail financial services by 2030. Seven recommendations follow. No new rules are immediately proposed, with frameworks (Consumer Duty, Senior Managers Regime) that follow principles based regulation judged to be suitable to adapt to the risks.
I think it's a smart long-term plan. But it doesn't fully address the immediate dangers posed by cyber crime, and it leaves traditional lenders carrying the risk for technology they do not control.
Key takeaways
|
What the review gets right, and where it leaves firms exposed
The FCA are not proposing a new AI specific rulebook, but suggest adapting current principles-based frameworks instead. Relying on existing rules like the Consumer Duty and the Senior Managers Regime is a sound call. The autonomy spectrum (judging AI by how much it does on its own) is a highly useful tool for firms to assess their own technology and how they use it.
We think that there are several issues that require further thought and action from both the regulator and firms:
- Proving a firm is treating customers fairly becomes incredibly difficult when decisions are automated
- Annual validation and testing will be insufficient — firms will need to monitor every AI decision continuously
- The review struggles to reach the technology firms that actually train and deploy the models, relying on indirect pressure on regulated lenders instead
- Because the FCA cannot regulate Big Tech, banks and lenders carry all the legal and compliance risks for systems they do not own, or fully control, and some of those risks may be genuinely outside their control
The most pressing issue is urgency. The FCA is planning for 2030, and structural change takes time, significant threats are happening today. The review notes that AI will increase fraud and cyber risks (indeed the notable events where the unreleased Anthropic model ‘Mythos’ autonomously identified a host of major security vulnerabilities, happened after the review was finalised), but its recommendations lack urgency in response. AI is already used to create synthetic identities and launch cyber attacks faster than firms can respond. Firms and consumers need defences now.
Recommendation 3: How far are firms from continuous model governance, really?
Recommendation 3 says point-in-time model validation won't be enough for AI systems, and that firms need more continuous approaches to model governance and monitoring. From what I see in client work, most firms are further from this than the review implies.
The knowledge problem
Many firms are aware they lack the knowledge and experience to manage the new risks AI brings. We see this when we walk model risk teams through what AI governance actually requires. This is understandable given the pace of change, the fundamental differences in the new technology, and how accessible and flexible AI is across a wide range of use cases meaning it can appear in all manner of business areas not typically associated with ‘models’.
MRM teams are experts in static, deterministic financial risk models, where fixed inputs produce consistent outputs. Many have limited experience managing the specific risks of probabilistic AI, including the fact that a third-party vendor can update a model's weights overnight, which would invalidate any previous validation testing and compliance checks.
New technology into legacy frameworks
We see risk functions considering how to push AI through governance structures designed to manage more linear problems. While the regulatory guidance from the FCA and PRA is that current principles-based approaches are appropriate, applying these frameworks without adaptation is insufficient. A point-in-time, annual validation applied to technology that requires more frequent (or even real-time) monitoring doesn't work.
Risk teams don't generally have the enterprise tooling to measure things like semantic drift or to deploy automated oversight, using AI to judge AI. They are governing unpredictable systems with static tools built for traditional risk models.
The market is moving at different speeds
A capability divide has opened up and is likely to widen before it improves. While some larger institutions have established dedicated AI model risk functions, the broader market is fragmented. For most lenders, there is a split between data science teams wanting to deploy live AI and traditional validation teams working within frameworks that require point-in-time sign-offs. Until these teams align on what good looks like, continuous monitoring will remain out of reach for the majority.

What the review misses on Consumer Duty compliance
The review confirms that the Consumer Duty applies and confirms the rules haven't changed. But it gives limited guidance on what constitutes acceptable evidence in a highly automated system. It confirms the preferred outcomes-based regulation but offers no real support on how to evidence it when the decision-making is automated. For firms trying to act now, that's a gap between intent and instruction.
Three problems that follow:
- Evidence. At Levels 4 and 5 on the autonomy spectrum, where AI is making decisions, AI is also generating the evidence that those decisions were fair. The review offers no direction on whether machine-generated evidence is suitable to prove Consumer Duty compliance.
- Accountability. The SM&CR explicitly forbids delegating ultimate accountability to a machine. If an automated system generates a biased compliance log, the FCA holds the Senior Manager accountable. Firms remain entirely responsible for outcomes regardless of how automated the process became.
- Bias. If the underlying training data is unbalanced, the AI's responses could be systematically biased against certain groups. Because the AI is also generating the compliance logs, that bias could be amplified and hidden from the human observer, creating a regulatory breach that the automated system is recording as compliant.
This may have the unintended consequence that risk-averse lenders stall their AI deployments for concern over breaching the rules.
Recommendation 6: What continuous supervision means for lenders
Recommendation 6 proposes an AI-enabled supervisory model where the FCA uses AI tools across authorisation, supervision and enforcement, including near real-time monitoring of outcomes and, over time, agent-to-agent communication between the regulator's systems and firms' own.
This requires lenders to greatly increase their capabilities, moving from the current periodic reporting to building and maintaining live data feeds.
Legacy versus challenger firms
Agent-to-agent communication requires high-frequency, automated API polling of structured data. That's a different proposition from pulling a manual data sample once a quarter. Newer fintechs and challengers are starting from a higher base and can adapt more easily. Retail banks and specialist lenders carry legacy system issues and may face resource constraints. FS firms are moving more cautiously than non-regulated firms, not out of ignorance, but because they are bound by existing model risk regulations and better understand the consequences of model uncertainty and error.
The marking-its-own-homework problem
If AI makes a lending decision, and AI generates the evidence and compliance log for measuring Consumer Duty outcomes, the system could be seen to be marking its own homework. We know that systems can be wrong, or can confidently present facts that have been hallucinated. Lenders must consider the need to maintain some independent, human-led checks between their decision engines and their reporting feeds so they do not pass automated errors to the regulator.
Regulatory pace
I’ve been impressed with the FCA’s work to engage with the industry, particularly in the testing and sandboxing offered to support firms explore use cases in tandem with them.. They have shown genuine openness to understanding and adopting the technology. However, there is a very real risk that the traditional regulatory cycle of assessment, consultation and review is too slow for the developments we are seeing in AI. The Mills Review itself demonstrates this where part of the response to the AI acceleration is to suggest more reviews. They must ask, is the advancement in the capability (and therefore risk) moving faster than the regulatory cycle can handle?
There is also a proportionality question Recommendation 6 doesn't resolve. What happens if a lender has a strict "zero AI" risk appetite but is delivering perfectly adequate customer outcomes? The FCA cannot force all firms to adopt agentic workflows to satisfy an automated reporting format. Good firms should not be penalised because they lack the infrastructure to connect to the FCA's reporting requirements.
The long-term gain
If the models can be effectively tuned for regulatory proportionality, the potential benefits are real. If routine compliance questions can be answered automatically by data and AI agents, the administrative burden on knowledge workers reduces on both sides. That would allow both lender and regulator to focus more resources on genuine, strategic risk management, which is where human judgement should sit.

What to do now (a quick checklist)
You cannot validate AI as if it were a traditional credit model. The underlying principles of MRM frameworks are sound, but their application for AI must be adapted. Three things worth starting now:
#1. Redefine risk appetite and ownership. Traditional credit risk produces predictable, deterministic outputs. Generative AI is entirely probabilistic, and the outcome is not completely explainable or repeatable. Establish who owns this new category of risk, then set a specific risk appetite that accepts qualitative uncertainty, for example setting a hard tolerance for factual hallucinations rather than demanding 100% static accuracy.
#2. Educate the entire risk function. We see that effective management of AI risk is reduced by a lack of understanding at oversight level. Train your staff and senior leadership on what machine learning and LLMs actually are, how they work, and why their output is inherently non-deterministic. If your validation teams expect fixed logic, they will reject every AI use case presented to them.
#3. Redefine the model. Focus on the LLM use case. You cannot validate a massive, third-party foundation model. The data and training methodology from the main providers are not accessible, and attempting to do so creates an impossible situation for current MRM policy. Adapt your framework so the "model" being validated is the specific use case. Governance should cover the LLM, the system instructions, and the data retrieval pipeline. Manage the probabilistic risk by enforcing rule-based boundaries, like "no advice" limits, to secure the commercial opportunities these models offer.
The speed and improvement in the technology brings with it an unprecedented level of new risk and opportunity. Forcing AI into legacy MRM frameworks will restrict your data science teams and risk making you uncompetitive. To take the opportunity, accept that the risk is genuinely new, and adapt how you manage it.
Further reading. You may also like: