Category: AI Regulation | Reading time: 9 minutes

Category: AI Governance | Reading time: 9 minutes

The scenario

A compliance officer at an FCA-regulated wealth manager is asked to approve a client communication. She wants to check it against the firm's own policy, so she asks the firm's Microsoft Copilot: can an adviser send this without additional compliance approval?

Copilot searches SharePoint. It finds a policy, quotes the relevant section, explains the rule and gives her a clear answer. The communication goes out.

The policy it relied on was superseded six months earlier. The current version introduced an additional approval requirement. Nobody noticed.

A complaint follows. The firm reconstructs the decision and finds that Copilot retrieved an obsolete document sitting alongside the approved version in the same library.

Nobody made a mistake in any sense that would have been visible at the time. The compliance officer asked a reasonable question. The tool answered it accurately from the document it found. The document was there because nobody removed it.

The defence that will not work

The instinctive response is that Copilot found the wrong document.

It did not, in any sense the tool would recognise. It searched an approved corporate source, found a document that matched the query, and reported what it said. There was no fabrication and no confident invention. None of the failure modes that dominate discussion of AI risk are present here.

The failure was earlier and duller. A superseded document remained retrievable, with nothing marking it as superseded, in a library the firm had designated as authoritative.

Your firm has probably had superseded documents on SharePoint for twenty years. What changed is that something now reads them and speaks with confidence.

So the governance question is not why Copilot got it wrong. It is why an obsolete document was still capable of being treated as authoritative in a regulated decision.

Where the accountability sits

Under the Senior Managers and Certification Regime, the compliance oversight function is SMF16. Where the failing control is a compliance control, that is the natural place to start looking.

But it is not automatic, and any article that says SMF16 is liable is overstating the position. Accountability depends on the firm's governance structure, the individual's Statement of Responsibilities, the activity concerned, and whether reasonable steps were actually taken. The FCA's design principle is that responsibility should sit with the most senior person responsible for the relevant activity, and that firms should be able to show clearly who does what.

The accurate framing is this. If compliance oversight sits with SMF16 at your firm, and the failing control is a compliance control, that is exactly the kind of failure they may be asked to explain. Whether anyone is personally liable turns on evidence of reasonable steps, not on job titles.

The question a supervisor would ask is not "why didn't you personally check this Copilot answer?" That misunderstands how senior management accountability works. It is closer to: what reasonable steps were taken to ensure the control environment could identify and manage this risk?

There is an awkward second question here, and it is the one firms find hardest. Document control usually sits with operations or IT. Retrieval quality sits with nobody. When a failure runs across a boundary like that, the Statement of Responsibilities is where you find out whether the boundary was ever drawn.

What the regulator has actually said

The FCA published its AI Update on 22 April 2024, setting out how existing rules apply to AI rather than creating a separate regime.1 It followed DP5/22, published jointly with the Bank of England and the Prudential Regulation Authority on 11 October 2022.2

The position was settled more recently and more directly. The FCA's Mills Review, published on 6 July 2026 and led by Executive Director Sheldon Mills, found that the existing principles and outcomes-based framework remains sound, including the Consumer Duty and the Senior Managers Regime, and recommended no AI-specific rules. It also recorded that firms remain responsible for the services they provide, including where they use third-party models or agentic tools.3

That last clause is the one that matters here. Copilot is a third-party tool. The Review does not treat that as a transfer of responsibility. It treats it as a reason the evidence becomes harder to assemble.

The Review's organising idea is an autonomy spectrum, along which the human role moves from operator, through collaborator, consultant and approver, to observer. Accountability does not move along with it. What changes is how much work it takes to demonstrate.

For firms in scope of the EU AI Act, there is a parallel expectation. Article 12 requires high-risk systems to allow automatic recording of events over the system's lifetime, and Article 72(2) requires post-market monitoring to include, where relevant, an analysis of the interaction with other AI systems.4 The regulation is already asking about the system's record and its connections, not only about its outputs.

SharePoint access is not governed knowledge

This is where AI projects most often go wrong. A firm connects Copilot to SharePoint and says the AI is grounded on our own documents.

That sounds reassuring until you consider what is actually in that environment. The current policy. The previous policy. A draft replacement. An old board paper. Regulatory guidance from two years ago. A downloaded copy of legislation. Someone's internal interpretation of that legislation. A training deck summarising the rule for new starters.

All of those are relevant to a search. They are not equally authoritative.

A retrieval system can find a document that is semantically relevant while still retrieving the wrong document for the decision being made. So the question is not whether Copilot can find your documents. It is whether the organisation can distinguish which documents should be relied on, which need qualification, and which should no longer influence a regulated decision at all.

Why "latest document wins" does not work

The obvious fix is to prefer the most recent document. It fails immediately, and two examples show why.

An employee uploads a training presentation yesterday explaining an FCA requirement. Beside it sits the authentic regulatory source published last year. The training deck is newer. It is not therefore the stronger authority.

A policy dated January is formally approved. A March draft contains proposed changes that have not yet been through the board. The March document is newer. It is not therefore current policy.

Recency is a weak proxy for authority, and in a regulated environment it is often an inverted one. What governed retrieval needs is not date sorting. It is source authority, recorded explicitly.

What a governed library records

There is no FCA-prescribed metadata schema for Copilot, and anyone offering one is inventing it. But if a firm wants AI-assisted decisions to be reconstructable, its knowledge environment has to be capable of distinguishing several things.

Document status. Current, amended, superseded, withdrawn, draft, or not yet in force. As a field, not as a folder name. Superseded documents move to an archive excluded from normal retrieval rather than sitting beside the current version.

Citation priority. Authentic legislation or binding rule, official consolidated reference, official non-binding guidance, official commentary, internal interpretation, training resource. When two documents both answer a query, something other than relevance scoring has to decide which wins.

Legal effect. Binding, guidance, internal policy, working document. A system that says "you must" when the correct phrasing is "guidance suggests" has made a material error, and nothing in the retrieval layer catches it unless the distinction is recorded.

Text role. The authentic source, an amending instrument, a consolidated working text, or an explanatory document. These are not interchangeable, and consolidated texts in particular often carry no independent legal effect.

Relationships. What amends what. What supersedes what. Marking a document superseded is not enough on its own if nothing records what replaced it.

Verification. When the source was last checked, against what, and by whom. A date field nobody updates is worse than no field, because it implies a check that did not happen.

These look like document management questions. Once AI begins retrieving organisational knowledge and presenting it as an answer, they become part of the control environment around the model.

A confident answer can still be a governance failure

This is the uncomfortable property of generative AI that most discussion misses.

An answer can be well written, logically structured, correctly formatted and apparently well sourced, and still be wrong for the decision being made. The model does not need to hallucinate for an AI incident to occur. It can accurately summarise the wrong source.

Most AI risk discussion focuses on fabrication, which is dramatic, checkable and relatively easy to design against. Retrieval introduces a quieter class of risk. The information exists. The system finds it. It interprets it correctly. And the organisation should never have allowed that information to determine the answer.

That is not a model intelligence problem. It is a governance architecture problem, and no improvement in the model will fix it.

Human in the loop is not enough

There is a phrase worth being careful with: a human always checks the output.

Good. But what does checks mean?

Does the reviewer see the underlying source? Do they know whether it is current? Can they tell regulatory authority from internal commentary? Are conflicting sources surfaced to them? Are superseded documents flagged? Does the system record that the review happened? Could the firm reconstruct it in eighteen months?

Human oversight becomes meaningful when the reviewer has enough information to exercise judgement and the organisation can demonstrate that the judgement occurred. Without both, human in the loop is a governance slogan rather than a control.

Three questions for firms using Copilot today

If your firm uses Microsoft Copilot, ChatGPT Enterprise or another AI assistant against internal knowledge, this is the test.

Can you prove which version of the policy the AI relied on? Not the document name. The approved version, and its status at the moment the answer was given.

Can you explain why that source was treated as authoritative rather than superseded? Current policy, binding regulation, official guidance, a draft, or somebody's interpretation. The system finding a document does not answer that question.

Can you identify who reviewed the answer before it influenced a regulated decision, and evidence what they checked? Not that someone approved it. What they verified.

If the answer to any of those is no, the weakness may not be the model. It is likely to be the governance around the knowledge the model is allowed to use.

What I am building around this

This is the design problem behind the SAFE™ Microsoft 365 Copilot Governance Agent I am currently developing.

The architecture separates governance knowledge into distinct SharePoint libraries rather than treating every organisational document as one undifferentiated pool: the SAFE™ methodology itself, governance standards, regulatory sources, sector-specific sources, organisation policies, organisation evidence, templates, and an archive that retrieval excludes.

Within the regulatory source layer, documents carry the metadata described above: document status, citation priority, legal effect, text role, relationship type, related instruments, source identifiers and last verified date.

The objective is not better search. It is the difference between a system that can say "I found a document" and one that can say "I found the current approved source, here is why it has authority, here is what superseded the previous version, and here is the evidence supporting the answer."

That is a considerably higher standard, and it is a design methodology rather than an FCA-prescribed architecture. No regulator endorses it, and any supplier claiming otherwise is overstating their position.The bottom line

The interesting failures in AI governance are rarely the ones that get written about.

A hallucinated regulation is dramatic and easy to design against. A correctly retrieved obsolete policy is neither. It looks right, it cites a real internal document, and it passes every review that asks whether the AI made something up.

Anyone can connect a conversational interface to a folder of documents. The difficult work is underneath it: identity, permissions, approved knowledge, source authority, version control, retrieval, validation, human oversight, logging and evidence.

Under SM&CR, accountability stays attached to people and defined responsibilities rather than to software. AI does not remove that requirement. It makes the evidence architecture behind it more important.

The firms that govern AI well will not simply know what their Copilot answered. They will be able to show why it was allowed to answer that way.

References

This article is general guidance, not legal advice. How SM&CR accountability applies in a specific firm depends on that firm's Statement of Responsibilities and the allocation of prescribed responsibilities. That is a question for your compliance function and legal counsel.

Footnotes

  1. Financial Conduct Authority, AI Update, published 22 April 2024, setting out how existing rules apply to AI: https://www.fca.org.uk/firms/innovation/ai-approach

  2. DP5/22: Artificial Intelligence and Machine Learning, published jointly by the Bank of England, the Prudential Regulation Authority and the Financial Conduct Authority on 11 October 2022.

  3. Financial Conduct Authority, AI and the future of retail financial services (The Mills Review), published 6 July 2026: https://www.fca.org.uk/publications/corporate-documents/mills-review

  4. Regulation (EU) 2024/1689, consolidated text as at 27 July 2026, CELEX 02024R1689-20260727. Article 12 on record-keeping and Article 72 on post-market monitoring by providers: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02024R1689-20260727