Use AI to fix the data

Use AI to fix the data
Use AI to fix the data
AI & Data · demystifAI
3 August 2026
Romain Thierry
AI & Data · demystifAI

Use AI to fix the data

The received wisdom says that banks must fix their data before building up AI solutions. Thirteen years of Basel Committee reporting suggest following this advice is a Sisyphean task, because the data never gets right. This piece argues for reversing the order — deploying AI at the front-end of the data lifecycle to fix the data for good — and for starting small, in one bounded domain, where the result can be measured.

The standard advice on artificial intelligence in banking is to fix the data first. The Basel Committee gave banks that instruction back in November 2023, telling them to treat sound data quality as the foundation of any digitalisation project, and it repeated the point this January. The advice sounds very reasonable, but the risk profile is far from being neutral. A model handling a duplicated counterparty record will usually not bat an eyelid: it rightly assumes it is working with clean data. Errors compound quietly, and stressed data analysts and developers then spend hours reverse-engineering what went wrong and working out the downstream implications. If you have ever been caught in such a situation, you will remember it.

Nevertheless, the advice rests on an assumption, which is that the data would be fixed in due course. Unfortunately, thirteen years of the Committee’s own reporting say it will not.

So why not turn the problem upside down, and invert the order: instead of using advanced data technology at the end of the chain, to patch the reports, the same smart tools should be deployed at the front of the data lifecycle, where they can start to reinforce the data foundations on which most banks run. Let’s call it the data inversion; it is what this series is about.

Why should it work this time? Because AI is extremely good at dealing with precisely the scenarios that let poor data into a mastering layer in the first place: duplicate entities, inconsistent reference data, undocumented feeds. Future articles on this blog will work through those scenarios one by one; for now, let us stay with what the inversion means and why the evidence supports it.

First, a few words about that ‘unfortunately: the data won’t get fixed’. The Committee published its principles for risk data aggregation in 2013 and expected the largest banks to comply by 2016. Its latest assessment, in 2023, found two banks out of thirty-one fully compliant, and no earlier assessment had found more. These programmes are not badly run so much as unfinishable: after each year that passes, the objective is still about a year away. This is understandable. The world does not stand still (think Covid, Ukraine), scope keeps widening, and legacy systems are more tangled than anyone budgets for. Worse, a bank cannot always tell when it has arrived: supervisors found institutions that had declared their data compliant while still running multi-year remediation programmes. In practice, waiting for the all-clear means waiting for a signal nobody can send.

Another interesting finding in the same assessment lists what banks would like AI to do to meet those targets: automate documentation, cut manual handling, map how their data actually flows, and maintain lineage. The common reason given for not doing it yet is that the data is not good enough. Look at that list again, though — documentation, manual handling, flows, lineage. That is the data-quality work. Banks are holding the tool back until the very condition it would create already exists.

This concept of inversion becomes more palatable when one reads the Committee’s own case studies: they describe a bank that rebuilt the production of its supervisory returns, put data-quality checks at source and at its data hub, and used machine learning to find and monitor its data problems. Of course the report does not say whether this was quick or cheap, so let us not oversell it; but a supervisor watched it happen and recorded it as good practice, in the same document that tells banks to fix their data before turning to the machines.

In any case, the market has stopped waiting. The Bank of England and the Financial Conduct Authority surveyed 118 financial firms in 2024 and found three-quarters already using AI, with another tenth expecting to within three years. None of them, on the survey’s evidence, held out for a declaration that their data was good. The tools are in; the open question is what to point them at.

None of this means the machines can be left alone with the problem. Run an entity-resolution model over badly governed reference data and it will produce confident, wrong mappings in bulk, so enough governance to trust the output is still needed; the inversion is a virtuous cycle rather than a cure. And where a number is attested — signed off by an officer who carries the consequence — a machine’s proposed mapping cannot be the last word. The ‘human in the loop’ remains an indispensable feature of any AI-based solution.

What does this mean in practice? This does not call for another enterprise programme; rather, pick one bounded domain where the value is high and the boundary is clear (client and party data is the usual candidate) and use AI to compress the time it takes to reach quality there, measured against a manual baseline. Nobody, so far as I can find, has published that measurement. The first bank that does will settle a thirteen-year argument with a number.

For thirteen years banks have been told to fix their data before turning to the machines, and for thirteen years the reports have recorded that they have not.

One of the two instructions has to give. It should be the order.

If your bank has made that measurement, I would like to hear about it. Tell me in the comments.

References

  1. Basel Committee on Banking Supervision, Progress in adopting the Principles for effective risk data aggregation and risk reporting, 28 November 2023. bis.org/bcbs/publ/d559.pdf
  2. Basel Committee on Banking Supervision, Principles for effective risk data aggregation and risk reporting, January 2013. bis.org/publ/bcbs239.pdf
  3. Basel Committee on Banking Supervision, Newsletter no. 36, Implementation of the Principles for effective risk data aggregation and risk reporting, 6 January 2026. bis.org/publ/bcbs_nl36.htm
  4. Bank of England and Financial Conduct Authority, Artificial intelligence in UK financial services — 2024, 21 November 2024. bankofengland.co.uk
demystifAI · Use AI to fix the data · AI & Data · 3 August 2026 demystifai.info AI, in control.