While much of the financial services industry races to deploy ever more sophisticated artificial intelligence models, Bank of America is delivering a sobering counterargument: the intelligence of a model is largely irrelevant if the data feeding it cannot be trusted. In a moment when banks are competing loudly over who has built the most capable AI system, America's second-largest bank by assets is shifting the conversation to a less glamorous but arguably more consequential problem — the state of institutional data itself.
The bank's position is direct and unambiguous. An AI feature, regardless of the sophistication of the underlying model, is only as reliable as the data on which it depends. That reliability, in turn, is what ultimately determines whether artificial intelligence can be trusted to perform consequential functions in a regulated banking environment — from credit decisioning to fraud detection to personalized client advisory. Without clean, well-governed, accurately structured data, even the most advanced large language model becomes an unreliable instrument, prone to errors that in a financial context can carry serious regulatory and reputational consequences.
This argument challenges a prevailing assumption in the industry's current AI arms race. The dominant narrative in banking technology circles has revolved almost entirely around model capability: which institution has assembled the largest training dataset, which has partnered with the most advanced foundation model provider, and which has deployed generative AI tools at the greatest scale. Bank of America's intervention is a reminder that this framing may be dangerously incomplete. Model sophistication and data integrity are not equivalent priorities, and treating them as such creates a false foundation for AI strategy.
The underlying concern is structural. Large financial institutions carry decades of legacy data infrastructure — systems built at different times, under different standards, often in organizational silos that were never designed to communicate with one another. Customer records exist in fragmented states across divisions. Transaction histories are stored in formats that predate modern data architectures. Risk data compiled for one regulatory regime may not map cleanly onto the taxonomies required for another. Layering generative AI or machine learning models on top of this kind of fragmented, inconsistently governed data environment does not resolve those underlying problems — it amplifies them, because AI systems operate at a speed and scale that can propagate data errors far faster than human oversight can catch them.
Bank of America's argument is therefore not an anti-AI position. It is a sequencing argument. Before institutions invest heavily in building or acquiring more powerful models, they should invest in ensuring that the data those models will consume meets a standard of accuracy, completeness, and governance that justifies the trust being placed in automated outputs. In practice, that means addressing data quality at the source, establishing consistent taxonomies and lineage tracking across business lines, and building the kind of data governance frameworks that allow an institution to audit what an AI system was told and why it reached a particular conclusion.
The regulatory dimension of this argument is significant and should not be overlooked. Supervisory bodies across major jurisdictions — including the European Banking Authority and the Federal Reserve — have been increasingly explicit that accountability for AI-driven decisions in financial services rests with the institution deploying the technology, not with the model provider. That accountability is impossible to discharge if an institution cannot demonstrate that the data used to train and run its AI systems was accurate, unbiased, and properly governed. Clean data is not merely a technical nicety — it is a prerequisite for regulatory defensibility.
There is also a competitive logic to Bank of America's stance that is worth examining. Institutions that invest early in data infrastructure create a durable advantage that compounds over time. A well-governed data estate does not just improve today's AI applications — it improves every application built on that data in the future. By contrast, institutions that rush AI deployment on top of poor data quality will find themselves continuously firefighting model errors, retraining on corrupted inputs, and facing escalating remediation costs. The short-term competitive gain of a flashy AI feature launch can quickly become a long-term liability if the data foundation is weak.
What This Means for the Industry
Bank of America's data-first framework represents a maturation signal for institutional AI strategy in banking. The industry is moving — slowly but perceptibly — from the question of whether to deploy AI to the harder question of how to deploy it responsibly at scale. That transition requires banks to confront unglamorous infrastructure work that rarely generates headlines but ultimately determines whether AI delivers durable value or merely the appearance of it. For boards, chief data officers, and technology leadership teams across the sector, the message from one of the industry's most systemically significant institutions is clear: fix the foundation before you build the skyscraper. Model quality is a competition worth having — but only once the data it depends on is genuinely fit for purpose.
Written by the editorial team — independent journalism powered by Codego Press.