Data Readiness: The Core Foundation for AI Transformation

by

AI & Data Cell Division

Last updated:

AI model performance is improving rapidly, and the cost of using these models continues to fall. There are more models to choose from than ever before, and AI is evolving beyond simple chatbots into technology that can perform tasks on its own. Yet why is it still difficult to find companies that have increased operational efficiency and improved actual profitability through AX (AI Transformation)?

When companies first began adopting AI, their biggest concern was which model to choose. However, performance gaps between models are narrowing rapidly. According to the Stanford Institute for Human-Centered Artificial Intelligence (Stanford HAI), the inference cost of models with GPT-3.5-level performance fell to 1/280 of its previous level or less between November 2022 and October 2024. [1] Simply using a good model is no longer enough to secure a lasting competitive advantage.

The scope of tasks AI can handle has also expanded. Beyond answering questions or summarizing documents, it can find information, query databases, and perform calculations and comparisons. Yet these uses do not immediately translate into results. In McKinsey’s 2025 survey, 88% of respondents reported that their organizations used AI in at least one business function, but only about one-third had scaled it across the enterprise. Just 23% reported that they were scaling AI agents in their operations. [2] Ultimately, data, rather than the model, is what sets companies apart. Data readiness is the foundation that bridges the gap between adoption and reliable use in day-to-day work.

 

What Data Readiness Means in AX

In machine learning, data readiness has primarily been understood as a concept for assessing the quality and suitability of training data. In enterprise AX, however, it needs to be viewed somewhat differently. Companies often use already-trained AI models by connecting them to internal documents, databases, and business systems. In this setting, data serves less as material for retraining AI and more as evidence it consults when generating answers that reflect the company’s business context.

A state in which the meaning, provenance, and access permissions of a company’s documents and data have been clarified so that AI can use them in its work

 

In an enterprise context, data readiness is the state in which AI can find the materials it needs and use them as reliable evidence. Figures must be traceable to their sources, and the latest versions must be distinguishable from earlier ones. Access permissions for each source must also be clear. With these basic conditions in place, it becomes easier to verify AI-generated results and use them in actual work.

This does not mean that all data must be perfectly organized. The level of preparation required depends on who uses the AI and for what business purpose. For a management dashboard, aligning metric definitions and measurement criteria comes first. For document retrieval, when the original document was written and whether it is the latest version may matter more. If AI is to perform calculations or prepare reports as well, business rules and access permissions must also be organized. Ultimately, data readiness starts with the work to be performed using AI, rather than with the data itself. [3]

Data Preparation Starts with the Business Purpose

AI can be put to use relatively quickly for finding and summarizing information. But the situation changes when it begins comparing or calculating across multiple sources on a consistent basis to produce actual deliverables. Simply connecting a large number of documents is not enough. The fields to be included in the output and the materials to serve as authoritative sources must be defined, and calculation bases and reference dates must be aligned.

Suppose, for example, that investment terms proposed by several institutions are to be compared in a single table. The fields to compare, such as interest rates, maturities, and collateral, must first be defined, along with whether values should be copied directly from the documents or calculated separately. Where units or reference dates differ, they must be brought onto a consistent basis. Rules are also needed to ensure that terms absent from the documents are not filled in arbitrarily and that authoritative original documents take precedence. Clear criteria are necessary to maintain consistency across repeated tasks.

Having more material does not necessarily produce accurate answers. Older and current versions may be mixed together, the same figure may be recorded with different dates and units, and the same term may have different meanings across projects. As data volumes grow, so does the possibility of drawing on the wrong evidence. The National Institute of Standards and Technology (NIST), part of the U.S. Department of Commerce, emphasizes data provenance, access controls, and human review procedures. [4] Gartner, a global IT research and advisory firm, also highlights the need to clarify relationships between data and business rules. [5]

Different Data Environments Require Different Approaches

Across the companies CIP has worked with on AX projects, the starting point for data readiness has varied by industry and how work is organized. Large volumes of data did not necessarily mean a company was well prepared, nor did having established systems mean it could immediately put AI to use. Some companies needed to identify the relevant scope within extensive historical records, while others first needed to clarify where scattered materials were located and the context behind them. The cases in which AX has been applied so far fall into five industry categories: PE firms, traditional manufacturers, B2C service businesses, B2B service businesses, and real estate development and investment firms. Each category accumulates data differently and has different priorities to address first.

At PE firms, investment review materials, due diligence documents, contracts, and external research accumulate as each deal progresses, while access permissions vary by project. Documents account for most of the data, rather than standardized metrics, and even the same field may be defined slightly differently from deal to deal. In this environment, distinguishing which materials are used at which stage matters more than the volume of documents. It must also be clear whether the latest or an earlier version takes precedence. James, the AI agent CIP developed and uses internally, began in this environment. Initially, its role was largely to retrieve materials for specific projects. Its scope expanded as output formats and priority reference materials were organized for use across multiple projects. In tasks such as investment review, due diligence, and contract review, where sources and context must be understood accurately, how clearly deals, documents, and business context were connected had a greater influence on usefulness than model performance.

Traditional manufacturers have accumulated production, quality, cost, and sales data over many years. Often, however, the connection between that data and business decisions has not been clearly defined. Setting priorities was also central at KAF, a textile materials manufacturer. Rather than connecting all data at once, the scope of use was defined according to the business purpose. For tasks involving recent business performance, the consistency of the latest data and the definitions of metrics were checked first. For topics where long-term trends mattered, such as quality issues or cost structures, the scope was extended to cover the necessary historical periods. For AI to use time-series data reliably, not only the query period but also units and classification criteria, such as product categories and processes, needed to be organized consistently. As more data organized for specific business purposes accumulates, the scope for examining production and costs on a common basis also expands.

B2C service businesses accumulate customer and transaction data in relatively orderly systems. South Springs, a golf course operator, also had structured customer, membership, and operational data. The key task was to define the fields in existing systems and the relationships between data, rather than undertake a large-scale effort to organize new data. Because personal information was involved, separation of permissions and the principle of least privilege also had to be considered. When data structures, intended uses, and access criteria were clear, existing data could be connected to actual work relatively quickly. There is also considerable scope to extend its use to forecasting demand and adjusting pricing and operations based on booking and usage histories.

B2B service businesses operate around orders and projects. Quotations, designs, and project execution records are often managed separately by individual staff members and projects, leaving them closer to personal files than company assets. SolidENG is still at an earlier stage, where even gathering scattered materials in one place is a challenge. The first step is to decide where materials should be accumulated and according to what criteria. Connecting an AI model comes next. Once the collection stage has been completed, past project records can be reused to inform quotation and design decisions.

Real estate development and investment firms work with parties of different kinds within a single project, including developers, contractors, financial institutions, and permitting authorities. At a specialist real estate investment firm, the dispersion of materials across stakeholders and stages of work was also an issue. Scattered materials were first gathered into a shared repository, and the folder structure and the context in which documents had been created were organized around project stages. Bringing files together in one place was not enough. To enable AI to reliably find the evidence it needed, it was necessary to identify which decisions each document related to and to clarify its relationships with other documents. Once this is in place, the decision-making history at each project stage can be carried forward as review criteria for subsequent projects.

Although the five categories had different starting points, they shared a clear common feature. Rather than organizing all data at once, they defined the outputs needed in actual work and first prepared the data and rules to support them. PE firms clarified the relationships among deals, documents, permissions, and versions, while traditional manufacturers defined the scope of data use and the criteria for analysis. B2C service businesses connected structured data in line with business purposes and access criteria, while B2B service businesses and real estate development and investment firms began by organizing where materials were located and their project context. Data readiness is less about how much data a company has and more about understanding its current state and finding the right sequence of preparation for its business purposes.

Conclusion

Model performance is likely to continue improving while costs fall further. As more companies use similar models, differences will increasingly depend on how they connect company data to their work. This does not mean that all materials must be perfectly organized before AI is introduced. A more practical starting point is to select tasks where AI can provide substantial value and first prepare the required outputs.

From an investor’s perspective, data readiness can also serve as a criterion for assessing a company’s foundation for executing AX. Even between two companies with similar business performance, having data definitions, sources, permissions, and business rules in order can make a difference in how quickly AI-based value creation initiatives can be implemented and validated. Although it does not appear directly in financial statements, data readiness matters as an operational foundation for translating AI into improvements in efficiency and profitability. The capabilities CIP has built through experience with different data environments across its portfolio companies reflect the same principle. As CIP gains experience in quickly identifying each company’s starting point and connecting the necessary data, it builds a stronger foundation for implementing AX reliably across companies. Over time, this experience yields a set of methods validated for each type of company. Rather than leaving this experience scattered across individual projects, CIP aims to bring it together as a shared service that operates and manages portfolio companies’ work and data within a single framework. Centroid AI Shared Service (CASS) is the direction it is pursuing. Instead of each company building its own AI team and infrastructure, the investment manager provides a common foundation while each company focuses on its own data and operations. At that point, AX becomes a value creation tool that can be presented from the investment stage, rather than a separate initiative at each portfolio company. Ultimately, the success or failure of AX depends less on how quickly an AI model is adopted than on how reliably company data is translated into actual improvements in operational efficiency and profitability.

 

Sources

[1] Stanford Institute for Human-Centered Artificial Intelligence, AI Index Report 2025, 2025.

[2] McKinsey & Company, The State of AI in 2025: Agents, Innovation, and Transformation, 2025.

[3] Gartner, Lack of AI-Ready Data Puts AI Projects at Risk, 2025.

[4] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024.

[5] Gartner, Lack of Semantics Causes Inaccurate AI Agents and Wasted Spending, 2026.

Insights

Read more articles