Data Strategy
AI-Ready? Why Your Data Foundation Matters More Than the Right Tool
Published:

Key Takeaways: Most companies that embark on AI invest in tools before their data is in order. Research shows that roughly 80% of time in AI projects is spent on data preparation, and that poor data quality is the primary reason AI initiatives fail. This article explains why a strong data foundation is the prerequisite for every successful AI application, how to lay that foundation in 8 weeks, and which subsidies can significantly reduce the investment.
Why most AI projects fail, and why it is not the technology
Approximately 80% of AI projects do not deliver the expected results, and the cause is almost never the algorithm or the tool. The problem starts with the data feeding the system. When information is scattered across dozens of spreadsheets, CRM fields are incomplete, and definitions differ between departments, even the most advanced AI model cannot produce reliable outcomes.
Gartner reports that organisations lose an average of USD 12.9 million per year due to poor data quality. For SMEs this translates into missed revenue, flawed decisions, and wasted project budgets. The pattern is recognisable: a company invests EUR 50,000 to 100,000 in an AI pilot, discovers halfway through that the underlying data is unfit, and must pause or write off the project. The tool works perfectly fine, but the input is worthless.
This phenomenon is known as "garbage in, garbage out," and it is as old as computing itself. Yet companies repeat the mistake because the AI market sells tools, not foundations. Vendors demonstrate impressive capabilities on clean datasets, but the gap between that demo and your own data reality is often larger than expected. A machine learning model that detects fraud on a standard dataset performs entirely differently when fed inconsistent, incomplete transaction data from your own accounting system.
The lesson is clear: invest in your data foundation first, then in AI tooling. Companies that respect this order report three times higher success rates with AI implementations. The difference is not budget or technical expertise, but the discipline to get the basics right before reaching for advanced applications.
What a data foundation actually is
A data foundation is the technical and organisational base that ensures all relevant business data is reliable, accessible, and usable for analysis and automation. It is not a product you buy, but a combination of architecture, processes, and agreements that together determine how well your organisation can work with data.
Concretely, a data foundation consists of four layers that build upon each other. The first layer is data sources and integration. This concerns the connection between all your operational systems: ERP, CRM, webshop, accounting, production registration, and any external data sources. A solid integration layer ensures data flows together automatically and in a standardised format, rather than being manually copied via exports and Excel files.
The second layer is storage and structure. A central data warehouse or data lakehouse stores all integrated data in a structured format with consistent definitions. Here, customer IDs are unified, date formats standardised, and duplicates cleaned. This layer literally forms the foundation on which everything rests. Research from McKinsey shows that companies with a centralised data warehouse make data-driven decisions 23% faster than organisations working with dispersed data sources.
The third layer is data quality and governance. This encompasses the rules, processes, and tooling that ensure data remains correct, complete, and current. Think of validation rules at entry, automated quality checks, ownership per data source, and procedures for correcting errors. Without governance, any data foundation degrades within months into the same chaos it was meant to resolve.
The fourth layer is accessibility and analysis-readiness. Data that is correctly stored but not accessible to the right people delivers no value. This layer includes dashboards, reporting tools, APIs, and the documentation that enables users to work with data independently. For AI applications, this means datasets are available in the right format, accompanied by metadata and version control.
Five signs your data is not AI-ready
If you recognise at least three of the following five signs, your data in its current state is not suitable as a basis for AI applications. These signals point to fundamental deficiencies that need to be resolved first.
The first sign is that the same question yields different answers depending on whom you ask or which system you query. When sales reports a different revenue figure than finance, or when the number of active customers differs between CRM and invoicing system, unified definitions are missing. AI models trained on inconsistent data produce unreliable predictions. Research from Harvard Business Review indicates that only 3% of data in the average company meets basic quality standards.
The second sign is that more than 15% of critical data fields are empty or invalid. Check your CRM: how many contacts lack an email address, job title, or industry code? How many orders are missing a source or campaign label? Missing values are not merely a cosmetic problem. They form gaps in the pattern that an AI model either ignores or misinterprets. A churn prediction model running on customer data with 30% missing interaction history will systematically misclassify customers.
The third sign is that your organisation spends more than two hours per week manually combining data from different sources. This often happens via Excel exports merged with VLOOKUP formulas. Every manual step introduces error risk and delay. Moreover, it does not scale: what works with a hundred rows fails at ten thousand. Organisations that recognise this typically spend an estimated 40% of their analysis time on data preparation rather than actual analysis.
The fourth sign is that no documented definitions exist for core concepts. What is an "active customer"? Someone who purchased in the last 12 months? Or someone with a running contract? If this definition is not formally established and consistently applied, every team interprets the concept in its own way. AI models require unambiguous labels and definitions to recognise patterns; ambiguity is their kryptonite.
The fifth sign is that your data lives exclusively in Excel files stored locally on laptops. According to research, 82% of SMEs still work primarily with spreadsheets for business-critical data. This is understandable but unsustainable once you want to deploy AI applications. Excel offers no version control, no automatic validation, no concurrent editing by multiple users, and no API access for AI systems. The step from spreadsheet to database is the most fundamental step in building a data foundation.
How to build a data foundation in 8 weeks
A data foundation need not be a lengthy and expensive undertaking. With a focused approach you can lay a working base in 8 weeks that is immediately usable for analysis and gradually expandable towards AI applications. The key is pragmatism: start with your most important data streams and expand incrementally.
In weeks 1 and 2 you conduct a data audit. You inventory all data sources in your organisation, from ERP and CRM to spreadsheets and manual registrations. Per source you document what data it contains, how current it is, who is responsible, and how the data is currently used. Simultaneously you identify the first AI use case you want to pursue. This use case drives the prioritisation of the foundation: which data needs to be addressed first? An audit of an average SME with 5 to 8 data sources typically takes 3 to 5 working days.
In weeks 3 and 4 you set up central data storage. This could be a cloud data warehouse such as BigQuery, Snowflake, or a lighter solution like PostgreSQL, depending on your scale and ambition. You define the data model, establish naming conventions, and record initial definitions in a data dictionary. Cloud storage costs are negligible for SME volumes: expect EUR 50 to 200 per month for the first terabytes.
In weeks 5 and 6 you build the first data pipelines that automatically extract data from your source systems and load it into the warehouse. Tools like Airbyte, Fivetran, or custom Python scripts ensure data is synchronised daily or in real time. This eliminates manual exports and guarantees your central storage is always current. On average, an SME has 4 to 6 critical data sources connected during this phase.
In weeks 7 and 8 you implement data quality checks and build the first dashboards. Automated checks flag when data is missing, duplicates emerge, or values fall outside expected ranges. Dashboards provide immediate insight into the data now centrally available. After these 8 weeks you have a working data foundation ready for the first AI experiments.
The total investment for this trajectory typically ranges between EUR 15,000 and 40,000, depending on the number of data sources and the complexity of your system landscape. This is a fraction of the cost of a failed AI project on a flawed data base.
The role of subsidies in building your data foundation
Building a data foundation qualifies for several Dutch innovation subsidies, enabling you to recover 30 to 60% of the investment. This makes the business case even more compelling and lowers the barrier to getting started.
The WBSO (R&D Tax Credit) is the most accessible subsidy for data foundation projects. When your project contains technical novelty, for example developing a domain-specific data model, building custom integrations, or implementing automated quality controls, you can partially deduct labour costs and outsourced R&D expenses. The WBSO effectively delivers 30 to 40% cost reduction on qualifying hours. In 2025 over 22,000 companies used the WBSO, demonstrating the scheme is broadly accessible.
The SLIM subsidy specifically targets learning and development in SMEs. When your data foundation project includes training employees in data literacy, dashboard usage, or data governance, you can receive up to 60% subsidy on these training costs. This is particularly relevant because a data foundation only delivers value when employees can work with it effectively.
For more extensive trajectories that continue into AI implementation, the MIT R&D subsidy offers possibilities. This subsidy supports technically innovative projects with a contribution of 35% up to a maximum of EUR 350,000. When your data foundation forms the basis for an innovative AI application, the entire trajectory can fall under this scheme.
Combining subsidies, also known as subsidy stacking, can bring your own contribution down to 40 to 50% of the total investment. A data foundation project of EUR 30,000 effectively costs EUR 15,000 to 20,000 after WBSO deduction and potential SLIM subsidy. For that investment you lay a base that not only makes AI possible, but also immediately improves your daily reporting, forecasting, and operational decision-making.
When is the right moment to start with AI?
The right moment to start with AI is not when the technology is perfect, but when your data foundation fulfils three basic conditions: your core data is centralised and standardised, you have at least 12 months of historical data available for your intended application, and there is a concrete business problem that demonstrably benefits from predictive or automating technology.
Do not wait for perfection. A data foundation is never "done"; it grows with your organisation and ambitions. The mistake many companies make is endlessly preparing without ever taking the step towards application. The goal of the 8-week approach is precisely to break this cycle: you build a workable base and immediately start a first AI pilot on that base.
The Dutch market is moving fast. According to research by Statistics Netherlands (CBS), 24% of companies with more than 10 employees now use AI applications, a doubling compared to two years ago. Companies that get their data foundation in order now will be ready for their first AI implementation within a quarter. Companies that wait risk competitors capturing that lead first.
The practical next step is clear. Start with a data audit to map your current situation. Identify your most valuable AI use case. Calculate the subsidy opportunities that reduce your investment. And start building. A solid data foundation is not only the prerequisite for AI, it is the prerequisite for every form of data-driven business. You choose the tool later. You lay the foundation now.
Get the AI-subsidy radar
1 email per month. New subsidies, deadlines, and what changed for SMEs. 5-minute read.
Unsubscribe with one click. No spam, ever.
Keep reading
Related articles

Data Strategy
Bad Data Costs You Money: Improve Data Quality in 5 Steps
Poor data quality costs businesses 15-25% of revenue on average. These 5 steps help you structurally improve data quality and save thousands of euros.
Read more →

Data Strategy
From Excel to Data Foundation: A Step-by-Step Guide for SMEs
Learn how to transition from Excel to a scalable data foundation in four stages, with concrete costs, timelines, and subsidies.
Read more →

Data Strategy
What is a data foundation? Definition, components and why you need it
A data foundation is the combination of data sources, integrations, storage and quality rules that lets an organisation reliably use data for reporting and AI.
Read more →
Let's talk business
Do you want to know how we can help you grow your business? Schedule free consultation with one of our experts and discover the possibilities.


