Skip to content
Stratalytic

DATA DRIVEN DECISIONS

Data Strategy

Bad Data Costs You Money: Improve Data Quality in 5 Steps

Published:

Dashboard with data visualisations showing the impact of data quality on business results

Key Takeaways: Poor data quality costs large organisations an average of $12.9 million per year according to Gartner. For SMEs, that translates to thousands or tens of thousands of euros wasted on lost time, missed opportunities and flawed decisions. This article describes five concrete steps to structurally improve your data quality, including tooling recommendations and subsidy options through the Dutch WBSO programme.

The hidden cost of bad data

Data quality is one of those topics where every business owner acknowledges its importance, yet few address it structurally. The numbers tell a sobering story. Gartner estimates that poor data quality costs organisations an average of $12.9 million per year. IBM places the total cost of bad data in the United States at $3.1 trillion annually. For the average Dutch SME, those figures are less astronomical but no less painful.

Consider a trading company with 30 employees and 5 million euros in revenue dealing with duplicates in their customer database, outdated contact details and inconsistent product codes. The impact is immediately measurable. Account managers call clients already served by a colleague. Invoices go to wrong addresses. Inventory counts are off because the same product is registered under three different codes. The Harvard Business Review calculated that knowledge workers spend up to 50% of their time hunting for and correcting data errors, or confirming data they suspect is unreliable.

The indirect costs are at least as significant. Management decisions based on unreliable data lead to wrong investments. Marketing campaigns that miss the mark because customer segmentation is flawed. Compliance risks because customer records are not current. According to MIT Sloan, only 16% of managers trust the data they receive for decision-making. That distrust is not irrational. It is the logical consequence of years of data management on autopilot.

Five common data quality problems

Most data problems in SMEs are not exotic. They are predictable, recognisable and therefore solvable. The first and most common problem is duplication. The same customer appears three times in the CRM, each time with a slightly different spelling. "Van der Berg", "V.d. Berg" and "Vandenberg" are three different entities to a computer. Research by Experian shows that 94% of businesses suspect their customer data contains duplicates. For most companies, it concerns 10-30% of all records.

The second problem is missing values. Fields left empty because they were not mandatory at input, or because the information simply was not available at the time of registration. A customer database without email addresses for 40% of records is useless for email marketing, but that only becomes apparent when the campaign is set up. Missing birth dates make age segmentation impossible. Empty fields for industry or company size undermine every attempt at B2B audience analysis.

The third problem is inconsistent formatting. Dates stored sometimes as "07-04-2026", sometimes as "2026/04/07" and sometimes as "7 April 2026". Phone numbers with and without country codes. Amounts with and without VAT. Addresses with and without apartment numbers. These inconsistencies make reliable filtering, sorting or matching impossible. According to research by Informatica, data teams spend an average of 40% of their time standardising formats instead of performing analysis.

The fourth problem is outdated records. Employees who have left but still appear in the system. Suppliers who have been inactive for years. Customers who have moved but whose addresses were never updated. Data degrades faster than most businesses realise: B2B contact data becomes outdated at a rate of approximately 30% per year according to Dun & Bradstreet.

The fifth problem is the absence of a single source of truth. The same information is maintained in the CRM, in Excel spreadsheets, in the accounting software and in the heads of employees. Nobody knows which source is correct. At 60% of SMEs there is no definitively established source system per data domain. The result is that every department works with its own version of the truth.

Step 1: audit your data

The first step towards better data quality is knowing where you stand. A data audit is not a months-long project. It is a structured inventory that you can complete in one to two weeks. Start by identifying your five most important data sources. For most SMEs these are the CRM system, accounting software, email marketing platform, ERP system and any Excel files functioning as shadow databases.

Measure four core indicators per source. Completeness: what percentage of required fields is populated? Uniqueness: what percentage of records is unique versus duplicate? Consistency: do values follow a fixed format? Timeliness: when were records last updated? A pragmatic approach is to take a sample of 200-500 records per source and assess them manually. This gives you a reliable picture of the current state within a few days.

Document your findings in a simple scoring table. Per data source, per quality dimension, a score from 1 to 5. This becomes your baseline measurement and the reference point for all improvements that follow. Companies that perform such a baseline measurement discover on average 3 to 5 times more problems than they expected beforehand, based on data management consultancy experience.

Step 2: define quality rules

After the audit you know where the problems are. The next step is documenting explicit rules that your data must satisfy. This sounds bureaucratic, but it need not be more than a shared document of two to three pages. Define per data field the expected format, permitted values and whether the field is mandatory.

An example: for the field "phone number" the format is +31-6-XXXXXXXX, the field is mandatory for all active customers and numbers without a country code are automatically supplemented with +31. For the field "customer status" the permitted values are "active", "inactive" and "prospect", free text is not allowed and every customer has exactly one status. These rules do not need to be perfect at first setup. What matters most is that they exist and that the entire team knows them.

Involve the people who work with the data daily when drafting these rules. The sales representative knows which customer fields are essential. The accountant knows which financial formats cause problems. The marketer knows which segmentation fields are needed for effective campaigns. By bundling this knowledge into a shared set of quality rules, you create buy-in and prevent the rules from existing only on paper. Companies with documented data quality rules report a 40-60% reduction in new data errors within the first year according to DAMA International.

Step 3: clean at the source

The most cost-effective way to improve data quality is preventing errors at the moment of entry. Every error that does not enter the system never needs to be traced and corrected later. The ratio is lopsided: correcting a data error costs 1 euro to prevent, 10 euros to correct and 100 euros in damage if the error goes undetected, according to David Loshin's 1-10-100 rule.

Make this concrete by revising your entry forms and processes. Make fields mandatory that should be filled according to your quality rules. Use dropdowns and selection lists instead of free-text fields where possible. Implement real-time validation that flags input errors immediately. A postcode field that automatically checks whether the entered combination of postcode and house number exists. An email field that performs syntax validation. A VAT number field that verifies against the VIES database whether the number is valid.

Train your staff on the importance of clean data entry. This does not need to be a multi-day training. A one-hour session showing the concrete consequences of sloppy entry has more effect than any policy document. Show how much time the team spends correcting errors. Show which customers were lost due to a wrong phone number. Make it tangible and personal.

Step 4: automate validation

Manual checking does not scale. Once your quality rules are defined and your entry processes improved, the next step is automating validation. This means that every time data enters the system or is modified, it is automatically checked against established rules.

For companies working with a modern data platform, tooling such as dbt (data build tool) offers built-in testing functionality. You write simple tests that check during every data transformation whether fields are not empty, whether values fall within expected ranges and whether references to other tables are valid. A dbt test verifying that every order has a valid customer ID is four lines of code. A test checking that revenue figures are not negative is three lines. The investment in setting up these tests is minimal compared to the hours saved on manual checking.

For more advanced validation, Great Expectations is a widely used open-source tool. You define "expectations" about your data: the column "age" contains values between 18 and 120, the column "email" always contains an @ sign, the total number of records does not deviate more than 10% from yesterday. Great Expectations automatically generates documentation of your data quality rules and reports violations. Companies implementing automated data validation report an average reduction of 60-80% in data errors reaching the production environment.

A simpler alternative for companies not yet working with a data platform is setting up automated checks in Google Sheets or Excel via scripts or add-ons. Even a weekly automated report checking for duplicates, empty fields and anomalous values is a major step forward compared to no checks at all.

Step 5: monitor continuously

Data quality is not a one-time project but an ongoing process. Data degrades constantly: employees leave, customers move, systems get integrated, new sources are added. Without continuous monitoring you inevitably slide back to the situation before the improvement efforts.

Set up a monthly data quality report tracking the four core indicators from your audit: completeness, uniqueness, consistency and timeliness. Define thresholds within which scores must remain and escalate when a score drops below the threshold. A pragmatic target for most SMEs is a data quality score of at least 85% across all four dimensions. That sounds like a high bar, but it is achievable with the preceding four steps as foundation.

Assign a data owner per domain. This does not need to be a full-time role. It is the person responsible for the quality of a specific dataset. The sales manager owns the customer database. The operations manager owns inventory data. The finance director owns financial data. By making ownership explicit, you prevent the situation where "everyone" is responsible and therefore "no one" takes action. Organisations with designated data owners score an average of 35% higher on data quality measurements than those without, according to the Data Governance Institute.

Implement a quarterly deep-dive in which you analyse data quality trends, identify new problem areas and update your quality rules. The data quality cycle is not linear but iterative. Each cycle yields new insights that make the next cycle better.

Tools that help

The good news is that you do not need to reinvent the wheel. Proven tools are available for every step in the data quality process, from open-source to enterprise-grade.

For data transformation and testing, dbt is the de facto standard in the modern data stack. It is open-source, well-documented and has an active community. dbt allows you to version-control, test and document your data transformations. The tests you write run automatically with every pipeline execution, catching errors before they reach your reports. More than 40,000 companies worldwide use dbt, from startups to Fortune 500 enterprises.

Great Expectations provides a framework for defining, documenting and validating data expectations. It integrates with all common data platforms and automatically generates data documentation. For SMEs serious about data quality but without the resources for an enterprise data governance platform, Great Expectations is an excellent starting point.

For deduplication there are specialised tools such as Dedupe.io (an open-source Python library) that uses machine learning to identify duplicates even when records do not match exactly. OpenRefine (formerly Google Refine) is a free desktop tool for cleaning and transforming data, particularly suitable for one-off cleanup actions on CSV and Excel files. For companies with data in a cloud warehouse, platforms like Snowflake and BigQuery offer built-in data quality monitoring functions.

The choice of tooling depends on your current technical maturity. A company managing its data primarily in Excel starts with OpenRefine and structured spreadsheet validation. A company with a cloud data warehouse implements dbt tests and Great Expectations. A company with an already mature data platform adds monitoring dashboards and automates the entire quality process.

WBSO: subsidy for data engineering R&D

This is where it gets interesting for SMEs seriously investing in data quality. Developing automated data quality processes, building data integration pipelines and implementing advanced validation logic can qualify for the WBSO (Wet Bevordering Speur- en Ontwikkelingswerk), the Dutch R&D tax credit scheme. The WBSO provides a reduction on payroll taxes or a deduction for self-employed professionals conducting technical-scientific research or technical development.

The key is the word "development". Standard implementation of existing software does not qualify for WBSO. But developing a company-specific data integration platform that connects multiple sources and automatically monitors data quality can qualify. Building machine learning models for deduplication based on your specific datasets is technical development. Designing a data governance framework with automated compliance checks can be classified as systematically organised development.

The WBSO provides starting entrepreneurs a payroll tax reduction of 40% on the first 350,000 euros in R&D salary costs and 16% above that threshold. For self-employed professionals a fixed deduction applies. This can make a significant contribution to financing your data quality project. Applications are submitted through RVO and must be filed before work begins. More information about conditions and the application process can be found on our WBSO page.

Does your data quality project have a research or collaboration component? The MIT scheme can provide additional funding for feasibility studies preceding the actual development. By combining subsidies you maximise the return on your data quality investment.

Conclusion: data quality as competitive advantage

Improving data quality is not a technical project for the IT department. It is a company-wide initiative with direct impact on your revenue, costs and decision-making. The five steps in this article, from audit to continuous monitoring, form an approach that scales from a one-person business to an organisation with hundreds of employees.

Companies that structurally maintain their data quality make better decisions, work more efficiently and are better prepared for deploying advanced technologies like machine learning and AI. Because it does not matter how sophisticated your algorithm is: if the data going in is wrong, the output is worthless. Or as the age-old wisdom in data goes: garbage in, garbage out. The difference is that you now know how to take the garbage out.

Get the AI-subsidy radar

1 email per month. New subsidies, deadlines, and what changed for SMEs. 5-minute read.

Unsubscribe with one click. No spam, ever.

Let's talk business

Do you want to know how we can help you grow your business? Schedule free consultation with one of our experts and discover the possibilities.

Rutger Geerlings, founder of Stratalytic

Rutger Geerlings

Solution Architect

Discover what data and AI can concretely deliver

Latest cases

All cases