Key Points
Definition: Data quality describes how well data fulfills its purpose—as measured by dimensions such as completeness, accuracy, consistency, uniqueness, timeliness, and validity.
The costs are well documented: Gartner estimates that poor data quality causes an average annual loss of at least $12.9 million per company.
Up to 25% of revenue: Data quality expert Thomas Redman shows that flawed data can wipe out 15 to 25% of revenue.
Six dimensions are key: The DAMA-UK framework defines completeness, uniqueness, consistency, correctness, validity, and timeliness as standard metrics.
Data quality is not a cleanup project: It is an ongoing governance process with clear responsibilities and measurable KPIs.
AI raises the bar: Without robust data quality, automation merely scales existing errors.
Goldright provides the foundation: With over 20 years of experience in master data management, Goldright makes data quality measurable and controllable.
What Is Data Quality? The Definition
Data quality describes how well a dataset fulfills its purpose—that is, how complete, accurate, consistent, unique, current, and valid the stored information is. It is not a binary property, but rather a multidimensional measure whose composition varies depending on the use case.
The DAMA UK framework has established itself as the standard, defining six core dimensions: completeness, unambiguity, consistency, correctness, validity, and timeliness. Since 2013, these dimensions have formed the basis for numerous data quality tools and have been adopted, among others, in the UK Government Data Quality Framework.
High data quality does not mean that every single data field must be perfect. It means that the data is sufficiently reliable for the specific business purpose—a billing address must be correct, whereas an optional free-text field for internal notes need not be. Data quality must therefore always be evaluated in the context of its intended use.
Data quality is particularly closely linked to master data: customer, supplier, product, and employee data are the entities most severely impacted by quality deficiencies because they are used simultaneously by numerous processes. Our guide to master data provides an in-depth introduction.
It is important to distinguish this concept from related terms. Data quality is not the same as data quantity—more data does not automatically lead to better decisions if its reliability is questionable. Nor is data quality synonymous with data security: A dataset can be extremely well protected and still contain incorrect information. Data quality refers exclusively to the accuracy of the content and the usability of the data itself.
The term “master data quality” is often used synonymously with “data quality,” but it is more precisely defined: It refers specifically to the quality of master data as a subset of all corporate data. Since master data are the most stable and widely used data objects, master data quality is often the most important lever in practice for tangibly improving the data quality of an entire company.
Why Is Data Quality So Important?
Data quality has evolved from a quiet IT metric to a strategic success factor. Four proven facts show why this topic belongs on every CDO’s agenda. The common thread: Whereas a data error used to result in an incorrect invoice, today the same error can derail an AI project or jeopardize regulatory approval.
1. Poor data quality has been proven to cost millions
Gartner estimates the average annual cost of poor data quality at at least $12.9 million per company—a figure cited as an industry benchmark that encompasses productivity losses, poor decision-making, and compliance costs.
2. Faulty data destroys revenue
Data quality expert Thomas Redman demonstrates in his widely cited analysis that poor data quality can cost companies 15 to 25% of their revenue—due to rework, poor decisions, and a loss of trust in their own figures. For a company with 500 million euros in revenue, that amounts to a potential loss of 75 to 125 million euros annually.
3. AI Projects Fail Because of the Data Set, Not the Algorithm
Gartner expects that by 2026, approximately 60% of AI projects that are not supported by AI-ready data will be abandoned. Data quality is therefore not only an operational risk but also a strategic risk for every automation and AI initiative.
4. Regulation Increases the Pressure for Clean Data
Article 10 of the EU AI Act (Regulation (EU) 2024/1689) explicitly requires relevant, representative, and largely error-free training, validation, and test datasets for high-risk AI systems. Data quality is therefore no longer optional for many companies, but rather a compliance requirement.
Where the Costs of Poor Data Quality Actually Arise
The costs of poor data quality are rarely visible in a single line item—they are spread across numerous processes and only add up to the magnitudes documented by Gartner and Redman when viewed in their entirety. Four cost drivers recur in nearly every company.
Operational Inefficiencies
Duplicate data maintenance across multiple systems, manual corrections, and rework tie up resources that are then lacking elsewhere. Every invoice that must be manually corrected due to an incorrect supplier address costs time, which adds up to a significant amount over thousands of transactions per year.
Poor Decisions at the Executive Level
Reports and analyses are only as reliable as the underlying data. When executives make decisions based on contradictory metrics, the resulting poor decisions are often the most costly consequences of poor data quality—even if they are harder to quantify than a single incorrect invoice.
Regulatory Risks
GDPR violations due to incomplete or incorrect personal data, back taxes resulting from erroneous filings, and audit findings stemming from inconsistent business data are direct, often underestimated follow-up costs. For companies in the DACH region, these requirements are becoming increasingly stringent.
Loss of Trust and Slowed Decision-Making
DThe effect that is hardest to measure but most costly in the long run: When employees and managers no longer trust their own data, every decision is slowed down by additional manual checks. This doesn’t just affect a single invoice—it slows down the entire company.
A Comparison of the Six Dimensions of Data Quality
To improve data quality, you must first break it down into its components. The following six-dimensional model, based on DAMA UK, has established itself as the standard in practice—it allows you to pinpoint quality issues precisely, rather than speaking in general terms about “bad data.”
Dimension | Key Question | Example of a Violation | Typical KPI |
Completeness | Have all required fields been filled in? | Missing phone number in the customer record | Percentage of complete records |
Accuracy | Do the values reflect reality? | Incorrect company name in the commercial register | Error rate per sample |
Consistency | Do the systems match? | Two addresses for the same customer | Discrepancy rate between systems |
Uniqueness | Is there only one record per object? | The same supplier was created twice | Duplicate rate |
Validity | Does the data comply with the format/rules? | Invalid tax ID, incorrect date format | Percentage of records compliant with rules |
Timeliness | Is the data up to date? | Outdated executive data | Average age of records |
Not every dimension is equally critical for every use case. For supplier payment data, accuracy is paramount because errors lead to incorrect payments. For customer master data, uniqueness is crucial, while for regulatory data, timeliness and validity are key. A good target framework specifies, for each use case, which dimension must meet which threshold—experience shows that blanket quality targets without this prioritization are ineffective.
In practice, the most business-critical dimension is usually uniqueness. Duplicates are the most costly and common quality defect because they directly result in duplicate payments, distorted revenue analyses, and contradictory customer histories. Industry benchmarks show that companies without formal data quality processes regularly have duplicate rates of 10 to 30%—a figure that can only be sustainably reduced through systematic validation.
Measuring Data Quality: KPIs and Methodology
Data quality can only be managed if it is measured. Without a baseline, it is impossible to demonstrate progress or justify a business case to the board. In practice, two categories of metrics have proven effective: quantitative metrics and qualitative indicators.
Quantitative Metrics
Quantitative metrics can be collected automatically and on a regular basis: the completeness rate of critical required fields, the number of identified and resolved duplicates, the average time to update changed data records, and the percentage of automatically validated data records. These metrics form the backbone of every data quality dashboard.
Qualitative Indicators
Qualitative indicators provide additional context: employee satisfaction with the data, the reduction in manual corrections in day-to-day operations, the perceived improvement in decision-making speed, and the assessment by external auditors during audits. While they are more difficult to automate, they are often crucial for the acceptance of a data quality initiative within the company.
The Right Frequency for Measurement and Reporting
Critical domains, such as customer or supplier master data, should be monitored continuously, with automated alerts triggered when thresholds are exceeded. Less critical domains can be checked on a quarterly basis. It is important that the measurement frequency matches the rate of change of the respective data—rigid, uniform review cycles for all domains waste resources in non-critical areas and overlook problems in critical ones.
The Four Root Causes of Poor Data Quality
Before a company invests in tools, it should understand the actual causes of its data quality problem. In practice, four recurring root causes emerge—they are rarely technical in nature, but mostly organizational.
Root Cause 1: Evolving System Landscapes Without Centralized Governance
Most companies have grown over decades—both organically and through acquisitions. The result is a patchwork of ERP instances, legacy systems, and cloud applications, each of which maintains its own version of the truth about customers, products, and suppliers. Without a central authority for data quality, inconsistency becomes systemic and inevitable.
Every new cloud application that a business unit introduces on its own creates yet another data silo. Without an overarching authority to define which system is the primary source for which domain, the number of conflicting versions grows faster than an organization can manually reconcile them.
Root Cause 2: Lack of Data Ownership
“Whose data is this?” — in many organizations, this question remains unanswered. If no one is responsible for a dataset, no one actively maintains it. IT manages systems but not content; business units maintain content in spreadsheets but not in the system. Without designated data owners, every quality initiative is doomed to fail.
Data ownership is therefore not a bureaucratic detail, but the lever with the greatest impact. As soon as a designated person from the business unit is responsible for the quality of a domain—with a clear mandate and the right tools—behavior changes noticeably. Anonymous data sets become well-maintained assets with a face behind them.
Root Cause 3: No Validation at the Source
Most data quality issues arise during initial data entry—due to missing required fields, missing format checks, or missing duplicate checks. If an erroneous data record is cleaned up only after the fact, it has already been incorporated into dozens of downstream processes and reports. Prevention is always less costly than correction.
This becomes particularly critical during manual initial data entry without any system support—for example, when a new supplier is added via an email attachment. Without automated validation at the time of entry, every typo, every incomplete entry, and every unnoticed duplicate makes its way directly into the production database.
Root Cause 4: Data Quality as a One-Time Project Rather Than an Ongoing Process
Many companies clean up their data once—as part of a migration or an audit—and then consider the issue resolved. Without continuous monitoring and established processes, however, data quality reverts to its original level within a few months because new employees, new systems, and new processes constantly introduce fresh inconsistencies.
Not a Data Cleansing Problem, but a Governance Problem
The most common misconception about data quality is treating it as a technical cleanup project—a one-time cleanup carried out by IT or an external service provider before returning to “actual” business operations. This perspective explains why so many data quality initiatives revert to their original state after a short time.
The sustainable perspective is different: Data quality is the result of governance—of clear responsibilities, binding rules, and continuous measurement. Technology can support and automate this process, but it does not replace it.
Incorrect Framing | Correct Framing |
“We’ll clean up the data once.” | “We’ll establish long-term data quality governance.” |
“Data quality is an IT issue.” | “Data owners in the business units are responsible.” |
“A tool solves our problem.” | “Governance and processes come before technology.” |
“We need better data.” | “We define measurable targets for each dimension.” |
“Quality is complete when the project ends.” | “Quality is continuously measured and managed.” |
This shift in perspective has practical implications for the budget, sponsorship, and roadmap. If data quality is treated as a governance issue, it receives C-level sponsorship and a permanent structure rather than a temporary project budget. Learn more about the role of governance in master data management in the article on the Golden Record.