Product Data Quality Metrics Every Catalog Team Should Track
Build a product data scorecard that measures completeness, validity, duplicates, classification, provenance, freshness, supplier quality, and business impact.
Completeness alone can hide a broken catalog. Every field can be populated and every one of them wrong. Effective product data quality metrics measure whether values are valid, consistent, current, attributable, and fit for their downstream purpose — not merely present.
A useful product data scorecard has three layers: the quality of individual fields, the integrity of product records, and the business outcomes affected by those records.
The core catalog quality KPIs
| Metric | Practical definition | Useful cut |
|---|---|---|
| Required-attribute completeness | Populated required fields ÷ required field opportunities | Category, supplier, channel |
| Identifier validity | Identifiers passing format, checksum, uniqueness, and ownership rules | Identifier type and source |
| Duplicate rate | Confirmed duplicate entities ÷ active product entities | Category and ingestion source |
| Classification accuracy | Reviewed products assigned to the correct class | Taxonomy level and category |
| Schema conformance | Values satisfying type, cardinality, vocabulary, range, and unit rules | Attribute and destination |
| Provenance coverage | Published values with an attributable source and capture time | Attribute and source tier |
| Freshness | Values inside their category-specific refresh window | Volatility class |
| Review rate | Records requiring human action ÷ processed records | Reason and workflow stage |
| Write-back success | Accepted updates confirmed in the destination ÷ attempted updates | System and error reason |
| Downstream error rate | Transactions or listings failing because of product data | Channel and error type |
1. Attribute completeness — with the right denominator
Measure required opportunities, not all possible fields. An ingress protection rating may be required for an enclosure and irrelevant for a desk. Define requirements by product class, market, channel, and workflow.
Report critical-field completeness separately. A record missing marketing copy is not equivalent to one missing pack quantity, voltage, or a compliance attribute. Weighted reporting is useful, but always retain the raw counts beneath the score.
2. Identifier validity, not identifier presence
A populated GTIN field can still contain an invalid checksum, a placeholder, a pack-level code attached to an each-level record, or a duplicate copied across products. Validate syntax, checksum where applicable, expected owner or prefix, uniqueness rules, and relationship to packaging level.
Apply equivalent rules to MPNs, supplier SKUs, and internal IDs. Their validation is contextual: an MPN needs manufacturer identity; a supplier SKU needs supplier identity.
3. Duplicate and false-merge rates
Duplicate rate measures unresolved fragmentation. False-merge rate measures distinct products incorrectly collapsed together. Track both: a team can reduce apparent duplicates by merging aggressively while making the catalog less trustworthy.
Use reviewed samples and confirmed operational cases to estimate rates. Segment by source, category, and identifier availability so a clean consumer-goods feed does not conceal poor matching in the industrial long tail. The real cost of duplicate products explains why this KPI belongs on an operating dashboard.
4. Classification accuracy and coverage
Coverage answers whether a product has a class. Accuracy answers whether it has the right class. Measure top-level and leaf-level accuracy separately, plus the percentage assigned to generic fallback categories. Track taxonomy version so a score change caused by a revised hierarchy is not mistaken for model drift.
5. Source and provenance coverage
Source coverage asks whether priority manufacturers, supplier files, and documents are represented. Provenance coverage asks whether each operational value can be traced to a source, capture time, and transformation.
Create a stricter KPI for high-risk fields: authoritative provenance coverage. A value copied from a marketplace page may be attributable yet still fail your source hierarchy. Filling missing attributes with provenance shows how to retain that distinction.
6. Freshness and change latency
Set refresh windows by value volatility. Inventory may age in minutes, price in hours or days, technical dimensions in years, and compliance documentation according to certificate validity. Then measure:
- percentage inside the freshness window;
- time from source change to detection;
- time from approved change to destination write-back;
- stale values still published after a source withdrawal.
One universal “last updated” target is not meaningful.
7. Review and exception metrics
Review rate is not automatically bad. It may indicate a cautious control on a new supplier. Pair it with auto-accept precision, median review time, backlog age, reviewer agreement, and exception reasons. The goal is to automate repetitive certainty and focus people on consequential ambiguity.
8. Supplier quality
For every supplier, track delivery timeliness, required-field completeness, validity, duplicate contribution, schema drift, correction response time, and the percentage of values accepted without remediation. A supplier scorecard turns catalog cleanup into a feedback loop rather than a recurring internal cost.
9. Write-back reliability
An approved correction that never reaches the ERP or PIM has produced no operational benefit. Track attempted, accepted, rejected, retried, and confirmed updates. Include idempotency failures, mapping failures, destination validation errors, and latency to confirmation.
10. Downstream errors and business impact
Connect leading quality indicators to outcomes: search zero-results, listing rejection, order correction, return reason, invoice dispute, compliance hold, supplier onboarding time, conversion, and margin leakage. Do not claim causation from correlation alone; use incident links, controlled fixes, or before-and-after cohorts where possible.
Build a scorecard that cannot be gamed
- Define rules by product class and use caseDocument required attributes, accepted values, authoritative sources, tolerances, and freshness windows.
- Publish numerator, denominator, and exclusionsA percentage without its population is easy to misread. Preserve zero denominators and excluded records explicitly.
- Segment before aggregatingShow category, supplier, source, channel, and workflow stage. Weight an overall score by risk, not convenience.
- Set a baseline and targetsRecord the starting distribution, set thresholds per metric, and assign an owner and remediation action.
- Link every exception to evidenceStore the failing rule, source value, normalized value, affected destination, and resolution.
Related resources
Article
Product Content Audit
Turn catalog evidence into a prioritized remediation plan.
Glossary
Supplier Scorecard
Measure incoming data quality and supplier improvement.
Guide
The Real Cost of Duplicate Products
Connect identity defects to operational and commercial impact.
Guide
Fill Missing Attributes with Provenance
Improve completeness without losing trust in the source.
Playbook
Product Deduplication
Measure and remediate duplicate product identities.
Guide
Product Classification
Classify inherited records and validate the result.
Guide
Seven Field-Level Errors
Find populated values that still fail catalog quality rules.
FAQ
What are the most important product data quality metrics?
How is product data quality scored?
How often should catalog quality KPIs be reviewed?
Claro
See where your catalog breaks — free
Claro runs this automatically: resolve identity, fill missing attributes, validate updates, and write clean records back into your PIM/ERP. Upload a sample supplier file for a free catalog audit.
Get a free catalog audit