AI in Service of Data Quality—and Why Most Distributors Cannot Do What Brickworks Did

Brickworks used AI agents to recommend missing master-data values. For distributors, the harder problem is keeping changing supplier data trustworthy.

published ai-ready-product-datadata-qualityproduct-enrichment

You cannot run AI on a catalog you cannot trust. The fastest fix is to use AI to repair the data itself—with provenance on every value and a human validating before it is written back. The hard part is not the one-time cleanup. It is keeping the data clean as suppliers keep changing it.

A brick manufacturer just said the quiet part out loud

AI projects rarely stall because the model cannot generate an answer. They stall because the item master has 40,000 records and 11,000 of them are missing the attribute the answer requires.

Brickworks offers a useful reference point—not a customer story or a template every company can copy. After investing in its data platform, governance, and stewardship, the building-products manufacturer used AI agents to recommend values for incomplete fields, with subject-matter experts validating the recommendations before application. Its chief data and digital officer described the approach as using “AI in service of data quality,” as reported by iTnews.

Those are the two facts that matter: agents proposed missing values, and accountable people approved them. The model did not replace data governance. It made the governed workflow faster.

Why the “AI in service of data quality” framing is correct

The familiar sequence—clean all the data, then begin the AI program—is economically unrealistic for a large catalog. By the time a team finishes the cleanup, new supplier files have already changed the baseline. AI should help perform the repair, provided it operates inside controls rather than writing plausible text straight into production.

Claro’s loop is: detect a change, resolve product identity, validate and enrich the record, write accepted values back, and monitor the source again. That loop also explains the difference between data cleansing and data enrichment: cleansing corrects what is already represented, while enrichment adds missing facts from evidence. Most AI-ready catalog work needs both.

The part that does not generalise: Brickworks only had to clean its own data

A manufacturer can define an internal master, put it on one platform, and assign named stewards to domains it controls. A distributor’s catalog is assembled from 40 to 140 suppliers it does not control. Inputs arrive as spreadsheets, PDF price lists, BMEcat exports, portals, and emails. Attribute names, pack structures, and document versions change independently.

That is the missing data layer between supplier documents and the PIM. The nominal steward of a supplier field is the supplier, but the distributor still bears the operational consequences when that field is blank, ambiguous, or stale. This recurring dependency is also why product-data debt accumulates in industrial distribution.

Project versus product: why one-time cleanup does not hold

A cleanup is accurate at a moment in time. It starts decaying when the next price file lands: a supplier renames an attribute, changes a pallet quantity, replaces a datasheet, launches a variant, or discontinues a line.

The meaningful choice is therefore one-time versus continuous enrichment. A project produces a cleaner snapshot. A product observes incoming change and preserves a trustworthy state. Teams can make that concrete with a repeatable process to detect and fix product-data drift.

What “AI recommends, human validates” needs to be safe

Three controls separate an evidence-backed proposal from faster guessing:

  1. A confidence score on every proposed value. The score should reflect source quality, identity-match strength, extraction certainty, and validation results—not merely model fluency. See confidence score.
  2. Provenance to the exact evidence. A reviewer must be able to open the source document, page, table, or feed cell behind the value. See data provenance and the guide to an AI enrichment source link.
  3. A reviewable, reversible write-back. Accepted changes need an audit trail; uncertain or conflicting values need a queue. The workflow should support correction without losing the prior value.

These controls are how teams pursue enrichment without hallucination, design human-in-the-loop product-data review, and set confidence thresholds for auto-merge. For a concise operating test, use the guide to trusting AI-enriched data.

Before and after: the difference is evidence

Record Fire class Compressive strength Water absorption Pack and coverage
Before: BRK-STD-65 · Brick
After: BRK-STD-65 · Brick A1 to EN 13501-1 20 N/mm² 7% 500 pcs/pallet; 48.5 pcs/m² at 10 mm joint

The after record is not trustworthy merely because it has more fields. Each value must carry its source—for example, supplier DoP PDF, page 2—plus extraction date, confidence, validation result, and approval status. If the source changes, the affected values can be found and re-evaluated.

What to do if you do not have a year and a Snowflake platform

Start with one supplier and one category, not the entire catalog.

  1. Measure completeness for the attributes that matter to transactions, search, and compliance.
  2. Collect the current supplier feed and source documents.
  3. Resolve each source to the correct product and variant.
  4. Let AI propose missing values, then validate format, units, allowed values, and category-specific ranges.
  5. Set a conservative auto-apply threshold and route every other proposal to review.
  6. Measure accepted values, rejected values, reviewer time, and supplier changes on the next refresh.

That small loop exposes whether the real constraint is extraction, identity, source quality, schema design, or review capacity. It also creates a defensible path to AI output validation for distributors without pretending the catalog will ever stop changing.

Claro

See where your catalog breaks — free

Claro runs this automatically: resolve identity, fill missing attributes, validate updates, and write clean records back into your PIM/ERP. Upload a sample supplier file for a free catalog audit.

Get a free catalog audit