Amazon Business at $60B: Agentic B2B Procurement Has a Product-Data Problem
Amazon Business reaching $60B in annualized gross sales shows where B2B procurement is going—and why distributors need trusted product identity before agentic commerce can work.
Amazon Business reaching $60 billion in annualized gross sales is more than another marketplace milestone. It is evidence that B2B buying is moving toward an environment where software can search, compare, approve, and eventually place orders through structured digital systems.
The usual conclusion is that distributors need better AI. The more useful conclusion is that they need better product data.
An agent can only purchase what it can identify. It cannot infer with acceptable confidence that four supplier records, two manufacturer part-number formats, and a holding company’s private label all refer to the same physical item. It cannot compare technical specifications that live in PDFs, interpret a blank pack quantity, or know whether EA, each, and 1/box describe compatible offers. The storefront may be ready for an agent while the catalog underneath it is not.
The $60B signal exposes a data-readiness gap
Amazon Business demonstrates what becomes possible when products, offers, buyers, policies, and transactions are accessible through operational digital infrastructure. At that scale, product data is not content added after the commerce system is built. It is part of the system.
Most industrial distributors are starting from somewhere else. Their source material arrives as supplier spreadsheets, PDF catalogs, emailed price lists, portal downloads, and inconsistent exports. One file calls an attribute nominal diameter; another uses DN; a third hides it in a description. GTINs may be missing, manufacturer part numbers may include different punctuation, and a single product may have separate identities for the manufacturer, regional operating company, and distributor SKU.
That creates a widening gap:
| AI-ready commerce expects | Many distributors actually have | Operational consequence |
|---|---|---|
| A stable product identifier | Missing GTINs and inconsistent MPNs | Agents cannot resolve an offer to the correct real-world product. |
| Structured, comparable specifications | Attributes embedded in PDFs, descriptions, and supplier-specific columns | Search and comparison depend on extraction or guesswork. |
| One current product record | Duplicates across suppliers, subsidiaries, and systems | Availability, price, and technical facts attach to competing identities. |
| Machine-accessible commercial data | Periodic spreadsheets and manual portal downloads | The agent sees stale pack, lead-time, price, or lifecycle information. |
| Evidence for important facts | Values copied forward without source or date | Neither the agent nor a reviewer can establish which value is trustworthy. |
Converting a spreadsheet to JSON does not close this gap. It makes the ambiguity accessible through an API. The record still has to be matched, normalized, validated, and kept current.
This is where Claro operates: between source documents and the systems that need dependable data. The job is to turn fragmented supplier inputs into canonical product records with source evidence, confidence, and an exception path for uncertain cases. The result can feed an ERP, PIM, ecommerce site, procurement platform, or AI agent without forcing any of those systems to solve supplier-data chaos on its own.
Supplier onboarding is the hidden bottleneck in agentic procurement
Agentic procurement is normally presented as a buying workflow: describe a need, let software find compliant options, compare them against policy, and complete the transaction. But every step assumes the item can be identified correctly.
Consider a maintenance team looking for a replacement bearing. The procurement agent encounters four records:
- the manufacturer’s current MPN;
- an older MPN retained in the buyer’s ERP;
- a supplier SKU with no GTIN;
- and a regional subsidiary’s record with dimensions written in free text.
A person may recognize that the four records describe the same bearing after opening a datasheet and calling the supplier. An agent sees four uncertain candidates. If it merges them carelessly, it may buy the wrong variant. If it refuses to merge them, it may miss the best available offer. If it selects the record with the richest marketing copy, it may optimize for completeness rather than compatibility.
That is not primarily an agent problem. It is an identity-resolution problem created during supplier onboarding.
The supplier-onboarding workflow therefore becomes commerce infrastructure. Before a new catalog is exposed to agents, a distributor needs to:
- 1Extract the source facts
Read spreadsheets, PDFs, price files, and feeds into structured fields without discarding the original evidence.
- 2Resolve product and supplier identity
Determine which rows are new products, variants, duplicates, replacements, or offers for items that already exist.
- 3Normalize and validate
Standardize identifiers, units, attribute names, pack structures, and allowed values, then flag conflicts rather than silently choosing one.
- 4Create a canonical record
Link every supplier and internal identity to one maintained representation of the real-world product while preserving the cross-references needed to transact.
- 5Monitor the next supplier change
Re-run the controls when a price list, datasheet, lifecycle status, or product specification changes.
This reframes why supplier onboarding takes weeks. The work is not merely loading rows into a destination. It is establishing whether each row describes the same thing the supplier, distributor, buyer, and agent believe it describes. The practical techniques are product matching, entity resolution, and a maintained canonical product record.
The contrarian lesson: stop upgrading systems until the inputs are trustworthy
Industrial distributors have spent a decade on digital transformation. Many have replaced an ERP, implemented a PIM, redesigned ecommerce, or added a data lake—and still maintain critical catalog fields by emailing spreadsheets between teams.
The projects did not necessarily fail because the software was wrong. They stalled because a destination system cannot manufacture trusted source data. A new PIM can enforce a required field, but it cannot decide which of three conflicting voltage values is correct. An ERP can store an MPN, but it cannot prove that two differently formatted MPNs refer to the same item. A polished storefront can display a specification table, but it cannot recover a specification that never left the supplier’s PDF.
Amazon’s advantage should not be reduced to a better storefront. The operational lesson is that product data, identity, offers, and transactions work together at scale. Industrial distributors do not need to imitate every part of Amazon Business. They do need to treat the data feeding their systems as infrastructure rather than as a cleanup task assigned after implementation.
Before approving another transformation project, ask four questions:
- What percentage of active products have a validated GTIN or manufacturer part number?
- How many apparent SKUs resolve to the same real-world product across suppliers and subsidiaries?
- Which transaction-critical attributes have a traceable source, date, and confidence level?
- How long does a supplier change take to reach every system and channel that depends on it?
If those answers are unknown, adding another platform creates one more place for unreliable data to travel. Fixing the upstream layer first improves the ERP, PIM, storefront, procurement workflow, and future agents at the same time. It also makes a later migration less risky because the organization moves resolved records rather than accumulated ambiguity.
The Brickworks connection: AI should repair the foundation before acting on it
This is the distributor version of the Brickworks data-quality lesson. Brickworks used AI agents to recommend missing master-data values inside a governed process, with subject-matter experts validating the recommendations. The model accelerated data repair; it did not bypass stewardship.
Distributors face the harder, recurring version because much of their product data originates outside the company. Dozens or hundreds of suppliers change their own documents and exports independently. That makes a one-time cleanup insufficient. The operating loop has to detect new inputs, resolve identity, propose corrections or enrichment from evidence, route uncertain changes to people, write accepted values back, and monitor again.
That loop is how AI becomes useful before the catalog is perfect. Use it first in service of data quality; then let buying agents act on the trusted result. Reversing the order only automates the consequences of unresolved data.
What distributors should do now
Do not begin with a company-wide agentic-procurement strategy. Begin with one high-value category and one representative set of supplier files.
- Measure identity coverage. Count valid GTINs and MPNs, duplicate candidates, orphaned supplier SKUs, and unresolved cross-references.
- Select transaction-critical attributes. Prioritize the facts that determine fit, compliance, pack quantity, price, and availability—not every possible marketing field.
- Build one evidence-backed canonical set. Attach each accepted value to its source and retain confidence and review status.
- Test real procurement questions. Ask an agent to find substitutes, compare compatible products, or select an approved item, then inspect where uncertainty blocks the task.
- Operationalize the refresh. Repeat the process when the next supplier feed arrives instead of treating the first clean result as finished.
The $60 billion milestone shows that digital B2B procurement is already large. The agentic shift raises the quality bar again: systems will make more decisions at machine speed and have less tolerance for catalog ambiguity. For distributors, trusted product identity is no longer preparation for some distant AI future. It is the admission ticket to the commerce infrastructure being built now.
Free catalog audit
See what an agent can—and cannot—identify
Send Claro a representative catalog extract. We will show you the missing identifiers, duplicate identities, conflicting specifications, and untraceable values blocking agent-ready commerce.
Product-data layer
Build trusted product identity
See how Claro resolves supplier inputs into canonical, evidence-backed records for the systems and agents you already use.
Source and related reading
Source
MarketScale: Amazon Business reaches $60B
The market signal behind this analysis: Amazon Business's annualized gross-sales milestone and the rise of agentic AI in B2B procurement.
Related article
Agentic commerce runs on machine-readable product data
Why machine-readable is only half the requirement and continuous trust is the durable infrastructure layer.
Related article
The missing layer between supplier documents and the PIM
Where fragmented source data must be extracted, matched, validated, and made ready for downstream systems.
Related guide
Add an AI layer without replacing ERP
How to improve catalog operations above existing systems of record rather than beginning with another replacement.
FAQ
What product data does agentic B2B procurement require?
Procurement agents need structured, current product records with stable identifiers, normalized specifications, pack and unit data, availability, pricing, and a clear link between supplier items and the buyer’s internal catalog. The records also need provenance and confidence so an agent can distinguish verified facts from uncertain matches.
Why is product identity a prerequisite for agentic commerce?
An agent cannot reliably compare, recommend, or purchase an item when the same product appears under multiple supplier and holding-company identities, or when identifiers and specifications are missing. Trusted product identity gives the agent one canonical record for the real-world product and preserves the cross-references needed to transact.
Should distributors replace their ERP or PIM before preparing for AI procurement?
Usually not. A new system will not correct unreliable supplier inputs by itself. Distributors should first measure and repair the product data feeding their existing systems, then decide whether a platform upgrade is still necessary. Clean, resolved data improves the current stack and makes any later migration safer.
Claro
See where your catalog breaks — free
Claro runs this automatically: resolve identity, fill missing attributes, validate updates, and write clean records back into your PIM/ERP. Upload a sample supplier file for a free catalog audit.
Get a free catalog audit