Case studies / Data foundations

Data foundations · Industrial distribution

Turning a 40-million-SKU item master into data buyers could actually find

A leading, PE-owned global parts distributor had built an enormous catalog — tens of millions of SKUs across hundreds of manufacturers — but its own website couldn’t reliably find them. Searches for in-stock parts came back empty, paid ads pointed shoppers to the wrong product, and the company was quietly running two separate, disconnected item masters for two parts of the business. We rebuilt the data underneath the catalog — sourcing, validating, and standardizing product identifiers at scale — and gave the client a field-by-field, dollar-quantified view of which data actually drives sales.

Outcomes

13% → 90%
Identifier coverage

GTIN/UPC coverage on the client’s single largest manufacturer line — worth over a quarter of total revenue — from a brand whose own website carries no GTIN data at all.

~60%
Paid-search click-through lift

Early testing showed click-through rate on paid ads rising by roughly this much once verified GTINs were populated into the product feed.

$258M
Sales matched to verified data

Of trailing two-year sales value, newly connected to a validated product identifier that didn’t exist in the item master before.

Figures reflect progress tracked jointly with the client against agreed data-quality metrics. The click-through figure comes from early, short-window tests flagged as preliminary at the time.

Sector
Industrial & factory-automation parts distribution (MRO / break-fix aftermarket)
Engagement
Item master data enrichment and data architecture optimization, delivered as a two-track program
Delivery
Built and operated by DataGrokr end-to-end; delivered to the client in partnership with Insight Factory
Cloud
Multi-cloud data pipelines (GCP, AWS) alongside the client’s existing Azure-hosted reporting warehouse

The challenge

A catalog too big to trust, split across two systems that couldn’t talk

The client’s item master held roughly 40 million records across hundreds of manufacturer brands, and it was quietly undermining the business it was supposed to support. Customers searching the site for an in-stock part by its manufacturer’s own barcode got zero results. Paid search ads matched to the wrong product. Whatever the precise frequency, it was consistent enough with the broader eCommerce metrics — search-to-click and click-to-cart rates below where the client wanted them — to justify treating item data as the root cause. And underneath it all, the company was running two independent item masters for two parts of its business, with no way to reconcile a record in one against a record in the other.

What made this hard wasn’t fixing any one field — it was that no single data source could be trusted on its own. Manufacturer websites were the obvious place to start, but coverage varied wildly: the client’s single largest product line, worth more than a quarter of total revenue, came from a manufacturer whose own site carries no GTIN or barcode data at all. And the two item masters weren’t just separate databases; they used fundamentally different rules for what made a record unique, so merging them outright would have meant a costly, disruptive rebuild neither business unit was asking for.

Before

  • Documented cases of searches on the client’s own site returning zero results for in-stock parts with a valid manufacturer barcode
  • Paid ads mismatched to the wrong part, consistent with search-to-click metrics running below target
  • Two independent item masters, for two parts of the business, with no way to reconcile records between them
  • No quantified view of which item-data fields actually moved sales — only anecdote and instinct

After

  • Verified, standardized product identifiers rebuilt at scale across the highest-revenue manufacturer lines
  • The client’s top product line’s identifier coverage lifted from 13% to 90% of its sales base
  • An automated “ItemMap” letting the two item masters interoperate for the first time, without merging either one
  • A field-by-field, dollar-quantified roadmap — which data actually lifts click-through, conversion, and margin, and which doesn’t

How we approached it

Enrich by revenue impact, not record count

01

Treat every data source as a hypothesis, not a source of truth

Manufacturer websites were the starting point, but never the final word — we cross-validated every candidate identifier against competitor listings and third-party product databases, because manufacturer data alone was inconsistent enough that trusting it outright would have introduced as many errors as it fixed.

02

Prioritize by revenue, not by record count

With 40 million records to work through, we ran the numbers before writing a line of enrichment logic: a Pareto analysis on two years of sales identified exactly which manufacturers and fields were worth the cost of enrichment — and just as importantly, which weren’t. One clear finding: 137 manufacturers, 11.6 million line items between them, had sold for less than ten cents a part over two years. That’s effort deliberately not spent.

03

Solve interoperability without a rip-and-replace merge

Rather than forcing the client’s two item masters into one system — expensive, disruptive, and not what either business unit needed — we designed a federated model: centralize a small set of core identity fields at the point of item creation, and layer an automated, monthly “ItemMap” job on top that keeps both systems resolving to each other, while each keeps running its own operations independently.

Every enrichment tied back to a dollar

We didn’t treat “cleaner data” as its own reward. Each data field we improved was tested against a real business metric — GTIN presence against paid-search ROAS and click-through rate, ad-title field ordering against ROAS, product lifecycle stage against realized margin and pricing power — so the client could see, even from early and preliminary reads, which data investments looked likely to pay for themselves, and prioritize what came next accordingly.

Under the hood

PythonLLM-based categorization & deduplicationWeb crawling & data-validation pipelinesGoogle BigQueryGCPAWSGoogle Ads / Merchant CenterGoogle AnalyticsSQL ServerAzureElasticsearch

Next step

Sitting on a product, customer, or item catalog too big to trust?

We rebuild the data underneath the systems that run your business — sourced, validated, and quantified against the outcomes that actually matter — without forcing a disruptive, all-or-nothing rebuild.

Start a conversation →