A product data quality assessment should take three to six weeks and end with three things. A scored picture of one part of your catalogue. A named mechanism behind each failure. A costed, sequenced remediation plan you can take to a budget holder. If it ends with a dashboard and no plan, it was an expensive way to confirm what everyone already suspected. This is the method we run, written out in full so you can run it yourself.
What a product data quality assessment is actually for
It is not for proving the data is bad. Everybody in the building knows that already.
It exists to answer a specific commercial question. Should we buy a PIM. Can we launch on this marketplace. Why are returns rising in this category. How much will it cost to get 40,000 SKUs ready for a new website. The assessment is the evidence base for a decision, and the decision determines what you measure.
We start every one of these by asking what changes as a result. If nobody can name a decision, we say so and stop. An assessment with no decision attached becomes a report that gets circulated once and never opened again.
Before you start: three things you need
A decision, named. See above. Write it in one sentence at the top of the brief.
A scope you can finish. One or two categories with real revenue behind them, ten to fifty thousand SKUs. Whole-catalogue assessments sound thorough and produce averages that hide everything.
A person who can answer category questions. Someone who knows what a bore diameter is, or which fastener grades you actually stock. Usually a merchandiser or technical product manager. Without them, stage three stalls and the whole thing drifts.
Stage 1: Write down the questions before you touch data
Three to five questions, no more. Each one has to be answerable with a number.
A bad question is how good our data is. A good question is what proportion of the top category can be published to Amazon today without manual intervention. The second has a testable definition and an obvious owner.
Get those questions agreed in writing by whoever asked for the assessment. This step takes an hour and prevents the most common failure, which is delivering a technically correct assessment that answers a question nobody asked.
Output: a one-page brief with the decision, the questions, the scope and the named category expert.
Stage 2: Pull four extracts, not one
Most assessments look at the PIM export alone. That tells you what you have, not how it got there or where it goes. Pull four things:
- The item export for the scoped categories. Every attribute, every language, every channel flag. Not the trimmed version marketing uses.
- The raw supplier data as it arrived, for the five largest suppliers in scope. Spreadsheets, XML, PDFs, whatever the format was. This is where root causes live.
- The published output, scraped from your own website and from any marketplace or trade portal you sell through.
- The schema itself. The attribute list, the types, the units, the controlled vocabularies, the category assignments.
Comparing extract one against extract three finds transformation and syndication faults nobody knew about. Comparing extract two against extract one finds onboarding faults. Those are different problems with different fixes, and one export cannot separate them. If the supplier side dominates, the answer is usually a change to supplier data onboarding rather than a cleansing project.
Output: four datasets in one place, with a note on how long each took to obtain. That duration is itself a finding.
Stage 3: Define the standard before you measure against it
You cannot score a catalogue against a standard that does not exist. In most assessments this stage is the longest, and it is the one clients least expect.
For each category in scope, produce a one-page definition:
- Mandatory attributes. Cannot publish without them.
- Commercially important attributes. Drive filtering, comparison and search in this category.
- Optional attributes. Improve the page, do not block it.
- Per attribute: type, unit, permitted values, and the source you consider authoritative.
Then get the category expert to sign it. Not approve it in a meeting, sign it. This document becomes the denominator for every completeness number in the assessment, and it will be challenged the moment the scores are unflattering.
If you already have a working attribute schema, this stage takes a day. If you do not, it takes two weeks, and that is a finding worth reporting on its own. Our taxonomy and attribution work exists mostly because this artefact is missing so often. The argument for why the schema comes first is set out in our piece on product attributes.
Output: a signed category standard per category in scope.
Stage 4: Run the automated rules
Now the cheap part. One testable rule per attribute per dimension, run over the whole extract.
Completeness rules count populated mandatory attributes per category. Validity rules test type, unit, format and membership of a controlled list. Uniqueness rules group by brand plus manufacturer part number, then by GTIN, then by a normalised signature of the key attributes. Consistency rules profile the distinct values of each attribute and flag near-duplicates such as “Zinc”, “ZP” and “BZP”.
Report pass rates per attribute, not per product. A 74 per cent product-level score tells you nothing. “Thread pitch is populated on 31 per cent of records and valid on 12 per cent” tells you exactly what to fix on Monday.
ISO 8000-8:2015 is useful here because it names the limit of this stage. It splits quality into syntactic, semantic and pragmatic. Automated rules test syntactic quality thoroughly. They cannot test whether the value matches the real product. That is stage five.
Output: a rule-by-rule pass rate table, and a list of every failing record so the results are traceable.
Stage 5: Sample by hand, because tools cannot check truth
This is the stage most assessments skip, and skipping it is what makes them worthless.
Take a stratified sample. We use around 100 SKUs per category, weighted towards high revenue, high return rate and high enquiry volume. For each one, open the manufacturer’s current datasheet and check the mandatory attributes value by value. Record which attribute was wrong, not just that the product was wrong.
Three things come out of this that nothing else finds. Values that are valid, complete and factually wrong. Records describing a superseded version of the product. Attributes that were populated by assumption rather than by reading a source.
Expect this to take a person two to three days per category. It is the most expensive stage per record and the highest value output in the assessment. It is the only number a commercial director will actually believe.
Output: an accuracy error rate per attribute, with the sample size stated.
Stage 6: Find the mechanism, not the symptom
Every finding needs a cause, and the cause is almost never carelessness.
ISO 8000-61:2016 sets out a process model for data quality management. It puts root cause analysis and process improvement alongside data cleansing, not after it. That ordering is the whole point. Cleansing without mechanism change buys you eighteen months.
In practice we see the same handful of mechanisms:
- No mandatory attribute list, so nobody knew what complete meant.
- No type or unit enforcement at the point of entry, so free text got in.
- No controlled vocabulary, so four spellings of the same finish coexist.
- Supplier data accepted in any format, so normalisation happens by hand every time.
- No source or verification date held, so nobody can tell fresh values from stale ones.
- Enrichment treated as a project, so the catalogue decayed from the day it finished.
Map each measured failure to one of these. When you do, a list of forty findings usually collapses into four or five mechanisms. That is a far easier thing to take to a board.
Output: a mechanism per finding, and a count of findings per mechanism.
Stage 7: Size the remediation and sequence it
The plan is what makes the assessment worth paying for. It needs three columns: what, how much, and in what order.
Size the work in the units the work is actually done in. Records to re-key. Attributes to populate. Datasheets to read. Images to match. Suppliers to renegotiate a feed with. Then apply your own throughput rates, or ours if you are buying the delivery too, to turn that into days and cost.
Sequence by two axes. Fix mechanisms before you fix records, always, or you will pay twice. Then order the record-level work by revenue exposure rather than by how bad the score is. The worst category in the catalogue is often the one nobody buys from.
Present three options, not one. Mandatory-only across the full scope. Full enrichment of the top categories. Everything. Give a cost and a duration for each. Budget holders approve a choice far more readily than a single number.
Output: a costed, sequenced remediation plan with three scope options.
How long a product data quality assessment takes
Three to six weeks of elapsed time for one or two categories, assuming the extracts arrive promptly. The distribution of effort is not what people expect.
Stage 3, defining the standard, is usually the longest if no schema exists. Stage 5, hand sampling, is the biggest fixed cost per category. Stages 1, 2 and 4 are quick once the data is in the room. Stage 7 takes two or three days if stages 5 and 6 were done properly, and much longer if they were not.
We run this as a fixed-scope engagement. A single category is priced at £4,000 to £6,000. A multi-category scope is priced at £12,000 to £25,000. The output is the same either way: scored findings, mechanisms, and a costed plan.
What the output should look like
Four artefacts. Anything else is packaging.
The scorecard. Pass rates per attribute per dimension, per category, with the sample-based accuracy figure shown next to the automated figures rather than blended into them.
The failing record lists. Every finding traceable to actual SKUs. Without this, the scores cannot be challenged or verified, and they will be.
The mechanism map. Findings grouped by cause, with a count against each.
The remediation plan. Costed, sequenced, three scope options.
We publish the scorecard and the plan together in the same document. Separating them is how assessments turn into shelfware. The engagements this leads into are visible across our case studies and our product data services.
Five ways a data quality assessment goes wrong
Scoring the whole catalogue at once. Produces averages. Averages hide the categories that matter and flatter the ones that do not.
Measuring completeness against every field in the PIM. Guarantees a terrible score that means nothing, because most fields belong to other categories.
Publishing an automated pass rate with no hand-checked accuracy sample. The green number gets believed, the assessment gets closed, and the accuracy problem surfaces six months later as returns.
Stopping at findings. A list of problems with no mechanism and no cost is not an assessment. It is a complaint.
Assessing the PIM export only. You cannot separate a supplier onboarding fault from a syndication fault with one extract. You will fix the wrong end of the pipe.
There is a sixth, which is running the assessment before anyone agreed what decision it feeds. We covered why measurement without a decision attached goes nowhere in our piece on product content performance.
Key takeaways
- A product data quality assessment is evidence for a decision. Name the decision first or do not start.
- Scope to one or two revenue-bearing categories. Whole-catalogue scoring produces useless averages.
- Pull four extracts: the item export, the raw supplier data, the published output and the schema.
- Define the standard per category and get it signed before measuring anything against it.
- Automated rules test syntax only. Hand sample around 100 SKUs per category for accuracy.
- Map every finding to a mechanism. Forty findings usually collapse into four or five causes.
- Deliver a costed, sequenced plan with three scope options, not a dashboard.
We run this on live catalogues, and we have published the method rather than a capability statement. A method you can copy is better proof. If you would rather not run it yourself, a thirty minute call is enough to scope it. Book a call, or read more about how the remediation work that follows is delivered in our product data services.