Skip to content
Home » Insight » Supplier Onboarding: A Product Data Process That Scales

Supplier Onboarding: A Product Data Process That Scales

Search for supplier onboarding and almost everything you find is about procurement. Vendor risk, bank details, compliance questionnaires, master data in the ERP. All necessary, none of it the reason your new range is not on the website. The part that takes the time is product data. Getting a specification, a description, images and a hundred structured attributes out of a supplier, in a state you can sell from. That is a different process with different owners, and most distributors run it as a series of one-offs. Below is the supplier onboarding stage model we use to make it repeatable.

Why supplier onboarding is really two processes

Procurement onboarding ends when the supplier exists in your ERP and you can raise a purchase order. Product data onboarding is only just starting at that point, and it runs on a different clock.

The two also fail differently. Procurement onboarding fails loudly, because you cannot pay someone. Product data onboarding fails quietly. The line gets created, the stock arrives, and the product sits on the site with a supplier part number for a title and no attributes. Nobody raises a ticket, because nothing is broken. It just does not sell.

The other structural difference is who holds the data. Your supplier’s finance team can answer every procurement question in an afternoon. The product data lives with their marketing team, their technical author, their ERP, three PDFs and a website. Often nobody at the supplier has ever assembled it in one place. You are not requesting a file, you are asking them to do work they have never done.

So a repeatable supplier data onboarding process gets designed around what suppliers can actually produce. Not around what you would like to receive.

Stage 1: Define what onboarded means

Almost nobody does this, and it is the stage that determines whether the other seven work.

Write down, per category, what a finished product record looks like. Which attributes are mandatory. How many images and at what specification. Whether a datasheet is required. What the description has to contain. Which identifiers must be present and which are optional.

Inputs: your attribute schema per taxonomy node, your channel requirements, your regulatory obligations for the category.

Outputs: a definition of done per category group, in writing, agreed by ecommerce, merchandising and the category buyer.

The common mistake: defining it once for the whole catalogue. A luminaire and a box of screws do not have the same definition of done. Pretending otherwise gives you a standard that is impossible for one and meaningless for the other. This is where attribute scoping from taxonomy and attribution earns its return. The definition of done is the mandatory set for a node.

Stage 2: Tier your suppliers by data capability

Suppliers get segmented by spend everywhere. Almost nobody segments them by what data they can send, and that is the variable that drives supplier onboarding effort.

Four tiers cover most distributor bases.

Tier one, structured publishers. They can send BMEcat, ETIM classified data, a GDSN entry or a clean API feed. Usually larger manufacturers in electrical, plumbing or FMCG. Your job is integration, not collection.

Tier two, spreadsheet capable. They can complete a well-built template accurately if you scope it properly. This is the largest tier for most distributors.

Tier three, document only. They have datasheets, catalogues and a website, and nothing structured behind them. You extract rather than collect.

Tier four, nothing usable. Small manufacturers, importers, own-brand factories. The data will be created by you or by an agency, from samples and photographs.

Inputs: supplier list, a sample of what each has previously sent, spend and line count.

Outputs: a tier per supplier, and a collection route per tier.

The common mistake: treating tier one suppliers like tier two. Sending a spreadsheet template to a manufacturer who publishes ETIM is the single most common waste in this whole process. They hand-key it from the same system that could have exported it. It comes back worse than the feed would have been.

Stage 3: Publish the specification, not just the form

The artefact you send is downstream of the specification. Write the specification first and generate the collection artefact from it.

For each category group that means field names, types, units, value lists, mandatory flags and an example of a good value. Not a blank template with a tab of instructions nobody reads. An example row filled in correctly does more work than two pages of guidance.

Inputs: the definition of done from stage one, the tier from stage two.

Outputs: a category-scoped template for tier two. A mapping specification for tier one, an extraction brief for tier three, a creation brief for tier four.

The common mistake: one form for everything. We wrote a whole piece on why the general purpose new line form fails, and every failure in it starts here.

Stage 4: Collect through the right channel

Now you actually ask for the data, and the channel follows the tier rather than the supplier’s size.

Tier one goes to integration. Map their feed once, then re-run it. The cost is front-loaded and then close to zero per line. That is why it is worth doing even at modest line counts.

Tier two gets the validated template. A named contact, a deadline tied to the launch date, and a chase schedule that exists before it is needed. Two chases, then escalation to the buyer. Not “we will follow up”.

Tier three gets extracted. Datasheets and catalogues go through a structured extraction process against your attribute list, then back to the supplier for confirmation rather than for completion. Asking someone to check twenty values is a far smaller ask than asking them to supply them.

Tier four gets created, and it needs to be priced into the range decision rather than absorbed. A line that requires photography and specification writing from scratch costs real money before it sells anything.

The common mistake: running every supplier through the template because it is the route you already have. That is how tier three suppliers end up returning empty forms and tier one suppliers end up hand-keying.

Stage 5: Validate at the gate

Everything arriving hits a gate. This is the stage that decides whether supplier onboarding is a process or a habit, and it needs a written test with a binary outcome.

The test has four layers, in this order. Structural: is the file the right shape, are the identifiers present and unique. Completeness: are the mandatory attributes for this node populated. Validity: do the values match the type, the unit and the value list. Plausibility: is the weight in kilogrammes, is the length a length, is the price within an order of magnitude of the rest of the node.

Plausibility is the layer people skip, and it catches the errors that are most expensive later. A pallet weight in the item weight field passes every other check.

Outputs: an accepted set, and a rejected set. Rejections go back with the failing rows and the reason, not a general request to try again.

The common mistake: a gate that never rejects. If nothing has failed the gate in six months, you do not have a gate. You have a receipt. Exceptions are fine when they have a named owner and a date attached, which is how you stop a temporary pass becoming permanent.

Stage 6: Map to your model

Accepted supplier data is still the supplier’s data. This stage makes it yours.

Three jobs happen here. Taxonomy mapping puts each item in your node, not theirs. Attribute mapping aligns their field names to yours, which is where a mapping table per supplier pays for itself the second time they send anything. Value normalisation converts their vocabulary to your value lists, units included. “Stainless steel A2”, “SS A2” and “A2 S/S” all land on one value.

Keep the mapping tables. That is the whole point. The first onboarding for a supplier is expensive and every one after it is not. That only holds if the mapping is a stored artefact, not a decision someone made in a spreadsheet.

The common mistake: normalising by hand each time, with no stored mapping. It feels faster on line one and costs you on every line after.

Stage 7: Enrich the gaps

There will be gaps. A supplier who supplies ninety per cent of what you asked for has done well, and the rest still needs filling before the product sells.

Sort the gaps by whether they can be derived, extracted or only requested. Derived values come from other attributes you already hold. Extracted values sit in a datasheet you already have. Requested values are the ones that genuinely require the supplier, and that list should be short by the time you get here.

Marketing content is usually the biggest gap, and it is also the one you can absorb yourself. Descriptions, feature bullets and search terms are your voice anyway, and waiting for a supplier to write them in your tone is a poor trade. Our product content enrichment work is mostly this stage. It runs behind a pipeline that has already done stages one to six.

The common mistake: holding the whole line for the content gap. Publish with the technical data complete and the description generated, then improve it. A product with a full specification and a plain description sells. A product held back for six weeks sells nothing.

Stage 8: Publish, then measure the right things

Publication is per channel, not global. Your website, your marketplace listings, your printed catalogue and your customer punchout each have a different minimum. Hold the readiness state per channel and publish where the item is ready rather than waiting for the strictest one.

Then measure four things, and only four. These are the supplier onboarding numbers worth putting in front of a board.

Cycle time, from range approval to first sellable record, split by tier. This is the number the commercial team cares about. First-pass rate at the gate, by supplier. It tells you which specifications are unclear and which suppliers need a different tier. Completeness against the definition of done, by category, not overall. An overall completeness percentage hides everything useful. Post-publication correction rate. Corrections made in the first ninety days after launch are the true measure of whether the gate is working.

The common mistake: measuring attribute fill rate across the whole catalogue and reporting it monthly. It moves slowly, it flatters, and it tells nobody what to do next.

The supplier onboarding timeline in practice

The shape matters more than the durations, and the durations vary enormously by tier.

Stages one and two are set up once and reviewed annually. They are a project, not a process, and they typically run over a few weeks with the category and ecommerce teams.

Stage three is per category group, done once and updated when the schema changes. Stages four to eight are the running process, repeated per supplier and then per range.

The first supplier through a new pipeline is slow, because you are still writing the mapping. The tenth is fast. That crossover is the entire argument for a repeatable process. The honest measure is not the first cycle time but the fifth. A sophisticated supplier with a structured feed takes two to four weeks on the first pass. A datasheet-only supplier takes four to eight weeks. A bespoke supplier with unstructured sources takes eight to fourteen. By the fifth supplier in a tier, the top of each range roughly halves.

Where supplier onboarding gets stuck

Three failure patterns account for most of it. We see all three across the industrial distributor catalogues we work on.

The first is starting at stage four. Somebody buys a portal or builds a template, sends it out, and discovers that nobody agreed what onboarded means. The collection tool is fine. It has been loaded with an undefined specification.

The second is treating onboarding as a project rather than an operation. A catalogue clean-up runs, the backlog clears, and eighteen months later the same backlog exists because nothing changed about how new suppliers arrive.

The third is ownership. Supplier onboarding for product data sits between buying, merchandising and ecommerce, and in many distributors it belongs to none of them. Whoever is most annoyed by the gap does the work. That does not scale, and it is usually the real reason the process gets rebuilt from scratch every time.

Key takeaways

  • Product data supplier onboarding is a separate process from procurement onboarding, with different owners and a different clock.
  • Define what onboarded means per category before you build any form or buy any tool.
  • Tier suppliers by data capability, not spend, and give each tier its own collection route.
  • The gate needs a written test with a binary outcome. A gate that never rejects is a receipt.
  • Store the mapping tables. First onboarding is expensive, every one after it should not be.
  • Measure cycle time, first-pass rate, completeness per category and post-publication corrections. Nothing else.

If supplier onboarding runs differently every time and nobody can say how long it takes, that is a process problem, not a tooling problem. We map the current process against this model, tier your supplier base and hand back the sequence to fix it in. Thirty minutes on a call is usually enough to see whether it is worth doing. Get in touch, or read how we run supplier data onboarding and what that looks like for B2B distributors.