Partner Selection

Distributor Performance Evaluation: Most Scorecards Grade the Wrong Thing

md.kim 2026. 9. 19. 23:47

Most distributor scorecards are borrowed from procurement. A partner can score full marks on every line and still be selling nothing.

The scorecard came back green. On-time ordering, payment behaviour, forecast submission, responsiveness, documentation, complaint handling — all at or above target, four quarters running.

The business was not moving.

Nothing on that sheet was wrong. Every number was real and every number had been collected honestly. The sheet simply could not see the thing I needed to know, because it had been designed to evaluate a supplier, and I was evaluating a distributor. Those are not the same relationship, and almost every template on the internet treats them as if they were.

Buying from someone and selling through someone are different problems

Search for a distributor performance scorecard and what you will actually find is a supplier scorecard — procurement software vendors and consultancies offering weighted KPI templates. They are competent documents. They are also built for the opposite direction of trade.

When you buy from a supplier, you know what you want and you can inspect whether you got it. Delivery is observable. Quality is testable. The metrics work because the outcome is in your hands.

When you sell through a distributor, you are buying an outcome you cannot observe — demand created in a market you are not standing in. Everything the procurement sheet measures happens on the way to the partner. The thing you care about happens after.

Buying from a supplierSelling through a distributoryou specify what you wantyou specify an outcome you cannot seedelivery is observabledemand creation is not observablequality can be inspectedeffort cannot be inspectedthe outcome lands in your handsthe outcome lands in their marketcompliance is close to performancecompliance says nothing about performanceThe template you downloaded was written for the left-hand column.It measures the journey to your partner. The business happens after that point.
Two relationships, one template. Procurement scorecards are good instruments pointed in the wrong direction. Nothing on them is false; everything on them is upstream of the question.

Every line can be green while nothing sells

This is the specific failure, and it is worth stating plainly because it is so easy to miss in a review: a distributor can satisfy every procurement metric without a single unit leaving a shelf.

They order on time — into their own warehouse. They pay on time — for stock that is not moving. They submit the forecast on the date it is due — the same forecast as last quarter. They respond within twenty-four hours, file the documents, and handle the two complaints that arrived. Green, green, green.

I wrote about the underlying arithmetic in The Month-Seven Test: your P&L is built on sell-in, the business is built on sell-out, and in cross-border distribution the gap between them can run most of a year. A procurement scorecard lives entirely on the sell-in side of that gap.

Four quarters of a distributor scorecardon-time ordering96%payment behaviour100%forecast submitted100%responsiveness92%documentation98%complaint handling94%units that left a shelf11%Everything above the line is a promise kept. The line below it is the business.
The compliance ceiling. Six metrics, four quarters, no complaint worth raising. None of them are downstream of the moment a customer decides to buy.

Kaplan and Norton made the general version of this argument when they introduced the balanced scorecard in 1992: financial and backward-looking measures alone give misleading signals, because measurement systems shape behaviour. A partner reviewed on ordering discipline becomes excellent at ordering discipline. That is not cynicism on their part. It is the system working exactly as designed.

The signals that actually arrive first are behavioural

Over enough appointments you notice that the numeric signals are always last. By the time sell-through misses, the decision was made three quarters earlier by people who stopped attending. The things that moved first were never on a scorecard, because they are awkward to quantify — which is a poor reason to ignore them.

Five signals, roughly in the order they appear
1. Seniority at the review drops. The person who signed the contract sends a deputy, then the deputy sends a coordinator. Nobody announces this. It is the earliest reliable signal there is, and it costs nothing to track — just write down who was in the room.
2. Bad news stops arriving before you ask for it. A partner who is still invested tells you about the delisting before the report does. When you start learning things only by asking, the relationship has already changed.
3. The forecast stops changing. A forecast that is identical three quarters running is not a forecast. It is a form being filled in. Movement — in either direction — is evidence that someone is looking at the market.
4. The lag between your question and their number gets longer. Not because they are hiding anything. Because the data is no longer at anyone's fingertips, which tells you who is looking at it.
5. They stop asking about next year. Questions about upcoming launches, next season's range, what is coming from the factory — these come from a partner planning to still be selling. Their absence is the quietest signal and usually the most advanced.
month 0month 3month 6month 9month 12seniority at the review dropsbad news stops arriving earlythe forecast stops changingthe reorder conversation slipssell-through misses targetbehavioural — free to collect, nobody records themnumeric — on every scorecard, arrives last
The cheap signals are the early ones. Who came to the meeting costs nothing to record and moves two to three quarters before the number does. Most review packs contain the opposite selection.

A scorecard is not a measurement device. It is a meeting agenda.

Here is the part that changed how I build these.

Whatever is on the sheet is what the quarterly review is about. The sheet does not describe the conversation — it causes it. Put six compliance metrics on it and you have designed a compliance audit: forty minutes of explaining a 92%, and no time left for the market. The partner prepares for the meeting you built.

Which means the real question when you add a line is not "can we measure this reliably." It is "do I want to spend twenty minutes talking about this every quarter."

The same partner, two scorecards, two meetingscompliance scorecardon-time ordering · paymentforecast filed · responsivenessthe meeting you getexplaining a 92%defending last quarteragreeing to file things soonernobody mentions the marketoutcome scorecardsell-through · door countwho attended · forecast movementthe meeting you getwhere the stock actually iswhich doors stopped reorderingwhat changed their forecastthe partner prepares for this oneYou are not choosing metrics. You are choosing what forty minutes a quarter is spent on.
What you score is what you discuss. Both sheets are honest. Only one of them produces a conversation that could change the outcome.

What to actually put on it

Short, and weighted by how early the signal arrives rather than how cleanly it can be collected.

Cumulative sell-through, against the month it is in. Not a flat target — the curve from the month-seven piece. One number, and the only one that is genuinely about the business.
Doors or accounts that reordered, as a proportion of doors that ever ordered. Catches decay that total volume hides when one large account is masking twenty small ones going quiet.
Forecast movement. Not accuracy — movement. Record whether the number changed and what changed it. A flat line is the finding.
Who attended the last two reviews, by title. Free to collect. Earliest signal you will get. Nobody records it.
One open line: what did they tell you before you asked? Written in words, not scored. If it is blank two quarters running, that is worth more than any percentage on the sheet.

Keep the compliance metrics if you like — they are cheap and occasionally a payment problem really is the story. But put them at the bottom, unweighted, where they belong: as hygiene, not as the answer.

And one caution, because this list can fail the same way the last one did. The moment any of these becomes a target, it starts being managed rather than measured — Goodhart's law, which does not spare a scorecard just because the metrics are better chosen. The defence is not a cleverer metric. It is that the fifth line stays open text, and that somebody reads it.

The five-line version

IF YOU READ NOTHING ELSE
01  Most distributor scorecards are supplier scorecards — built for buying from someone, not selling through them.
02  A partner can score full marks on every compliance line while nothing leaves a shelf.
03  The early signals are behavioural: who attends, whether bad news arrives unasked, whether the forecast ever moves.
04  A scorecard does not describe the quarterly meeting. It causes it. Choose lines you want to spend twenty minutes on.
05  Keep one line as open text, or the good metrics become targets too.

Your turn

Take your current scorecard. How many lines are downstream of a customer deciding to buy? If the answer is zero or one, you have a supplier sheet.
For your largest partner: who attended the last two reviews, by title? If you cannot answer, start recording it this quarter — it costs nothing.
Pull the last three forecasts they sent. Did the numbers change? If not, you have been receiving a form, not a forecast.
Think back to your last quarterly review. What did the forty minutes get spent on — and was that the scorecard's doing?

Sources

  • Robert S. Kaplan & David P. Norton, "The Balanced Scorecard—Measures that Drive Performance", Harvard Business Review, January–February 1992 — for the argument that backward-looking measures give misleading signals, and that measurement systems shape behaviour.
  • Goodhart's law — Charles Goodhart, "Problems of Monetary Management: The U.K. Experience", 1975. The widely quoted formulation is Marilyn Strathern, "Improving Ratings: Audit in the British University System", European Review 5(3), 1997.
  • The five behavioural signals are drawn from practice rather than literature, and are offered as such. They are cheap to record and easy to disprove on your own data, which is the only test that matters here.

Cases are drawn from real appointments with countries, companies, brands, individuals and proprietary figures removed; only the structure remains. The scorecard figures in the charts are illustrative. This is not legal advice.

Written from the brand side of this table — twenty-two years appointing and managing distribution partners across Asian markets, across categories and price tiers.

Previously in this series: How to Evaluate a Distributor Before You Sign · Exclusive or Not · When to Stop Selling Cross-Border and Set Up Locally · What to Do When Your Distributor Asks for a Lower Price · When Is a Market Ready for a Local Launch? · Marketing Money in a Distribution Agreement · When Your Distributor Stops Paying · The Month-Seven Test · How to Manage the Distributors You Never Visit · Parallel Imports Are a Price Gap, Not a Piracy Problem · Stop Pricing One Export Market at a Time · Your Marketplace Store Is Four Rights, Not One Decision