← Back to blog

The Vendor Scorecard Metrics Retailers Actually Track

vendor scorecard metricssupplier scorecardOTIFfill rateretailer compliancechargebacksSQEP

The UNFI line review opens with a slide the brand has never seen before: its own scorecard. Fill rate 93%, under the 95% bar for most of the year, service-level fines accruing, and a corrective-action flag on two SKUs. The fines on the slide total $9,450. The two flagged SKUs carry $180,000 of annual distribution.

That meeting happens, with those numbers, to Cinderhaven Provisions, a fictional $25M specialty food brand whose figures are a synthetic dataset built to illustrate a scorecard structure that is entirely real. The structure has one property worth the whole article: the fines scale smoothly, and the consequences arrive as cliffs.

The scorecard that matters is the one kept on you

Search this keyword and the pages that rank teach the opposite lesson: how to build a weighted scorecard for your own suppliers, with a 1-to-5 scale and a quarterly review. Useful for grading a co-packer. Useless for the score that moves revenue, because every retailer and distributor a food brand ships to already runs a scorecard on the brand, and it is not on a 1-to-5 scale. It is denominated in percentages and deducted in dollars.

The metric set is consistent across programs: on-time percentage, in-full or fill-rate percentage, ASN accuracy, lead-time reliability, and, increasingly, data quality itself. Walmart prices that last category by the defect. Under its Supplier Quality Excellence Program, an ASN that never downloads bills at $25 per purchase order for a non-DSDC supplier, and any other PO-accuracy defect at $200 plus $1 per impacted case. One unreadable label template, repeated across a year of shipments, runs that schedule to $68,000. A misdeclared shipment is scored before anyone opens a carton, which is one reason OTIF fines are a data problem before they are a logistics problem.

Your procurement scorecard is optional. Theirs arrives as deductions.

Five programs, and no two measure the same thing

The bars and the fees are published, mostly through supplier-education channels rather than the retailers' own marketing:

| Retailer | The bar | The fine | |---|---|---| | Walmart | 98% on time for collect, 90% for prepaid, 95% in full for both | 3% of cost of goods on every case that missed; early counts as late | | Target | 95% fill rate, scored in Greenfield's supplier performance dashboard | 5% of the cost of goods not received, $150 minimum per violation | | UNFI | 95% fill rate per PO | 3% of shorted value after two consecutive weeks under the bar | | KeHE | 98% inbound fill, 92% on time | 3% of shorted value, plus on-time fines billed quarterly | | Kroger | Thresholds vary by category and distribution center, and are not published | $500 a day past a two-day delivery window, with most other penalties charged as flat dollars per incident |

Read down the middle column and the bars scatter: 98, 95, 92, 90, each measured at its own level, and one program that will not say. Read down the right column and the fees scatter too. Walmart, UNFI, and KeHE take 3%, Target takes 5% and applies it to cost of goods rather than to shorted billing value, and Kroger does not use a percentage at all.

That non-comparability is the finding. A brand cannot rank its accounts by scorecard pressure, because no two programs measure the same event in the same unit: per case at Walmart, per PO-location at Target, per PO at UNFI, per shipment day at Kroger. Managing all five to one internal service-level number is managing to a number none of them keeps.

Five bars, five units, five fee schedules, and one consequence behind all of them.

The fine is the small number on the scorecard

Price Cinderhaven's bad year at UNFI, where 18% of its volume runs: $4.5M in annual billing. Fill rate stuck at 93%, so 7% of ordered value ships short. The service-level fine is 3% of the shortage: $4.5M × 7% × 3% = $9,450. A CFO scanning the deduction ledger sees a nuisance line, smaller than the freight surcharges. No fee schedule in the table above changes that verdict; the fine is a rounding error on a $4.5M account under any of them.

The scorecard's other output is the flag. UNFI's published escalation for persistent shortfalls is a corrective-action plan, then delisting. Cinderhaven carries all fifty of its SKUs at UNFI, so two flagged items at roughly average velocity represent $4.5M ÷ 50 × 2 = $180,000 of annual distribution. The fine and the flag come off the same 93%, and the flag costs nineteen times the fine. The gap is wider than the slide shows, because the two sides are not even measuring the same fill rate. In the case behind a brand reading 99% against a retailer reading 84%, the fifteen-point spread was entirely definitional. The scorecard runs on their measurement, not yours.

This is the shape short-ship costs take everywhere: the number on the fill-rate report is under 60% of the real cost, and the rest sits in three systems nobody joins. Lailara's short-ship model prices all four layers from a brand's own order data.

The fine is a fee. The cliff is the business.

Their scorecard is reproducible from your own files

Nothing on the retailer's slide is secret, and none of it is theirs alone. On-time and in-full come from POs, ASNs, and receiving data the brand already holds. The scored defects arrive itemized as deduction codes on every remittance. Store-level velocity, the number behind delist flags, arrives weekly in the 852 feed. A brand that assembles those files into one table per retailer, each account scored on that account's own bar rather than a house average, is reading the same scorecard the buyer will present, a quarter earlier and without the meeting.

Few do. The retailer grades weekly. Most brands check the grade at contract renewal.

Your two biggest accounts, scored on their own bars

Take your two largest accounts. A quarter of their purchase orders, ASNs, and remittance detail is enough for me to rebuild the scorecards those retailers are keeping on you, bar by bar, each on its own measurement rather than your house average. Two accounts, one quarter. A flag read early is a corrective-action plan; read late it is a line review.