Supplier performance scorecards: the metrics, the weights, and the arithmetic
Seven metrics, the default weights, target and floor bands, and one supplier scored end to end. No gated download — the tables on this page are the template. Plus the two parts every template skips: getting the data out of three systems, and running the review without it turning into a spreadsheet argument.
On this page
What a supplier scorecard is, and the seven metrics that matter
A supplier scorecard is a weighted score, normally out of 100, that grades one supplier on delivery, quality, and responsiveness over a fixed period. Seven metrics carry almost all the signal: on-time delivery (25%), in-full line fill (20%), lead time variance (15%), ASN accuracy (10%), document accuracy (10%), responsiveness (10%), and defect rate (10%). Everything else is a subset of those or a vanity column.
Most scorecards fail for one of two reasons. Either they carry twenty metrics, so no supplier knows what to fix, or they carry two — on-time and quality — so the metrics that actually drive your cost stay invisible. Seven is the number that survives a real quarterly review: enough to be fair, few enough that a supplier's operations manager can hold the whole thing in their head and act on it.
Note what is deliberately absent. Price is not on the scorecard. Price is negotiated separately, and mixing it in lets a cheap supplier buy their way out of a delivery problem. Total cost of ownership belongs in the sourcing decision, not the performance score. Sustainability and diversity metrics, if you track them, belong on a separate compliance page, because they move on a different timescale and are not something a plant manager can fix inside a quarter.
Measure everything at PO line level, not order level. An order with nine lines received complete and one line short is a 90% line fill and a 0% order fill. Line level is what your production schedule and your customer promise actually feel, and it is the level at which a supplier can act on the feedback.
Scorecard total = Σ (metric points × metric weight). Seven metrics, 100 points, one number per supplier per period.
| Metric | What it measures | Formula | Default weight |
|---|---|---|---|
| On-time delivery | PO lines received inside the agreed window against the first confirmed date | (lines on time ÷ total lines) × 100 | 25% |
| In-full (line fill) | Lines received complete on the first receipt, no back-order tail | (lines complete on first receipt ÷ total lines) × 100 | 20% |
| Lead time variance | Consistency of actual lead time, not its length | standard deviation of (receipt date − PO release date), in days | 15% |
| ASN accuracy | Whether the advance ship notice matched what actually arrived | (ASNs matching receipt on SKU, quantity and carton count ÷ ASNs sent) × 100 | 10% |
| Document accuracy | Complete, correct paperwork on first submission | (shipments with all required docs correct first time ÷ shipments) × 100 | 10% |
| Responsiveness | Time to a substantive answer on a PO change or an issue | median hours to acknowledge, plus % of issues closed inside the agreed SLA | 10% |
| Defect rate | Units rejected at receipt, in process, or at your customer | (rejected units ÷ units received) × 1,000,000, expressed as PPM | 10% |
Who owns the scorecard
Settle one thing before you settle the metrics: who owns the scorecard. Procurement owns the supplier relationship and the commercial leverage. Supply chain owns the consequence when the line lands late. Quality owns the defects. Receiving owns the timestamps that decide the on-time number, and usually doesn't know it. If that is unresolved before the first scored period, the scorecard becomes an interdepartmental argument conducted in front of a supplier, which is the worst possible venue for it. The arrangement that works at most companies: procurement owns the conversation and the consequence, supply chain owns the definitions and the data, and the two publish one number jointly. Whoever else builds it will eventually stop, because assembling it every month is nobody's actual job.
How to score each metric: targets, floors, and the points formula
Raw percentages are not comparable across metrics. 96% on-time and 96% document accuracy are not the same quality of performance, because the useful range differs. Nobody runs 60% document accuracy and stays approved, but 88% on-time is depressingly common. So convert every metric to points on a shared 0 to 100 scale using two numbers: a target that earns full marks, and a floor that earns zero.
For metrics where higher is better, points are the distance from the floor as a share of the floor-to-target range. For inverted metrics where lower is better — lead time variance in days, response time in hours, defects in PPM — swap the terms.
points = clamp( (actual − floor) ÷ (target − floor) × 100, 0, 100 )
points = clamp( (floor − actual) ÷ (floor − target) × 100, 0, 100 )
Set the floor at roughly the level where you would open a resourcing conversation, and the target at the level a genuinely good supplier in that category hits.
Two calibration checks. If more than a third of your suppliers score full marks on a metric, the target is too soft and the metric has stopped discriminating. If nobody clears half marks, the floor is fantasy and you will spend the review defending the scorecard instead of discussing performance.
Avoid step bands of the sort where 95.0 to 97.9 percent scores 85 points. Bands create cliff effects that suppliers optimize against: a supplier at 94.9% has every incentive to reach exactly 95.0% and none to reach 96.5%, and the review turns into an argument about two late lines rather than a discussion of the pattern. A linear ramp between floor and target means every real improvement scores something.
One category caveat on defects. Automotive tier-one programs commonly run targets under 50 PPM with best-in-class under 10. General industrial fabrication and non-critical commodity parts sit in the hundreds to low thousands. Set the floor from your own category and your own rejection history, not from a benchmark slide.
| Metric | Full marks (target) | Zero (floor) | Direction | Weight |
|---|---|---|---|---|
| On-time delivery | 98% | 85% | Higher is better | 25% |
| In-full, first receipt | 99% | 90% | Higher is better | 20% |
| Lead time variance (σ) | 1.0 day | 5.0 days | Lower is better | 15% |
| ASN accuracy | 98% | 85% | Higher is better | 10% |
| Document accuracy | 99% | 90% | Higher is better | 10% |
| Responsiveness (median acknowledgement) | 4 hours | 48 hours | Lower is better | 10% |
| Defect rate | 100 PPM | 2,000 PPM | Lower is better | 10% |
Worked example: scoring one supplier for one quarter
Halvorsen Extrusions supplies aluminum profiles on a 20-day quoted lead time. Over Q3 you released 412 PO lines to them. Here is the quarter, scored with the default weights and bands above.
Column width = weight · Filled height = points · Filled area = weighted contribution
Halvorsen Q3: 54.8 of 100 — 0.2 points below the Conditional line. Scroll to see all seven metrics →
54.8
of 100, weighted
38.5
points from a 90% on-time headline
8.8
of a possible 25 from ASN accuracy + lead time variance
Read the result carefully, because this is where most scorecards get argued with. The headline that procurement would normally report is “90% on-time, 97% complete”, which sounds survivable. Converted to points, 90% on-time is 38 of 100, because the target is 98% and the floor is 85%. That is the whole purpose of the target and floor: they anchor the score to what good actually looks like rather than to how the number feels.
Now look at the two lowest weighted contributions: ASN accuracy at 2.8 and lead time variance at 6.0. Those are almost always the two metrics nobody was tracking before the scorecard existed, and they are the two doing the most damage downstream. A wrong ASN forces receiving to do a blind count, which typically takes three to five times as long per pallet depending on pack configuration and carton labelling. A lead time standard deviation of 3.4 days on a 20-day lead time is what is driving your safety stock, which is the next section but one.
At 54.8 the supplier lands in At risk — 0.2 points below the Conditional line, which is two late lines short of a corrective action plan and into active resourcing. That is uncomfortable, and it is also the argument for the linear ramp: the score is close enough to the boundary that the review will want to argue about rounding, and the only defense is that every metric moved continuously and nothing was banded. The score did not tell you what to do. It told you where to look, and the two lowest weighted contributions told you what to put on the agenda.
Halvorsen scores 54.8 of 100. A 90% on-time headline converts to 38 of 100 points, and the two metrics nobody was tracking — ASN accuracy and lead time variance — contribute just 8.8 of a possible 25.
| Metric | Q3 result | Points calculation | Points | Weight | Weighted |
|---|---|---|---|---|---|
| On-time delivery | 371 of 412 lines = 90.0% | (90.0 − 85) ÷ (98 − 85) | 38.5 | 25% | 9.6 |
| In-full, first receipt | 398 of 412 lines = 96.6% | (96.6 − 90) ÷ (99 − 90) | 73.3 | 20% | 14.7 |
| Lead time variance | σ = 3.4 days | (5.0 − 3.4) ÷ (5.0 − 1.0) | 40.0 | 15% | 6.0 |
| ASN accuracy | 88.6% | (88.6 − 85) ÷ (98 − 85) | 27.7 | 10% | 2.8 |
| Document accuracy | 97.1% | (97.1 − 90) ÷ (99 − 90) | 78.9 | 10% | 7.9 |
| Responsiveness | median 19 hours | (48 − 19) ÷ (48 − 4) | 65.9 | 10% | 6.6 |
| Defect rate | 640 PPM | (2,000 − 640) ÷ (2,000 − 100) | 71.6 | 10% | 7.2 |
| Total | 100% | 54.8 |
Nobody at this supplier had a disastrous quarter — 90% on time, 96.6% in full, no crisis anyone escalated — and the weighted arithmetic still lands two-tenths of a point on the wrong side of 55.
Adjusting the weights for your operation
The default weighting is built for a manufacturer or distributor with a mix of domestic and imported inbound. It is a reasonable starting point and a bad finishing point. The weights should reflect what a failure on that metric actually costs you, and that is entirely a function of your operation.
If you run sequenced or just-in-time assembly, lead time variance and on-time dominate, because a two-day swing consumes buffer you do not carry. Line stoppage cost is the reason. Published figures for automotive line-down run into the tens of thousands of dollars per minute at OEM final assembly, and typically the hundreds to low thousands at a tier-one plant, driven by takt time, unit contribution margin, and whether the shift can be recovered on overtime. Use your own number — your finance team already has it, and the point is only that whatever it is, it makes variance expensive.
If you supply a graded retailer, in-full weighs as heavily as on-time, because you are being scored on both and penalised for either. Retailer penalty programs commonly run 1 to 3 percent of order value per non-compliant order, and a short-shipped line from your supplier becomes a short-shipped line to your customer. That is the same arithmetic covered in the OTIF guide, one tier upstream.
If you operate under regulatory scope — pharma, defense, aerospace, food — document accuracy and defects carry the weight, because a missing certificate of analysis or certificate of origin does not slow a shipment down, it stops it. A customs hold on incorrect paperwork typically costs two to five days plus per-diem charges, and demurrage and detention at North American ports commonly runs $150 to $400 per container per day after free time, escalating in tiers.
And if you are the 3PL or contract manufacturer rather than the shipper, run two scorecards from the same record: one on the vendors you buy from, and one per client on the inbound they direct to you. The second is the one that ends the argument about whose miss it was, because a client-nominated supplier landing late is a fact you can produce at line level instead of a claim you have to defend.
Pick the profile closest to your operation and adjust from there. The only hard rule: the weights must sum to 100, and you must publish them to your suppliers before the period they apply to.
Default — mixed inbound manufacturer
Baseline
- On-time
- 25
- In-full
- 20
- LT variance
- 15
- ASN
- 10
- Docs
- 10
- Responsiveness
- 10
- Defects
- 10
JIT / sequenced assembly
- On-time
- 30
- In-full
- 15
- LT variance
- 25
- ASN
- 10
- Docs
- 5
- Responsiveness
- 5
- Defects
- 10
Retail or CPG supplying a graded customer
- On-time
- 30
- In-full
- 30
- LT variance
- 10
- ASN
- 10
- Docs
- 5
- Responsiveness
- 5
- Defects
- 10
Regulated — pharma, defense, aerospace
- On-time
- 20
- In-full
- 15
- LT variance
- 10
- ASN
- 5
- Docs
- 25
- Responsiveness
- 5
- Defects
- 20
Long-lead import program
- On-time
- 20
- In-full
- 15
- LT variance
- 30
- ASN
- 10
- Docs
- 15
- Responsiveness
- 5
- Defects
- 5
MRO and spares
- On-time
- 25
- In-full
- 20
- LT variance
- 10
- ASN
- 5
- Docs
- 5
- Responsiveness
- 25
- Defects
- 10
Every profile sums to 100. Publish yours to your suppliers before the period it applies to.
Lead time variance is the metric you are not scoring, and it costs the most
Almost every supplier scorecard in circulation grades on-time delivery and ignores lead time variance. That is backwards, because variance, not length, is what your inventory pays for. The safety stock formula makes this unarguable.
Safety stock = Z × √(LT × σD2 + D2 × σLT2)
Z = service factor (1.65 at 95%) · LT = mean lead time in days · σ_D = daily demand σ · σ_LT = lead time σ
Notice that σ_LT is multiplied by average demand D, and both are squared. On any high-volume part, that second term swamps the first.
Run it on Halvorsen. Take one representative SKU: 60 units per day average demand, demand standard deviation of 18 units, 20-day average lead time, and the measured lead time standard deviation of 3.4 days. Unit cost is $85.
Cut the variance
−163
units of safety stock · σ_LT 3.4 → 1.5 days
$13,855 working capital released on one SKU · ~$3,050/yr carrying at 22%
Cut the lead time
−6
units of safety stock · quoted lead time 25 → 20 days
Same variance, five days of negotiation, almost no buffer released
At σ_LT = 3.4 days, safety stock is 1.65 × √(20 × 18² + 60² × 3.4²) = 1.65 × √(6,480 + 41,616) = 362 units. Get the supplier to σ_LT = 1.5 days, changing nothing else, and safety stock falls to 1.65 × √(6,480 + 8,100) = 199 units. That is 163 units, or 45%, off one SKU. At $85 per unit that is $13,855 of working capital released, and roughly $3,050 a year of carrying cost at a 22% carrying rate. Carrying rates realistically run 18 to 30 percent depending on capital cost, storage, obsolescence risk, and insurance. On a 40-SKU program with a similar demand profile, that is about $554,000 of working capital and $122,000 a year.
Now the comparison that should change how you negotiate. Suppose the quoted lead time were 25 days instead of 20, variance unchanged at σ_LT = 3.4. Safety stock would be 1.65 × √(25 × 18² + 60² × 3.4²) = 1.65 × √(8,100 + 41,616) = 368 units. Win those five days back — 25 down to 20, variance untouched — and you land at 362. Six units. Five days off the quoted lead time bought you six units of buffer. Making the same lead time predictable bought you 163.
That is why lead time variance earns 15% of the score, and why the corrective action you ask for is “tell me your real lead time and hit it” rather than “give me a shorter lead time”. A supplier who quotes 26 days and delivers on day 26 every time is worth substantially more to you than one who quotes 20 and lands anywhere between 17 and 27.
Safety stock = Z × √(LT × σ_D² + D² × σ_LT²). On the worked SKU: five days off the quoted lead time released 6 units of buffer. Cutting lead time variance from 3.4 days to 1.5 released 163.
The number has to survive being checked by the supplier it grades
Every supplier who receives a score does the same thing with it first. They reconcile it against their own shipping records, line by line, and they open the review with the handful of lines where you are wrong. A handful of bad lines out of 412 barely moves the score. It still ends the program, because from that point on every number on the page is negotiable.
So the bar is not accurate enough to manage by. The bar is whether the supplier can rebuild your number from documents they already hold — the purchase order, the packing list, the proof of delivery. That is settled in the plumbing rather than in the metric design, and almost nobody budgets for it.
Where the data actually lives, and which joins break
The metrics above are not hard to define. They are hard to populate. A single PO line's story is spread across at least three systems, and the joins between them break in specific, predictable ways — the same seven ways, at almost every company.
The promise date lives in the ERP. The arrival lives in the WMS. The transport event lives in the TMS or the carrier's feed. The quantity exists in three places — PO line, receipt, invoice — and they disagree. The ASN arrives by EDI and is never diffed against what turned up. Defects surface in the QMS weeks after the receipt was closed.
The single most common defect in supplier on-time reporting is this: most ERPs overwrite the confirmed delivery date on the PO schedule line when a supplier reschedules. By the time the goods arrive, the PO says the supplier was on time, because the PO says the promise was the date they eventually hit. Unless you are capturing the change log and snapshotting the first confirmed date, your on-time number is measuring nothing. Track reschedules as a separate metric — schedule change rate — so the supplier cannot quietly move the goalposts and score full marks.
The second most common defect is unit of measure. The PO is in cases, the receipt is keyed in eaches, and the conversion factor is maintained in a third place. A 96% line fill becomes 100% or 12% depending on which direction the conversion goes wrong, and it will not be obvious from the summary.
The third is receipt timestamps. A WMS goods receipt records when a receiver keyed the transaction, not when the truck hit the gate. A Friday 18:00 arrival receipted Monday 08:15 reads as a three-day miss. If you have a gate or yard log, join to that. If you do not, at minimum agree a business-day rule with your suppliers and apply it consistently.
And one that is structurally unfixable if you get the terms wrong: under Ex Works or FCA terms, the supplier's obligation ends at their dock. If you grade them on delivery date, you are grading your own forwarder. Measure ready-to-ship date and booking notification accuracy instead, and grade the carrier separately. The same logic applies in reverse under DDP or DAP, where delivery date is fair game.
| Data you need | System of record | Why the join breaks |
|---|---|---|
| First confirmed delivery date | ERP purchase order schedule line | Overwritten on every reschedule; no history unless the change log is captured and snapshotted |
| Actual arrival at gate | Yard gate log, or carrier POD | WMS receipt timestamp records keying time, not arrival time; weekend arrivals read as multi-day misses |
| Received quantity by line | WMS goods receipt | Multiple partial receipts per line; unit-of-measure conversions between each, case, and pallet |
| Ready-to-ship or pickup date | EDI 856 supplier ASN or forwarder booking | Frequently absent on Ex Works and FCA moves, which is exactly where it is the only fair measure |
| ASN content versus actual | EDI 856 EDI 856 versus WMS receipt | The ASN is keyed against the PO, the receipt is keyed against the ASN, and the difference is never stored |
| Transit exception ownership | TMS TMS or carrier event feed | No link from a carrier event back to the PO line, so supplier fault and carrier fault cannot be separated |
| Defect and reject quantity | QMS QMS or WMS inspection record | Defects found line-side weeks later rarely post back against the original receipt or period |
Seven joins. Every one of them decides a number you are about to publish.
The six definitions you must settle before you publish a number
Write these down, get them signed off internally, and send them to every supplier before the first scored period. Every one of them is a live argument waiting to happen in a quarterly review, and every one of them is cheap to settle in advance and expensive to settle in the room.
1. Which date is the promise
Requested date, first confirmed date, or latest confirmed date. Requested date grades the supplier on your wishful thinking; latest confirmed grades them on nothing, because it moves every time they move it.
2. The tolerance window
Day-exact, minus two to zero, or plus or minus one day. Early is not free: it consumes dock labor, receiving capacity, and storage you planned for something else.
3. Line level versus order level
Order-level scoring is punishingly binary and gives the supplier no signal about which line failed. An order with nine lines complete and one short is a 90% line fill and a 0% order fill.
4. Partial receipts
Does a line receipted 80% on the due date and 20% four days later count as on time? Whatever you decide here decides whether a back-order tail scores as success.
5. Incoterm-adjusted responsibility
Under Ex Works or FCA terms the supplier's obligation ends at their dock, so a delivery-date score is really a score on your own forwarder. Mixing terms inside one on-time percentage produces a number nobody can act on.
6. Exclusions
A PO released inside the quoted lead time, a customer-driven engineering change, a credit hold, or your own forecast collapse are your failures, not theirs. If exclusions are decided case by case in the review, every review becomes a negotiation about exclusions.
One more reason to settle these in writing: you are on the other side of a scorecard too. Whatever room your customer's definitions leave — the tolerance window, the promise date, how partials are counted — you have already found it. Assume your suppliers are exactly as good at reading yours. Every definition you leave loose is a definition someone will optimize against, and you will not find out until the quarter is closed.
The score has to survive the walk from the account manager to the plant
The people in your quarterly review are rarely the people who can fix anything. You get an account manager and, on a good day, an ops director. The schedule that keeps missing you is built by a planner two buildings away who has never seen your scorecard and does not know your part numbers are the ones under review.
That is what the line-level detail is for. A single score travels badly inside somebody else's business. Forty-one late PO lines with dates, part numbers and confirmed-versus-actual travels well, because it survives being forwarded. Ask in the review who inside their organization receives it — and if the answer is nobody, you have found why the last three corrective actions changed nothing.
Running the review without a spreadsheet argument
The scorecard is the easy part. The review is where these programs live or die, and they die in a very specific way: forty minutes of the meeting spent disputing whether line 4417 was actually late, and five minutes on what changes next quarter. Four rules stop that.
First, the pre-read rule. Publish the score and the full line-level detail five business days before the meeting. Disputes must be filed in writing at least two business days before, citing PO line numbers. Anything not disputed in advance is accepted for the period and can be corrected in the next one. This single rule reclaims most of the meeting, because a supplier who has to reconcile line numbers in writing usually finds you were right.
Second, publish your own performance back to them. Average PO lead time actually given versus the quoted lead time, forecast accuracy at the horizon they plan against, schedule change rate, and payment on time. If your average PO gives 8 days of notice against a 21-day quoted lead time, their on-time score is your problem and you are about to pay a corrective-action consultant to discover that. This reciprocal page is the fastest way to turn an adversarial review into a working one.
Third, bring cost, not percentages. “You were 90% on time” is a debate. The Halvorsen quarter, costed out: 41 late lines drove 6 expedited air moves at an average $4,200 premium over the planned ocean lane, which is $25,200. Premiums vary widely by lane and consignment, from a few hundred dollars for a regional truck upgrade to five figures for an intercontinental charter, so use your own actuals. Add roughly 3 hours of planner rework per exception — 41 lines, 123 hours — at a $65 fully loaded hourly rate (planner rates realistically run $45 to $85 loaded depending on geography and seniority), which is $7,995. Add the $13,855 of working capital locked in extra safety stock on one SKU from the variance calculation. That is $33,195 of expensed cost plus $13,855 of tied-up capital, from one supplier, in one quarter. Nobody argues with that the way they argue with a percentage.
$25,200
expedite premiums, 6 air moves
$7,995
planner rework, 123 hours at $65
$13,855
working capital, one SKU
$33,195 expensed + $13,855 tied up. One supplier. One quarter.
Fourth, structure the corrective action. Containment this week, root cause within two weeks, permanent fix with a named owner, a dated verification, and the specific metric that will prove it. Closed only after three consecutive months at target. If your organization has quality DNA, run it as an 8D. If it does not, the four fields above are enough. What is not enough is “supplier to improve on-time performance”, which is what most corrective action plans say and why most of them fail.
Expect one specific answer in that conversation: their supplier is the problem. Sometimes it is true. Treat it as a scope question rather than an excuse — ask which tier-2 part, on which lines, and require the recovery plan to name that supplier and its date. If a tier-2 constraint is real, it will appear on the same lines every month, and it changes what you should be asking for: buffer, a second source qualified at their tier, or a longer quoted lead time you can actually plan against.
Then attach consequences to the band, publish the consequences table in advance, and apply it without exception. A scorecard with no consequence is a newsletter.
41 late lines: $25,200 in expedite premiums, $7,995 in planner rework, and $13,855 of working capital locked in extra safety stock on a single SKU. Take those numbers to the review, not “90% on time”.
| Band | Score | Share of wallet | Commercial posture | Review cadence |
|---|---|---|---|---|
| Preferred | 85 to 100 | Eligible for growth and new programs | Longer terms, forecast sharing, joint cost-reduction work | Quarterly business review |
| Approved | 70 to 84 | Hold current volume | Standard terms, standard scorecard | Quarterly scorecard, semi-annual review |
| Conditional | 55 to 69 | Frozen, no new part numbers | 90-day corrective action plan; dual-source qualification begins in parallel | Monthly until two consecutive quarters above 70 |
| At risk | Below 55 | Actively resourcing | Exit plan with dated resourcing milestones; cost recovery claim opened | Monthly, with procurement leadership in the room |
One honest caveat on the bands
The table assumes you can resource, and for a real share of your base you cannot: sole-source tooling, a customer-directed supplier your OEM told you to buy from, a casting with an eighteen-month requalification, one approved vendor on the drawing. Scoring those suppliers still matters — the score is your evidence in the commercial conversation and your justification for the buffer you are already carrying — but the consequence has to change. For a supplier you cannot exit, At risk means escalation to their leadership rather than yours, a joint recovery plan with your engineering team in the room, buffer stock priced into the part and charged back where the contract allows, and dual-source qualification funded as a project instead of waved as a threat. Publishing a consequence you cannot execute is worse than publishing none, because the first time you don't execute it, every band above it becomes advisory.
The first 90 days: how to stand this up without getting the first scorecard rejected
Do not launch a scored program across your whole supply base. Every supplier scorecard that gets rejected in its first quarter was rejected on data quality, not on metric design, and the rejection is usually correct.
Before day one
Find the last one. Most companies asking this question have already tried. There is a workbook somewhere, built by an analyst who has since left, last updated eight months ago, that two or three people still quote in meetings. Open it. It will tell you which definitions were already agreed, which suppliers have already seen a score, and which join broke first. It will also tell you the real reason it stopped, which is almost never metric design — it is that assembling it every month was one person's unpaid second job, and that person moved teams.
Days 1–15
Settle the definitions, pick three pilots
Sign the six definitions off with procurement, planning, quality, and receiving in one room. Then pick three pilot suppliers spanning your range: one strategic and generally good, one problem supplier, one long-lead importer. Send them the definitions, the weights, and the bands.
Days 16–45
Build the joins, then hand-audit 50 lines
Snapshot the first confirmed date, resolve the unit-of-measure conversions, decide the arrival timestamp rule, and wire the ASN diff. Then do the step everyone skips: hand-audit 50 PO lines end to end against the systems and compare to what your automated calculation produced. If the two disagree by more than a couple of lines, you have found the reason your program would have been rejected. Fix it before anyone outside the company sees a number.
Days 46–60
Shadow run a full period
Calculate everything, publish nothing. Look at the score distribution. A healthy first distribution is a spread: a couple of suppliers in the 80s, most in the 60s and 70s, and one or two below 55 that nobody in the room is surprised by. If your three pilots all land in the same band, your targets and floors are not discriminating and need recalibration.
Days 61–75
Share unscored data with the pilots
Line-level detail, no score, no consequence. Invite them to dispute the data, and expect them to find real errors in it. Fixing those before a score is attached is what buys you the right to attach one.
Days 76–90
First scored period, consequences published
Run the first scored period with the consequences table published alongside it. Then expand in tiers, by spend and criticality, rather than all at once.
One rule to carry through all of it: never let the first scorecard a supplier ever sees carry a consequence. The first one is a calibration exercise for both sides. The second one has teeth.
Where this gets hard, and what Orkestra actually does about it
Section six is the bottleneck, and it is not a method problem. One PO line's story sits in four places: the promise and its change history in the ERP, the receipt and the received quantity in the WMS, the transport events in the TMS or the carrier feed, the ASN in the EDI stream. Assembling that by hand is a month-end exercise, which is why most scorecard programs end up grading last quarter and spending the review arguing about the data instead of the supplier.
Orkestra is a layer over the ERP, TMS, and WMS you already own. It does not replace any of them. What it does for this problem specifically is hold one record per PO line — first confirmed date and its change history, ready-to-ship, transport events, arrival, received quantity by line, ASN versus actual, document status — so the number assembles continuously instead of being rebuilt in a workbook the week before the review. Order management holds the PO-to-receipt lifecycle. Analytics runs your metric definitions against that record, so when a supplier disputes line 4417 you open the line, not the workbook. And because the record is live rather than retrospective, the Exception Monitoring Agent can flag an inbound line drifting past its confirmed date while there is still a week to expedite, split, or re-promise — which over a quarter is worth more than the score itself.
Supplier on-time in-full is the same measurement problem as customer OTIF, one tier upstream, and it breaks the same way: the data existed in time to act, but it lived in four systems with no owner and no workflow. The customer-facing half of it is in the OTIF guide.
DBW, a global automotive supplier, replaced spreadsheet-driven coordination with real-time visibility and cut supply chain costs 18%. Matalco unified supply chain data across its systems and partners. Neither replaced an ERP, TMS, or WMS to do it.
What Orkestra does not do
We are not an ERP, a TMS, a WMS, or an SRM suite, and we do not replace any of them. We do not hold your supplier master, run sourcing events, manage contracts, or issue purchase orders. We cover the delivery half of this scorecard — on-time, in-full, lead time variance, ASN accuracy, document status — and we do not cover the quality half: defect and reject data lives in your QMS and stays there, and if anyone tells you otherwise ask to see the inspection record. We do not rate your suppliers for you, and we do not set your weights, targets, or floors, because those are commercial decisions that belong to your procurement organization. And clean, live data does not run the conversation. No software will hold the quarterly business review, write the corrective action plan, or make the resourcing decision. It removes the argument about whether the number is right, which is the part that was wasting the meeting.
