How to Build a Carrier Scorecard in 7 Steps
Build a carrier scorecard with real TMS and EDI data instead of spreadsheets. Steps, KPI formulas, weighting, and a quarterly review template.
Why most carrier scorecards fail before the first review
Most carrier scorecards die the same way. Someone builds a spreadsheet in Q1, asks each carrier to submit their own on-time delivery numbers, and by Q3 nobody trusts the tab anymore because two carriers report 97% OTIF while your customers keep complaining about late deliveries. A carrier performance scorecard is only worth building if it runs on a structured, consistent set of metrics pulled from data you control, not numbers a carrier hands you in a PDF once a quarter. This post is a build guide, not another list of KPIs to consider. By the end you'll have a working scorecard fed by your TMS and EDI feeds, not by carrier goodwill.
What you need before you start
You need four things in place before you open a spreadsheet or a BI tool. Skip any of these and the scorecard will be built on sand.
- Six to twelve months of shipment-level data per carrier: pickup and delivery timestamps, ordered versus delivered quantity, damage and claims records, and invoice line data.
- A defined, current list of active carriers per lane, split by mode (road FTL, LTL, parcel, whatever mix you run).
- Access to your own execution data: TMS exports, EDIFACT IFTSTA milestone messages, proof-of-delivery scans, or WMS receiving records. If your carrier connections aren't sending structured milestone data yet, that's a separate project worth solving first, since an EDI feed without proper IFTMIN/IFTSTA integration will leave you back to manual PODs.
- Internal agreement on what "on time" and "in full" actually mean at your company: arrival at the dock, unloading complete, or goods available for putaway. Get this wrong and every carrier conversation turns into a dispute over definitions instead of performance.
Start in a spreadsheet or BI tool if that's what you have. The end goal is TMS-native reporting, but you don't need to wait for that to get value from the first version.
Building the scorecard, step by step
Here's the sequence that gets you from nothing to a working scorecard your carriers can't argue with.
- Pick three or four metrics tied to business outcomes. OTIF, damage or claims ratio, billing accuracy, and tender acceptance rate cover most of what matters for a shipper. Resist the temptation to track fifteen KPIs. A scorecard with four metrics gets reviewed every month. A scorecard with fifteen gets opened twice a year.
- Write the exact formula for each metric before you build anything. For OTIF specifically, the standard method is to multiply the percentage of on-time deliveries by the percentage of complete deliveries, since this multiplication method rather than averaging ensures that both components must be satisfied for an order to count as successful. Worth knowing: some practitioners argue this shortcut distorts the real number when on-time and in-full aren't statistically independent, and recommend counting orders that clear both conditions directly instead, as counting orders directly is more reliable than multiplying a separate on-time rate by a separate in-full rate. Either method works as long as you pick one, document it, and use it consistently across every carrier and lane, so comparisons stay fair.
- Map each metric to a specific data source, not a carrier report. OTIF pulls from your TMS execution record cross-referenced with the IFTSTA delivery milestone. Damage and claims come from your own claims log and POD condition notes, not the carrier's incident summary. Billing accuracy compares the carrier's invoice line against the agreed rate card in your TMS, flagging any variance automatically. If a metric's only source is something the carrier sends you, it doesn't belong on the scorecard yet.
- Set scoring bands and weights per metric. Green, amber, red thresholds work fine for a first version, say above 95% is green, 90-95% is amber, below 90% is red for OTIF. Weight the metrics according to what actually costs you money. A retailer running JIT replenishment weights OTIF heavier than a bulk chemicals shipper who cares more about damage ratio.
- Score per lane, not just per carrier. A national road carrier that's excellent on domestic trunk routes can be mediocre on a cross-border corridor through a different depot network. Blending everything into one carrier-level number hides exactly the information you need to act on. If you're renegotiating a lane, you need the lane-level score, not the average.
- Automate the pull instead of chasing PDFs every month. This is the step most shippers skip, and it's the one that determines whether the scorecard survives past two quarters. Multi-carrier connectivity and TMS platforms including Cargoson, MercuryGate, project44, Transporeon (now part of Alpega), and nShift can consolidate carrier execution data and TMS records into a single feed that updates the scorecard automatically. Manual entry is where scorecards go to die, usually around month four when whoever owns the spreadsheet gets busy with something else.
- Set a review cadence and tie the score to a decision. Monthly for an internal ops check, quarterly for a formal business review with the carrier, and annually feeding into tender and contract renewal criteria. A scorecard nobody acts on is a filing exercise, not a management tool.
How you'll know it's working
Two consecutive review cycles producing numbers the carrier doesn't dispute is the first real signal. If a carrier still argues about the OTIF figure in month six, either your data source or your definition of "on time" needs fixing, not the carrier's excuse.
The second signal is more concrete: the scorecard actually gets used in a live tender decision or renegotiation, not just archived after the QBR. And the real shift you're after is moving from reactive to proactive. Instead of discovering a carrier's performance dropped from a monthly summary three weeks after the fact, a properly automated scorecard flags the drop as it happens, giving you time to shift volume before a customer SLA breach lands on your desk.
Failure mode: relying on carrier self-reported data
The most common trap in carrier performance management is trusting the carrier's own on-time report instead of your TMS or POD timestamp. Carriers aren't lying, usually. They're measuring against a different clock, often "dispatched on time" rather than "delivered on time," or they exclude exceptions they consider their customer's fault. The fix is straightforward: cross-check the carrier-declared OTD figure against your own delivery timestamps every review cycle. A persistent gap, say five or more points, is the signal to switch your source of truth entirely, not to keep splitting the difference in the carrier's favor.
Cadence matters here too. Weekly checks catch exceptions and SLA breaches while there's still time to act on a specific shipment. Monthly reviews track trends and carrier mix shifts. Quarterly sessions are where you bring the evidence to the table for a contract conversation. Trying to run all three off the same monthly spreadsheet is how carriers end up grading their own homework.
A worked example for European road freight
Here's how the same template applies whether you're scoring a national road carrier or a pan-European network carrier on a cross-border lane.
| Metric | Formula | Weight | Data source | Target band |
|---|---|---|---|---|
| OTIF | On-time % × In-full % | 40% | TMS execution record + IFTSTA milestone | Green ≥95%, Amber 90-95%, Red <90% |
| Damage/claims ratio | Claims / total shipments | 25% | Claims log + POD condition notes | Green <0.5%, Amber 0.5-1%, Red >1% |
| Billing accuracy | Correct invoice lines / total lines | 20% | Invoice vs. TMS rate card | Green ≥98%, Amber 95-98%, Red <95% |
| Tender acceptance rate | Accepted tenders / total tenders offered | 15% | TMS tender log | Green ≥90%, Amber 80-90%, Red <80% |
The scores only mean something against a market backdrop. With EU road freight contract rates at 140.1 index points, a solid 3.2-point increase quarter on quarter in Q1 2026, carriers landing in the red band on OTIF or claims are exactly the candidates you flag first for renegotiation, since you're already paying more for the lane and getting less reliability in return.
Where this fits in your broader TMS setup
A carrier scorecard isn't a standalone spreadsheet project, it's one module of how your transport management system should already be operating. Multi-carrier platforms like Cargoson, MercuryGate, Transporeon, nShift, and project44 increasingly build scorecarding in natively, pulling from live shipment data rather than a sheet someone updates by hand every Friday afternoon. That's the direction worth building toward even if you start with Excel.
Before your next QBR, run through this: four metrics maximum, formulas documented and agreed internally, data sourced from your TMS and EDI feeds rather than carrier PDFs, scores broken out by lane, and a clear owner for the automation step. Get those five things right and the scorecard stops being a quarterly chore and starts being the thing that decides who keeps your freight next year.