A service level agreement is not a trust exercise or a legal trophy; it is a shared measurement system that tells both parties whether the operation is healthy and what happens when it is not. This article shows how to design an SLA with a supply or fulfillment partner: which metrics can actually be measured, how to set thresholds that change behavior, and the review mechanics that keep the document alive after signature.
What an SLA is actually for
An SLA does three jobs that goodwill cannot. It converts vague expectations — "fast dispatch", "good quality" — into definitions both sides can compute the same way. It creates consequences proportionate to misses, so underperformance surfaces as a number instead of as a feeling accumulated over months. And it creates a forum: a scheduled review where the numbers are read together, which is where most operational problems actually get solved. What an SLA is not is a substitute for choosing a competent partner — a weak operation with a strict SLA produces disputes, not performance. The document formalizes a working relationship; it cannot create one. That is why serious partners often propose SLAs themselves: measurement discipline is a feature of operations that expect to be measured.
Only metrics with an owner and a data source
Most SLA failures are definition failures. "Orders ship quickly" collapses under the first dispute; "orders placed before 14:00 local time dispatch the same business day, measured from the warehouse management system's dispatch-scan timestamp, excluding carrier-confirmed collection delays" does not. Every metric in an SLA needs three things in writing: a formula, a data source both parties accept, and an owner who answers for it. The metrics that earn their place in most ecommerce supply SLAs, with definitions and commonly used target ranges — ranges vary by category, lane and program, so treat them as starting points for negotiation rather than as standards:
| Metric | Definition (formula and source) | Common target range | Notes |
|---|---|---|---|
| Dispatch on-time rate | Orders dispatched within the agreed handling window after cut-off, per WMS scan, divided by all orders | 97–99% (varies) | Define cut-off time and business-day handling explicitly |
| Oversell rate | Oversold orders per 1,000 orders, from order and stock system logs | Near zero; under 1 per 1,000 (varies) | Depends on sync cadence; define the protocol, not just the number |
| Pick accuracy | Wrong-item or wrong-quantity shipments per 1,000 orders, from claims and QC audits | Under 2 per 1,000 (varies) | Verify by sampling, not only by complaints received |
| Transit in-range rate | Parcels delivered within the published lane range, from carrier scans | 90–95% (varies by lane) | Carrier-owned; remedy design must reflect limited partner control |
| Customer-reported defect rate | Defect-verified complaints per 100 orders, from coded support tickets | 0.5–2% (varies by category) | Verify against inspection records; code reasons, do not paraphrase |
| Exception first response | Hours to first substantive response on exception tickets, from the ticket system | Within 4–24 business hours (varies) | "Substantive" needs defining; auto-replies do not count |
| Inbound dock-to-stock | Business days from receipt at dock to sellable inventory, from WMS timestamps | 1–3 business days (varies) | Matters most before peaks and campaigns |
| Stock count accuracy | Cycle-count matches versus system records | 98–99.5% (varies) | Agree count frequency and method in the SLA itself |
Five to eight metrics is the working maximum. A forty-line SLA is a document nobody reads and everybody resents; the discipline of choosing which five to eight metrics carry the program is itself most of the design work.
Thresholds, tiers and measurement windows
A single pass/fail threshold wastes the information in near-misses. Working SLAs use three bands — a target, an acceptable range and a critical line — so that a month at 96.8% dispatch on-time reads as "watch", not as "breach". Set the bands from data rather than aspiration: take the last three months of actual performance where it exists, add the improvement you are actually paying for, and set the critical line where real damage begins. Agree the measurement window explicitly — monthly rolling windows smooth out single bad weeks, while weekly windows surface problems faster but generate noise in low-volume programs. And set a minimum volume below which a metric is reported but not enforced; percentage targets on thirty orders a month measure luck, not performance.
Remedies that change behavior
The remedies section of an SLA exists to change incentives, not to extract rent, and the design shows. Service credits — a percentage of fees for the affected period, scaled by how far below the critical line performance fell — are the standard instrument, and they work when they are proportionate and automatic: computed from the agreed data source, applied without a negotiation. Two design rules keep remedies healthy. First, match remedy control to actual control: a partner cannot owe credits for carrier-caused transit misses beyond the duty to have booked a competent carrier and escalated visibly. Second, pair every credit with a root-cause obligation — a written answer to "what failed and what changes" — because the credit is compensation, and the root cause is prevention. Chronic misses across consecutive windows justify an escalation ladder that ends in an exit right; a well-built SLA includes the door, which paradoxically is what makes both parties invest in never using it. The exception-handling workflows these clauses sit on top of are covered in our guide to order exception handling.
The review cadence that keeps SLAs honest
An SLA without a scheduled review is a PDF. The cadence that works is weekly for operations — a thirty-minute call reading the operational metrics, the exception log and the coming fortnight's volumes — and monthly or quarterly for the SLA itself: scores against bands, credits applied, root causes closed, thresholds recalibrated against reality. Both parties should see the same scoreboard generated from the agreed data sources; dueling spreadsheets are how measurement disputes begin. When the data plumbing comes from the partner's systems, the visibility commitments to look for are the ones described on our fulfillment service page, and the quality-side metrics — defect rates, inspection coverage — draw on the practices in our quality and compliance process. Recalibration is part of the design: as volumes grow and lanes stabilize, yesterday's target bands should tighten, and the scheduled review is where that happens by agreement rather than by ambush.
Frequently asked questions
What if a prospective partner refuses to sign an SLA?+
Ask what they would measure instead. Refusal to define and share performance numbers is itself diagnostic, and the safest reading is that the operation does not measure itself. Some partners resist specific thresholds while accepting definitions and transparent reporting — that is a workable negotiation; refusing measurement outright is not.
Should remedies be penalties or service credits?+
Service credits computed from the agreed data are the workable instrument: automatic, proportionate and boring. Pure penalties invite dispute over every incident, and remedies that require litigation to activate protect nobody. The stronger lever is usually the escalation ladder and exit right, not the size of the credit.
How many metrics should a first SLA contain?+
Five or six. Dispatch on-time, oversell or pick accuracy, transit in-range, defect rate, exception response, and stock accuracy cover the failure modes that actually damage programs. Add metrics when the operation can honestly measure the first set — a short SLA that is read weekly beats a long one that is filed.
How often should SLA targets be renegotiated?+
Review scores monthly, recalibrate bands quarterly, and renegotiate the full agreement annually. Between those points, targets should move only by written agreement — targets adjusted verbally in month nine are disputes scheduled for month twelve. The recalibration rhythm is a feature: it lets both parties tighten standards as performance earns it.
