The 3PL Evaluation Checklist: 42 Questions Before You Sign
The best way to evaluate a 3PL is to ask the same 42 written questions of every provider, require evidence for each answer, and test the highest-risk claims in a controlled pilot. A polished sales presentation is not proof. BondJet recommends comparing the operating process, data, contract language, and exit readiness before comparing headline rates.
The cheapest quote can become the most expensive choice once dimensional weight, special handling, inventory variance, claims, and rework appear. This 3PL evaluation checklist gives operations leaders a practical way to expose those costs and risks before signing. It covers pricing, SLAs, systems, inventory, packaging, claims, peak season, and exit terms, then turns the answers into a weighted decision.
Key Takeaways
- Use all 42 questions across eight contract areas and require a document, system record, report, or live demonstration for every material claim.
- A credible 3PL should request at least 11 data sets before giving a dependable quote or onboarding plan.
- Run a representative pilot with written baselines, thresholds, exception tests, and a go/no-go review; do not treat a few perfect orders as validation.
- Weight evidence quality and operational fit above sales confidence, and make missing data ownership or exit rights automatic disqualifiers.
- BondJet can review the completed scorecard against the needs of high-value, fragile, and complex-SKU cross-border products.
Download the 3PL Evaluation Scorecard Template
Use the table below as a working 3PL scorecard template. Copy one row for each of the 42 questions and keep evidence beside the answer. A verbal “yes” should score no higher than an undocumented claim.
| Field | What to record |
|---|---|
| Question ID | 1-42 |
| Provider answer | The exact written response |
| Evidence link | Contract clause, report, screenshot, invoice, SOP, or demo recording |
| Internal owner | Finance, operations, IT, legal, or customer experience |
| Score | 0-5 using the evidence scale below |
| Risk | Low, medium, high, or disqualifying |
| Follow-up | Open item, owner, and due date |
Evidence score: 0 = no answer; 1 = verbal claim; 2 = generic document; 3 = relevant sample; 4 = customer-specific proof; 5 = customer-specific proof validated in the pilot.
Recommended category weights are pricing 15%, SLA 15%, systems and data 15%, inventory 15%, packaging 10%, claims 10%, peak season 10%, and exit 10%. Adjust them before issuing the request for proposal, not after seeing providers' scores.
3PL Evaluation Checklist: Pricing Questions 1-6
1. What is included in each receiving, storage, pick, pack, and shipping charge?
Red flag: Bundled labels such as “fulfillment fee” with no activity definition.
Evidence to request: An itemized rate card, billing glossary, and annotated invoice for an order mix similar to yours.
2. Which billing units and rounding rules apply?
Red flag: Weight, storage, labor, or order increments are omitted from the proposal.
Evidence to request: Written formulas and three worked examples, including a split shipment and a multi-line order.
3. How are actual weight and dimensional weight calculated?
Red flag: The provider cannot show the divisor, measurement point, or carton dimensions used for billing.
Evidence to request: Carrier rules, warehouse measurement SOP, and sample shipment records. Compare the method with published carrier guidance such as the UPS dimensional weight explanation.
4. What minimums, deposits, implementation fees, and recurring platform fees apply?
Red flag: Minimum monthly charges appear only in the contract appendix.
Evidence to request: A complete fee schedule and 12-month cost model for low, expected, and peak volumes.
5. Which surcharges can be passed through, marked up, or added manually?
Red flag: “At cost” is used without defining source documentation or markup.
Evidence to request: A surcharge list, carrier invoice example, approval rules, and audit rights.
6. How can rates change during the term?
Red flag: The provider may change rates immediately or without objective triggers.
Evidence to request: Escalation formula, notice period, historical change example, and termination rights after a material increase.
Pricing should be modeled at order level. BondJet starts with product, destination, packaging, and service requirements because a responsible quote must reflect the actual operating profile rather than a single attractive per-order number.
3PL Selection Checklist: SLA Questions 7-12
7. How does the SLA define an order received, accepted, released, shipped, and delivered?
Red flag: “Same-day shipping” has no clock, time zone, or system event attached.
Evidence to request: A data dictionary and timestamped order journey from the warehouse system.
8. Where are cutoffs measured, and what happens to orders received after cutoff?
Red flag: Sales-channel time, warehouse time, and carrier handoff time are mixed.
Evidence to request: Cutoff table by service and facility, including weekends and holidays.
9. Which exclusions stop the SLA clock?
Red flag: Broad exclusions let the provider classify ordinary exceptions as customer-caused.
Evidence to request: An exhaustive exclusion list with reason codes and sample reports.
10. What remedy applies when the provider misses an SLA?
Red flag: The SLA has targets but no service credit, correction plan, or escalation duty.
Evidence to request: Contract language defining eligibility, calculation, claim window, and corrective action.
11. How will SLA performance be reported and reconciled?
Red flag: Monthly percentages cannot be traced to individual orders.
Evidence to request: A sample dashboard, raw export, metric owner, and dispute workflow.
12. How are SLA definitions or targets changed?
Red flag: The provider can revise metrics through a portal notice.
Evidence to request: Formal change-control clause requiring impact assessment and mutual written approval.
Ask BondJet for a customer-specific SLA working session when evaluating high-value cross-border fulfillment. The useful output is not a generic promise; it is a written definition of receiving, inspection, SKU checks, packaging, dispatch, exceptions, and reporting that both teams can verify.
3PL Due Diligence Checklist: Systems and Data Questions 13-17
13. Which platforms, marketplaces, carriers, and APIs are supported, and who maintains each integration?
Red flag: “We integrate with everything” without version, ownership, or support boundaries.
Evidence to request: Integration catalog, architecture diagram, implementation owner, and a live demonstration.
14. What happens when an order, inventory update, or tracking event fails?
Red flag: Failed messages rely on someone noticing an inbox.
Evidence to request: Retry logic, alert thresholds, exception queue, escalation path, and incident example.
15. How are order cutoffs, holds, edits, cancellations, and duplicate orders handled?
Red flag: The operating rules differ between the user interface, API, and warehouse floor.
Evidence to request: State diagram, test cases, permission matrix, and audit logs.
16. How is customer and operational data protected?
Red flag: Security answers consist only of a logo or unsupported certification claim.
Evidence to request: Access-control policy, encryption statement, incident-response plan, backup test, and current independent assessment where applicable. The NIST Cybersecurity Framework 2.0 is a useful common reference.
17. Who owns the data, and can we export complete records on demand?
Red flag: Exports omit event history, images, reason codes, or identifiers needed to migrate.
Evidence to request: Contract clause, full export schema, sample file, API limits, retention schedule, and deletion process.
Inventory Questions 18-22
18. How is inbound inventory identified, counted, and accepted?
Red flag: Receiving closes against cartons without SKU-level variance handling.
Evidence to request: Receiving SOP, sample discrepancy record, photo policy, and timestamp report.
19. How is inventory accuracy defined and measured?
Red flag: The provider quotes a percentage but excludes quarantined, damaged, or unlocated stock.
Evidence to request: Metric formula, denominator, location-level report, and recent reconciliation example.
20. What cycle-count program applies to our SKUs?
Red flag: Counts occur only after a customer reports a shortage.
Evidence to request: ABC rules, count frequency, blind-count method, variance thresholds, and correction approvals.
21. How are damaged, returned, restricted, or suspect units quarantined?
Red flag: Non-sellable stock can be allocated to live orders.
Evidence to request: Status codes, physical segregation method, release permissions, and audit trail.
22. How are inventory variances investigated and financially reconciled?
Red flag: Adjustments can be posted without root cause, approval, or supporting evidence.
Evidence to request: Investigation workflow, adjustment log, liability clause, and sample closure report.
For complex SKU assortments, BondJet combines inbound identification, inspection records, SKU management, and outbound checks. That process is particularly relevant when sets, accessories, condition, or product variants affect whether the buyer receives the correct item.
Packaging Questions 23-27
23. Who approves the packaging specification for each SKU or product family?
Red flag: Packers choose materials from memory or personal preference.
Evidence to request: Version-controlled pack-out instructions, approval owner, and training record.
24. What evidence proves that the approved pack-out was followed?
Red flag: There is no link between an order and its packing record.
Evidence to request: Scan events, material consumption records, workstation checks, and sample photos where appropriate.
25. Can packaging materials or carton sizes be substituted without approval?
Red flag: Substitutions are allowed whenever stock runs low.
Evidence to request: Approved alternatives, change threshold, notification rule, and exception log.
26. How are carton dimensions controlled before carrier handoff?
Red flag: No one compares the packed carton with the expected dimensional profile.
Evidence to request: Cartonization rules, measurement process, variance report, and billing feedback loop.
27. How are fragile, high-value, or presentation-sensitive products handled?
Red flag: “Fragile” means only adding a sticker or more void fill.
Evidence to request: Product-specific risk assessment, inspection scope, drop or pack test where relevant, photo evidence, and escalation rules.
BondJet's figure seller fulfillment case shows why packaging should be evaluated as an operating control, not merely a consumable charge. For high-value goods, inspection, accessory checks, protective materials, and traceable pack-out decisions reduce avoidable damage and condition disputes. Specific protection still needs to be assessed by product, route, and carrier requirements.
Claims Questions 28-32
28. What are the reporting windows for loss, damage, shortage, and mis-shipment?
Red flag: Different deadlines exist in the contract, carrier terms, and support portal.
Evidence to request: Consolidated claims matrix by event type, service, and responsible party.
29. What evidence must each party provide for a claim?
Red flag: Evidence requirements are disclosed only after a claim is filed.
Evidence to request: Claim checklist covering invoices, photos, scans, packaging, delivery proof, and timelines.
30. How is liability divided among the merchant, 3PL, and carrier?
Red flag: Every failure is automatically treated as a carrier issue.
Evidence to request: Responsibility matrix and contract clauses for warehouse error, inadequate pack-out, carrier loss, and customer instruction.
31. Who escalates carrier claims and communicates with the recipient?
Red flag: Ownership shifts between teams while deadlines continue running.
Evidence to request: Named owner, escalation ladder, communication templates, and response targets.
32. How can we see claim status, aging, outcome, and recovered amount?
Red flag: Claims are tracked in private spreadsheets with no merchant access.
Evidence to request: Sample dashboard or export with case ID, dates, owner, status, reason, and settlement.
Do not accept an unverified statement about compensation as a substitute for terms. Coverage, declared value, exclusions, documentation, and settlement rules should be confirmed in writing for the actual product and service.
Peak Season Questions 33-37
33. What forecast format, horizon, and update cadence do you require?
Red flag: The provider asks for a forecast but does not define how it affects staffing or capacity.
Evidence to request: Forecast template, variance bands, submission calendar, and capacity response.
34. What proof shows that labor, space, equipment, and carrier capacity can support our peak?
Red flag: Capacity is described as “scalable” without numbers or constraints.
Evidence to request: Site capacity model, staffing plan, utilization limits, carrier allocation, and contingency assumptions.
35. Which cutoffs, SLAs, fees, or receiving rules change during peak?
Red flag: Temporary restrictions can be announced after inventory arrives.
Evidence to request: Peak calendar, notice requirement, historical example, and contract precedence.
36. How are backlog and aging reported each day?
Red flag: Only completed volume is reported, hiding orders waiting in queues.
Evidence to request: Daily queue report by age, cause, priority, and expected clearance.
37. What recovery plan starts when backlog exceeds the agreed threshold?
Red flag: The response is simply to “add people” with no trigger or owner.
Evidence to request: Trigger levels, incident lead, prioritization rules, extra-shift plan, communication cadence, and closure review.
Exit Terms Questions 38-42
38. What notice period, termination rights, and exit fees apply?
Red flag: Exit charges are open-ended or calculated only after notice.
Evidence to request: Fee schedule, notice mechanics, termination-assistance clause, and worked example.
39. How will the final stock count be performed and disputed?
Red flag: The provider's final count is binding without merchant observation or reconciliation.
Evidence to request: Joint count procedure, cutoff time, blind-count rules, variance process, and sign-off record.
40. How will inventory be prepared and transferred to the next location?
Red flag: There is no committed sequence for holds, labeling, palletization, documents, and pickup.
Evidence to request: Transfer plan, unit and pallet manifest format, condition record, schedule, and chain of custody.
41. Which complete data set will be exported, in what format, and by when?
Red flag: Only current on-hand quantity is available after termination.
Evidence to request: Export containing SKU master, lots or serials where applicable, locations, orders, returns, claims, adjustments, images, tracking, invoices, and event history.
42. When will remaining copies be deleted, and how will deletion be confirmed?
Red flag: Data remains indefinitely in operational tools, partner systems, or user accounts.
Evidence to request: Retention map, legal exceptions, deletion schedule, access revocation record, and written confirmation.
An orderly exit is part of service quality. Inventory, records, images, identifiers, and event history should remain usable through the transition. If a provider resists defining the exit while pursuing the sale, record that resistance as current evidence.
What Data Must a 3PL Request Before Quoting or Onboarding?
A credible provider cannot design a dependable operating plan from monthly order volume alone. At minimum, it should request:
- SKU master: identifiers, descriptions, variants, kits, barcodes, values, and handling attributes.
- Dimensions and weights: product, retail pack, case, and expected outbound cartons.
- Order history: line count, units per order, split orders, cancellations, and daily volatility.
- Destination mix: countries, regions, postal zones, residential share, and remote-area exposure.
- Service levels: delivery promises, order cutoffs, carrier preferences, and restricted services.
- Inventory profile: average and peak units, pallets, turns, seasonality, aging, lots, or serials.
- Packaging rules: approved materials, presentation standards, inserts, labels, and fragile-item needs.
- Returns and claims history: reasons, rates, evidence, disposition, and financial impact.
- Forecasts: promotions, launches, peak periods, inbound schedule, and expected variance.
- Integration map: commerce platform, marketplace, ERP, support tools, data owners, and event flows.
- Compliance constraints: product restrictions, dangerous goods status, customs data, tax model, and destination requirements.
A provider that asks better questions is more likely to uncover hidden work before launch. BondJet uses this discovery stage to map inspection, SKU management, customized packaging, warehousing, international dispatch, customs support, and tracking to the merchant's actual product and lane requirements.
How to Design a Representative 3PL Pilot
A pilot should test the future operation in miniature, including exceptions. Use the following acceptance framework.
| Pilot element | Required decision |
|---|---|
| Scope | Representative fast, slow, fragile, high-value, multi-line, and return-prone SKUs |
| Lanes | Major destinations plus at least one operationally difficult lane |
| Baseline | Current cost, cycle time, accuracy, damage, claims, and support workload |
| Sample | Enough orders and inbound activity to exercise each required workflow |
| Thresholds | Written pass/fail level for every critical metric |
| Exception tests | Hold, edit, cancellation, stock variance, damaged inbound, failed integration, return, and claim |
| Evidence cadence | Daily exceptions and weekly raw-data review |
| Decision | Named go/no-go owners, unresolved-risk list, and remediation deadline |
Measure from system events rather than recollection. Reconcile the provider's report against commerce, warehouse, and carrier data. A pilot acceptance criteria document should also define whether one critical failure overrides the average score.
For high-value or fragile goods, include inbound photos, SKU and accessory checks, packaging authorization, carton measurement, handoff tracking, and exception communication. BondJet can help shape a pilot around these controls, including a small initial trial where appropriate, without treating a successful sample as a blanket guarantee for all future products or routes.
How to Score Providers and Apply Quality Gates
First, score each answer from 0 to 5 using the evidence scale. Then calculate each category's weighted result:
Category result = average question score / 5 x category weight
Do not let a high total hide a fundamental control gap. Apply automatic disqualifiers before ranking providers:
- Refusal to provide a complete fee schedule or sample invoice.
- No auditable inventory adjustment history.
- No merchant ownership and usable export of operational data.
- No written security or incident-response process.
- No defined exit inventory and data handover.
- Contract language that contradicts the proposed SLA without an agreed correction.
- Material pilot failure with no validated corrective action.
Finally, compare the risk profile, not just the total. A provider scoring 82% with strong evidence in your critical categories may be safer than one scoring 90% through polished but generic documents.
FAQ About How to Evaluate a 3PL
How many 3PL providers should we score?
Three to five serious candidates usually gives a useful comparison without overwhelming the team. Apply the same data pack, questions, scoring weights, and deadlines to each provider.
Who should own the 3PL evaluation?
Use a cross-functional owner group: operations leads the workflow; finance validates total cost; IT checks integrations and data; customer experience reviews exceptions and returns; legal confirms liability, SLA, security, and exit language.
Is a warehouse tour enough evidence?
No. A tour confirms that a facility and process exist at one moment. Pair it with customer-specific documents, raw reports, system demonstrations, samples, contract clauses, and pilot results.
Should the lowest-cost 3PL win?
Only when cost is measured against the same service, packaging, destination, billing, claims, and risk assumptions. Compare total landed fulfillment cost and expected exception cost, not a headline pick-and-pack fee.
When should legal review begin?
Begin before the operational finalist is selected. The contract may reveal exclusions, unilateral changes, data restrictions, liability gaps, or exit costs that change the score.
Make the Decision From Evidence
A sound 3PL decision rests on four things: complete operating data, 42 consistent questions, customer-specific evidence, and a representative pilot. Price still matters, but it belongs beside measurable SLAs, system controls, inventory accuracy, packaging discipline, claims ownership, peak readiness, and an executable exit.
Start by copying the scorecard, assigning internal owners, and marking every unsupported answer as an open risk. Then test the most consequential claims before moving inventory at scale.
BondJet supports growth-stage cross-border sellers that need inspection, SKU management, custom packaging, warehousing, and international fulfillment connected in one traceable process. Review BondJet's fulfillment approach, learn more about BondJet, or contact the team to assess your completed scorecard and define evidence-based next steps.