Every warehouse-management talk opens with "you can't manage what you don't measure." Fair enough - except that in practice the opposite problem is more common: companies track forty metrics that nobody draws any conclusion from, because the report lands in an inbox at 4pm on Friday and no one opens it. A sensible warehouse KPI set is five, maybe seven numbers - but ones that a specific person looks at on a specific cadence and knows how to act on. Below are the ones that, after several hundred rollouts, we consider genuinely worth measuring, along with the ranges you can expect.
Inventory accuracy - the number-one metric
Inventory accuracy - how well the system's stock figures match what's physically on the shelf. If this one is broken, everything else stops mattering - picking productivity is irrelevant when every tenth line the operator hits an empty location that "the system says" is holding stock. A healthy WMS-run warehouse sits at 99.5% or better. Drop below 98% and the operational pain starts: picking stalls, emergency moves, and the "sorry, turns out we don't have it" call to the customer.
How you count it is what matters. Accuracy measured on total units across the whole warehouse lies - a surplus of an SKU in one location masks a shortage in another, and on paper "it all adds up." The honest metric is the percentage of locations where physical stock equals the system count down to the unit, measured via cycle counts. At 99% accuracy per location, one in a hundred operator visits ends in a surprise - that's livable. A "99.8% accuracy" figure counted on totals can be hiding hundreds of locations that don't reconcile.
Lines per labor hour - picking productivity
The core productivity measure: how many order lines the team picks per hour worked. The ranges are wide because they depend on the operation profile: full-pallet picking with a forklift runs 15-30 lines/h, case picking from shelf racking 40-80, piece picking in a well-organized warehouse 60-120, and with automation (conveyors, pick-to-light) 150-250. Benchmarking yourself against another industry is pointless - what counts is your own week-over-week trend.
Two rules for counting it honestly. First, the denominator is the picking team's entire time on the clock, not just "terminal-in-hand" time - otherwise the metric climbs from shuffling people onto other tasks rather than from any process improvement. Second, lines are not units: an order for 200 pieces of a single SKU is one line and one trip, an order for 5 pieces of five SKUs is five lines. Warehouses bragging about "units per hour" usually just have a wholesale order profile.
Dock-to-stock - how long goods sit on the dock
The time from physical unloading to the moment goods are received, put away, and available to sell in the system. In retail and e-commerce this is a metric you can see directly in money: stock sitting a full day in the receiving buffer is stock you can't sell for a day, even though the capital has already been paid for it. Good operations keep dock-to-stock under 4 hours for standard deliveries and under 24 hours for deliveries that need a quality check.
It's worth measuring two segments separately: unloading → receipt completed in the system, and receipt → put-away into a location. The first segment breaks down over missing advance notices (with no ASN, every receipt is a manual count), the second over put-away bottlenecks, usually at peak season. Splitting it out immediately shows which side the problem is on - without the split, the discussion at the stand-up comes down to "receiving is slow."
OTIF and fill rate - the metrics your customer grades you on
OTIF (on time, in full) - the percentage of orders delivered on time and complete. Fill rate - the percentage of lines filled from available stock on the first attempt. These are the metrics the warehouse gets judged on from the outside: retail chains write OTIF into contracts and charge penalties below the threshold (typically 95-98%, with the biggest players able to demand 98.5%).
The trap is that each side counts OTIF differently. The warehouse counts "shipped from the warehouse on time," the carrier "delivered within the window," the customer "received at my dock, complete and undamaged." Between the first and third definition there can be 5-8 percentage points of difference - and a row at the quarterly business review. If the customer counts OTIF their way, you have to measure it in-house against that same definition, otherwise your "99% here" figure is just decoration.
Mispick rate - picking errors counted in ppm
The share of lines picked wrong: wrong SKU, wrong quantity, wrong batch. In paper-based warehouses, 1-3 errors per hundred lines is typical (10,000-30,000 ppm). Scanning every line brings that down to 0.05-0.1% (500-1,000 ppm), and scanning with quantity validation and a check-weigh at packing - below 200 ppm. This is one of the strongest arguments for a WMS at all, because every picking error costs real money: return freight, re-shipment, a claim, sometimes a lost customer. The cost of a single error in B2C is typically 30-80 PLN; in B2B with contractual penalties it can run into the thousands.
It's worth measuring broken down by operator, zone, and SKU. The distribution is almost never even: usually 20% of the errors come from one person (needs retraining), and another 30% are generated by two or three "problem" SKUs - near-identical product variants sitting next to each other that slotting should have separated.
Cost per line handled - the hardest and the most important
The full operating cost of the warehouse (wages plus overheads, floor space, utilities, equipment with depreciation, materials, the system) divided by the number of lines handled. In Polish conditions in 2026 this comes out at anywhere from 2-4 PLN per line in efficient piece-picking operations to 15-25 PLN for full-pallet picking with a narrow volume. This metric does two things no other one will: it lets you honestly price serving a customer or channel (fulfillment for e-commerce vs. replenishing a retail chain) and it shows whether growing volume is actually spreading your fixed costs or just piling on overtime.
You calculate it once a month, not daily - and that's fine. Day-to-day variation in cost per line is noise; what matters is the monthly trend and the year-over-year comparison for the same season.
Five, not forty - how to build a metric set
A metric only works when it has an owner, a review cadence, and a possible response. A practical setup that doesn't die after a month:
- Daily, at the 10-minute shift stand-up: yesterday's lines/h, this morning's picking backlog, the number of errors caught at packing. Three numbers on the board, two minutes of discussion, a "who goes where" decision.
- Weekly, warehouse manager: inventory accuracy from cycle counts, dock-to-stock, mispick rate broken down by zone. This is where decisions about retraining, re-slotting, and procedure changes get made.
- Monthly, management/owner: cost per line, OTIF in the customer's definition, productivity month-over-month and year-over-year. The level for decisions about headcount, investment, and pricing.
Anything beyond this set should justify itself by answering the question "who does what differently when this number moves." If there's no answer, the metric is report decoration and you can delete it without loss.
Where the WMS gets these numbers
The whole mechanism rests on one thing: every operation in the WMS has a timestamp, an operator, and a location. A pick scan is a ready-made record of "who, what, from where, when" - lines/h falls out of that with no extra work. Receipt and put-away have their own timestamps, so dock-to-stock comes out of the system on its own. Cycle counts record variances per location - that's the source of inventory accuracy.
What the WMS won't measure on its own: costs (it needs data from accounting), OTIF in the "delivered" definition (it needs statuses from the carrier via integration), and root causes - the system will show that dock-to-stock jumped on Tuesday from 3 to 9 hours, but that two agency workers didn't show up in the morning is something you still have to add yourself. A KPI from the system is a thermometer; the diagnosis stays with people.
Benchmarks - go easy on the comparisons
Industry ranges, including the ones quoted above, are treated as a reference point for the first measurement - not as a target. A warehouse with 30,000 SKUs of small items won't hit the lines/h of a full-pallet operation, and shouldn't try. The real value of comparison is internal: the same operation week over week, shift over shift, season over season. A metric that has crept up 15% over three months says more than any table from an industry report.
And one warning from the field: tying bonuses straight to a single KPI corrupts the data faster than it improves the result. Bonus-driven picking productivity goes up - and so do errors, because people stop reading labels. Bonus-driven OTIF goes up - because planners quietly stretch the promised dates. If you're going to pay a bonus, tie it to a pair of metrics that keep each other in check: productivity alongside quality, on-time alongside in-full.
Summary
The minimum set for a warehouse that wants to know what's going on inside it: inventory accuracy per location, lines per labor hour, dock-to-stock, mispick rate, and once a month cost per line plus OTIF counted the way the customer counts it. Five or six numbers, three review cadences, every number with an owner. The rest is optional.
Weaver WMS calculates these metrics from data that gets created at every scan anyway - productivity reports per operator and zone, receiving times, cycle-count results, a picking-error log. The data also goes out through a REST API, so if a company has its own Power BI dashboard or a data warehouse, the warehouse feeds its numbers into it automatically. From experience: just putting a board with yesterday's three numbers in front of the team can lift productivity by a few percent within a month - before anyone has optimized a single thing.