You can measure a robotic system's throughput in a pilot. The question that keeps a site manager up at night is a different one: what happens the day a robot stops mid-peak - does shipping stop with it? This article treats reliability not as a promise but as what it is in operation: a question of accuracy, availability, and fallback. Failure modes explicitly included.
Why the reliability question is the real ops question
By the time a site manager weighs in on robotic picking, the in-principle question is usually settled. Whether robotics pays off against manual picking at all - throughput, cost per pick, warehouse fit - is what the comparison of robotic vs. manual picking answers. What comes after that is no longer a spreadsheet problem. It is an operational objection: shipping now depends on machines, and machines fail.
The objection is fair. A manual warehouse has an obvious buffer - if one picker is out, another steps in. Once robots take over the core of picking, that buffer looks lost at first glance, and that impression alone blocks projects whose economics were proven long ago. The site manager wants to lower operational risk, not trade it for a new one.
The answer starts by taking the word apart. In a warehouse, "reliability" means three different things that too often get lumped together:
- Accuracy - am I picking the right thing? The quality of each individual pick.
- Availability - is the system up when I need it? The uptime of the installation.
- Continuity - what happens when it isn't? The fallback when something fails.
These three layers have different answers, different evidence, and different risk profiles. Treating them separately is the difference between an honest assessment and a marketing claim. The best-evidenced one first.
Layer 1 - accuracy: more reliable than manual
Accuracy is the layer where robotics wins most clearly, and the only one of the three with defensible numbers behind it. A system-guided pick using pick-by-light or display guidance does not make the mistakes that are routine in manual work: the wrong bin, the swapped variant, the transposed digit in a quantity.
Manual picking in industry runs at an error rate of 0.1 to 0.3 percent (industry data). That sounds low, but at 5,000 picks a day it adds up to five to fifteen mispicks daily - each carrying returns, rework, and trust costs. Robotic goods-to-person picking reaches, per NEO product data, a picking accuracy of 99.9 percent at an error rate below 0.01 percent. That is roughly an order of magnitude better.
For the ops lead, this means one thing: on this layer, robotics is not the risk but the mitigation of it. The mispick that does not happen is the return that does not arise.
Layer 2 - availability: the honest treatment of failure modes
Availability is the layer where the honest answer matters. Robots fail - every machine does, eventually. The relevant question is not whether a failure happens but what it means for total throughput. This is where a fleet parts ways with a single machine.
What the failure modes are
"The robot fails" is not one fault but a category with very different consequences. The failure of a single robot - a flat or faulty battery, a jammed wheel, a sensor error - is the most common and the most harmless. It affects one unit out of many. The failure of the pick station, where goods are presented to the person, weighs more, because throughput converges there; a warehouse with several stations spreads that risk, one with a single station does not. Disruptions to the WMS connection - the interface that feeds orders into the system - do not disable the hardware but the supply of orders. And then there is the aisle traffic bottleneck: not a defect, but a throughput ceiling that appears under dense occupancy or peak overload.
Two more conditions cost availability without being machine faults. A Wi-Fi or network outage hits fleet coordination, since robots draw their travel orders over the network; a well-built deployment absorbs brief dropouts through local buffering, a longer outage it does not. And maintenance windows are planned but real non-availability of individual units. A reliable system does not talk these cases away - it makes sure none of them stops shipping.
Redundancy through the fleet
The structural difference between a robot fleet and a conventional automation installation lies in what a single failure triggers. When the crane in an automated storage and retrieval system fails, the entire aisle stops - one machine, one aisle, one single point of failure. In a fleet of many autonomous units, the failure of one robot lowers throughput proportionally, but operation does not stop. The faulty unit is pulled out and replaced while the rest keep working.
That is the core of the availability argument, and it is mechanical, not rhetorical: a system whose output is spread across many identical units degrades gracefully on a single failure instead of failing hard. How far throughput dips depends on fleet size and the buffer built in - a fleet sized to average volume with reserve absorbs the loss of one unit without a noticeable effect. That sizing is fixed before full call-off.
The scaling bottleneck as a reliability topic
There is a limit on availability that is not a failure and grows more important as the fleet grows: traffic management in the aisles. Additional robots integrate quickly, but past a certain density the bottleneck shifts from the number of robots to the physical aisle capacity. In narrow aisles, densely stocked zones, and at critical handover points, a traffic bottleneck forms that control software can ease but not dissolve at will. Aisle width is a hard limit no algorithm computes away.
For the reliability question, this deserves to be stated plainly: a fleet does not scale linearly without end in the same footprint. Anyone planning for the next five years' volume should know which of the three bottlenecks they hit first - the robot count, traffic management, or pick-station capacity. The second lever beside more robots is to add another station, which raises throughput without filling the aisles further. That is exactly what the peak test in the pilot is for: run the real peak volume before deciding on full call-off, not during ramp-up.
Layer 3 - continuity: falling back to manual picking
The third layer works structurally differently in an AMR retrofit than in any other automation architecture - and it is the real reason the residual risk here is smaller than it first appears.
The shelving stays, so the manual path stays open
In an AMR retrofit, the robots move into the existing shelving without rebuilding it. The shelving stays unchanged. That single structural fact has a large operational consequence: if the automation fails - hardware, software, or interface - the team keeps picking manually at the same shelves while the fault is fixed. The fallback process is the same process the warehouse ran before automation. You deactivate the robots, and manual picking continues.
This is not a stopgap but a property of the architecture. Two mechanisms absorb a failure: parallel operation before go-live, which lowers the chance that a fault first surfaces live, and the simple fallback afterwards, because the manual infrastructure was never torn out.
Why this is the structural difference to AS/RS and shuttle
In an automated storage and retrieval system or a shuttle installation, that way back is blocked. To install these systems, the existing shelving is replaced - goods then sit in a grid or in shuttle-served channels a person cannot reach without the machine. If the installation goes down, there is no manual parallel path, because the footprint for manual access no longer exists. Falling back to manual is difficult to impossible once the infrastructure is rebuilt; in a retrofit with unchanged shelving, it is the default.
That distinction is the core of the continuity argument. Not "the automated system fails less often" - that cannot be claimed without defensible availability figures - but "when it fails, a tested fallback stays in place". That lowers the risk through the way the system is built, not through a promise.
Mixed operation as a standing property, not just a transition
In the collaborative and bin-to-person approach, mixed operation runs from day one: part of the shelf rows is served by robots, another keeps running manually, with the split configurable in software. That shipping does not stop during the switchover is a rollout advantage and a topic of its own.
For reliability in steady-state operation, the flip side of the same property matters: manual capacity does not disappear after go-live. It stays as a safety net, because people can keep working the same shelves. A warehouse running 70 percent automated and 30 percent manual has a different range of response in a failure than a fully automated installation with no manual way back. That the shelving stays unchanged is therefore a standing reliability reserve, not just a rollout advantage.
The service model - who carries the availability risk
The technical view ends at a commercial question: when a robot fails, who deals with it, and at whose cost? Here the business model decides alongside the technology.
With pay-per-pick the hardware stays the vendor's problem
In the pay-per-pick model and under robotics-as-a-service, the vendor stays the owner of the robots. That shifts the risk allocation at a point that is decisive for availability: hardware availability, maintenance, and spare-parts supply become the vendor's problem, not the operator's. A unit going down is not a fixed asset the operator has to service, stock spares for, or depreciate, but a service the vendor owes. The NEO platform supplies robots, station, and software; the operator pays per pick actually completed and buys availability as an outcome, not as an asset.
The pilot phase as a reliability proof on real data
The most defensible reliability evidence is not one from a data sheet but one from your own warehouse. A pilot with a few robots over typically four weeks before full call-off tests the installation under real conditions - including how it behaves at peak volume and how quickly a failure is absorbed. A project that scales up after measurable success has a different risk profile than one that demands the full installation on day one. For the reliability question, the pilot is less a delay than the actual proof.
Reliability is a system property, not a machine property
The flaw in the opening question is a single word. "What happens when a robot fails?" treats reliability as a property of one machine. It is a property of the operating model.
No serious vendor can promise that a robot never fails. What a well-designed operating model does is absorb the failure before it reaches shipping. Three layers work together: the fleet absorbs the single failure through redundancy, so throughput dips instead of breaking off; the unchanged shelving keeps manual operation available if the automation as a whole stalls; and the service contract shifts the hardware and availability risk to the vendor, who can carry and guarantee it. What is reliable is not the individual robot but the system that makes it replaceable.
Next step: check the operating reality in detail
How fleet scaling, traffic management, and the operating flow actually play out across the different AMR concepts is covered in depth in the AMR-Tiefenvergleich 2026 whitepaper (in German) - the right basis for the ops reliability question.
If you want to check how a retrofit behaves in your own warehouse under failure and peak conditions, we answer the question against your real setup in a free session: See NEO in action.