Spirit AI: factory success is not proof of household autonomy
An industrial robot completing a defined operation and a household helper coping with an unpredictable environment are different propositions. Spirit AI provides a useful example of why the two should not be assessed with the same headline metric.
In a September 18 Reuters interview, Spirit AI co-founder Gao Yang discussed robots at CATL and JD.com and a potential software breakthrough by mid-2027. That is an entrepreneur’s forecast, not a verified delivery date for a general-purpose home robot. [1]
Define the task before comparing success rates
An earlier Spirit AI account describes XiaoMo working on battery-pack testing at CATL’s Zhongzhou facility. The company reports more than 99% success for connector insertion. That is a result for a specific operation, not a measurement of general reliability in arbitrary environments. [2]
Assessing a deployment should start by defining success. Did the robot merely insert a connector, or also interpret the test result correctly and handle an exception? How many attempts were counted? How were retries and human interventions recorded? Without those details, comparing percentages from different demonstrations can produce a misleading conclusion.
Human assistance does not automatically invalidate a deployment
A practical assessment should ask how much exception handling costs and whether interruptions occur predictably. A narrowly focused system may be useful even with periodic supervision. Buying a home robot on the promise of complete independence would require a different standard of evidence and a different approach to failure handling.
Instead of asking only whether a robot is autonomous, buyers can ask: for which task, in which environment, for how long and with whose assistance? These are questions that can be tested in a pilot and incorporated into acceptance criteria.
Factories and homes need different acceptance criteria
This is not a simple split between an “easy factory” and a “difficult home.” Either environment can contain demanding exceptions. A useful pilot separates conditions that are fixed in advance from capabilities the robot is supposed to demonstrate. If a supplier prepares every object, pickup position and movement sequence, that result does not yet establish independent handling of changes to those conditions.
| Pilot question | What to record |
|---|---|
| Task success | Completion of the whole agreed operation, not one successful movement. |
| Interventions | Every instance of human assistance, its cause and recovery time. |
| Pace and interruptions | Full cycle time including waiting and retries. |
| Environmental changes | New objects and positions tested without additional training. |
This is an editorial evaluation framework, not a table of published Spirit AI results. It helps suppliers and customers compare the same problem instead of arguing about the meaning of “autonomy” after deployment.
An example: why a 99% figure cannot determine the cost
Consider a purely illustrative assumption: one unsuccessful attempt in a hundred. The percentage alone cannot tell us whether the robot retries moments later or stops a station until a technician arrives. One outcome creates a short delay; the other may create a lengthy interruption. This is not a XiaoMo measurement. It shows why recovery time and recovery method need to be recorded separately.
That is why a success rate needs the cost of exception handling alongside it. Buyers do not necessarily need a robot that never makes a mistake. They need to understand whether failures fit an acceptable risk level, whether they are detected and whether the next action is predictable.
What should an EU customer evaluate?
A pilot on the customer’s own process, with records of completed operations, interruptions and operator effort, would be more informative than a generic demonstration. Another customer’s deployment is a reference point, not an automatic financial result for a new site. Hardware, integration and ongoing support also need separate budget lines.
Our humanoid total-cost-of-ownership guide develops that approach. The overview of commercially available humanoids offers a separate check: distinguishing hardware with a purchasing pathway from a project that remains an announcement.
The takeaway: Spirit AI is a useful case for defining robotic applications precisely. It is not evidence that general-purpose household autonomy is a solved problem.
Looking at a different application? The RoboMorrow robot database offers a way to begin with the task rather than the appearance of the machine.
XiaoMo at a CATL station. Archival Spirit AI photograph; it does not document a new September 19–20 deployment. Image source.
Editorial analysis based on the September 18 Reuters interview and Spirit AI’s archival deployment account. Company metrics are not an independent RoboMorrow test. The CATL station photograph comes from earlier company documentation.