Tidewell Robotics

A success rate without a denominator

The best-documented robot releases in the field publish success rates from 32 to 92 percent with no trial count behind them, and a 0.54-million-parameter policy nearly matches a 4.1-billion-parameter one on the benchmark everyone quotes. At the trial counts this field actually runs, a published success rate cannot separate a real improvement from noise. Here is what to ask a robotics vendor for instead, and the ten rules we have bound ourselves to before we publish a number of our own.

Insight · 4 September 2026 · Updated 11 September 2026 · 12 min read · Tidewell Article Crew, edited by Timothy Mo

On 30 July 2026 Google DeepMind published eleven success rates for Gemini Robotics 2 [single source]. They run from 32 percent, sweeping with a dustpan, to 92 percent, unscrewing a light bulb.

The chart's caption carries two sentences of methodology: "Each bar represents the average success rate over multiple tasks within the same skill category. For multifinger tasks we show individual task performance."

Read both. The second answers half of what we came to ask: the five multifinger figures are individual named tasks, and they are both ends of the range — 92 percent is unscrewing the bulb, 32 percent is the dustpan. The other six bars are averages over an unstated number of tasks. So DeepMind does tell you what each bar is.

What neither sentence tells you, for any of the eleven, is out of how many attempts. Nothing else on the page tells you either. The charts carry error bars, drawn with no values and no N behind them, from the best-resourced robotics laboratory in the world. Ninety-two percent of how many?

Physical Intelligence's π0.7, 16 April 2026, is thinner still [single source]. Read end to end, it states no numerical success rate in its body text at all; success appears only in bar charts labelled "Success Rate (%)", running roughly 60 to 100 percent [single source — read off the charts, not stated in prose]. Its most striking claim is a comparison: "The success rate of π0.7 on this task actually matches the 'zero shot' success rate of expert human teleoperators…" — no number on either side of the match.

These are two of the best-documented releases in the field. That is the point: if the best disclosure is this thin, the problem is the convention and not the companies. Of six outlets we mapped on 11 September 2026 — those two, read in full, plus Skild, Boston Dynamics, LionsBot and Hivebotics [background] — not one publishes an intervention rate or an uptime figure.

So a published success rate, on its own, is not procurement evidence. Not because vendors lie. Because at the trial counts this field runs, the number cannot separate a real improvement from noise, and almost nobody publishes the denominator that would let a buyer check. What to ask for instead is below, with the standard we bind ourselves to before we have a number of our own to defend.

What a denominator buys you

The counter-example is a paper that attacks the field's favourite benchmark, scores narrowly below the model it measures itself against, and publishes its rollout count, its protocol and its own seed band so anyone can see how little that gap is worth. MINERVA — Sendai, Matsushima and Iwasawa, arXiv 2609.03715, 3 September 2026 — asks how small a manipulation policy can be and still solve LIBERO. The answer is 0.54 million parameters, for 95.05 percent across the four standard suites (the abstract rounds to 95.1). That is 2.4 points below the reported LeRobot π0.5 port, which Table I names outright: 4.1 billion parameters, four-suite average 97.50 percent, a ratio of 7,700× fewer parameters. Do the division yourself: 4.1 × 10⁹ ÷ 0.54 × 10⁶ is 7,593 [inference]. The four figures are consistent, which is why we print all four.

Now the part that matters more. MINERVA publishes its denominator — "standard-LIBERO results use 50 episodes per task × 10 tasks × 4 suites = 2,000 rollouts under fixed seed" — and its own noise floor: "a single training seed moves the 4-suite average by ±1 point (the baseline spans 94.60–96.75; its long-suite score alone spans 87.6–92.8)."

The two numbers in that headline do not share a denominator. MINERVA's 95.05 percent is 2,000 rollouts; Table I's π0.5 row is the LeRobot implementation's reported result, and the table says what it rests on — "10 episodes/task, 400 rollouts". Five times fewer. The paper discloses it, which is the only reason this paragraph can exist.

The gap is 2.4 points. The declared seed band is ±1 point, and the baseline's own four-suite average moves across a 2.15-point range from the training seed alone [inference]. So the reported difference between a 0.54M policy and a 4.1B one is barely more than double the noise a random seed introduces. It survives at all because MINERVA aggregated 2,000 rollouts. At the 50-rollouts-per-task granularity the field normally reports, a cell at 95 percent carries a Wilson interval 13.3 points wide [inference], and every individual cell in that table is statistically mute.

We only know any of this because MINERVA said so. That is the case for the denominator, made by the most disciplined paper in the set rather than the least — and the same discipline turns up two further findings. A task-ID permutation probe changes only the task-ID mapping, nothing about the scene, the robot or the task, and it "reduces success to near chance": LIBERO instruction conditioning is largely selecting among memorised tasks. Under LIBERO-Plus perturbations the same policy drops to 46–56 percent.

Four orders of magnitude of parameters, 2.4 points of LIBEROFour orders of magnitude of parameters, 2.4 points of LIBEROLeRobot π0.5, 4.1B parametersas reported in Table I: 400 rollouts97.5 %MINERVA, 0.54M parameters2,000 rollouts: 50 × 10 × 4 suites95.05 %The same policy, LIBERO-Plusperturbed test distribution46–56 %097.5MINERVA, 0.54MTest set perturbedThe 4.1B comparatorPublished range
  • LeRobot π0.5, 4.1B parameters — 97.5 % — as reported in Table I: 400 rollouts
  • MINERVA, 0.54M parameters — 95.05 % — 2,000 rollouts: 50 × 10 × 4 suites
  • The same policy, LIBERO-Plus — 46–56 % — perturbed test distribution
  • MINERVA, 0.54M
  • Test set perturbed
  • The 4.1B comparator
  • Published range
Parameters and LIBERO success, from Table I of MINERVA — Sendai, Matsushima and Iwasawa, arXiv 2609.03715, 3 September 2026. Four orders of magnitude of parameters, 4.1 billion against 0.54 million, buy 2.4 points on the four standard suites; perturbing the test distribution costs between thirty-nine and forty-nine points. MINERVA names the denominator behind its own row — 2,000 rollouts, 50 episodes × 10 tasks × 4 suites at a fixed seed — and behind the comparator's, which is the LeRobot port's reported result at 10 episodes per task, 400 rollouts. It declares a ±1-point seed band on the four-suite average, which is the only reason anyone can see that the 2.4-point headline is barely more than double the paper’s own noise. RIPL’s 0.09-billion-parameter probe belongs on this axis and cannot be drawn: the audit places it “at or near the best reported result” and publishes no figure for it.

Some benchmarks are much better than others

On 2 June 2026 Jiang, Tan, Wheeler, Sun, Ayalew and Walter of RIPL (arXiv 2606.04233) audited five manipulation benchmarks and asked what they actually measure. They name four failure modes with one diagnostic each — shortcut solvability, lack of statistical significance, creeping overfitting, data-source dependence — under the cleanest principle in the paper: "A benchmark score is evidence of capability only if any policy achieving it has the capability." On LIBERO, "a simple 0.09B probe with no language encoder and no large-scale robotics pretraining scores at or near the best reported result." On CALVIN, "randomizing block poses within the training range drops performance for every tested policy" — within the training range, not outside it.

Then the number a buyer can use. "On LIBERO and SimplerEnv, only 19.8% and 19.7% of SOTA claims, respectively, are provably significant. On RoboCasa and RoboTwin 2.0, the shares rise to 53.3% and 73.7%, respectively." [single source]

That is an asymmetry, not a condemnation. Four out of five claims on the two most-quoted benchmarks cannot be shown to be real; on two others the majority can, and those two appear far less often in progress claims. The question to put to a vendor deck is therefore which benchmark, and what share of claims on it survive a significance test. Not whether benchmarks are meaningless: that LIBERO can be solved at 0.54 million parameters is a useful fact in itself.

One limit we are obliged to state. Nowhere we can read does the audit name the test behind "provably significant", or the trials per evaluation. We can quote 19.8 percent; we cannot say how it was computed.

The arithmetic, and who said it first

Toyota Research Institute said this before we did, in public, in plain English, with more evidence behind it than anyone else has assembled. TRI's Large Behavior Model study — arXiv 2507.05331, 7 July 2025, since published in Science Robotics, 2026 — ran 1,800 real-world evaluation rollouts and over 47,000 in simulation, at 50 rollouts per real-world task and 200 per simulation task. Its §4.1.5 describes the method: a sequential hypothesis testing framework that "sequentially compares paired outcomes of each evaluation trial until either a decision is reached or all the trials are consumed."

Their conclusion is on the LBM project page at toyotaresearchinstitute.github.io/lbm1/: "It's easy for noise due to experimental variation to dwarf the effects being measured, and many robotics papers may be measuring statistical noise due to insufficient statistical power." The same page quantifies it: "With 50 rollouts, for example 10 rollouts each on 5 behaviors, the resulting CI width is generally 20% to 30% absolute success rate." Badithela and colleagues (arXiv 2510.04354) say the same of practice: policies "evaluated on a small number of hardware trials without any statistical assurances."

What no source we could find supplies is the number a buyer needs: how many trials it takes to separate the effect sizes this field routinely claims. So we computed it, and print the formula so it can be checked on a calculator. For x successes in n trials, with p̂ = x/n and z = 1.96, the 95 percent Wilson score interval has half-width (z/(1 + z²/n))·√(p̂(1−p̂)/n + z²/4n²); the widths below are twice that. At an observed 90 percent, the interval width runs 38.6 percentage points at n=10, 27.3 at n=20, 17.0 at n=50 and 2.6 at n=2,000 [inference].

Two things fall out. First, a check on TRI: at n=50 our arithmetic gives 26.7 points at p̂=0.50 [inference], squarely inside TRI's published band and reached independently. Second, the number nobody prints. On the standard two-proportion test, two-sided α = 0.05 at 80 percent power, separating 90 percent from 92 percent takes 3,213 trials per arm; separating MINERVA's 95.1 percent from π0.5's 97.5 percent takes 970 [inference].

The field runs 20 to 50.

Interval width against trial count, at an observed 90 percentInterval width against trial count, at an observed 90 percentn = 1038.6 ppn = 20SO-101, 20 episodes per model-task27.3 ppn = 3022.2 ppn = 50TRI per real task; MINERVA per task17 ppn = 10011.9 ppn = 320SO-101, all four models and tasks6.6 ppn = 1,0003.7 ppn = 2,000MINERVA’s LIBERO aggregate2.6 ppA two-point improvement, the effect routinely claimed038.6Counts almost nobody runsWhere this field runs, 20 to 502 points, the claim
  • n = 10 — 38.6 pp
  • n = 20 — 27.3 pp — SO-101, 20 episodes per model-task
  • n = 30 — 22.2 pp
  • n = 50 — 17 pp — TRI per real task; MINERVA per task
  • n = 100 — 11.9 pp
  • n = 320 — 6.6 pp — SO-101, all four models and tasks
  • n = 1,000 — 3.7 pp
  • n = 2,000 — 2.6 pp — MINERVA’s LIBERO aggregate
  • A two-point improvement, the effect routinely claimed — 2 pp
  • Counts almost nobody runs
  • Where this field runs, 20 to 50
  • 2 points, the claim
Width of the 95 percent Wilson score interval against the number of trials, at an observed 90 percent, z = 1.96. The arithmetic is ours [inference], computed from the formula in the paragraph above and reproducible on a calculator. At the counts this field actually runs, a published 90 percent means somewhere between 69.9 and 97.2 percent at n = 20, and between 78.6 and 95.7 percent at n = 50. Even 2,000 trials leave 2.6 points of interval, still wider than the two-point improvement papers routinely claim. The annotations mark where real protocols sit on the trial axis; every width here is computed at 90 percent rather than at that study’s own rate, so MINERVA’s actual aggregate, 95.05 percent over 2,000 rollouts, is narrower at 1.9 points.

This article is not exempt from its own rule. Three of our best sources quote a percentage with no denominator. The RIPL audit's 19.8 percent names no test and no n. Moritz Reuss's survey of 164 VLA submissions to ICLR 2026 — the sharpest public critique of benchmark saturation there is, the one that says "LIBERO is basically solved and showing 99% vs 98% is not very helpful…" — contains one numeric proportion, "90% of papers mentioned in this post all test in either LIBERO, SIMPLER or CALVIN" [single source], whose denominator is a set the reader cannot count. MEMOBench (arXiv 2609.07047) reports "31.9% average success rate" for the strongest memory-module baseline [single source; brand new, one team, simulation] with no trial count behind it. None is dishonest; the convention is that deep — deep enough to have caught us twice inside the writing of this piece.

A summarisation of the RIPL page handed our researcher a tidy significance figure for CALVIN that exists nowhere on it, caught only by re-fetching with a narrower question. The MINERVA comparison above stood in this article, and in its chart, as a clean head-to-head until review: we had not read Table I's footnote closely enough to notice that one side rests on 2,000 rollouts and the other on 400. It reads as it does now because someone went back to the source.

What to ask for

Some people already publish well, and in a way anyone could copy. Yu and Qiu's SO-101 study (arXiv 2606.08881, 7 June 2026, two authors on low-cost hardware) states its denominator in the methods — "For each model-task pair, we conduct 20 independent real-world evaluation episodes, resulting in a total of 320 evaluation episodes across four models and four benchmark tasks" — names a failure taxonomy, and defines Recovery Rate as "N_successful recovery / N_recovery opportunity", both terms written out. A metric defined as a fraction with its denominator inside the definition is exactly what this article asks for. It reports no confidence intervals, so it is half the protocol: at n=20 with 95 percent observed, the Wilson interval runs 76.4 to 99.1, 22.7 points wide [inference]. The cheapest credible contribution anyone could make is this study with intervals.

Two more. RoboArena (arXiv 2506.18123) ran over 600 pairwise real-robot episodes across seven policies and seven institutions, where evaluators "are required to perform double-blind evaluations over pairs of policies": the contribution is the protocol, not the compute. Tommoro's Habilis-β (arXiv 2602.18813) makes interventions a headline metric, 124 tasks per hour against 137.4 seconds mean time between interventions, simulation figures kept separate.

The best disclosure of the lot is a vendor's. Dyna Robotics' DYNA-1 page leads with "a 99.4% success rate—zero interventions, full shift reliability" across 850-plus napkins folded in a 24-hour run [single source], then publishes the distribution behind it in the next paragraph: "While 98% of folds reach near-perfect quality (grade ≥3), only 75% hit our rigorous quality bar." Most companies stop after the first paragraph.

So, seven questions, short enough to paste into a tender.

  1. How many trials, and of what? The denominator behind every percentage, per task.
  2. What is the confidence interval, and by which method? Named, so it can be reproduced.
  3. The per-condition cells, not the aggregate. Which task, which site, which day.
  4. What counts as a failure, and what as a recovery? Defined in writing before the run.
  5. What is the intervention rate? Per hour, with the teleoperated hours beside it.
  6. Simulation or hardware? For each number, separately.
  7. Who ran it, and were they blind to which system was under test?

A vendor who cannot answer these is not necessarily selling a bad robot. They are selling one whose performance cannot yet be evaluated, and those are different purchases at different prices.

The rules we sign before we have a number

We have no fleet data and no published results. Every cell of our Scoreboard is empty on purpose across all six product rows, and the one target on it, taskSuccessRate: 95, is a target on an empty board — not a result, not a near-result, and not a claim until a denominator sits next to it. The promise on our Brain page, that our target for any task we sell is 95 percent success without intervention and that we publish the number, has had no protocol behind it. This is that protocol, in force from today. It governs what we may publish rather than what we must achieve, which is why every rule can be honoured with no measurements in hand.

  1. No success rate without its denominator. Every percentage Tidewell publishes carries the number of trials it was computed from, in the same sentence or the adjacent cell. A rate with no denominator does not go on the Scoreboard, and does not go in a deck.
  2. No point estimate without an interval. Every published success rate carries a 95 percent Wilson score interval with the method named, so it can be reproduced. Wilson is for proportions, and four of the five columns on our Scoreboard are not proportions: any other published quantity that is an estimate rather than a count — interventions per hour among them — carries an interval with its own method named. Where the interval is wider than the claimed effect, the interval is printed first.
  3. Per-condition, before aggregate. Rates are published per task, per site and per condition. An aggregate is published only alongside the cells behind it, never instead of them.
  4. A written failure definition, fixed before the run. What counts as a failure, what counts as a recovery and what counts as an intervention are written down before the run starts, and published with the result unchanged. We adopt the SO-101 shape: a named failure taxonomy, and recovery reported as a fraction with its denominator inside the definition.
  5. Interventions counted, never netted out. Interventions per hour is published beside every success rate, and the teleoperated hours are published with it. A task completed after a human took the controls is not a success.
  6. Simulation and hardware are never mixed in one number. Every figure is labelled simulation or hardware. A simulation figure is never published as evidence of field performance, with or without a transfer argument.
  7. Who ran it, and were they blind. Every result names the party that ran it and states whether the evaluator knew which policy or machine was under test — including when the answer is "not blinded".
  8. A Limitations section on every claim. Adopted from CoRL 2026: limiting assumptions, failure modes, and what we could not test, published with the result and not on request.
  9. Video at 1x, uncut. Adopted from Holson's rule for his Humanoid Olympic Games: "a 1x speed video with no cuts… running autonomously." Any clip that is cut or sped up is labelled an illustration and carries no number. Where a clip carries a number, the clip is unedited and the intervention count is stated with it.
  10. The rule that costs something. When a number comes out worse under this protocol than it would have without it — when the interval swallows the claim, when the per-condition breakdown exposes a cell we would rather have averaged away, when an intervention count turns a 95 percent into an 80 percent — we publish the worse number, under this protocol, and we publish the better one nowhere. We do not retire a measurement because we dislike it. A published figure is corrected in place, dated, with the previous value left visible. This rule binds the same measurement: what we suppress is a laxer method applied to the run we actually did — not the aggregate, which rule 3 requires us to publish beside the cells behind it, and not a genuinely better result from a later run, which is published like any other under this same protocol, with its own denominator and interval.

The five columns behind that board — task success rate, teleoperated hours, autonomous hours, interventions per hour and fleet hours, set out on our Data page — will each carry a trial count and an interval when they are filled, across all six product rows.

Rule 10 is the only one of the ten that costs us anything, which is why it is the only one that proves the other nine. It is also, today, free: we have nothing to suppress. The bill is signed and the cost has not been paid, and the board is public, so anyone can check whether we pay it.