Tidewell Robotics

How we designed a test we can fail

A cleanliness score graded by the company selling the robot is not evidence. So we wrote the experiment down before the machine existed: the score computed on the robot and hash-committed to a sealed record before the swabber walks in, the swabber blinded, the scoring done by the hospital's own infection-prevention team and never by us, across at least two sites so one site cannot satisfy the gate. Here is the protocol, where each part of it came from, the claim about the regulator we had to narrow on our own page along the way, what a pre-commitment does not buy, and why the threshold itself stays inside the agreement — which is also where the world's trial-registration standard leaves it.

Insight · 14 September 2026 · 12 min read · Tidewell Article Crew, edited by Timothy Mo

A cleanliness score graded by the company selling the robot is not evidence, and everyone in the room knows it. Hospitals have measured environmental cleaning for years with ATP swabs and fluorescent markers. The instruments are not the hard part. The hard part is the arrangement around them: who computed the score, who took the swab, and whether anyone fixed what counts as a pass before seeing both.

We are building a restroom cleaning robot, C1, and nothing of it is built: no part on order, no site secured, no swab taken, nothing measured on a machine. That is why the experiment that can kill it is already written down, and why its shape is published on the Cleaning page and the roadmap.

The case for publishing a design now, rather than beside a result, is put better by the body that enforces it than by us. The International Committee of Medical Journal Editors asks editors to require registration of clinical trials "at or before the time of first patient consent for enrollment", with a minimum 24-item data set lodged "at the time of registration and before enrollment of the first participant". It states why in one sentence: "The purpose of clinical trial registration is to prevent selective publication and selective reporting of research outcomes, to prevent unnecessary duplication of research effort, to help patients and the public know what trials are planned or ongoing into which they might want to enroll, and to help give ethics review boards considering approval of new studies a view of similar work and data relevant to the research they are considering" [read 11 September 2026; the page carries no revision date]. Then the sentence this article is built on:

"Retrospective registration, for example at the time of manuscript submission, meets none of these purposes."

None of the four. A protocol produced after the data has arrived describes what was done; as a commitment it is worth nothing. That applies to a vendor's acceptance trial at least as hard as to a drug: the vendor has a commercial interest the trialist does not.

What we wrote down, in the order it happens

From the Cleaning page today, in the section that opens "Nothing is built yet":

"The test is blinded, and we do not score it. The per-fixture score is computed on the robot and hash-committed to a sealed record before the swabber enters the room; the swabber is blinded to the robot's score; scoring is done by the hospital's infection-prevention team, or by an independent contractor engaged by us but reporting to them, and never by Tidewell engineers. The swab and marker data are the hospital's own. The two tables are locked before they are joined. We receive a de-identified paired table under a data-use agreement, and no site is named without written permission; the hospital's infection-prevention lead signs an acceptance statement, or does not."

The order is the content: the two tables are locked before they are joined, and the infection-prevention lead signs an acceptance statement, or does not.

Two structural commitments sit around that sequence, and both cost us something. We do not score at any stage, so the number the programme lives on is produced by people we cannot instruct. And the test "spans at least two sites by design, so one site cannot satisfy the gate" — one enthusiastic ward cannot carry the product.

Acceptance is not permanent either. A manual control sample continues, and if disagreement between the robot's record and the manual audit exceeds a rate agreed in advance, acceptance lapses automatically: manual audit resumes at its previous frequency, no negotiation, no decision by us. The default is the unusual part. Most procurement acceptance arrangements require the customer to actively withdraw, so a bad month leaves acceptance standing [inference — our reading of procurement practice, not a sourced survey]. Here a bad month ends it, and nobody has to win an argument with the vendor first.

Fixed in writing before anything runs

The outcome, by namethe name of the outcome, not an abbreviation
The metricthe metric or method of measurement used, as specific as possible
The timepointthe timepoint or timepoints of primary interest
The masking schemewhether masking is used and, if so, who is masked
Sealed — the thresholdagreed in writing with the hospital's infection-prevention team; not published. The envelope is inside this region, and it is empty.

The shape is the registration standard's, not ours. The first four boxes are what the WHO Trial Registration Data Set requires, and what ICMJE requires lodged before the first participant is enrolled; the fifth is what the data set does not require, and what we do not publish.

  1. Computethe per-fixture score is computed on the robot
  2. Commithash-committed to a sealed record before the swabber enters the room
  3. Blindthe swabber is blinded to the robot's score
  4. Runthe swab and marker data are the hospital's own
  5. Scoreby the hospital's infection-prevention team, or a contractor engaged by us but reporting to them, and never by Tidewell engineers
  6. Lockthe two tables are locked before they are joined
  7. Reveal, and the acceptance statementwe receive a de-identified paired table under a data-use agreement; the infection-prevention lead signs an acceptance statement, or does not

The order is the content: each step is closed before the next one can see anything.

  • Fixed before the run, in writing, and not revisited afterwards.
  • Our side: the robot, and the record it seals.
  • A step we do not perform and cannot instruct.
  • The human gate: the bar, and the acceptance statement. Both belong to the hospital’s infection-prevention team.
A design, not a record: nothing here has been run. C1 is not built, no part is on order, no site is secured and no swab has been taken. All seven steps are quoted from our own Cleaning page. The region above it is not our invention: the WHO Trial Registration Data Set fixes the outcome, the metric, the timepoint and the masking (the items ICMJE requires lodged before the first participant is enrolled), and fixes no success threshold at all, which is why the sealed envelope sits inside the region rather than missing from it. The envelope is empty by design and stays empty: the bar is the hospital infection-prevention team's to agree, in writing, before any test runs, and no value for it appears anywhere on this site. No number appears in this figure.

Where each part came from

The bar this gate is aimed at is Singapore's. The National Cleaning Standards for Acute Healthcare Facilities 2024 was developed by the National Infection Prevention and Control Committee, commissioned by the Ministry of Health, and it says what it is: "This document serves as a checklist for self-assessment of the cleaning standards and environmental hygiene plan in an acute healthcare facility." Its introduction says "selected standards will eventually be incorporated into the relevant regulations (e.g. Acute Hospital Regulations) or licensing conditions" — future tense, in the document's own words. It defines two grades of element: "Core elements define activities fundamental for environmental cleaning. Expected elements identify good-to-have activities…"

Two elements reach our machine. Element 3.3.4 is a Core element: "Technical audits including visual assessment and at least one of the following tools: residual bio burden or environmental marking should be undertaken regularly." Element 3.3.6 is an Expected element, quoted whole: "Institutions take into consideration audit technologies that use objective evidence-based methodology to support the subjective measurement and efficacy of the cleaning process."

Our own page overstated both of those until 11 September 2026. It opened that paragraph "MOH already mandates instrument evidence" and closed it by calling the gate a bar the regulator had already written. A self-assessment checklist is not a mandate, a good-to-have element is not a requirement, and "will eventually be incorporated" is not a rule in force. The page was corrected that day and carries an inline note saying what it used to say. We point at it rather than stepping over it: this article's subject is a claim narrowed to fit its evidence, and the nearest example is ours.

Element 3.3.5, itself an Expected element, defines what an audit process should contain: "technical audit (checks and scores cleanliness outcomes against the safe standard), efficacy audit… and external audit." Our gate is a technical audit in that three-way taxonomy and nothing else. The C1 page asks whether the robot's cleanliness score correlates with ATP and fluorescent-marker audits "closely enough for a hospital infection-control team to accept it as the visual assessment element of their audit. If the answer is no, the programme stops." MOH's own wording asks for technology that supports the subjective measurement rather than replacing it. The bio-burden and marking element is not replaced: it continues at the frequency the hospital sets.

Registered Reports go further than registration itself. The Center for Open Science names the two features a journal policy must have to qualify: "Peer review occurs prior to observing the outcomes of the research", and an in-principle acceptance "that will not be revoked based on the outcomes, but only on failings of quality assurance, following through on the registered protocol…" Over 300 journals offer it, either as a regular submission option or as part of a single special issue [single source; the page is undated, read 11 September 2026, and the count moves — a search summary in the same pass still said 260]. We submit nothing to a journal. What we took is the structure: whoever decides that a result counts commits to the criterion before the result exists, and cannot withdraw because they dislike the outcome.

What the commitment does not buy

The hash is the part that sounds strongest and is the weakest. A commitment proves that a value existed at or before a time, and that the value revealed later is the same one. It removes exactly one kind of cheating: adjusting the robot's score once the swab result is known. It proves nothing else [inference]. Not that the value was computed correctly — a hash over a wrong number is a valid commitment to a wrong number. Not that the scoring algorithm was fixed between rounds; if the method can change, every commitment is honest and the sequence is not. Not that the fixture set was fixed: choosing which fixtures enter the paired table after the early rounds defeats the arrangement without breaking a hash. And not that every committed value was disclosed rather than the best of several, unless the count of commitments is fixed in advance.

The work is done by a written agreement, made with people who do not work for us, before any data exists.

That agreement will have a witness who is not us; the hospital keeps its copy. The published protocol, today, has none. The evidence that it was fixed before any data existed is that it appears on a website we own, edit and deploy, and the section above reports us correcting a paragraph of that same page on 11 September 2026. A commitment whose only witness is the committer can be revised by the committer, and the git history is ours too. The remedy is a copy lodged where we cannot edit it — with a registry, or with the hospital that signs the agreement — and as of 14 September 2026 no such copy exists.

That leaves the weakest joint in our own independence. The published protocol offers two scoring arrangements, and they are not equivalent. The hospital's own infection-prevention team has no commercial relationship with us; a contractor engaged by us but reporting to them does, and "reporting to them" is a reporting line, not the source of the money. The hospital's own team is preferred. The fallback exists because an infection-prevention team's time is the scarcest resource in a hospital, and a gate that depends on their unpaid labour may never run at all. It is the weaker of the two, for the obvious reason [inference].

Two sites are two sites, not a sample. And a pre-registered gate can still be a badly chosen gate: fixing the question early does not make it the right question. What it cannot be is a question chosen after the answers were in.

Nor does it buy novelty. Robotics pre-registers a great deal, in the corner of it that runs clinical and human-subjects research, where journals make registration a condition for clinical trials: on 11 September 2026 there were 854 OSF registrations with "robot" in the title and 1,993 ClinicalTrials.gov studies matching the term, and the titles we inspected are clinical and human-subjects research run by universities and hospitals. What we could not find on those two registries, on that date, is a vendor pre-registering the acceptance criterion for its own product — "disinfection robot" returns nothing on ClinicalTrials.gov, and "cleaning robot" returns a single endotracheal airway-cleaning device. The two calls for papers we read for the ACM/IEEE International Conference on Human-Robot Interaction, 2026 and 2027, contain no pre-registration or Registered Reports track at all [inference — our count of the fetched pages].

The bounds belong in the same breath as the claim. AsPredicted cannot be searched for an absence at all. Its own description says "Pre-registration remains private until an author makes it public", and public lookup is by number only, from a published paper. ACM Transactions on Human-Robot Interaction's author guidelines returned HTTP 403 to us, and that is the likeliest place in robotics for a Registered Reports policy to exist. We did not read the ICRA, IROS, RSS or CoRL guidelines, and searched no tender documents, no other registry and no non-English source. The idea is also live in the field next door: Preregistration for Experiments with AI Agents, arXiv 2606.11217, Michelle Vaccaro, 3 May 2026, argues for carrying pre-registration into experiments run on AI agents. That is experiments on agents, not product acceptance, so it is not a counterexample — but pretending nobody had the thought would be worse than having it second.

WhoScore computedValue committedSwab takenManual scoringTables lockedTables joinedAcceptance statement
The robotcomputes the per-fixture scorewrites it to the sealed record
The Tidewell engineerit is our machine's output, and the protocol does not blind us to itknows the commitment exists; cannot alter itthe swab readingthe manual score — and does no scoring, at any stageeither tablereceives a de-identified paired table, under a data-use agreementreceives it, or its absence; cannot issue it
The swabberthe robot's scorethe committed valuetakes the swab and marker readings
The hospital's infection-prevention teamthe swab and marker data are theirsscores it — or an independent contractor engaged by us, reporting to themsigns it, or does not
Whoever joins the tablesboth tables are locked firstjoins them

A design, and nothing in it has been run: no swab has been taken and no cell contains a value — this is who knows what, not what anyone found. A struck-through cell is something the published protocol blinds that role to. An em dash is not a claim that the role knows nothing; it means the published protocol does not speak to it, and the difference matters in the last row: the written protocol locks the two tables before they are joined and names nobody as the party who joins them, so the row is drawn because the step exists, not because a role has been assigned. Every cell is from the Cleaning page's protocol as it stands today, the locking step and the acceptance statement included.

What stays inside the agreement, and what we publish

There is no threshold in this article and there will not be one. The reason is already on our own diagram: "No numeric ATP or marker threshold exists in the Singapore standard, so the bar is set with the hospital's own infection-prevention team, in writing, before any test runs." The document sets cleaning-time and staffing norms in its annexes, but nowhere in its twenty-four pages does it set an ATP reading, a marker-removal criterion, or any other pass mark for a cleanliness audit [inference — our reading of the document]. There is no national number to fall back on. The bar is somebody else's to agree, and a vendor publishing one in advance would be announcing a number the party who has to accept it has not seen.

That would still be a dodge if the practice we borrow from did the opposite. It does not. The WHO Trial Registration Data Set, version 1.3.1, is the world's reference for what "registered" means. For each primary outcome it requires "The name of the outcome (do not use abbreviations) / The metric or method of measurement used (be as specific as possible) / The timepoint(s) of primary interest", and under study type it requires "Masking (is masking used and, if so, who is masked)". What you will measure, how, when, and who is blinded, all fixed before the first participant. It requires no success threshold: twenty-four items, and not one of them is the value at which the trial will be called a success. Our protocol fixes those same four things and leaves the threshold inside the written agreement, which is the shape the registration standard itself takes rather than a departure from it [inference — the comparison is ours; the items are verbatim].

What publishes, then, in future tense and with no date attached: the protocol as committed, with its statistic, blinding scheme and roles; the paired table as received from the hospital, de-identified, under the data-use agreement, no site named without written permission; and the result, pass or fail, with the infection-prevention lead's acceptance statement or its absence.

A success rate without a denominator governs how we report a number once we have one. This piece sits upstream of it: that one asks whether a number can mean anything, this one fixes the question before there is a number.

The gate is at month 10. If the correlation is not there, our pages already say what happens, and they say it flatly because it was decided before there was anything to lose. "If the answer is no, the programme stops." Failing it "ends the product rather than delaying the phase". Those two sentences are the only thing that makes this a test we can fail, and they are worth exactly as much as the dates on which they were written.