Wiping down bottles
contributor capture · 1080p source
Household robots are not held back by their hands. They are held back by never having watched anyone do the work. Every clip of an ordinary cleanup is a lesson they cannot get from the internet, from a simulator, or from each other.
Drag to turn it. This is the thing your twenty seconds is teaching.
You already make a mess. You already clean it up. The same two minutes, filmed, is what teaches a robot how.
The hardware problem is largely solved — joints, grip, balance, battery. What a household robot still lacks is the ordinary judgement a person spends no thought on at all. That part is not engineered. It is learned, from examples, one at a time.
What to gather before what. Sweep the rice up first or the water spreads it further — nobody ever writes this down.
Order of operationsWhich thing is a tool, which is rubbish, which goes back in the cupboard.
Semantic placementWhat to do when the first attempt smears it instead of lifting it.
Failure dataWhen a surface is actually clean, rather than merely wiped once.
Terminal stateFour clips from our own capture set, with the decision the agent published for each. Two were kept. Two were not, and the rules that failed are printed on them.
Read the full ledger ↗contributor capture · 1080p source
contributor capture · 1080p source
contributor capture · 2160p source
contributor capture · 2160p source
One in two of these did not make it, which is close to the discard rate the whole industry runs at. The difference is that the two refusals are on this page rather than deleted from a bucket, each one naming the rule that stopped it. A contributor can see what to change. A lab can see what it is not being sold. The duration on each card is the clip as submitted; where that runs long, the player shows a trimmed excerpt of it.
Every figure below is something you can check on this page today. We are not going to print a total we cannot evidence — the whole point of the ledger is that the numbers hold up when somebody looks.
Named checks, each answered separately in writing by the agent before any verdict is computed.
From upload to a published verdict with the reasoning attached. No human reviewer in the loop.
Accepted or refused, every decision goes on the record with the rule that produced it.
One open for contributions now, five more specified and waiting in the queue behind it.
When there are ten thousand hours in the set, that number goes here and the ledger will back it. Until then this row stays small and true.
Eight named checks. Your clip is judged against words you can read before you film.
Rubric v1The agent answers every rule in writing. It never gets to announce the verdict.
BRAIDAccepted or refused, the reasoning goes on the record where anyone can read it.
Open ledgerEvery clip keeps its contributor record. The lab that trains on it knows where it came from.
ProvenanceA language model learned from text a person would need a hundred thousand years to read. A robot arm has nothing comparable to learn from. Controlling joints is harder than producing sentences, and no simulator gets the physics of a wet cloth on a counter right. The only source is a person doing the thing, on camera, once.
That footage barely exists — so the largest robotics companies on earth are now paying to create it from scratch.
Committed over twelve months to data and compute, after launching a platform that pays people to film household work.
Forbes, August 2026Videos uploaded to that one platform so far, at a rate of thirty minutes of footage every second, from 108 countries.
Forbes, August 2026What robotics companies already spend annually buying real-world manipulation data from third parties.
MIT Technology Review, 2026Footage gathered by a single data contractor. Others run thousands of workers across more than fifty countries.
MIT Technology Review, 2026
Filming is not training. Footage becomes training only when something can tell a good demonstration from a bad one — and say why.
The part every closed collection app leaves outBetween twenty and thirty per cent of collected footage never becomes usable training data. Hands drift out of frame, the task is half finished, the light goes. Someone pays for those hours anyway.
One folding policy went from eight per cent success to eighty-three on the same two hundred hours, simply reweighted by a model that could judge quality. The bottleneck is not volume. It is judgement.
Grading failed attempts instead of deleting them took real-world success from thirty per cent to ninety in two hundred episodes. Most pipelines throw that data away by default.
Nothing here asks you to stage anything. You make the mess you were going to make, you clean it the way you always clean it, and the recording is the only thing you do differently.
See the full brief ↗On a stand, or propped against a jar. It must not move while you film, and both your hands need to stay inside the frame.
The one you were going to make anyway — the dal that goes over, the flour, the tea. Then clear it exactly the way you always do. No demonstrating, no slowing down. Twenty seconds is usually the whole thing.
Send the clip straight from your camera roll. The agent samples frames, reads it against all eight rules, and comes back in under a minute.
Accepted, and your cleanup joins the training set with your name on it. Refused, and you get the exact rules that failed and the sentence that decided each one. Nothing is left unexplained.
It is the first task because it is the one a robot fails hardest at: a mess has no fixed shape, no fixed place, and no instruction manual. Watching a person decide what to gather first is the whole lesson.
Spilled, contained, clear. Three states the agent has to tell apart in your clip — and so will the robot.
Every one of these is answered separately, in writing, before anything is decided. Seven of them can refuse a clip. The eighth only ever raises its grade.
Containment first is scored rather than enforced. Cleaning it the other way round is still useful — how people actually sequence a mess is a large part of what we are here to collect, and a rule that punished the unusual order would quietly delete it.
All four sit on a fixed surface inside one frame. Mopping and vacuuming are deliberately excluded — they break the framing rule, and a clip that breaks framing teaches a robot nothing.
See the setup ↗The middle one is why this exists. A closed collection app tells you your clip was rejected. This one tells you which rule, in whose words, and what measurement overruled the model.
clip a41c8e07 · 38.0s
All seven required checks cleared. Rice was gathered toward the centre before the cloth came out, so containment scored too.
→ added to the training set
clip 7b20d914 · 11.0s
→ not accepted · film it again
clip c93f1a55 · 41.0s
Every rule cleared, and it was still refused. Too close to clip a41c8e07 — distance 0.11 against a threshold of 0.40.
→ not accepted · try a different mess
SERV Reasoning answers each rule separately with a written reason. Code computes accept or refuse from those answers. If the model volunteers an opinion that contradicts its own answers, the code wins and the disagreement is recorded.
Containment never causes a refusal — it sets a grade. Reweighting an identical dataset by a quality model once took a folding policy from 8% to 83% success, so a graded set is worth more to a lab than a merely admitted one.
Duration comes from the file, not from an opinion about the file. When the two disagree, the file wins and the override appears on the record.
If the reasoning pass fails to return a usable verdict, your clip is queued for another look. Nobody is refused because our software misbehaved.
Recent copyright litigation has separated the training activity, often defensible, from how the data was acquired — which is where the liability actually sits. Buyers now ask for a rights record per asset rather than a catalogue-level claim. Ours is built in.
What gets picked up, in what order, where each thing belongs, whether a mess is contained before it is absorbed, when the tool changes. Sequencing and semantic placement — the part long-horizon household policies are weakest at, and the part no force sensor records.
Not force data. Surface-wiping datasets are specified around 6-axis force/torque at 500Hz, collected by physically guiding a robot arm, because wiping means regulating normal force in a 2–15 newton band. A phone cannot record newtons and we do not pretend otherwise. This is pre-training data, not a teleoperation substitute.
Each opens as the one before it fills. They get harder, longer and less scriptable in that order — which is exactly the direction household data is shortest in.
Long-horizon and semantic: what goes to the sink, what goes to the bin, what goes back in the cupboard. A robot that gets this wrong throws away your dinner.
Teaches · object destinationNo correct action sequence exists, so it tests planning rather than motion. Two people will do it two different ways and both are worth keeping.
Teaches · planningA large deformable object under tension, handled with two hands. Simulation is close to hopeless at this and always has been.
Teaches · deformablesObject categorisation, placement and two-handed coordination, all inside one fixed frame with a legible end state.
Teaches · categorisationTool use with an unambiguous before and after, and a tool that has to be re-gripped halfway through.
Teaches · tool useIf you think a household task is under-represented, say so. Tasks are JSON files in a public repo; adding one is a pull request, not a roadmap meeting.
Open · MIT repoA dataset nobody can audit is a dataset nobody can buy. So here are the gaps, before anyone has to ask.
The robots arriving in kitchens over the next decade will have learned from ordinary people doing ordinary things on camera. The only question is whether anyone kept a record of who taught them, and whether the lesson was any good. The agent is open source, the rules are written down, and every decision it has made is public.