Metric 3D environment
Dense coloured point cloud in real metres, with a pose for every frame. Global metric scale, held across the whole session.
A robot foundation-model team
We trained robot foundation models before we built this pipeline, so we know what data has to do — and we do not ship hours of footage and wish you luck. Every batch is trained on, run on a real robot, and delivered with the checkpoint and the eval that prove it worked. That loop is the product.
Every hour we deliver is shot in the wild — in real homes, on real beds piled with real laundry, in the light and clutter people actually live with. No lab bench, no staged set, nothing arranged for the camera. That is where robots are going to work, so it is the only place worth mining.
What comes out of a real room is ore. The value is in there, and so are the idle minutes, the drifting sensors and the hours nobody labelled. Most robot data is sold in exactly that state, with the refining left to whoever buys it. We refine it before anything leaves.
both ego views, the approved instruction, its action steps and the grounded moments
Our own feedforward VIO tracks the rig from the first frame to the last — continuously, with no reset and no external tracker. It holds global metric scale for the whole session: a metre in the first minute is a metre in the sixth hour, in one room or across a whole building.
It keeps that accuracy through the conditions that end a classical run — people crossing the frame, the gripper filling the view, a doorway blowing out the exposure, a light switched off mid-take. No markers, no motion capture, nothing installed in the room.
20 recordings replayed against a reference trajectory, with the per-recording error
Every sensor on the rig lands on a single time base — both camera bodies, the IMU, the gripper encoder, the microphones. Our synchronisation works across separate devices that never share a cable and across sensors running at completely different rates.
Theoretical alignment error is under a millisecond: a twentieth of a camera frame at 50 fps, a fifth of an IMU sample at 200 Hz. It holds for the length of the session rather than only at the start, and every session is checked before it leaves the pipeline.
Dense coloured point cloud in real metres, with a pose for every frame. Global metric scale, held across the whole session.
4K per lens at 50 fps, back to back — the full sphere, including whatever is behind the operator.
Accelerometer and gyroscope, calibrated per device and shipped with the extrinsic to the camera it sits behind.
Position and orientation of the gripper at 50 Hz in the same metric frame as the map, with the jaw angle alongside — a demonstration you can replay.
Stereo sound for contact events and speech, on the same time base as every other stream.
Who shot it, on which rig, when, where it got to in the pipeline — carried with the clip instead of living in a spreadsheet.
Every capture ships with its own calibration — per lens, per body, per rig. Nothing is assumed shared.
In-the-wild footage doesn't fit a fixed vocabulary. Every clip is decomposed by a vision-language model into a task, ordered minitasks, and per-moment affordances grounded back onto the frame — then a reviewer confirms it. Below is one real capture with its real annotation, playing.
—
How far through the task this moment is — not just what is happening in it.
Where to act and which way the action goes, tracked live against the clip.
Free-form targets written by the model, not picked from a closed class list.
There is no closed form between a data distribution and how a policy behaves on a real robot. It can only be measured. So the measurement sits inside the line rather than at the end of it: a batch trains the moment it is complete, only the parameters that changed go back out to the rigs, and the operator sees what the last hour did to the policy before the next hour starts.
The gap between "the model fails at this" and "we are out collecting exactly that" closes in hours instead of weeks. Nobody collects blind, and nothing is judged three weeks late by a training run that diverged.
Engineering compresses how many rounds of exploration a capability takes. It cannot compress that to one — the last stretch has no analytic solution, only measurement. What it can do is make every round fast, cheap and repeatable.
A batch that will not help is caught on site, at the moment it is shot, by the model it was meant to improve. The rework, the re-shoots and the discarded hours go away with it.
The dataset, the training recipe, the checkpoint trained on it and that checkpoint's real-robot eval report — so what arrives is a measured result, not hours of footage that should help.



Access
Datasets, rigs, or the whole pipeline running inside your own infrastructure — start with the task you're stuck on.