Sinewgrid Robotics trains policies that move real hardware.
Sinewgrid Robotics trains vision-language-action policies for manipulation and whole-body control. Pretraining draws on video and simulation at scale. Post-training uses demonstrations and reinforcement learning on tasks that can be measured on real robots.
Illustrative ordering of data volume. Not measured values.
- PretrainVideo and simulation
- ImitateTeleoperated episodes
- Post-trainReinforcement learning
- EvaluateHeld-out real tasks
Why robot learning is a data problem.
A language model can learn from a large share of the written web. A robot has no such source. Demonstrations recorded by a human operator are exact but slow and costly to collect, human video is broad but carries no record of the actions taken, and simulation is plentiful but only an approximation of real physics. Sinewgrid Robotics treats these three sources as a pyramid and uses each for what it does well. Pretraining draws on video and simulation at scale to give the policy a general sense of objects and motion. Imitation on teleoperated episodes then teaches precise actions, and reinforcement learning in post-training improves the policy on tasks that can be measured. The result is a vision-language-action policy for manipulation and whole-body control.
What matters at the end is what happens on real hardware. Each task family has a fixed held-out set, a checkpoint is scored on real robots, and the gap between simulation and the real result is reported next to it, so a good simulated score is never mistaken for a real one. The compute is uneven: simulation rollouts can run on interruptible capacity, training needs a pool with fast interconnect, and evaluation runs on its own pool so it never competes with training. We work in open formats, with logs in MCAP, scenes in OpenUSD and demonstrations in a LeRobot-style dataset layout, so a partner can bring data in and take it out. Every policy has defined stop conditions and is tested on held-out tasks before it moves hardware outside the lab. The first body we target is chosen with the first partner.
Tasks that can be scored on a real robot.
Each task family has a fixed held-out set. A checkpoint is judged on real hardware, and the gap to simulation is reported alongside.
| Task family | Example | How it is scored |
|---|---|---|
| Tabletop manipulation | Pick and place objects of varied shape. | Success rate over a fixed set of held-out trials. |
| Articulated objects | Open a drawer, a door or a lid. | Success rate and number of recoveries after a failed attempt. |
| Bimanual assembly | Fit two parts together with two arms. | Task completion within a time limit. |
| Mobile manipulation | Fetch an item from a shelf and return it. | Success rate in layouts not seen in training. |
More than one body
Sinewgrid Robotics targets arms, bimanual setups and whole-body platforms. The first embodiment is chosen with the first design partner.
Open formats end to end
Robot logs in MCAP, scenes in OpenUSD, and demonstrations in a LeRobot-style dataset format, so a partner can bring data in and take it out.
Simulation as the base of the pyramid
Large simulated rollouts, for example in NVIDIA Isaac Sim and Isaac Lab, supply the volume that real robots cannot.
Steady, and limited by rollouts and bandwidth
Simulation rollouts can run on interruptible capacity. Training needs a pool with fast interconnect. Evaluation runs on its own pool, so it never waits behind training.
See the Robotics reference designA safety case before any deployment
Every policy has defined stop conditions and is tested on held-out tasks before it moves real hardware outside the lab.
Read the commitmentsPlanning a training run?
Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.