All domains / Exploring

Sinewgrid Autonomous trains world models for robots and autonomous systems.

Video-scale models that predict how a scene evolves can generate training scenarios for robots and vehicles, especially rare ones that logs do not contain. This domain is at the exploring stage.

FIG. · A generated scene with predicted paths that widen with time. Schematic, not a real recording.
Overview

Generating the cases the logs do not contain.

Robots and vehicles fail on the rare cases: an unusual obstacle, odd lighting, a sensor fault, a situation that happens once in a million miles. Logged runs rarely contain enough of them to train on, and collecting them for real is slow and unsafe. A world model learns to predict how a scene evolves from video and three-dimensional sensor data, and once trained it can generate new scenarios, including hard ones, for a policy to practice on. Sinewgrid Autonomous would build these models. The compute demand is large on both sides. Training on video moves far more data than training on text, so storage and interconnect matter, and generating a long, consistent scene takes many iterative steps for every clip.

The work is a loop that does not stop. Logs the partner is entitled to use are ingested and auto-labeled, with a sample checked by people, the world model is trained, rare and hard scenarios are generated and checked for realism, and policies trained with them are scored on real held-out cases. We keep a hard line between generated and real: generated scenes are labeled as such, and no policy is judged on generated data alone. We would report a realism check, the change in a policy’s score on real held-out cases after training with generated scenarios, and the cost per generated hour. Training runs on a pool with fast interconnect, generation on interruptible capacity and evaluation on its own pool. This domain is at the exploring stage.

Why it needs compute

Where the capacity goes.

Video is large

Training on video and three-dimensional sensor data moves far more data than text, so storage and interconnect matter.

Generation is expensive too

Producing long, consistent scenes needs many iterative steps for every clip.

A data loop that never stops

Logged runs are auto-labeled, rare cases are generated, and the results feed the next training run.

Data

What the models learn from.

Logged runs

Drive and robot logs the partner is entitled to use, in open formats such as MCAP.

Simulation

Simulated scenes in OpenUSD, for example with NVIDIA Isaac Sim, to cover what logs do not.

Held-out real tests

A fixed set of real-world test cases that models never train on.

How a program runs

From a question to a checked result.

  1. Collect and labelLogs are ingested and auto-labeled, with a sample checked by people.
  2. Train the world modelThe model learns to predict how scenes evolve.
  3. Generate scenariosRare and hard cases are generated and checked for realism.
  4. Evaluate on held-out dataPolicies trained with them are scored on real held-out cases.
Compute shape

Heavy, bandwidth-bound training

A pool with fast interconnect for training, interruptible capacity for generation, and a separate pool for evaluation.

See the architecture
Limits we hold

Generated is not real

Generated scenes are labeled as such, and no policy is judged only on generated data. Real held-out tests decide.

Read the commitments
Where we would start

A first engagement, scoped small.

First engagement

One environment, such as a warehouse floor or a road segment, and one kind of hard case to generate.

What the partner brings

Logged runs they may use, the target platform, and a held-out set of real test cases.

What we bring

The world model, scenario generation and the evaluation against real held-out data.

What we would report

Measures we commit to publishing.

Each result is reported against these measures, which are defined before the first run.

Realism check

How often reviewers or a trained classifier can tell generated scenes from real ones.

Real held-out gain

The change in a policy's score on real held-out cases after training with generated scenarios.

Cost per generated hour

Capacity used per hour of generated scene, by purchasing mode.

Questions we expect

Plain answers.

Is generated data a substitute for real data?

No. It fills gaps. A policy is judged on real held-out cases, never only on generated ones.

Do you build vehicles or robots?

No. We build models and the compute around them. Hardware comes from partners.

Which formats do you use?

Open ones, such as OpenUSD for scenes and MCAP for logs, so partners can bring data in and take it out.

Public work such as NVIDIA Cosmos and Isaac Sim shows the direction. They are referenced as context, not as partners. Related: Sinewgrid Robotics: shares the simulation and held-out evaluation pipeline. Status: Exploring

Planning a training run?

Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.

[email protected]