Approach

Two models with different compute shapes.

Sequence and omics models are limited by data curation and bursty experiments. Robot policies are limited by rollout throughput and interconnect bandwidth. A scheduler that treats both as generic GPU jobs wastes capacity on each. Sinewgrid plans capacity, storage and scheduling across clouds around these two shapes.

Why vertical compute

Match the compute to the problem.

General-purpose GPU scheduling treats every job as the same kind of thing: some number of accelerators for some number of hours. Real programs are not like that. A biology program alternates between a long, steady pretraining run and short bursts of fine-tuning and ranking after each lab batch, and its bottleneck is the quality of its data and the speed of its experiments. A robotics program is limited by how many simulated rollouts it can run, how fast its accelerators can talk to each other, and how many hours of real hardware it can get for evaluation. Give both the same generic queue and each wastes capacity in a different way. Our approach is to start from the shape of the problem and then plan capacity, storage and scheduling around it, buying steady capacity where demand is steady and flexible capacity where demand arrives in bursts. The second idea is the loop. Every real-world result, whether an assay, a robot trial, a verified simulation or a held-out test, goes back into training, so that the compute buys learning and not only throughput. The same approach carries over to the other domains we study and to the enterprise services we run on the same platform.

Cell & mRNARobotics
Primary dataSingle-cell expression matrices, perturbation screens, sequence libraries, assay readouts.Multi-camera video, joint and force logs, simulator rollouts, teleoperation episodes.
Training patternLarge pretraining on sparse, high-dimensional tables and sequences, then fine-tuning per program.Vision-language pretraining, imitation learning, then reinforcement learning on continuous rollouts.
What limits throughputData curation, input pipelines for very large sparse matrices, and demand that follows experiment schedules.Rollout generation, simulation throughput, and sustained interconnect bandwidth for large batches.
How progress is checkedWet-lab validation. Results return to training as new data.Held-out tasks on real robots. The sim-to-real gap is measured per task.
Safety reviewBiosecurity screening of generated sequences and tiered model access.Physical safety cases with defined stop conditions before real-world use.
Closed loops

Real-world results return to training, and the capacity stays busy.

Each lane sends measurements back into the next training run. The shared control plane handles scheduling, checkpoints, evaluation and telemetry for both, across whichever clouds hold the capacity.

Wet-lab assaysperturb · transfect Single-cell data.h5ad · scRNA-seq Train & fine-tunecell state · mRNA Design candidatesUTR · CDS · LNP CELL & mRNAclosed loop
SimulationOpenUSD · physics Rollouts & demosMCAP · teleop Train policyVLA · RL post-train Sim-to-real evalheld-out tasks ROBOTICSclosed loop
Shared control plane SchedulerCheckpointsEvaluationTelemetryAccess control

Planning a training run?

Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.

[email protected]