Architecture without owning the hardware.
The major clouds already run some of the largest GPU clusters in the world. Our work is to choose the right capacity for each job, design the network, storage and keys around it, and keep a run moving when one pool runs short.
Designed around failure, cost and change.
An architecture is a set of decisions about what will go wrong and what will change. We assume that nodes will fail during long runs, that capacity prices and availability will shift from quarter to quarter, and that the best model for a task today will not be the best model next year. So the design separates what the cloud supplies, which is compute, storage and the cluster network, from what we design and operate above it: scheduling, the reliability and cost layer, the experiment and data layer, and the workloads themselves. That line matters. It means a training job is defined once and behaves the same on AWS, Google Cloud, Microsoft Azure, Oracle Cloud or a GPU cloud, and that moving work is an operating choice and not a rewrite. It means data is placed where moving the work is cheaper than moving the data, keys stay with the customer, and failure is rehearsed in drills, with recovery times reported along with the method. The model layer follows the same logic: tokens from several providers are routed by task, measured per completed task and kept behind one gateway, so that no single provider becomes a dependency. The references further down are the providers’ own public designs. We build on those patterns and do not claim them as ours.
Provider names appear as design references. No provider has endorsed or partnered with Sinewgrid unless an agreement is listed on this site.
Match the way we buy compute to the shape of the job.
No single purchasing mode fits both lanes. The mix below is the starting design. Target mix and commitment levels are set per program, once agreements are executed.
| Mode | Best for | Trade-off |
|---|---|---|
| Reserved or committed | Steady robotics rollouts and long pretraining runs. | Lowest unit cost, but the commitment is a fixed obligation. |
| Short-term blocks | Scheduled pretraining or fine-tuning windows with a known start and end. | Capacity is guaranteed for the window only. |
| On-demand | Bursty biology experiments that follow lab schedules. | Flexible, but price and availability vary by region and day. |
| Spot or preemptible | Simulation rollouts, hyperparameter sweeps and other jobs that checkpoint often. | Cheapest, but jobs can be interrupted and must resume cleanly. |
The leading AI platforms under the architecture.
GPUs train and serve our own models. Agents, quant code generation, evaluation, data labeling and research assistants run on frontier models from ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google) and Grok (xAI), and that work is billed in tokens. We treat their compute and tokens as capacity to plan, route and report, the same way we treat GPU hours, and the architecture is sized for a steady, large volume of both.
| Purchasing mode | Fits | How we manage it |
|---|---|---|
| Pay as you go | Exploration and low-volume tasks | Per-project budgets and alerts |
| Committed spend | Steady, predictable usage | Sized from measured demand, then reviewed each quarter |
| Batch processing | Large offline jobs such as labeling and evaluation | Queued behind a deadline instead of run on demand |
| Priority capacity | Event-driven bursts with firm deadlines | Reserved only where the cost of a miss justifies it |
Available modes differ by provider and change over time. We confirm each against the provider's current terms before relying on it.
Route by task
Each task goes to the model that meets its quality bar at the lowest cost, with a second provider as fallback if one is rate-limited or unavailable.
Cut repeated tokens
Prompt caching, shorter context and structured outputs reduce spend, and verification steps are used where accuracy requires them, not by default.
Measure cost per task
Spend is reported per completed task and per project, so token cost sits next to GPU cost in every run report.
Protect customer data
Customer and partner data is sent to an outside model only where the agreement allows it, under the provider's data-retention terms.
No lock-in
Prompts, evaluations and tool definitions are kept provider-neutral, so switching or adding a provider is a configuration change.
Clear boundaries
Outputs from outside models are used only as each provider's terms permit. Buying tokens from a provider does not imply a partnership or an endorsement.
One architecture for each lane.
Data-heavy, bursty, tied to the lab
- Data. Omics matrices and sequence libraries live in object storage, in the same region as the compute that reads them.
- Compute. A steady pretraining pool, plus burst fine-tuning pools sized to the experiment calendar.
- Loop. Assay results return through a validated ingest pipeline and become the next training set.
Throughput-heavy, steady, tied to the robot
- Data. Teleoperation episodes and robot logs land in object storage, with scenes and assets versioned beside them.
- Compute. A simulation pool that leans on spot capacity, a training pool placed for fast interconnect, and a separate evaluation pool.
- Loop. Policies are checked on held-out real-robot tasks, and the failures become new data.
Six decisions we make the same way everywhere.
Network and placement
Keep tightly coupled jobs inside one high-bandwidth domain, and place data and compute in the same region.
Storage and data gravity
Tier storage for long sequential reads and for millions of small robot files, and move compute to the data before moving data to compute.
Identity and keys
Isolated tenants, private networking and customer-held keys for weights and proprietary assay data.
Resilience
Health checks, automatic node replacement and job resume, so a failed node costs minutes and not a run.
Cost and goodput
Report useful training time and cost per run in the same terms on every provider, so capacity can be compared fairly.
Portability
Containers, open checkpoint formats and infrastructure as code, so a run can move between providers without a rewrite.
Seven layers, and which ones we own.
The cloud supplies the bottom three. We design and operate the four above them, so the same training job behaves the same way on any provider.
Workloads
Training, fine-tuning, simulation rollouts and evaluation jobs for Cell, Robotics and customer models, packaged as containers.
SinewgridExperiment and data layer
Dataset versions, checkpoint registry, experiment tracking and lineage, so any result can be traced to its data and code.
SinewgridScheduling
Slurm for tightly coupled training and Kubernetes for pipelines and services, with queues that match each job to a capacity pool.
SinewgridReliability and cost
Health checks, node replacement, resume from checkpoint, goodput reporting and per-run cost.
SinewgridCluster and network
GPU nodes on a high-bandwidth fabric, with placement that keeps coupled jobs close together.
CloudStorage
Parallel file systems for training reads, object storage for datasets and checkpoints, local disk for scratch.
CloudCompute
GPU instances bought as reserved, short-term, on-demand or spot capacity under the provider's terms.
CloudSame job, different building blocks.
What each provider documents publicly for large training clusters. We use the differences to place work, and we keep the job definition the same.
| Provider | Managed cluster option | Scheduler | Network and storage, as documented |
|---|---|---|---|
| AWS | SageMaker HyperPod; AWS ParallelCluster | Slurm or Amazon EKS | Elastic Fabric Adapter for the fabric; Amazon FSx for Lustre for training storage |
| Microsoft Azure | CycleCloud Workspace for Slurm | Slurm | InfiniBand on ND-series GPU clusters; Azure Managed Lustre option |
| Google Cloud | Cluster Director | Slurm or GKE | Topology-aware placement; Filestore, Managed Lustre or Cloud Storage attached to the cluster |
| Oracle Cloud | OCI Supercluster | Not specified in the source we used | RDMA cluster network over RoCE v2; bare metal instances; Lustre-based file storage |
| GPU clouds | For example CoreWeave SUNK | Slurm on Kubernetes | Slurm control plane runs as Kubernetes workloads on the provider's GPU fabric |
Summaries of public provider documentation. Features and names change, and we verify each against current documentation before a design is used.
When a node fails during a long run.
The system is designed to respond in four steps. Recovery times from failover drills are reported with the method used.
- DetectHealth checks and job telemetry flag a failing node or a straggler.
- ContainThe scheduler drains the node and pauses the affected job.
- ReplaceA spare node in the same placement group takes its place, or the job moves to another pool.
- ResumeThe job restarts from the latest checkpoint, and the lost time is recorded in goodput.
Checkpoint cadence
Set from the cost of a checkpoint against the cost of lost work, per job, not as one global number.
Portable state
Checkpoints are written in open formats to object storage, so a resumed job can start in a different cloud.
Drills
Failover drills are part of the operating plan, with resume times reported by provider and the method.
The patterns we build on, in the providers' own words.
Each card summarizes a public document. Sinewgrid is not affiliated with these companies, and the summaries are ours.
Resilient training clusters
Amazon SageMaker HyperPod offers Slurm or Amazon EKS orchestration, monitors instances and automatically replaces faulty hardware. We hold our designs to the same resilience bar on every cloud.
Read the source →Slurm in your own subscription
Azure CycleCloud Workspace for Slurm deploys a Slurm cluster, shared storage and networking into the customer's own subscription, with optional Azure Managed Lustre. That ownership model suits customer-held keys.
Read the source →Scale-up and scale-out limits
Azure's ND GB300 v6 clusters join 72-GPU NVLink racks with 800 Gb/s of InfiniBand per GPU in a non-blocking fat tree. Job placement is planned around topology limits like these on any cloud.
Read the source →Hybrid by workload
Recursion runs its image data pipeline on Google Kubernetes Engine and uses cloud GPUs and TPUs for inference, alongside its own training supercomputer. Pipelines and inference are natural places to use cloud capacity.
Read the source →One interface across clouds
SkyPilot runs workloads across Kubernetes, Slurm and cloud VMs, fails over when a pool lacks capacity and picks the cheapest available zone. We evaluate open schedulers like it before building our own.
Read the source →A cloud pipeline for robot learning
AWS's published reference fine-tunes NVIDIA Isaac GR00T N1.5 on GPU instances through AWS Batch, keeps checkpoints on shared storage and evaluates in Isaac Lab. Our robotics lane follows the same shape.
Read the source →Managed Slurm or GKE clusters
Cluster Director provisions Slurm or GKE clusters from validated templates, places nodes by topology and runs continuous health checks with straggler detection. We copy the idea of qualifying a cluster before a job lands on it.
Read the source →RDMA fabric on bare metal
OCI documents an RDMA cluster network over RoCE v2 on bare metal instances for large GPU clusters. Fabric and virtualization overhead are the first things we compare across providers.
Read the source →Slurm on Kubernetes
CoreWeave describes SUNK, which runs Slurm on Kubernetes so scheduled training jobs and Kubernetes services share one cluster. It is the closest public model for our two-scheduler design.
Read the source →Plain answers.
Do you own GPUs or data centers?
No. We source capacity from the major clouds and GPU clouds, and our work is the architecture, scheduling and data layer on top.
Which cloud will you use?
Whichever fits the job, the region and the commitment. We design for AWS, Google Cloud, Microsoft Azure, Oracle Cloud and GPU clouds, and keep workloads portable so no single provider is a dependency.
Can we run in our own cloud account?
That is the intended deployment model. Running the control plane in your subscription keeps your data, weights and keys under your control, in the same way the providers' own reference designs deploy into the customer's account.
What happens if a pool runs out of capacity mid-run?
The scheduler checkpoints, moves the job to another pool and resumes. How fast that happens is something we will measure in drills and publish, not promise in advance.
Is Sinewgrid affiliated with any cloud provider?
No. Providers are named here as design references and as sources of public documentation. No provider has endorsed or partnered with Sinewgrid.
Why both Slurm and Kubernetes?
Slurm is the common scheduler for tightly coupled training, and Kubernetes suits pipelines, simulation and services. Running both is a pattern the providers above also document, so we design for it from the start.
How is customer data kept apart?
By isolated tenants, private networking and keys the customer holds. Training data does not move between customers or into shared models unless the agreement says so.
Will you run on our existing capacity?
Yes, where it fits. If you already hold reserved GPUs on a cloud, the control plane can schedule onto them, and we plan around your commitment instead of buying duplicate capacity.
Planning a training run?
Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.