Cloud architecture

Architecture without owning the hardware.

The major clouds already run some of the largest GPU clusters in the world. Our work is to choose the right capacity for each job, design the network, storage and keys around it, and keep a run moving when one pool runs short.

Why this design

Designed around failure, cost and change.

An architecture is a set of decisions about what will go wrong and what will change. We assume that nodes will fail during long runs, that capacity prices and availability will shift from quarter to quarter, and that the best model for a task today will not be the best model next year. So the design separates what the cloud supplies, which is compute, storage and the cluster network, from what we design and operate above it: scheduling, the reliability and cost layer, the experiment and data layer, and the workloads themselves. That line matters. It means a training job is defined once and behaves the same on AWS, Google Cloud, Microsoft Azure, Oracle Cloud or a GPU cloud, and that moving work is an operating choice and not a rewrite. It means data is placed where moving the work is cheaper than moving the data, keys stay with the customer, and failure is rehearsed in drills, with recovery times reported along with the method. The model layer follows the same logic: tokens from several providers are routed by task, measured per completed task and kept behind one gateway, so that no single provider becomes a dependency. The references further down are the providers’ own public designs. We build on those patterns and do not claim them as ours.

AWSGoogle CloudMicrosoft AzureOracle CloudGPU clouds

Provider names appear as design references. No provider has endorsed or partnered with Sinewgrid unless an agreement is listed on this site.

Capacity strategy

Match the way we buy compute to the shape of the job.

No single purchasing mode fits both lanes. The mix below is the starting design. Target mix and commitment levels are set per program, once agreements are executed.

ModeBest forTrade-off
Reserved or committedSteady robotics rollouts and long pretraining runs.Lowest unit cost, but the commitment is a fixed obligation.
Short-term blocksScheduled pretraining or fine-tuning windows with a known start and end.Capacity is guaranteed for the window only.
On-demandBursty biology experiments that follow lab schedules.Flexible, but price and availability vary by region and day.
Spot or preemptibleSimulation rollouts, hyperparameter sweeps and other jobs that checkpoint often.Cheapest, but jobs can be interrupted and must resume cleanly.
Model tokens

The leading AI platforms under the architecture.

GPUs train and serve our own models. Agents, quant code generation, evaluation, data labeling and research assistants run on frontier models from ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google) and Grok (xAI), and that work is billed in tokens. We treat their compute and tokens as capacity to plan, route and report, the same way we treat GPU hours, and the architecture is sized for a steady, large volume of both.

Purchasing modeFitsHow we manage it
Pay as you goExploration and low-volume tasksPer-project budgets and alerts
Committed spendSteady, predictable usageSized from measured demand, then reviewed each quarter
Batch processingLarge offline jobs such as labeling and evaluationQueued behind a deadline instead of run on demand
Priority capacityEvent-driven bursts with firm deadlinesReserved only where the cost of a miss justifies it

Available modes differ by provider and change over time. We confirm each against the provider's current terms before relying on it.

Route by task

Each task goes to the model that meets its quality bar at the lowest cost, with a second provider as fallback if one is rate-limited or unavailable.

Cut repeated tokens

Prompt caching, shorter context and structured outputs reduce spend, and verification steps are used where accuracy requires them, not by default.

Measure cost per task

Spend is reported per completed task and per project, so token cost sits next to GPU cost in every run report.

Protect customer data

Customer and partner data is sent to an outside model only where the agreement allows it, under the provider's data-retention terms.

No lock-in

Prompts, evaluations and tool definitions are kept provider-neutral, so switching or adding a provider is a configuration change.

Clear boundaries

Outputs from outside models are used only as each provider's terms permit. Buying tokens from a provider does not imply a partnership or an endorsement.

Reference designs

One architecture for each lane.

Sinewgrid Cell

Data-heavy, bursty, tied to the lab

  • Data. Omics matrices and sequence libraries live in object storage, in the same region as the compute that reads them.
  • Compute. A steady pretraining pool, plus burst fine-tuning pools sized to the experiment calendar.
  • Loop. Assay results return through a validated ingest pipeline and become the next training set.
Reference design
Sinewgrid Robotics

Throughput-heavy, steady, tied to the robot

  • Data. Teleoperation episodes and robot logs land in object storage, with scenes and assets versioned beside them.
  • Compute. A simulation pool that leans on spot capacity, a training pool placed for fast interconnect, and a separate evaluation pool.
  • Loop. Policies are checked on held-out real-robot tasks, and the failures become new data.
Reference design
On every cloud

Six decisions we make the same way everywhere.

Reference design

Network and placement

Keep tightly coupled jobs inside one high-bandwidth domain, and place data and compute in the same region.

Reference design

Storage and data gravity

Tier storage for long sequential reads and for millions of small robot files, and move compute to the data before moving data to compute.

Reference design

Identity and keys

Isolated tenants, private networking and customer-held keys for weights and proprietary assay data.

Reference design

Resilience

Health checks, automatic node replacement and job resume, so a failed node costs minutes and not a run.

Reference design

Cost and goodput

Report useful training time and cost per run in the same terms on every provider, so capacity can be compared fairly.

Reference design

Portability

Containers, open checkpoint formats and infrastructure as code, so a run can move between providers without a rewrite.

The reference stack

Seven layers, and which ones we own.

The cloud supplies the bottom three. We design and operate the four above them, so the same training job behaves the same way on any provider.

Workloads

Training, fine-tuning, simulation rollouts and evaluation jobs for Cell, Robotics and customer models, packaged as containers.

Sinewgrid

Experiment and data layer

Dataset versions, checkpoint registry, experiment tracking and lineage, so any result can be traced to its data and code.

Sinewgrid

Scheduling

Slurm for tightly coupled training and Kubernetes for pipelines and services, with queues that match each job to a capacity pool.

Sinewgrid

Reliability and cost

Health checks, node replacement, resume from checkpoint, goodput reporting and per-run cost.

Sinewgrid

Cluster and network

GPU nodes on a high-bandwidth fabric, with placement that keeps coupled jobs close together.

Cloud

Storage

Parallel file systems for training reads, object storage for datasets and checkpoints, local disk for scratch.

Cloud

Compute

GPU instances bought as reserved, short-term, on-demand or spot capacity under the provider's terms.

Cloud
How the clouds differ

Same job, different building blocks.

What each provider documents publicly for large training clusters. We use the differences to place work, and we keep the job definition the same.

ProviderManaged cluster optionSchedulerNetwork and storage, as documented
AWSSageMaker HyperPod; AWS ParallelClusterSlurm or Amazon EKSElastic Fabric Adapter for the fabric; Amazon FSx for Lustre for training storage
Microsoft AzureCycleCloud Workspace for SlurmSlurmInfiniBand on ND-series GPU clusters; Azure Managed Lustre option
Google CloudCluster DirectorSlurm or GKETopology-aware placement; Filestore, Managed Lustre or Cloud Storage attached to the cluster
Oracle CloudOCI SuperclusterNot specified in the source we usedRDMA cluster network over RoCE v2; bare metal instances; Lustre-based file storage
GPU cloudsFor example CoreWeave SUNKSlurm on KubernetesSlurm control plane runs as Kubernetes workloads on the provider's GPU fabric

Summaries of public provider documentation. Features and names change, and we verify each against current documentation before a design is used.

When something breaks

When a node fails during a long run.

The system is designed to respond in four steps. Recovery times from failover drills are reported with the method used.

  1. DetectHealth checks and job telemetry flag a failing node or a straggler.
  2. ContainThe scheduler drains the node and pauses the affected job.
  3. ReplaceA spare node in the same placement group takes its place, or the job moves to another pool.
  4. ResumeThe job restarts from the latest checkpoint, and the lost time is recorded in goodput.

Checkpoint cadence

Set from the cost of a checkpoint against the cost of lost work, per job, not as one global number.

Portable state

Checkpoints are written in open formats to object storage, so a resumed job can start in a different cloud.

Drills

Failover drills are part of the operating plan, with resume times reported by provider and the method.

Public references

The patterns we build on, in the providers' own words.

Each card summarizes a public document. Sinewgrid is not affiliated with these companies, and the summaries are ours.

AWS

Resilient training clusters

Amazon SageMaker HyperPod offers Slurm or Amazon EKS orchestration, monitors instances and automatically replaces faulty hardware. We hold our designs to the same resilience bar on every cloud.

Read the source →
Microsoft Azure

Slurm in your own subscription

Azure CycleCloud Workspace for Slurm deploys a Slurm cluster, shared storage and networking into the customer's own subscription, with optional Azure Managed Lustre. That ownership model suits customer-held keys.

Read the source →
Microsoft Azure

Scale-up and scale-out limits

Azure's ND GB300 v6 clusters join 72-GPU NVLink racks with 800 Gb/s of InfiniBand per GPU in a non-blocking fat tree. Job placement is planned around topology limits like these on any cloud.

Read the source →
Google Cloud

Hybrid by workload

Recursion runs its image data pipeline on Google Kubernetes Engine and uses cloud GPUs and TPUs for inference, alongside its own training supercomputer. Pipelines and inference are natural places to use cloud capacity.

Read the source →
Open source

One interface across clouds

SkyPilot runs workloads across Kubernetes, Slurm and cloud VMs, fails over when a pool lacks capacity and picks the cheapest available zone. We evaluate open schedulers like it before building our own.

Read the source →
AWS

A cloud pipeline for robot learning

AWS's published reference fine-tunes NVIDIA Isaac GR00T N1.5 on GPU instances through AWS Batch, keeps checkpoints on shared storage and evaluates in Isaac Lab. Our robotics lane follows the same shape.

Read the source →
Google Cloud

Managed Slurm or GKE clusters

Cluster Director provisions Slurm or GKE clusters from validated templates, places nodes by topology and runs continuous health checks with straggler detection. We copy the idea of qualifying a cluster before a job lands on it.

Read the source →
Oracle Cloud

RDMA fabric on bare metal

OCI documents an RDMA cluster network over RoCE v2 on bare metal instances for large GPU clusters. Fabric and virtualization overhead are the first things we compare across providers.

Read the source →
GPU cloud

Slurm on Kubernetes

CoreWeave describes SUNK, which runs Slurm on Kubernetes so scheduled training jobs and Kubernetes services share one cluster. It is the closest public model for our two-scheduler design.

Read the source →
Questions we expect

Plain answers.

Do you own GPUs or data centers?

No. We source capacity from the major clouds and GPU clouds, and our work is the architecture, scheduling and data layer on top.

Which cloud will you use?

Whichever fits the job, the region and the commitment. We design for AWS, Google Cloud, Microsoft Azure, Oracle Cloud and GPU clouds, and keep workloads portable so no single provider is a dependency.

Can we run in our own cloud account?

That is the intended deployment model. Running the control plane in your subscription keeps your data, weights and keys under your control, in the same way the providers' own reference designs deploy into the customer's account.

What happens if a pool runs out of capacity mid-run?

The scheduler checkpoints, moves the job to another pool and resumes. How fast that happens is something we will measure in drills and publish, not promise in advance.

Is Sinewgrid affiliated with any cloud provider?

No. Providers are named here as design references and as sources of public documentation. No provider has endorsed or partnered with Sinewgrid.

Why both Slurm and Kubernetes?

Slurm is the common scheduler for tightly coupled training, and Kubernetes suits pipelines, simulation and services. Running both is a pattern the providers above also document, so we design for it from the start.

How is customer data kept apart?

By isolated tenants, private networking and keys the customer holds. Training data does not move between customers or into shared models unless the agreement says so.

Will you run on our existing capacity?

Yes, where it fits. If you already hold reserved GPUs on a cloud, the control plane can schedule onto them, and we plan around your commitment instead of buying duplicate capacity.

Planning a training run?

Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.

[email protected]