Research on running AI at scale, and on the models we train.
Sinewgrid is an operating company, and our research comes from the problems we meet in operation. This page sets out what we study and why, the agenda we intend to publish on, and the rules we hold ourselves to. It describes plans: nothing on it reports a result.
The hard problems are in the operation.
Most public work on large models starts after the hardware exists: a cluster is assumed, a scheduler is assumed, and the question is how to train well on it. We start earlier. We do not own data centers. We buy GPU capacity from several clouds, and we buy model tokens from several providers, and we have to make that mixed supply behave like one dependable system. That is a research problem in its own right, and the field has few open answers to it.
Three questions come up on almost every program. How do you measure useful work when jobs of very different shapes share capacity? What does it really cost to move a run, or its data, from one provider to another? And when a service makes thousands of model calls a day, how do you compare the real cost of finishing a task across providers, once retries, checking steps and caching are counted? Vendor figures rarely answer these, because they describe one machine or one price list, not a task from start to finish.
We think these questions deserve plain, reproducible answers, and that customers and other builders are better served when the methods are public. So we write them down as we go, and we release what the customers and partners involved allow us to release.
Systems, models and services.
Systems. This line covers the control plane and the capacity under it: how a training run is scheduled across clouds, how checkpoints move, where very large datasets should sit so that moving the work is cheaper than moving the data, and how to account for goodput, the share of paid time that produces training progress. The aim is a set of definitions and drills that another team could repeat on its own capacity and get comparable numbers.
Models. Here we study the models we train for biology and robotics, and how they are judged. In biology, the central questions are how many lab cycles a model needs before its proposals beat a strong baseline, and how generated sequences are screened for biosecurity risk before anyone sees them. In robotics, the central question is how to measure the gap between simulation and the real world with a fixed set of held-out tasks, so that one checkpoint can be compared with another without argument.
Services. This line covers the enterprise AI services we offer. We look at how to route work across frontier models, how to compare their cost per finished task, and how to test an assistant well enough to decide whether it should leave a pilot. The emphasis is on evaluation that a business owner can read: a test set written with the customer, a report of failures as well as successes, and costs that include everything that was spent.
Questions we intend to publish on.
Each item is a question we think the field needs answered in the open. Results are released as papers, datasets or reports, and each card names what the release will contain.
Goodput accounting for mixed workloads
How to report useful training time when bursty biology jobs and steady robotics jobs share capacity pools.
Portable checkpoints across clouds
Moving a training run between providers without losing progress, and what it costs in time and egress.
Egress-aware data placement
Where to keep very large omics matrices and robot logs so that moving work is cheaper than moving data.
Screening generated sequences
Methods for checking sequence-design outputs against biosecurity risk lists before release.
Measuring the sim-to-real gap
A fixed set of held-out real-robot tasks used to compare policy checkpoints.
Closed-loop design with a lab
How many lab batches a model needs before its proposals beat a strong baseline, and how to pick informative batches.
Cost per task across model providers
How to compare a routed set of frontier models on one task, counting retries, checking steps and cached tokens.
Evaluating enterprise assistants
Test sets, review steps and failure reports that show whether an assistant is ready to leave a pilot.
Four rules for every release.
Reproducible. Code, configuration and evaluation sets are released with the result, wherever licenses and partners allow, so that a reader can run the same test and compare. Where something cannot be released, we say what is withheld and why.
Negative results included. When a method does not beat its baseline, we say so and show the comparison. A field that publishes only wins learns slowly, and customers deserve to know what did not work as much as what did.
Reviewed for risk. Anything that touches sequence design or physical robots is reviewed for misuse before it is released, and we hold back what a review does not clear. This applies to our own work first.
Partner data stays private. We publish methods and aggregate findings, never a partner’s proprietary data without written consent. Customer and partner names appear only with written permission. See Trust for the commitments behind these rules.
Planning a training run?
Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.