Sinewgrid Protein trains protein and small-molecule design models.
Models that propose proteins, antibodies and drug-like molecules, rank them with structure-prediction and binding models, and learn from what a partner lab measures. This domain is at the exploring stage.
Designing proteins is a search problem.
The number of possible proteins and drug-like molecules is vastly larger than anything a lab can make and test, so progress depends on proposing a small number of good candidates and learning quickly from each result. Sinewgrid Protein is the domain where we would build models that generate designs for a target, such as a protein, a binding site or a property, and rank them with structure-prediction and binding models, with an estimate of how uncertain each ranking is. The compute demand has three parts: a large first pretraining run on sequence and structure collections that needs fast interconnect, many scoring and fine-tuning passes after every lab batch, and inference across thousands of candidate designs for each target. Together they make demand uneven and tied to the lab’s calendar.
A program runs as a loop. We define the target and the assay that will judge it, the model generates and ranks designs, a partner lab makes and tests a small, informative batch, and the measured results update the model before the next batch is chosen. We would buy steady capacity for pretraining and on-demand capacity for the scoring that follows each batch. We report three measures that are fixed before the first run: the hit rate against a baseline the partner already trusts, the number of rounds needed to reach a hit, and the compute spent per round. Biosecurity comes first. Design outputs are screened against risk lists, access is tiered, and capability evaluations are reviewed by outside experts before any release. This domain is at the exploring stage.
Where the capacity goes.
Large pretraining runs
Protein language and structure models learn from very large sequence and structure collections, so the first training run is long and needs fast interconnect.
Scoring after every batch
Each lab batch triggers many scoring and fine-tuning passes before the next batch is chosen, so demand arrives in bursts.
Many candidates per target
Ranking thousands of designs for one target multiplies inference, even when each design is small.
What the models learn from.
Public and licensed
Public sequence and structure databases, subject to a license review for each source.
Partner assays
Binding, activity and stability measurements from a partner lab, kept in the partner tenant under partner keys.
Generated here
Every tested design becomes a labeled example for the next round.
From a question to a checked result.
- Define the targetA protein, a binding site or a property, and the assay that will judge it.
- Generate and rankThe model proposes designs and ranks them, with an uncertainty estimate.
- Test in the labA partner lab makes and tests a small, informative batch.
- Return and retrainMeasured results update the model before the next batch.
Bursty, tied to the lab calendar
Steady capacity for pretraining, on-demand capacity for scoring and fine-tuning after each batch.
See the architectureBiosecurity first
Design outputs are screened against biosecurity risk lists, access is tiered, and capability evaluations are reviewed by outside experts before any release.
Read the commitmentsA first engagement, scoped small.
First engagement
One target and one assay. We agree what a successful design looks like before any compute is bought.
What the partner brings
A target, a lab able to make and test a small batch, and the right to use the resulting data.
What we bring
Candidate generation and ranking, capacity planning for each round, and the screening step.
Measures we commit to publishing.
Each result is reported against these measures, which are defined before the first run.
Hit rate against baseline
The share of tested designs that meet the target, compared with a simple baseline the partner already trusts.
Rounds to a hit
How many lab batches it took, so cost per result is visible.
Compute per round
Capacity used for training, scoring and screening, by purchasing mode.
Plain answers.
Do you run wet-lab experiments?
No. A partner lab makes and tests the designs. Our job is the models and the compute behind them.
Who owns the designs?
That is set in the agreement with each partner. We do not assume rights to a partner's sequences or data.
Can the model design anything?
No. Outputs are screened against biosecurity risk lists, and access is tiered by use.
Public work such as AlphaFold and ESM shows what is possible in this area. They are referenced as context, not as partners. Related: Sinewgrid Cell: shares the lab-in-the-loop pattern and the screening. Status: Exploring
Planning a training run?
Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.