All domains / Exploring

Sinewgrid Materials trains atomistic models for materials and chemistry.

Learned models of how atoms interact can stand in for slow quantum-chemistry calculations, so a team can screen far more candidate materials for batteries, catalysts and semiconductors. This domain is at the exploring stage.

FIG. · A model scanning an atomic lattice, and a screen of candidates. Schematic, not real data.
Overview

Replacing slow simulation with a learned model.

Finding a better battery electrolyte, catalyst or semiconductor usually means evaluating a very large number of candidate structures, and the standard way to evaluate one is a quantum-chemistry calculation that can take hours or days. Learned atomistic models offer a shortcut. Trained on the results of many such calculations, a model can estimate the energy and forces for a new structure in a fraction of the time, which lets a team screen far more candidates before spending on the expensive calculation or a lab test. Sinewgrid Materials would build and run these models. The work has an unusual compute profile: the training data is itself made by physics calculations, which is heavy compute, and the screening that follows is a very large parallel job.

A program starts with one property, such as ionic conductivity or catalytic activity, together with a stability filter. The model scores a large candidate set and flags the structures it is least sure about. Those, and a selection of the best-scoring ones, get a high-accuracy simulation and then a lab test where possible, and the verified results go back into the training set for the next sweep. Simulation and screening run on interruptible capacity, and training uses a steady pool with fast interconnect. A screening hit is a hypothesis until it is verified, so we report the verified hit rate, the speed-up over the calculation the model replaces, measured on the same candidates, and the compute spent per verified hit, not only the model’s own score. This domain is at the exploring stage.

Why it needs compute

Where the capacity goes.

Data comes from simulation

Training sets are built from very large numbers of physics calculations, which is itself a compute-heavy step.

Retrain as results arrive

As new calculations and measurements come in, the model is retrained and the screen is repeated.

Large screening sweeps

Scoring millions of candidate structures is a parallel job that fits interruptible capacity.

Data

What the models learn from.

Public databases

Open computed-materials databases, such as the Materials Project, subject to a license review.

Simulation we run

High-accuracy calculations chosen to fill gaps where the model is least certain.

Partner measurements

Lab measurements from a partner for the materials they care about, kept in their tenant.

How a program runs

From a question to a checked result.

  1. Define the propertyA target such as ionic conductivity or catalytic activity, and a stability filter.
  2. Train and screenThe model scores a large candidate set and flags the uncertain ones.
  3. VerifySelected candidates get high-accuracy simulation, then a lab test where possible.
  4. Feed backVerified results are added to the training set for the next sweep.
Compute shape

Interruptible simulation, steady training

Simulation farms run on interruptible capacity. Training needs a steady pool with fast interconnect.

See the architecture
Limits we hold

Verified before it is believed

A screening hit is a hypothesis. We report how many hits held up under high-accuracy simulation and in the lab, not only the model score.

Read the commitments
Where we would start

A first engagement, scoped small.

First engagement

One property and one material family, for example an electrolyte or a catalyst, with a stability filter.

What the partner brings

A property target, a way to measure it, and access to the relevant lab or measurement service.

What we bring

The learned model, large screening sweeps on interruptible capacity, and the verification plan.

What we would report

Measures we commit to publishing.

Each result is reported against these measures, which are defined before the first run.

Verified hit rate

The share of model-selected candidates that held up under high-accuracy simulation, then in the lab.

Speed-up over simulation

How much faster the model is than the calculation it replaces, measured on the same candidates.

Compute per verified hit

Capacity spent on training, screening and verification for each confirmed candidate.

Questions we expect

Plain answers.

Does the model replace quantum-chemistry calculations?

It stands in for them during screening. Selected candidates are still verified with the accurate method.

How do you handle uncertainty?

The model reports where it is least certain, and those cases go to high-accuracy simulation first.

Is the data public?

We start from open databases where the license allows and add calculations and partner measurements as needed.

Public work such as Google DeepMind GNoME shows what large-scale screening can look like. It is referenced as context, not as a partner. Related: Sinewgrid Robotics: shares the simulation-at-scale pipeline. Status: Exploring

Planning a training run?

Tell us the model, the data and the schedule. We reply with a capacity plan and the evidence behind it.

[email protected]