Skip to content

Python API

The functions most scripts need are importable from the package itself:

from torchnep import train_nep, train_nep_sharded, predict_dataset, predict_dataset_sharded, export_valid_split

Training

train_nep

train_nep(
    config_file: str,
    data_file: str,
    output_dir: str = ".",
    device: str = None,
    precision: str = "float32",
    use_autograd_forces: bool = False,
    use_swa: bool = False,
    swa_start: int = None,
    use_compile: bool = None,
    print_interval: int = 1,
    restart: bool = True,
    checkpoint_interval: int = 100,
    prediction_interval: int = 100,
    finetune_from: str = None,
    resume_from: str = None,
    recompute_q_scaler: bool = False,
    slim_types: bool = False,
    energy_key: str = "energy",
    use_gpumd_qscaler: bool = False,
    run_seed: int = None,
    valid_file: str = None,
    valid_ratio: float = None,
    valid_strategy: str = "stratified",
    neighbor_mode: str = "auto",
)

Train a NEP model on a single device (GPU / CPU / MPS).

neighbor_mode: how the training data is held in host memory — cached (full neighbor lists with displacement vectors, fastest), compact (int32 indices + int8 image shifts, ~4x smaller, rij rebuilt on the device per batch), on_the_fly (positions + cells only, the neighbor search runs on the device per batch — fits any dataset). auto (default) takes the first of these whose estimated peak footprint fits in 75% of this process's memory budget.

Hyperparameters (epoch / batch / lr / lambda_e,f,v / stage2* / …) come from config_file only. See README for the full nep.in reference.

Launch: python run_train.py

Parameters:

Name Type Description Default
config_file paths.
required
data_file paths.
required
output_dir paths.
required
device "cuda" | "cpu" | "mps" — auto-detected if omitted.
None
precision "float32" (default) or "float64".
'float32'
use_autograd_forces True -> autograd-through-rij forces (slower, gold

standard); False (default) -> analytical chain rule. The type-pair contraction backend is chosen automatically (see torchnep.ops.resolve_backend).

False
use_swa True -> maintain an averaged model during stage 2 and save it

as nep_average.txt at the end.

False
swa_start first epoch included in the SWA average (default: the

last 100 epochs). Averaging all of stage 2 degrades energies.

None
use_compile None (default) = auto — compile on GPU, eager on

CPU; True/False force it. Missing Triton (GPU) or C++ toolchain (CPU) degrades to eager with a log note. epoch after a one-time first-epoch compilation cost; needs Triton, which ships with the CUDA PyTorch build). With autograd forces the first-order gradient is materialized via make_fx and compiled too.

None
print_interval log a line to screen every N epochs (all epochs still

land in output.log).

1
restart on fresh output_dir, write new log; otherwise resume from

checkpoint.pt if present.

True
checkpoint_interval save checkpoint.pt every N epochs.
100
prediction_interval every N epochs, run predict_from_store with the

current-epoch weights and overwrite {energy,force,virial}_train.out in output_dir — lets you watch the parity-plot converge live. The end-of-training predict overwrites them with the nep_best.txt model. Set to 0 or a negative value to disable.

100
finetune_from path to an existing .pt or nep.txt to load weights from

(weights only — a NEW training starts from them: epoch 1, nep.in lr, fresh optimizer). The q_scaler stored with the weights is kept.

None
resume_from path to a checkpoint (.pt) to CONTINUE from — restores

model, optimizer (incl. lr), scheduler, epoch and best losses, so the run picks up exactly where that checkpoint left off. Use checkpoint_stage1.pt to redo stage 2 with edited stage2_* settings. Mutually exclusive with finetune_from; takes precedence over the automatic checkpoint.pt pickup.

None
recompute_q_scaler only with finetune_from. Default False — the

q_scaler stored with the loaded weights is kept (it is part of the potential they were trained as). True recomputes it from the new dataset, which rescales the NN inputs: the loaded model starts off wrong and must re-adapt (see bench_test/restart_rework/t8).

False
slim_types drop element types not present in data_file before training.
False
energy_key name of the comment-line tag read as the reference energy

(default "energy"). Set to "atomization_energy" to train against atomization energies instead of totals.

'energy'
use_gpumd_qscaler Default False — torch's default init with the

self-consistent q_scaler (computed from the model's actual init coefficients). On a 600-epoch 4-seed PdCuNiP benchmark this converges to clearly better minima than the GPUMD-style start (~12% lower E/V RMSE, ~3% lower F, train and validation alike). True reproduces GPUMD's initialization instead: every parameter re-initialised uniform(-1, 1) (SNES mu init) and the q_scaler computed with all coefficients = 1.0 (GPUMD's generation-0 initial_para) — useful for GPUMD-comparison runs. Either way the saved nep.txt is fully GPUMD-compatible (the scaler is stored in the file). Only applies to fresh training (ignored under finetune_from).

False
run_seed master RNG seed for this run. None (default) -> a fresh random

seed each run, so repeated runs differ (independent weight init AND per-epoch batch shuffle) — the stochastic-testing behaviour. Pass an int to make a run fully reproducible: the same seed reproduces both the initial weights and the batch ordering bit-for-bit. The seed is saved into checkpoint.pt and restored on resume, so a resumed run continues along the very same shuffle stream regardless of what seed (or None) is passed at resume time.

None
valid_file path to a validation .xyz. Evaluated (frozen weights, full

set) every epoch: nep_best is the epoch with the lowest validation loss and the plateau LR scheduler steps on the validation loss — both anti-overfitting. Writes GPUMD-style *_test.out files.

None
valid_ratio alternative to valid_file — hold out this fraction of

data_file (e.g. 0.1) as the validation set. The split is drawn from run_seed, so it is reproducible for a given seed and is preserved exactly on resume (the checkpoint's seed wins). Mutually exclusive with valid_file.

None

train_nep_sharded

train_nep_sharded(
    config_file: str,
    data_file: str,
    output_dir: str = ".",
    precision: str = "float32",
    use_autograd_forces: bool = False,
    use_swa: bool = False,
    swa_start: int = None,
    use_compile: bool = None,
    print_interval: int = 1,
    restart: bool = True,
    checkpoint_interval: int = 100,
    prediction_interval: int = 100,
    finetune_from: str = None,
    resume_from: str = None,
    recompute_q_scaler: bool = False,
    slim_types: bool = False,
    energy_key: str = "energy",
    use_gpumd_qscaler: bool = False,
    run_seed: int = None,
    valid_file: str = None,
    valid_ratio: float = None,
    valid_strategy: str = "stratified",
    neighbor_mode: str = "auto",
)

Data-sharded NEP training. Launch via torchrun (or any launcher that sets RANK / LOCAL_RANK / WORLD_SIZE / MASTER_ADDR / MASTER_PORT).

All hyperparameters (epoch / batch / lr / lambda_e,f,v / stage2* / ...) come from config_file — see README for the nep.in reference.

Each rank loads structures[rank::world_size] only, so data-store memory scales as 1/world_size. Gradients are all-reduced by DDP; q_scaler and epoch metrics are all-reduced explicitly.

torchrun --nproc_per_node=N run_train.py

Parameters mirror train_nep (same restart semantics: resume is an exact continuation — lr/optimizer/scheduler/SWA/best gates from the checkpoint, nep.in's lr never re-applied on resume; resume_from continues from a specific checkpoint such as checkpoint_stage1.pt; finetune_from keeps the source model's q_scaler unless recompute_q_scaler=True). Runtime-exclusive differences are the DDP launch, CUDA/gloo backend auto-select, and distributed aggregation of q_scaler / epoch metrics / the frozen-weight best evaluation.

run_seed mirrors train_nep: None -> a fresh random seed each run (rank 0 draws it and broadcasts, so all replicas agree); pass an int for a fully reproducible run. It seeds the weight init and the per-epoch shuffle, and is saved into / restored from the checkpoint for exact resumption.

valid_file / valid_ratio mirror train_nep: a validation set (own .xyz, or a run_seed-drawn holdout fraction of data_file) evaluated every epoch — nep_best and the plateau LR schedule then follow the validation loss. The validation frames are sharded across ranks like the training frames; the error sums are all-reduced, so every rank sees the identical validation loss (schedulers stay in lock-step).

Each rank keeps its shard in host memory and streams only the current batch to its GPU (basis computed on the fly, CPU assembly prefetched one batch ahead) — per-GPU memory scales with batch, not shard size. DDP collectives are untouched: same step count per rank, gradients all-reduced as usual.

export_valid_split

export_valid_split(
    data_file: str,
    valid_ratio: float,
    run_seed: int,
    output_dir: str = "split",
    strategy: str = "stratified",
    min_stratum: int = 20,
)

Write GPUMD-ready train.xyz / test.xyz with train_nep's split.

Reproduces exactly the validation split that train_nep(data_file, valid_ratio=r, run_seed=s, valid_strategy=...) uses internally, so the exported pair can train the SAME data partition in GPUMD (or any other code) and loss curves stay comparable. Frames are copied verbatim (raw text, untouched fields and precision), in input-file order.

strategy: "random" (default) or "stratified" — see :func:stratified_split_indices.

Returns (train_path, test_path, n_train, n_valid).

parse_nep_in

parse_nep_in(filename: str) -> Dict

Parse nep.in parameter file.

Parameters:

Name Type Description Default
filename str

Path to nep.in file.

required

Returns:

Type Description
dict

Dictionary of NEP parameters.

Prediction

predict_dataset

predict_dataset(
    model_file: str,
    xyz_file: str,
    output_dir: str = ".",
    dtype: str = "float32",
    device: str = None,
    batch_size: int = None,
    verbose: bool = True,
    energy_key: str = "energy",
    output_descriptor: int = 0,
    chunk_atoms: int = None,
)

Run streamed, batched prediction on a full dataset and save GPUMD-format outputs.

The file is first indexed (byte offsets + atom counts, one I/O pass, no parsing), then processed in consecutive chunks of roughly chunk_atoms atoms: each chunk is read, neighbor-listed (multi-process), predicted in device-sized batches and appended to the output files before the next chunk is touched. Host memory is therefore bounded by the chunk size and device memory by the batch size — any dataset finishes on any machine, only the wall time differs. Outputs are identical to the former whole-dataset path.

Outputs (per-atom for energy and virial; per-atom raw for forces): - energy_train.out: e_pred e_target (eV/atom, per frame) - force_train.out: fx fy fz fx_t fy_t fz_t (eV/A, per atom) - virial_train.out: xx yy zz xy yz zx (pred, ref) (eV/atom, per frame) - stress_train.out: same layout in GPa (GPUMD sign convention) - descriptor.out (only when output_descriptor != 0): mode 1 — per-frame averaged scaled descriptor, one row per frame mode 2 — per-atom scaled descriptor, one row per atom Matches GPUMD's output_descriptor / descriptor.out schema.

The format mirrors GPUMD's *_train.out files, so the two can be diffed column by column. When reference labels are present, a summary of the energy/force/virial RMSE and MAE is printed at the end. The contraction backend is chosen automatically (see torchnep.ops.resolve_backend).

Parameters:

Name Type Description Default
batch_size int or None

Frames per compute batch. None (default) sizes the batch automatically from the device's free memory and the first chunk's average pair counts (CUDA/ROCm; other devices fall back to 1000), with an OOM-halving retry.

None
chunk_atoms int or None

Atoms per streamed chunk (host-memory bound; ~1 GB per 200k atoms of dense metal at cutoff 6/5). None -> TORCHNEP_PREDICT_CHUNK_ATOMS env var, else 200000.

None
output_descriptor int

0 — disabled (default). 1 — write per-frame averaged q * q_scaler to descriptor.out. 2 — write per-atom q * q_scaler to descriptor.out.

0

predict_dataset_sharded

predict_dataset_sharded(
    model_file: str,
    xyz_file: str,
    output_dir: str = ".",
    dtype: str = "float32",
    batch_size: int = None,
    verbose: bool = True,
    energy_key: str = "energy",
    output_descriptor: int = 0,
    chunk_atoms: int = None,
)

Multi-rank (multi-GPU / multi-node) version of :func:predict_dataset.

Launch one process per GPU with torchrun or srun (the launcher sets RANK / LOCAL_RANK / WORLD_SIZE / MASTER_ADDR / MASTER_PORT). Rank 0 indexes the file once and broadcasts the frame index; the frames are then split into WORLD_SIZE contiguous ranges balanced by atom count, every rank streams its own range exactly like predict_dataset (same chunking, auto batch size and OOM retry, bounded memory) into a per-rank part directory, and rank 0 finally concatenates the parts in rank order into the usual *_train.out files and prints the global E/F/V RMSE/MAE table. The output is identical to a single-process predict_dataset. Rank 0 shows the progress of its own share; the other ranks are silent. With WORLD_SIZE == 1 (or no launcher) this is plain predict_dataset.

NEPCalculator

NEPCalculator(
    model_file: str, dtype=torch.float64, device="cpu"
)

NEP4 calculator: loads a trained model and computes atomic properties.

Parameters:

Name Type Description Default
model_file str

Path to nep.txt (GPUMD NEP4 format).

required
dtype dtype

Precision (default: float64 for accuracy).

float64
device str or device

Compute device (default: 'cpu').

'cpu'

compute

compute(
    species: list,
    positions: ndarray,
    cell: ndarray,
    compute_descriptor: bool = False,
    return_components: bool = False,
) -> Dict[str, torch.Tensor]

Compute energy, forces, and per-atom virial.

Parameters:

Name Type Description Default
species list of str

Element symbols for each atom.

required
positions (N, 3) array

Atomic positions in Angstrom.

required
cell (3, 3) array

Lattice vectors (row-major).

required
compute_descriptor bool

Also return scaled descriptors.

False
return_components bool

Also split the result into the NEP (neural-network) part and the ZBL repulsive part. Adds energy_nep / forces_nep / virial_nep and energy_zbl / forces_zbl / virial_zbl to the output; the plain energy / forces / virial keys remain the NEP+ZBL total (their exact sum). For models without ZBL the ZBL part is zero.

False

Returns:

Type Description
dict with 'energy', 'forces', 'virial', optionally 'descriptor' and
(if ``return_components``) the per-contribution breakdown above.

get_descriptor

get_descriptor(species, positions, cell)

Compute scaled descriptors. Returns (N, dim) numpy.

compute_tiled

compute_tiled(
    species: list,
    positions: ndarray,
    cell: ndarray,
    block_size="auto",
    query_sub_chunk: int = 4000,
    backend: str = "auto",
    compile: bool = False,
) -> Dict[str, torch.Tensor]

Energy / forces / virial for large cells with bounded peak memory.

This is the MD-oriented counterpart of :meth:compute. Instead of building the whole neighbor list and back-propagating through it (which scales the peak memory with the total pair count — infeasible for a 500k-atom, high-density cell), it:

  • builds the spatial bin table once (:class:~torchnep.neighbor.CellList);
  • streams over blocks of block_size centre atoms, querying only that block's neighbors at a time;
  • computes per-atom energy with the analytical force/virial path (ops.compute_analytical_forces) — no autograd graph is retained, so peak memory is set by a single block, not the whole system. (This is the same analytical kernel training/prediction use, hence also torch.compile-friendly.)

Each directed pair is owned by exactly one block (the block holding its centre atom i); contributions scatter into global per-atom accumulators, so the result is identical to :meth:compute to round-off.

For cells too small for the linked-cell decomposition this transparently defers to :meth:compute (the autograd path is cheap there).

Parameters:

Name Type Description Default
species as in :meth:`compute`.
required
positions as in :meth:`compute`.
required
cell as in :meth:`compute`.
required
block_size int or auto

Number of centre atoms processed per tile. Larger = faster but more peak memory (block-independent N-sized scratch + this block's pairs). "auto" (default) sizes it from the cell density, model dims, dtype and the device's free memory (~35% budget); pass an int to override.

'auto'
query_sub_chunk int

Inner chunk for the neighbor query (bounds the candidate tensor).

4000
backend "auto" | "loop" | "bmm" type-pair contraction (see ops).
'auto'
compile bool

torch.compile the descriptor + analytical-force kernels (dynamic shapes, so the pair-count changes between MD steps do NOT trigger recompilation — one warm-up compile, then a single dynamic graph). Forces the bmm backend (compile-friendly; the per-type Python loop graph-breaks). Worth it for many steps / large cells.

False

Returns:

Type Description
dict with 'energy' (N,), 'forces' (N, 3), 'virial' (N, 9) — same keys and
conventions as :meth:`compute` (sans the component split / descriptor).

NEP

NEP(
    model_file,
    dtype="float64",
    device="cpu",
    tiled="auto",
    block_size="auto",
    compile=False,
    **kwargs,
)

Bases: Calculator

ASE calculator backed by a torchnep NEP4 model (nep.txt).

Parameters:

Name Type Description Default
model_file str

Path to a GPUMD-format nep.txt (NEP4).

required
dtype str or dtype

Compute precision; "float64" (default) reproduces GPUMD to round-off, "float32" is faster.

'float64'
device str or device

Torch device, e.g. "cpu" (default) or "cuda".

'cpu'
**kwargs

Forwarded to ase.calculators.calculator.Calculator (e.g. label).

{}
Notes

stress is only reported for fully periodic cells (a finite volume is required); for molecules / clusters it is omitted. Non-periodic systems are handled by embedding the atoms in a vacuum box large enough that no atom sees a periodic image within the model cutoff.

get_components

get_components(atoms=None)

Full NEP / ZBL / total breakdown of energy, forces, and stress.

Returns {'nep': {...}, 'zbl': {...}, 'total': {...}} where each inner dict has 'energy' (float, eV), 'forces' ((N, 3) eV/Å) and, for periodic cells, 'stress' (6-vector Voigt, eV/ų). The total block equals what get_potential_energy / get_forces / get_stress return.

get_energy_components

get_energy_components(atoms=None)

Return the NEP / ZBL / total potential-energy split (eV).

Useful to inspect how much of the energy comes from the ZBL repulsive baseline versus the neural-network NEP part::

{'nep': ..., 'zbl': ..., 'total': ...}

For a model trained without ZBL, 'zbl' is 0 and 'nep' equals 'total'. See :meth:get_components for forces and stress too.

Models

slim_model

slim_model(
    model: NEPModel, keep_type_names: List[str]
) -> NEPModel

Return a new NEPModel containing only the specified element types.

Mathematically equivalent to the original for structures that contain only the kept element types. Parameters for removed types are discarded:

  • fitting_nets: only the nets for kept types are copied
  • c_param_2 / c_param_3: rows and columns for removed types are dropped
  • b1, q_scaler: copied unchanged (independent of num_types)

Parameters:

Name Type Description Default
model NEPModel

Source model (any number of types).

required
keep_type_names list[str]

Subset of model.type_names to retain, in any order. The output model uses this order.

required

Returns:

Type Description
NEPModel

Smaller model on the same device/dtype as the source.

Plotting

NEPPlotter

NEPPlotter(
    font="Arial",
    fontsize=7,
    dpi=300,
    cmaps=None,
    max_points=300000,
    colors=None,
    panel_labels="abcdefghijkl",
    label_format="{}",
    label_weight="bold",
    frame=False,
    rc=None,
    font_dir=None,
    cmap_range=(0.1, 0.9),
    cmap_reverse=True,
)

Figures for torchnep / GPUMD training and prediction outputs.

Parameters:

Name Type Description Default
font str

Font family (default "Arial"); matplotlib's default when absent.

'Arial'
fontsize float

Base font size in points (labels; ticks and legends one point smaller).

7
dpi int

Figure / saved-file resolution (default 300).

300
cmaps dict

Colormaps of the density (hexbin) panels per set, overrides of DEFAULT_CMAPS (train: Blues, valid: Reds).

None
cmap_range (float, float)

Part of each density colormap used (default (0.1, 0.9)), kept clear of the colormap's white end.

(0.1, 0.9)
cmap_reverse bool

True (default): density colormaps run reversed, so cells with few points (the outliers) take the dark end and crowded cells the light end; False: the usual direction (crowded cells dark).

True
max_points int

Scatter panels draw at most this many points (a fixed random subsample); metrics always use every point.

300000
colors dict

Overrides of DEFAULT_COLORS — keys train, valid (parity sets) and energy, force, virial, stress (curves).

None
panel_labels str or sequence

Labels of the panels of multi-panel figures (default "abcd..."); label_format wraps them ("{}" -> a, "({})" -> (a)), label_weight is their font weight. panel_labels=None for none.

'abcdefghijkl'
frame bool

False (default): only the left and bottom spines; True: the full box with ticks on all four sides.

False
font_dir str

Folder of .ttf/.otf files to register before looking up font (default: the TORCHNEP_FONT_DIR environment variable).

None
rc dict

Extra matplotlib rcParams applied on top of the built-in style.

None

dashboard

dashboard(
    path,
    out=None,
    virial=False,
    title=None,
    stage2="auto",
    figsize=None,
)

One figure per training run: (a) the training / validation RMSE curves of E, F, V from loss.out and (b)(c)(d) the energy / force / stress parity plots (scatter) with both sets overlaid — 2 x 2 panels of equal size, or one row of three when the data carry no stress labels. virial=True: the virial (eV/atom) instead of the stress. stage2: epoch where stage 2 started (dashed line); "auto" reads it from output.log / nep.in in path, None for none. figsize in cm.

loss

loss(
    path,
    out=None,
    stage2="auto",
    stress=False,
    figsize=None,
)

Training curves from loss.out: the E / F / V RMSEs against the epoch on a log scale, training solid and validation faint (same colour), a dashed line at the start of stage 2 — the panel of the dashboard on its own. stress=True adds the stress RMSE. path: the run directory or the loss.out file. figsize in cm.

parity

parity(
    path,
    kind="scatter",
    margins=False,
    virial=False,
    out=None,
    quantities=None,
    size=4.0,
    shift_energy=None,
    xyz=None,
    title=None,
    exclude=None,
    natoms=None,
    bins=80,
    cell="hex",
)

Parity plots of the run in path: energy, force and stress (the last only with stress labels) with training and validation overlaid and R^2 / RMSE / MAE per set. kind="density": log-count 2-D histograms (training Blues, validation Reds) with a small horizontal colorbar under every panel. margins=True adds a strip on top of each panel with the error NEP − DFT against the DFT value. The DFT and NEP axes share limits and ticks. virial=True shows the virial (eV/atom) instead of the stress in the third panel; quantities overrides the panels altogether (any of E F V S); shift_energy ("mean" / "element", needs xyz) removes a reference offset from the predicted energies; size is the side of one panel in cm. exclude: frame indices (of the training split) to leave out, e.g. an outlier list — needs natoms (atoms per frame, or the xyz) to drop their force rows too. bins: density cells across the x axis (default 80; fewer = larger cells); cell: "hex" (default) or "square".

errors

errors(
    path,
    out=None,
    quantities=("E", "F", "V"),
    shift_energy=None,
    xyz=None,
    force_magnitude=True,
    bins=120,
    figsize=None,
)

Error distributions of the run in path: one panel per quantity with the histogram of NEP − DFT on a log count axis, training and validation overlaid in the parity colours and annotated with RMSE / MAE / max, plus a last panel with the force error against the force magnitude.

quantities: any of E F V S. shift_energy ("mean" / "element", the latter needs xyz) removes a reference offset from the predicted energies first — for predictions of a set labelled with other DFT settings. force_magnitude=False drops the last panel, bins sets the histogram bins and figsize (cm) the figure size. Per-element errors have their own figure, :meth:periodic_table.

periodic_table

periodic_table(
    values=None,
    path=None,
    xyz=None,
    split="train",
    quantities=("E", "F"),
    out=None,
    cmap="YlOrRd",
    vmax=None,
    title=None,
    size=0.62,
    families=False,
)

Periodic table coloured by a per-element number: either the per-element RMSEs of a run (path + xyz -> :func:element_errors, one table per entry of quantities) or your own values ({label: {element: value}}). Elements without a value are grey. vmax: colour-scale top (default: the largest value); size: cell side in cm; families=True outlines the chemical families of the elements that carry a value (PT_FAMILIES) with a legend below.

element_errors

element_errors(path, xyz, split='train')

Per-element RMSEs of the *_<split>.out outputs of path for the frames of xyz: {"E": {el: meV/atom}, "F": {el: meV/A}, "n_atoms": {el: count}, "n_frames": {el: count}}. The force RMSE is over the atoms of the element; the energy RMSE over the frames that contain it (per-atom energies, every frame weighted once).

Extrapolation grade

build_active_set

build_active_set(
    model_file: str,
    xyz_file: str,
    output_file: str = "active_set.pt",
    rcond: float = 0.0001,
    tol: float = 1.01,
    sample_frames: int = 50000,
    init_rows: int = 4,
    max_passes: int = 10,
    device: str = None,
    precision: str = "float32",
    asi_file: str = None,
    chunk_atoms: int = None,
    energy_key: str = "energy",
    seed: int = 0,
    verbose: bool = True,
)

Build the MaxVol active set of a NEP model from its training set.

Streams xyz_file:

  1. over sample_frames random frames (all frames if fewer): the Gram matrix B^T B of every element, whose eigenvectors with singular value above rcond x the largest span the subspace of the grade, and a random sample of rows that seeds the active set (LU pivoting on init_rows x r rows, then MaxVol over the whole sample);
  2. over all frames: MaxVol swaps of every row whose grade exceeds tol;
  3. check passes (at most max_passes): rows that exceed tol against the final active set (a row passed early can, after later swaps) are swapped in, until a pass finds none — then every training atom has a grade <= tol. The log (and meta) reports whether it converged and the largest training grade.

The descriptors of pass 2 are kept in host memory for the check passes when they fit in 25% of the free memory (TORCHNEP_GAMMA_CACHE_GB sets another budget, 0 disables it).

precision="float32" (default) computes the descriptors and screens the rows of passes 2 and 3 in float32, regrading in float64 those near or above the threshold; "float64" does everything in float64. The Gram matrices, subspaces and MaxVol always run in float64. Either way every training atom ends with a grade <= tol; the active sets of the two precisions differ slightly, as any two MaxVol runs over rounded data.

Saves the :class:ActiveSet to output_file (and, with asi_file, in GPUMD's format, see :meth:ActiveSet.save_gpumd) and returns it.

select_structures

select_structures(
    model_file: str,
    active_set,
    xyz_file: str,
    output_xyz: str = None,
    max_frames: int = None,
    grade_min: float = None,
    grade_max: float = None,
    block_rows: int = None,
    output_active_set: str = None,
    device: str = None,
    precision: str = "float32",
    chunk_atoms: int = None,
    energy_key: str = "energy",
    verbose: bool = True,
)

Greedy D-optimal choice of structures that extend the training set.

A frame's grade is the largest over its atoms of max(gamma, gamma_res). Frames of xyz_file whose grade lies in (grade_min, grade_max] are candidates (grade_min defaults to the tol of the active set). They are visited from the highest grade down; a frame is taken when its grade, recomputed after the frames taken before it were added (their atoms enter the active set by MaxVol swaps, and their out-of-subspace directions extend the subspace), still exceeds grade_min — so near-duplicates of a taken frame are skipped. Frames with an element that has no active set grade inf. Stops after max_frames. The frames are graded here, no :func:compute_gamma call is needed first.

precision="float32" (default) computes the descriptors and the first grades of all frames in float32; the choice itself runs in float64. Writes the chosen frames verbatim to output_xyz and, if given, the active set extended by them to output_active_set (MaxVol within the subspace; the subspace itself changes only when the active set is rebuilt). Returns a dict with index (chosen frames, in order of choice), grade_at_choice and grade (grades of all frames against the original active set).

compute_gamma

compute_gamma(
    model_file: str,
    active_set,
    xyz_file: str,
    output_file: str = None,
    per_atom: bool = False,
    device: str = None,
    precision: str = "float32",
    chunk_atoms: int = None,
    energy_key: str = "energy",
    verbose: bool = True,
)

Extrapolation grades of every frame of xyz_file.

Grades only, to inspect structures; :func:select_structures grades the candidates itself when choosing. active_set: an :class:ActiveSet or the path of a saved one. Returns a dict of numpy arrays, per frame the maximum over its atoms of gamma / gamma_res / grade (the larger of the two, what :func:select_structures ranks by), atom (the atom with the largest grade), natoms, and with per_atom=True also gamma_atoms / gamma_res_atoms (every atom, file order). Saved as .npz when output_file is given.

precision="float32" (default) computes the descriptors and grades in float32 — gamma within 1e-4 and gamma_res within a few 1e-3 relative of "float64".

build_active_set_sharded

build_active_set_sharded(
    model_file: str,
    xyz_file: str,
    output_file: str = "active_set.pt",
    rcond: float = 0.0001,
    tol: float = 1.01,
    sample_frames: int = 50000,
    init_rows: int = 4,
    max_passes: int = 10,
    precision: str = "float32",
    asi_file: str = None,
    chunk_atoms: int = None,
    energy_key: str = "energy",
    seed: int = 0,
    verbose: bool = True,
)

Multi-GPU / multi-node :func:build_active_set: one process per GPU, launched with torchrun or srun (which set RANK / LOCAL_RANK / WORLD_SIZE / MASTER_ADDR / MASTER_PORT).

Every rank streams its own contiguous share of the frames. The Gram matrices are summed on the rank that owns each element (elements are dealt round-robin), which computes its subspace and seed active set; the ranks then run MaxVol on their shares, the owners merge the ranks' active rows and the check passes run as in :func:build_active_set, so the same guarantee holds: every training atom has a grade <= tol. Rank 0 writes output_file (and asi_file); every rank returns the same :class:ActiveSet. With one process this is :func:build_active_set.

select_structures_sharded

select_structures_sharded(
    model_file: str,
    active_set: str,
    xyz_file: str,
    output_xyz: str = None,
    max_frames: int = None,
    grade_min: float = None,
    grade_max: float = None,
    block_rows: int = None,
    output_active_set: str = None,
    precision: str = "float32",
    chunk_atoms: int = None,
    energy_key: str = "energy",
    verbose: bool = True,
)

Multi-GPU / multi-node :func:select_structures (torchrun / srun, one process per GPU): the ranks grade contiguous shares of the candidate frames, rank 0 makes the (sequential) greedy choice and writes the outputs; every rank returns the same dict. The choice is identical to :func:select_structures.

compute_gamma_sharded

compute_gamma_sharded(
    model_file: str,
    active_set: str,
    xyz_file: str,
    output_file: str = None,
    per_atom: bool = False,
    precision: str = "float32",
    chunk_atoms: int = None,
    energy_key: str = "energy",
    verbose: bool = True,
)

Multi-GPU / multi-node :func:compute_gamma (torchrun / srun, one process per GPU). Every rank grades a contiguous share of the frames; rank 0 assembles them in file order, writes output_file and returns the dict — the other ranks return None. The result is identical to :func:compute_gamma.

ActiveSet

ActiveSet(calc, desc_calc, rcond, tol)

MaxVol active sets of every element of a NEP model.

Built by :func:build_active_set, reloaded with :meth:load; grades of descriptors with :meth:gamma, GPUMD's format with :meth:save_gpumd. calc holds the float64 weights, desc_calc computes the descriptors in precision (float32 or float64).

load classmethod

load(path, model_file, device=None, precision='float32')

Load an active set for model_file (the model it was built with); precision: of the descriptors and grades computed with it.

save

save(path)

save_gpumd

save_gpumd(path, digits=12)

Write the active set in the ASI format read by GPUMD's compute_extrapolation (plain text; per element its symbol, K, K and a K x K matrix M, row by row). GPUMD grades an atom with row b as max |b M|; M = [V diag(scale) A^-1, 0] (K x r, zero-padded to K columns) makes that exactly gamma (not gamma_res). Elements without an active set are left out, with a warning: GPUMD cannot grade their atoms, so the MD must not contain them. Returns the elements written.

gamma

gamma(types, q, precision=None)

Per-atom grades of scaled descriptors q (N, D) with types (N,).

Returns (gamma, gamma_res) (N,) float64 tensors; atoms of elements without an active set get inf. precision (default: the active set's) "float32" grades in float32 — gamma within 1e-4, gamma_res within a few 1e-3 relative of the float64 grades.