Python API¶
The functions most scripts need are importable from the package itself:
from torchnep import train_nep, train_nep_sharded, predict_dataset, predict_dataset_sharded, export_valid_split
Training¶
train_nep ¶
train_nep(
config_file: str,
data_file: str,
output_dir: str = ".",
device: str = None,
precision: str = "float32",
use_autograd_forces: bool = False,
use_swa: bool = False,
swa_start: int = None,
use_compile: bool = None,
print_interval: int = 1,
restart: bool = True,
checkpoint_interval: int = 100,
prediction_interval: int = 100,
finetune_from: str = None,
resume_from: str = None,
recompute_q_scaler: bool = False,
slim_types: bool = False,
energy_key: str = "energy",
use_gpumd_qscaler: bool = False,
run_seed: int = None,
valid_file: str = None,
valid_ratio: float = None,
valid_strategy: str = "stratified",
neighbor_mode: str = "auto",
)
Train a NEP model on a single device (GPU / CPU / MPS).
neighbor_mode: how the training data is held in host memory —
cached (full neighbor lists with displacement vectors, fastest),
compact (int32 indices + int8 image shifts, ~4x smaller, rij
rebuilt on the device per batch), on_the_fly (positions + cells
only, the neighbor search runs on the device per batch — fits any
dataset). auto (default) takes the first of these whose estimated
peak footprint fits in 75% of this process's memory budget.
Hyperparameters (epoch / batch / lr / lambda_e,f,v / stage2* / …) come
from config_file only. See README for the full nep.in reference.
Launch: python run_train.py
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config_file
|
paths.
|
|
required |
data_file
|
paths.
|
|
required |
output_dir
|
paths.
|
|
required |
device
|
"cuda" | "cpu" | "mps" — auto-detected if omitted.
|
|
None
|
precision
|
"float32" (default) or "float64".
|
|
'float32'
|
use_autograd_forces
|
True -> autograd-through-rij forces (slower, gold
|
standard); False (default) -> analytical chain rule. The type-pair contraction backend is chosen automatically (see torchnep.ops.resolve_backend). |
False
|
use_swa
|
True -> maintain an averaged model during stage 2 and save it
|
as |
False
|
swa_start
|
first epoch included in the SWA average (default: the
|
last 100 epochs). Averaging all of stage 2 degrades energies. |
None
|
use_compile
|
None (default) = auto — compile on GPU, eager on
|
CPU; True/False force it. Missing Triton (GPU) or C++ toolchain (CPU) degrades to eager with a log note. epoch after a one-time first-epoch compilation cost; needs Triton, which ships with the CUDA PyTorch build). With autograd forces the first-order gradient is materialized via make_fx and compiled too. |
None
|
print_interval
|
log a line to screen every N epochs (all epochs still
|
land in output.log). |
1
|
restart
|
on fresh output_dir, write new log; otherwise resume from
|
checkpoint.pt if present. |
True
|
checkpoint_interval
|
save checkpoint.pt every N epochs.
|
|
100
|
prediction_interval
|
every N epochs, run predict_from_store with the
|
current-epoch weights and overwrite {energy,force,virial}_train.out in output_dir — lets you watch the parity-plot converge live. The end-of-training predict overwrites them with the nep_best.txt model. Set to 0 or a negative value to disable. |
100
|
finetune_from
|
path to an existing .pt or nep.txt to load weights from
|
(weights only — a NEW training starts from them: epoch 1, nep.in lr, fresh optimizer). The q_scaler stored with the weights is kept. |
None
|
resume_from
|
path to a checkpoint (.pt) to CONTINUE from — restores
|
model, optimizer (incl. lr), scheduler, epoch and best losses, so
the run picks up exactly where that checkpoint left off. Use
|
None
|
recompute_q_scaler
|
only with finetune_from. Default False — the
|
q_scaler stored with the loaded weights is kept (it is part of the potential they were trained as). True recomputes it from the new dataset, which rescales the NN inputs: the loaded model starts off wrong and must re-adapt (see bench_test/restart_rework/t8). |
False
|
slim_types
|
drop element types not present in data_file before training.
|
|
False
|
energy_key
|
name of the comment-line tag read as the reference energy
|
(default |
'energy'
|
use_gpumd_qscaler
|
Default False — torch's default init with the
|
self-consistent q_scaler (computed from the model's actual init
coefficients). On a 600-epoch 4-seed PdCuNiP benchmark this
converges to clearly better minima than the GPUMD-style start
(~12% lower E/V RMSE, ~3% lower F, train and validation alike).
True reproduces GPUMD's initialization instead: every parameter
re-initialised uniform(-1, 1) (SNES mu init) and the q_scaler
computed with all coefficients = 1.0 (GPUMD's generation-0
|
False
|
run_seed
|
master RNG seed for this run. None (default) -> a fresh random
|
seed each run, so repeated runs differ (independent weight init AND per-epoch batch shuffle) — the stochastic-testing behaviour. Pass an int to make a run fully reproducible: the same seed reproduces both the initial weights and the batch ordering bit-for-bit. The seed is saved into checkpoint.pt and restored on resume, so a resumed run continues along the very same shuffle stream regardless of what seed (or None) is passed at resume time. |
None
|
valid_file
|
path to a validation .xyz. Evaluated (frozen weights, full
|
set) every epoch: nep_best is the epoch with the lowest validation loss and the plateau LR scheduler steps on the validation loss — both anti-overfitting. Writes GPUMD-style *_test.out files. |
None
|
valid_ratio
|
alternative to valid_file — hold out this fraction of
|
data_file (e.g. 0.1) as the validation set. The split is drawn from run_seed, so it is reproducible for a given seed and is preserved exactly on resume (the checkpoint's seed wins). Mutually exclusive with valid_file. |
None
|
train_nep_sharded ¶
train_nep_sharded(
config_file: str,
data_file: str,
output_dir: str = ".",
precision: str = "float32",
use_autograd_forces: bool = False,
use_swa: bool = False,
swa_start: int = None,
use_compile: bool = None,
print_interval: int = 1,
restart: bool = True,
checkpoint_interval: int = 100,
prediction_interval: int = 100,
finetune_from: str = None,
resume_from: str = None,
recompute_q_scaler: bool = False,
slim_types: bool = False,
energy_key: str = "energy",
use_gpumd_qscaler: bool = False,
run_seed: int = None,
valid_file: str = None,
valid_ratio: float = None,
valid_strategy: str = "stratified",
neighbor_mode: str = "auto",
)
Data-sharded NEP training. Launch via torchrun (or any launcher that sets RANK / LOCAL_RANK / WORLD_SIZE / MASTER_ADDR / MASTER_PORT).
All hyperparameters (epoch / batch / lr / lambda_e,f,v / stage2* / ...)
come from config_file — see README for the nep.in reference.
Each rank loads structures[rank::world_size] only, so data-store memory scales as 1/world_size. Gradients are all-reduced by DDP; q_scaler and epoch metrics are all-reduced explicitly.
torchrun --nproc_per_node=N run_train.py
Parameters mirror train_nep (same restart semantics: resume is an
exact continuation — lr/optimizer/scheduler/SWA/best gates from the
checkpoint, nep.in's lr never re-applied on resume; resume_from
continues from a specific checkpoint such as checkpoint_stage1.pt;
finetune_from keeps the source model's q_scaler unless
recompute_q_scaler=True). Runtime-exclusive differences are the DDP
launch, CUDA/gloo backend auto-select, and distributed aggregation of
q_scaler / epoch metrics / the frozen-weight best evaluation.
run_seed mirrors train_nep: None -> a fresh random seed each run
(rank 0 draws it and broadcasts, so all replicas agree); pass an int for a
fully reproducible run. It seeds the weight init and the per-epoch shuffle,
and is saved into / restored from the checkpoint for exact resumption.
valid_file / valid_ratio mirror train_nep: a validation set
(own .xyz, or a run_seed-drawn holdout fraction of data_file) evaluated
every epoch — nep_best and the plateau LR schedule then follow the
validation loss. The validation frames are sharded across ranks like the
training frames; the error sums are all-reduced, so every rank sees the
identical validation loss (schedulers stay in lock-step).
Each rank keeps its shard in host memory and streams only the current
batch to its GPU (basis computed on the fly, CPU assembly prefetched
one batch ahead) — per-GPU memory scales with batch, not shard
size. DDP collectives are untouched: same step count per rank,
gradients all-reduced as usual.
export_valid_split ¶
export_valid_split(
data_file: str,
valid_ratio: float,
run_seed: int,
output_dir: str = "split",
strategy: str = "stratified",
min_stratum: int = 20,
)
Write GPUMD-ready train.xyz / test.xyz with train_nep's split.
Reproduces exactly the validation split that
train_nep(data_file, valid_ratio=r, run_seed=s, valid_strategy=...)
uses internally, so the exported pair can train the SAME data partition
in GPUMD (or any other code) and loss curves stay comparable. Frames
are copied verbatim (raw text, untouched fields and precision), in
input-file order.
strategy: "random" (default) or "stratified" — see
:func:stratified_split_indices.
Returns (train_path, test_path, n_train, n_valid).
parse_nep_in ¶
Parse nep.in parameter file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
filename
|
str
|
Path to nep.in file. |
required |
Returns:
| Type | Description |
|---|---|
dict
|
Dictionary of NEP parameters. |
Prediction¶
predict_dataset ¶
predict_dataset(
model_file: str,
xyz_file: str,
output_dir: str = ".",
dtype: str = "float32",
device: str = None,
batch_size: int = None,
verbose: bool = True,
energy_key: str = "energy",
output_descriptor: int = 0,
chunk_atoms: int = None,
)
Run streamed, batched prediction on a full dataset and save GPUMD-format outputs.
The file is first indexed (byte offsets + atom counts, one I/O pass, no
parsing), then processed in consecutive chunks of roughly
chunk_atoms atoms: each chunk is read, neighbor-listed (multi-process),
predicted in device-sized batches and appended to the output files before
the next chunk is touched. Host memory is therefore bounded by the chunk
size and device memory by the batch size — any dataset finishes on any
machine, only the wall time differs. Outputs are identical to the former
whole-dataset path.
Outputs (per-atom for energy and virial; per-atom raw for forces):
- energy_train.out: e_pred e_target (eV/atom, per frame)
- force_train.out: fx fy fz fx_t fy_t fz_t (eV/A, per atom)
- virial_train.out: xx yy zz xy yz zx (pred, ref) (eV/atom, per frame)
- stress_train.out: same layout in GPa (GPUMD sign convention)
- descriptor.out (only when output_descriptor != 0):
mode 1 — per-frame averaged scaled descriptor, one row per frame
mode 2 — per-atom scaled descriptor, one row per atom
Matches GPUMD's output_descriptor / descriptor.out schema.
The format mirrors GPUMD's *_train.out files, so the two can be diffed
column by column. When reference labels are present, a summary of the
energy/force/virial RMSE and MAE is printed at the end. The contraction
backend is chosen automatically (see torchnep.ops.resolve_backend).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
batch_size
|
int or None
|
Frames per compute batch. |
None
|
chunk_atoms
|
int or None
|
Atoms per streamed chunk (host-memory bound; ~1 GB per 200k atoms of
dense metal at cutoff 6/5). |
None
|
output_descriptor
|
int
|
0 — disabled (default).
1 — write per-frame averaged |
0
|
predict_dataset_sharded ¶
predict_dataset_sharded(
model_file: str,
xyz_file: str,
output_dir: str = ".",
dtype: str = "float32",
batch_size: int = None,
verbose: bool = True,
energy_key: str = "energy",
output_descriptor: int = 0,
chunk_atoms: int = None,
)
Multi-rank (multi-GPU / multi-node) version of :func:predict_dataset.
Launch one process per GPU with torchrun or srun (the launcher
sets RANK / LOCAL_RANK / WORLD_SIZE / MASTER_ADDR / MASTER_PORT). Rank 0
indexes the file once and broadcasts the frame index; the frames are then
split into WORLD_SIZE contiguous ranges balanced by atom count, every
rank streams its own range exactly like predict_dataset (same chunking,
auto batch size and OOM retry, bounded memory) into a per-rank part
directory, and rank 0 finally concatenates the parts in rank order into
the usual *_train.out files and prints the global E/F/V RMSE/MAE
table. The output is identical to a single-process predict_dataset.
Rank 0 shows the progress of its own share; the other ranks are silent.
With WORLD_SIZE == 1 (or no launcher) this is plain predict_dataset.
NEPCalculator ¶
NEP4 calculator: loads a trained model and computes atomic properties.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_file
|
str
|
Path to nep.txt (GPUMD NEP4 format). |
required |
dtype
|
dtype
|
Precision (default: float64 for accuracy). |
float64
|
device
|
str or device
|
Compute device (default: 'cpu'). |
'cpu'
|
compute ¶
compute(
species: list,
positions: ndarray,
cell: ndarray,
compute_descriptor: bool = False,
return_components: bool = False,
) -> Dict[str, torch.Tensor]
Compute energy, forces, and per-atom virial.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
species
|
list of str
|
Element symbols for each atom. |
required |
positions
|
(N, 3) array
|
Atomic positions in Angstrom. |
required |
cell
|
(3, 3) array
|
Lattice vectors (row-major). |
required |
compute_descriptor
|
bool
|
Also return scaled descriptors. |
False
|
return_components
|
bool
|
Also split the result into the NEP (neural-network) part and the
ZBL repulsive part. Adds |
False
|
Returns:
| Type | Description |
|---|---|
dict with 'energy', 'forces', 'virial', optionally 'descriptor' and
|
|
(if ``return_components``) the per-contribution breakdown above.
|
|
get_descriptor ¶
Compute scaled descriptors. Returns (N, dim) numpy.
compute_tiled ¶
compute_tiled(
species: list,
positions: ndarray,
cell: ndarray,
block_size="auto",
query_sub_chunk: int = 4000,
backend: str = "auto",
compile: bool = False,
) -> Dict[str, torch.Tensor]
Energy / forces / virial for large cells with bounded peak memory.
This is the MD-oriented counterpart of :meth:compute. Instead of
building the whole neighbor list and back-propagating through it (which
scales the peak memory with the total pair count — infeasible for a
500k-atom, high-density cell), it:
- builds the spatial bin table once (:class:
~torchnep.neighbor.CellList); - streams over blocks of
block_sizecentre atoms, querying only that block's neighbors at a time; - computes per-atom energy with the analytical force/virial path
(
ops.compute_analytical_forces) — no autograd graph is retained, so peak memory is set by a single block, not the whole system. (This is the same analytical kernel training/prediction use, hence alsotorch.compile-friendly.)
Each directed pair is owned by exactly one block (the block holding its
centre atom i); contributions scatter into global per-atom
accumulators, so the result is identical to :meth:compute to round-off.
For cells too small for the linked-cell decomposition this transparently
defers to :meth:compute (the autograd path is cheap there).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
species
|
as in :meth:`compute`.
|
|
required |
positions
|
as in :meth:`compute`.
|
|
required |
cell
|
as in :meth:`compute`.
|
|
required |
block_size
|
int or auto
|
Number of centre atoms processed per tile. Larger = faster but more peak memory (block-independent N-sized scratch + this block's pairs). "auto" (default) sizes it from the cell density, model dims, dtype and the device's free memory (~35% budget); pass an int to override. |
'auto'
|
query_sub_chunk
|
int
|
Inner chunk for the neighbor query (bounds the candidate tensor). |
4000
|
backend
|
"auto" | "loop" | "bmm" type-pair contraction (see ops).
|
|
'auto'
|
compile
|
bool
|
torch.compile the descriptor + analytical-force kernels (dynamic
shapes, so the pair-count changes between MD steps do NOT trigger
recompilation — one warm-up compile, then a single dynamic graph).
Forces the |
False
|
Returns:
| Type | Description |
|---|---|
dict with 'energy' (N,), 'forces' (N, 3), 'virial' (N, 9) — same keys and
|
|
conventions as :meth:`compute` (sans the component split / descriptor).
|
|
NEP ¶
NEP(
model_file,
dtype="float64",
device="cpu",
tiled="auto",
block_size="auto",
compile=False,
**kwargs,
)
Bases: Calculator
ASE calculator backed by a torchnep NEP4 model (nep.txt).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_file
|
str
|
Path to a GPUMD-format |
required |
dtype
|
str or dtype
|
Compute precision; |
'float64'
|
device
|
str or device
|
Torch device, e.g. |
'cpu'
|
**kwargs
|
Forwarded to |
{}
|
Notes
stress is only reported for fully periodic cells (a finite volume is
required); for molecules / clusters it is omitted. Non-periodic systems
are handled by embedding the atoms in a vacuum box large enough that no
atom sees a periodic image within the model cutoff.
get_components ¶
Full NEP / ZBL / total breakdown of energy, forces, and stress.
Returns {'nep': {...}, 'zbl': {...}, 'total': {...}} where each
inner dict has 'energy' (float, eV), 'forces' ((N, 3) eV/Å) and,
for periodic cells, 'stress' (6-vector Voigt, eV/ų). The total
block equals what get_potential_energy / get_forces /
get_stress return.
get_energy_components ¶
Return the NEP / ZBL / total potential-energy split (eV).
Useful to inspect how much of the energy comes from the ZBL repulsive baseline versus the neural-network NEP part::
{'nep': ..., 'zbl': ..., 'total': ...}
For a model trained without ZBL, 'zbl' is 0 and 'nep' equals
'total'. See :meth:get_components for forces and stress too.
Models¶
slim_model ¶
Return a new NEPModel containing only the specified element types.
Mathematically equivalent to the original for structures that contain only the kept element types. Parameters for removed types are discarded:
fitting_nets: only the nets for kept types are copiedc_param_2 / c_param_3: rows and columns for removed types are droppedb1,q_scaler: copied unchanged (independent of num_types)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
NEPModel
|
Source model (any number of types). |
required |
keep_type_names
|
list[str]
|
Subset of |
required |
Returns:
| Type | Description |
|---|---|
NEPModel
|
Smaller model on the same device/dtype as the source. |
Plotting¶
NEPPlotter ¶
NEPPlotter(
font="Arial",
fontsize=7,
dpi=300,
cmaps=None,
max_points=300000,
colors=None,
panel_labels="abcdefghijkl",
label_format="{}",
label_weight="bold",
frame=False,
rc=None,
font_dir=None,
cmap_range=(0.1, 0.9),
cmap_reverse=True,
)
Figures for torchnep / GPUMD training and prediction outputs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
font
|
str
|
Font family (default |
'Arial'
|
fontsize
|
float
|
Base font size in points (labels; ticks and legends one point smaller). |
7
|
dpi
|
int
|
Figure / saved-file resolution (default 300). |
300
|
cmaps
|
dict
|
Colormaps of the density (hexbin) panels per set, overrides of
|
None
|
cmap_range
|
(float, float)
|
Part of each density colormap used (default |
(0.1, 0.9)
|
cmap_reverse
|
bool
|
|
True
|
max_points
|
int
|
Scatter panels draw at most this many points (a fixed random subsample); metrics always use every point. |
300000
|
colors
|
dict
|
Overrides of |
None
|
panel_labels
|
str or sequence
|
Labels of the panels of multi-panel figures (default |
'abcdefghijkl'
|
frame
|
bool
|
|
False
|
font_dir
|
str
|
Folder of .ttf/.otf files to register before looking up |
None
|
rc
|
dict
|
Extra matplotlib rcParams applied on top of the built-in style. |
None
|
dashboard ¶
One figure per training run: (a) the training / validation RMSE
curves of E, F, V from loss.out and (b)(c)(d) the energy / force /
stress parity plots (scatter) with both sets overlaid — 2 x 2 panels
of equal size, or one row of three when the data carry no stress
labels. virial=True: the virial (eV/atom) instead of the stress.
stage2: epoch where stage 2 started (dashed line); "auto"
reads it from output.log / nep.in in path, None for
none. figsize in cm.
loss ¶
Training curves from loss.out: the E / F / V RMSEs against the
epoch on a log scale, training solid and validation faint (same
colour), a dashed line at the start of stage 2 — the panel of the
dashboard on its own. stress=True adds the stress RMSE.
path: the run directory or the loss.out file. figsize in cm.
parity ¶
parity(
path,
kind="scatter",
margins=False,
virial=False,
out=None,
quantities=None,
size=4.0,
shift_energy=None,
xyz=None,
title=None,
exclude=None,
natoms=None,
bins=80,
cell="hex",
)
Parity plots of the run in path: energy, force and stress (the
last only with stress labels) with training and validation overlaid
and R^2 / RMSE / MAE per set. kind="density": log-count 2-D histograms
(training Blues, validation Reds) with a small horizontal colorbar
under every panel. margins=True adds a strip on top of each panel
with the error NEP − DFT against the DFT value. The DFT and NEP axes
share limits and ticks. virial=True shows the
virial (eV/atom) instead of the stress in the third panel;
quantities overrides the panels altogether (any of E F V S);
shift_energy ("mean" /
"element", needs xyz) removes a reference offset from the
predicted energies; size is the side of one panel in cm.
exclude: frame indices (of the training split) to leave out, e.g.
an outlier list — needs natoms (atoms per frame, or the xyz)
to drop their force rows too. bins: density cells across the x
axis (default 80; fewer = larger cells); cell: "hex" (default)
or "square".
errors ¶
errors(
path,
out=None,
quantities=("E", "F", "V"),
shift_energy=None,
xyz=None,
force_magnitude=True,
bins=120,
figsize=None,
)
Error distributions of the run in path: one panel per quantity
with the histogram of NEP − DFT on a log count axis, training and
validation overlaid in the parity colours and annotated with
RMSE / MAE / max, plus a last panel with the force error against the
force magnitude.
quantities: any of E F V S. shift_energy ("mean" /
"element", the latter needs xyz) removes a reference offset
from the predicted energies first — for predictions of a set labelled
with other DFT settings. force_magnitude=False drops the last
panel, bins sets the histogram bins and figsize (cm) the
figure size. Per-element errors have their own figure,
:meth:periodic_table.
periodic_table ¶
periodic_table(
values=None,
path=None,
xyz=None,
split="train",
quantities=("E", "F"),
out=None,
cmap="YlOrRd",
vmax=None,
title=None,
size=0.62,
families=False,
)
Periodic table coloured by a per-element number: either the
per-element RMSEs of a run (path + xyz -> :func:element_errors,
one table per entry of quantities) or your own values
({label: {element: value}}). Elements without a value are grey.
vmax: colour-scale top (default: the largest value); size: cell
side in cm; families=True outlines the chemical families of the
elements that carry a value (PT_FAMILIES) with a legend below.
element_errors ¶
Per-element RMSEs of the *_<split>.out outputs of path for the
frames of xyz: {"E": {el: meV/atom}, "F": {el: meV/A},
"n_atoms": {el: count}, "n_frames": {el: count}}. The force RMSE is
over the atoms of the element; the energy RMSE over the frames that
contain it (per-atom energies, every frame weighted once).
Extrapolation grade¶
build_active_set ¶
build_active_set(
model_file: str,
xyz_file: str,
output_file: str = "active_set.pt",
rcond: float = 0.0001,
tol: float = 1.01,
sample_frames: int = 50000,
init_rows: int = 4,
max_passes: int = 10,
device: str = None,
precision: str = "float32",
asi_file: str = None,
chunk_atoms: int = None,
energy_key: str = "energy",
seed: int = 0,
verbose: bool = True,
)
Build the MaxVol active set of a NEP model from its training set.
Streams xyz_file:
- over
sample_framesrandom frames (all frames if fewer): the Gram matrixB^T Bof every element, whose eigenvectors with singular value abovercondx the largest span the subspace of the grade, and a random sample of rows that seeds the active set (LU pivoting oninit_rowsx r rows, then MaxVol over the whole sample); - over all frames: MaxVol swaps of every row whose grade exceeds
tol; - check passes (at most
max_passes): rows that exceedtolagainst the final active set (a row passed early can, after later swaps) are swapped in, until a pass finds none — then every training atom has a grade <=tol. The log (andmeta) reports whether it converged and the largest training grade.
The descriptors of pass 2 are kept in host memory for the check passes
when they fit in 25% of the free memory (TORCHNEP_GAMMA_CACHE_GB sets
another budget, 0 disables it).
precision="float32" (default) computes the descriptors and screens
the rows of passes 2 and 3 in float32, regrading in float64 those near
or above the threshold; "float64" does everything in float64. The
Gram matrices, subspaces and MaxVol always run in float64. Either way
every training atom ends with a grade <= tol; the active sets of the
two precisions differ slightly, as any two MaxVol runs over rounded data.
Saves the :class:ActiveSet to output_file (and, with asi_file,
in GPUMD's format, see :meth:ActiveSet.save_gpumd) and returns it.
select_structures ¶
select_structures(
model_file: str,
active_set,
xyz_file: str,
output_xyz: str = None,
max_frames: int = None,
grade_min: float = None,
grade_max: float = None,
block_rows: int = None,
output_active_set: str = None,
device: str = None,
precision: str = "float32",
chunk_atoms: int = None,
energy_key: str = "energy",
verbose: bool = True,
)
Greedy D-optimal choice of structures that extend the training set.
A frame's grade is the largest over its atoms of max(gamma, gamma_res).
Frames of xyz_file whose grade lies in (grade_min, grade_max]
are candidates (grade_min defaults to the tol of the active set).
They are visited from the highest grade down; a frame is taken when its
grade, recomputed after the frames taken before it were added (their
atoms enter the active set by MaxVol swaps, and their out-of-subspace
directions extend the subspace), still exceeds grade_min — so
near-duplicates of a taken frame are skipped. Frames with an element
that has no active set grade inf. Stops after max_frames. The frames
are graded here, no :func:compute_gamma call is needed first.
precision="float32" (default) computes the descriptors and the first
grades of all frames in float32; the choice itself runs in float64.
Writes the chosen frames verbatim to output_xyz and, if given, the
active set extended by them to output_active_set (MaxVol within the
subspace; the subspace itself changes only when the active set is
rebuilt). Returns a dict with index (chosen frames, in order of
choice), grade_at_choice and grade (grades of all frames against
the original active set).
compute_gamma ¶
compute_gamma(
model_file: str,
active_set,
xyz_file: str,
output_file: str = None,
per_atom: bool = False,
device: str = None,
precision: str = "float32",
chunk_atoms: int = None,
energy_key: str = "energy",
verbose: bool = True,
)
Extrapolation grades of every frame of xyz_file.
Grades only, to inspect structures; :func:select_structures grades the
candidates itself when choosing. active_set: an :class:ActiveSet or
the path of a saved one. Returns a dict of numpy arrays, per frame the
maximum over its atoms of gamma / gamma_res / grade (the
larger of the two, what :func:select_structures ranks by), atom
(the atom with the largest grade), natoms, and with per_atom=True
also gamma_atoms / gamma_res_atoms (every atom, file order).
Saved as .npz when output_file is given.
precision="float32" (default) computes the descriptors and grades in
float32 — gamma within 1e-4 and gamma_res within a few 1e-3 relative of
"float64".
build_active_set_sharded ¶
build_active_set_sharded(
model_file: str,
xyz_file: str,
output_file: str = "active_set.pt",
rcond: float = 0.0001,
tol: float = 1.01,
sample_frames: int = 50000,
init_rows: int = 4,
max_passes: int = 10,
precision: str = "float32",
asi_file: str = None,
chunk_atoms: int = None,
energy_key: str = "energy",
seed: int = 0,
verbose: bool = True,
)
Multi-GPU / multi-node :func:build_active_set: one process per GPU,
launched with torchrun or srun (which set RANK / LOCAL_RANK /
WORLD_SIZE / MASTER_ADDR / MASTER_PORT).
Every rank streams its own contiguous share of the frames. The Gram
matrices are summed on the rank that owns each element (elements are
dealt round-robin), which computes its subspace and seed active set; the
ranks then run MaxVol on their shares, the owners merge the ranks' active
rows and the check passes run as in :func:build_active_set, so the same
guarantee holds: every training atom has a grade <= tol. Rank 0 writes
output_file (and asi_file); every rank returns the same :class:ActiveSet. With one
process this is :func:build_active_set.
select_structures_sharded ¶
select_structures_sharded(
model_file: str,
active_set: str,
xyz_file: str,
output_xyz: str = None,
max_frames: int = None,
grade_min: float = None,
grade_max: float = None,
block_rows: int = None,
output_active_set: str = None,
precision: str = "float32",
chunk_atoms: int = None,
energy_key: str = "energy",
verbose: bool = True,
)
Multi-GPU / multi-node :func:select_structures (torchrun /
srun, one process per GPU): the ranks grade contiguous shares of the
candidate frames, rank 0 makes the (sequential) greedy choice and writes
the outputs; every rank returns the same dict. The choice is identical to
:func:select_structures.
compute_gamma_sharded ¶
compute_gamma_sharded(
model_file: str,
active_set: str,
xyz_file: str,
output_file: str = None,
per_atom: bool = False,
precision: str = "float32",
chunk_atoms: int = None,
energy_key: str = "energy",
verbose: bool = True,
)
Multi-GPU / multi-node :func:compute_gamma (torchrun / srun,
one process per GPU). Every rank grades a contiguous share of the frames;
rank 0 assembles them in file order, writes output_file and returns
the dict — the other ranks return None. The result is identical to
:func:compute_gamma.
ActiveSet ¶
MaxVol active sets of every element of a NEP model.
Built by :func:build_active_set, reloaded with :meth:load; grades of
descriptors with :meth:gamma, GPUMD's format with :meth:save_gpumd.
calc holds the float64 weights, desc_calc computes the
descriptors in precision (float32 or float64).
load
classmethod
¶
Load an active set for model_file (the model it was built with);
precision: of the descriptors and grades computed with it.
save_gpumd ¶
Write the active set in the ASI format read by GPUMD's
compute_extrapolation (plain text; per element its symbol, K, K and
a K x K matrix M, row by row). GPUMD grades an atom with row b as
max |b M|; M = [V diag(scale) A^-1, 0] (K x r, zero-padded to K
columns) makes that exactly gamma (not gamma_res). Elements without an
active set are left out, with a warning: GPUMD cannot grade their
atoms, so the MD must not contain them. Returns the elements written.
gamma ¶
Per-atom grades of scaled descriptors q (N, D) with types (N,).
Returns (gamma, gamma_res) (N,) float64 tensors; atoms of elements
without an active set get inf. precision (default: the active
set's) "float32" grades in float32 — gamma within 1e-4, gamma_res
within a few 1e-3 relative of the float64 grades.