English | δΈζ
MLFCS is an ASE-first Python library for calculating symmetry-reduced force constants from atomic forces. It supports second and arbitrarily higher orders through one order-parameterized pipeline, with third- and fourth-order calculations as the primary production-validated paths. Fifth and higher orders can also be calculated and exported through the generic sparse HDF5 format; practical size is determined by cluster count, cutoff, supercell size, and available memory.
The third-order lineage of MLFCS explicitly references and draws on the algorithms, periodic-
image conventions, and sow/reap workflow of the GPL-licensed
thirdorder project. MLFCS develops these ideas into an
order-parameterized ASE/JAX architecture with new sparse, constraint-solving, acceleration, and
interoperability layers. See Third-party provenance for attribution and scope.
Numerical execution supports both CPU and GPU. CPU mode handles ordinary calculations and
large sparse linear algebra, while a CUDA-enabled JAX installation can move high-rank Cartesian
tensor rotations and batched transformations to a GPU. Memory is controlled through symmetry
reduction, contiguous sparse arrays, lazy dense materialization, matrix-free tensor actions,
small Gram null spaces, and sparse LSMR. JAX JIT, vmap, batched contractions, and displacement
deduplication improve throughput. Actual gains depend on the system, order, and hardware; GPU
execution does not replace cluster enumeration or sparse solvers that still run on the CPU.
The base package does not prescribe how forces are generated. Structures can be evaluated with any user-owned ASE Calculator or dispatched to an external workflow. An independent optional module uses phonopy and symfc to calculate temperature-dependent effective second-order force constants with a stochastic self-consistent harmonic approximation (SSCHA).
MLFCS provides a Python API only; it has no CLI.
An order-n force constant is the nth derivative of potential energy with respect to atomic
displacements, or equivalently an (n-1)-fold derivative of force. MLFCS computes it as follows:
ASE primitive structure
β
βΌ
deterministic supercell and neighbor cutoff
β
βΌ
space-group and permutation-reduced cluster orbits
β
βΌ
recursive central-difference displacement plan
β
βΌ
user-provided forces
β
βΌ
sparse symmetry reconstruction and optional strict ASR
β
βΌ
ForceConstants β HDF5 / NumPy / ShengBTE / phonopy
Space-group symmetry, force-constant index permutations, and stabilizer constraints reduce each cluster tensor to independent components. Only those components are sampled. The reconstructed result remains sparse until a dense representation is explicitly requested.
The acoustic sum rule (ASR) is imposed as a constrained projection in the independent orbit parameter space:
sum over one atom index of Phi(i1, ..., in) = 0
Small constraint systems use a Gram-matrix null space followed by sparse LSMR refinement; large systems use a sparse LSMR projection directly.
- One API and one numerical pipeline for
order >= 2. - Production-tested third- and fourth-order force constants.
- Analytic FC4 validation against an independently differentiated FCC Morse energy, including second-order finite-difference step convergence.
- End-to-end second- and fifth-order validation.
- ASE
Atomsand ASECalculatorat the public boundary. - External, checkpoint-friendly
sow()/reap()workflow. - Stable configuration IDs, plan hashes, and explicit atom-order mappings.
- Joint periodic-image cluster cutoff geometry.
- Recursive central-difference stencils.
- JAX-accelerated high-rank tensor transformations with CPU/GPU selection.
- JAX JIT,
vmap, and batched contractions for high-rank tensor throughput. - Contiguous sparse arrays, matrix-free actions, and lazy materialization to reduce peak memory.
- Displacement-key deduplication to reduce expensive calculator evaluations.
- Strict translational ASR using Gram null spaces and sparse LSMR.
- Generic sparse HDF5 for any order.
- ShengBTE output for orders 3 and 4.
- Full dense phonopy text output for order 2.
- Optional phonopy/symfc SSCHA with arbitrary ASE calculators.
MLFCS requires Python 3.12 or newer. Install and run it with uv:
uv syncRunnable API examples are available in examples/:
basic_fc2.pyruns FC2 directly with ASE's built-in EMT calculator;vasp_external_fc3.pyimplements a complete external VASPsow/ force collection /reapworkflow;nep89_orders.pyevaluates one or more orders with a user-supplied NEP89 model through calorine's ASE calculator.
Install the optional SSCHA dependencies when needed:
uv sync --extra sschaCalculator packages such as calorine or MACE are intentionally not base dependencies. Install the calculator required by your application separately.
| Quantity | Unit |
|---|---|
| Cell, positions, and displacements | Γ |
| Forces | eV/Γ |
Order-n force constants |
eV/Γ βΏ |
| Positive cutoff | Γ |
| Negative integer cutoff | Neighbor shell |
JAX numerical kernels use 64-bit floating point.
sow() does not write files or run a DFT program. It returns an ordered Python list of displaced
ASE Atoms objects. The user chooses how to serialize and calculate them: ASE can write each
structure as POSCAR-xxx for VASP or in another calculator's input format, after which the jobs
can be submitted by any local scheduler. When the calculations finish, use ASE to read each
result, extract its forces, restore the sow order (or key the forces by configuration ID), and
pass only those forces to reap(). If this positional order is guaranteed, configuration IDs
and a plan hash are not required.
For example, a positional VASP workflow is:
from pathlib import Path
import numpy as np
from ase.io import read, write
from mlfcs import ForceConstantCalculation
calculation = ForceConstantCalculation(
read("POSCAR"),
order=3,
supercell=(2, 2, 2),
cutoff=-5,
displacement=0.01,
symprec=1e-5,
jax_platform="auto", # "auto", "cpu", or "gpu"
)
# 1. sow(): obtain the displaced ASE structures in the exact reap order.
structures = calculation.sow(atom_order="grouped")
Path("vasp-jobs").mkdir(exist_ok=True)
for configuration_id, atoms in enumerate(structures):
job = Path("vasp-jobs") / f"POSCAR-{configuration_id + 1:03d}"
job.mkdir(exist_ok=True)
write(job / "POSCAR", atoms, format="vasp", direct=True, vasp5=True)
# 2. The user supplies INCAR, KPOINTS, POTCAR and submits every directory.
# MLFCS does not launch or configure VASP.
# 3. Read completed results with ASE in the same filename/order convention.
forces = []
for configuration_id in range(len(structures)):
job = Path("vasp-jobs") / f"POSCAR-{configuration_id + 1:03d}"
completed = read(job / "vasprun.xml", index=-1)
forces.append(completed.get_forces())
forces = np.asarray(forces)
# 4. reap(): the force at forces[i] must belong to structures[i].
fc3 = calculation.reap(
forces,
atom_order="grouped",
acoustic_sum_rule=True,
)
fc3.write("fc3.h5", format="hdf5")For Quantum ESPRESSO, ABINIT, CP2K, or another external program, replace only the ASE write()
and read() formats and provide that program's required input parameters. The sow/reap contract
is unchanged. Positional reap() needs no metadata when file names and returned forces preserve
the exact sow order. File formats such as POSCAR do not preserve the Python atoms.info metadata,
so a manifest containing the filename-to-configuration-ID relation and plan hash is recommended
for out-of-order jobs, restarts, long-term archives, and accidental-dataset detection. The complete
vasp_external_fc3.py example implements this optional safety
layer, force collection, missing-result checks, and final export; see the
external VASP workflow guide.
The force array must have shape:
(len(calculation.sow()), len(calculation.supercell), 3)
While the structures remain as ASE objects, every displaced structure also contains optional audit metadata:
atoms.info["mlfcs_configuration_id"]
atoms.info["mlfcs_plan_hash"]
atoms.info["mlfcs_atom_order"]
atoms.arrays["mlfcs_displacement"]When jobs return out of order, read them in any order and pass a mapping keyed by the original zero-based configuration ID:
fc3 = calculation.reap(
forces_by_configuration_id,
atom_order="grouped",
plan_hash=calculation.plan.hash,
)Missing or extra IDs, invalid shapes, non-finite values, and plan-hash mismatches are rejected.
calculator = make_my_ase_calculator()
fc3 = calculation.run(
calculator,
progress=lambda done, total: print(f"{done}/{total}"),
)Calculator evaluation is serial by design to avoid multiplying the memory used by large machine
learning potentials. Use sow() / reap() when external parallelism or checkpointing is needed.
For explicit checkpointing:
forces = calculation.evaluate(calculator)
np.savez_compressed("forces.npz", forces=forces, plan_hash=calculation.plan.hash)
fc3 = calculation.reap(forces, plan_hash=calculation.plan.hash)Stage reporting is enabled by default for sow(), reap(), and direct ASE-calculator runs.
It reports symmetry, cluster, displacement-plan, ASR, and force-evaluation progress using values
already computed by the calculation. Pass verbose=False to silence all stage and cutoff output.
report_cutoff=False suppresses only the detailed neighbor-shell lines.
The same constructor is used for all supported orders:
fc2_calculation = ForceConstantCalculation(atoms, order=2, cutoff=-6)
fc4_calculation = ForceConstantCalculation(atoms, order=4, cutoff=-3)
fc5_calculation = ForceConstantCalculation(atoms, order=5, cutoff=-1)A positive cutoff is a radius in Γ
. A negative integer selects a one-based neighbor shell. For
example, cutoff=-8 selects the eighth shell. MLFCS reports both the supercell capacity and the
selected radius:
Supercell neighbor limit: maximum shell = 33, maximum cutoff radius = 15.7504983443 Γ
Selected neighbor cutoff: shell = 8, cutoff radius = 7.5419604204 Γ
The first line is a capacity diagnostic for the finite supercell. The second line is the cutoff
actually used. Requests beyond the enumerable capacity are rejected. Use report_cutoff=False
to suppress both lines.
Higher orders grow combinatorially through cluster combinations, tensor components, permutations, and finite-difference signs. Use small cutoffs first and monitor configuration count and memory.
The canonical internal supercell order is:
z β y β x β primitive_atom
The primitive-atom index changes fastest. This is the default order used by sow() and reap().
For primitive-atom-grouped data:
structures = calculation.sow(atom_order="grouped")
force_constants = calculation.reap(forces, atom_order="grouped")Explicit mappings are available as:
calculation.index.grouped_permutation
calculation.index.internal_from_grouped
calculation.index.group_atoms(atoms)MLFCS performs the required grouped-order conversion automatically at the phonopy output boundary.
ASR is enabled by default:
constrained = calculation.reap(forces, acoustic_sum_rule=True)
raw = calculation.reap(forces, acoustic_sum_rule=False)The constrained result is the nearest solution in independent parameter space that satisfies translational invariance. Permutation symmetry supplies equivalent constraints on the other atom axes.
The output format is always explicit:
fc2.write("FORCE_CONSTANTS", format="phonopy")
fc2.write("fc2.hdf5", format="phonopy_hdf5")
fc3.write("fc3.h5", format="hdf5")
fc3.write("fc3.hdf5", format="phono3py_hdf5")
fc3.write("fc3.npz", format="numpy")
fc3.write("FORCE_CONSTANTS_3RD", format="shengbte")
fc4.write("FORCE_CONSTANTS_4TH", format="shengbte")| Format | Orders | Representation |
|---|---|---|
hdf5 |
Any | Sparse cluster tensors or dense arrays |
numpy / npz |
Any | Materialized NumPy arrays |
shengbte |
3 and 4 | Symmetry-closed translation-based text blocks |
phonopy |
2 | Full dense supercell FC2 text |
phonopy_hdf5 |
2 | Phonopy-compatible full-supercell force_constants HDF5 |
phono3py_hdf5 |
3 | Phono3py-compatible full-supercell fc3 HDF5 |
ShengBTE output is faithful by default: it writes exactly the symmetry-closed cluster support carried by the reconstructed sparse result. To reproduce the legacy thirdorder secondary joint-image filtering and block order, request compatibility explicitly:
fc3.write(
"FORCE_CONSTANTS_3RD",
format="shengbte",
compatibility="thirdorder",
)The phonopy and phono3py HDF5 writers use primitive-atom-grouped supercell order and stream one
first-atom slab at a time. They therefore do not materialize the full FC3 in memory. The native
hdf5 format remains the compact, order-parameterized MLFCS representation.
Sparse HDF5 is recommended for high orders. Dense materialization is explicit and emits a warning above the default 2 GB advisory budget:
fc5.write("fc5.h5", format="hdf5")
dense = fc5.materialize(5)
dense = fc5.materialize(5, max_bytes=None) # explicitly disable the warning budgetThe independent mlfcs.sscha module fits a temperature-dependent effective harmonic FC2 from
thermally sampled forces:
from mlfcs.sscha import SSCHA
sscha = SSCHA(
atoms,
supercell=(3, 3, 3),
temperature=300,
snapshots=1000,
max_iterations=10,
random_seed=42,
)
sscha.run(calculator)
sscha.use_average(last=5)
sscha.write("fc2-300K.hdf5", format="hdf5")Iteration zero fits an initial FC2 from small random Cartesian displacements. Each subsequent
iteration uses phonopy to sample the canonical harmonic ensemble and symfc to refit full FC2.
Thus max_iterations=10 performs one initialization and ten updates.
External per-iteration execution is also supported:
structures = sscha.sow()
result = sscha.reap(
forces_by_configuration_id,
energies=energies_by_configuration_id,
reference_energy=equilibrium_supercell_energy,
)Forces are sufficient for FC2 fitting. Energies are required only for the free-energy estimate.
Completed iterations are stored in sscha.history, and sscha.phonopy exposes the underlying
Phonopy object for meshes, bands, DOS, and thermal properties.
This is a stochastic effective-harmonic method, not an explicit FC3 bubble or FC4 loop calculation. See the SSCHA guide for details.
- Third and fourth order are the main production-tested finite-difference paths.
- Higher orders use the same implementation but can become prohibitively expensive.
- ShengBTE output is limited to orders 3 and 4.
- Explicit FC3 bubble and FC4 loop self-energies are not implemented.
- Non-analytic electrostatic corrections are not part of the base reconstruction pipeline.
- SSCHA stopping criteria are controlled by the caller; no universal automatic convergence threshold is imposed.
- Documentation index (δΈζ)
- External VASP workflow
- Technical overview
- Numerical validation and CI
- SSCHA guide
- Detailed old/new implementation comparison
Implementation comparisons, compatibility decisions, benchmark counts, and measured memory figures are intentionally kept in the comparison and technical documents rather than this user introduction.
All commands use uv and tests run serially:
uv sync --extra sscha
uv run pytest
uv run pytest -m "not reference"
uv run ruff check src tests reference_tools examples
uv run ruff format --check src tests reference_tools examples
uv buildhiphive and phono3py are development-only validation dependencies. The CI reference compares AlN FC3 values against an independent phono3py finite-difference result after hiphive converts both atom orderings and tensor representations to the same full-supercell form. The test hierarchy and independent reference commands are documented in tests/README.md.
The current release is 3.0.0. See CHANGELOG.md for release notes and
CONTRIBUTING.md for the development workflow.
MLFCS is distributed under the GNU General Public License v3.0 or later. Its third-order lineage and the boundary between borrowed ideas and new MLFCS development are documented in Third-party provenance.