Skip to content

Pipeline

From a directory of runs to a prediction

Four stages carry every run: ingest and cache the data once, train one config, predict on a structure that has never been computed, and evaluate what came out.

Pipeline stages

Ingest & cache

A VASP run directory, a flat archive of standalone CHGCARs from the Materials Project, or an already-prepared cache — all three normalize through one MaterialSource contract before anything downstream has to care which it was. Every material is then spectrally downsampled once to the training resolution and written to a per-material cache directory, so a second run over the same data costs nothing. Each cache remembers exactly what built it — resolution, potential source, every ingestion path — so two runs with different settings can never silently share stale fields under the same name.

Train

One training loop serves both tasks: AdamW, a cosine schedule, complex-aware gradient clipping for the spectral weights, and early stopping that restores the best-epoch weights rather than the last-epoch ones. K-fold cross-validation is one flag away, for a generalisation estimate with a spread rather than a single held-out score.

Predict

poraque-inference takes a bare structure directory, builds Vext analytically from its POTCAR, chains ext2chg into chg2tau, and writes CHGCAR-format ρ and τ — plus, on request, the Hartree potential (solved exactly, not predicted) and the full Kohn–Sham energy decomposition.

Evaluate & report

Every run writes a PDF report — loss curves, field-slice comparisons, parity plots — and a JSON metrics file archived beside the exact resolved config that produced it, so a result is reproducible without having to remember what changed.

Around the four stages

Sourcing data, spending a DFT budget wisely, and getting a number the pipeline itself can check.

Query-by-committee calibration

poraque-committee scores a labelled dataset and correlates disagreement against measured error — the free half of the active-learning loop, and the one to run first.

Active learning

poraque-active-learning ranks an unlabelled pool by that same disagreement measure and promotes the top-k structures into the training set, by move, copy or symlink.

Materials Project ingestion

poraque-mp turns a chemical space into a local dataset. Size is estimated with S3 HEAD requests before a byte transfers, so --estimate is a true dry run.

Symbolic distillation

An optional PySR extra fits a closed-form expression to the trained τ residual, with a physics-constrained objective and a Pareto front between accuracy and complexity.

Total-energy integration

poraque.physics.energy integrates the predicted fields into the full Kohn–Sham total energy, validated against exact Madelung constants and uniform-electron-gas limits.

Fast CPU inference

An optional C kernel for the spectral contraction — the one part of a Fourier layer PyTorch runs poorly at batch 1, which is every single-structure prediction.

Reading a DFT code's own output

One ingestion contract — geometry, parameters, pseudopotentials, fields — with each backend exposing exactly what it actually has.

Data ingestion
  • VASP 6.6+
Coming soon (in planning)
  • FHI-aims
  • Quantum ESPRESSO
  • GPAW

How a run stays reproducible

Config over code, one bundle over two files, ragged grids without padding — the decisions between pressing train and trusting the result.

Config over code

Every YAML key is optional and CLI flags override it in turn, so a run is fully described by one small file plus whatever it changed on the command line — and that resolved config is archived beside the results.

One bundle, both operators

ext2chg and chg2tau are saved into a single .pfno file rather than two, so the two halves of the chain cannot drift apart, be copied individually, or be mixed across training runs.

Ragged grids, one model

Materials keep whatever grid shape their ENCUT and cell size imply. The dataset buckets samples by shape for real batching, and a shape seen only once still trains — just in a batch of one.

How it is put together

fields/
The shared-grid field model: external potential, charge density, kinetic energy density, all on one mesh. VASP and FHI-aims I/O, plus the code-agnostic ingestion contract.
ml/
The Fourier neural operators, the differentiable DFT operators used to constrain them, and the training loop that serves both tasks.
physics/
Kohn–Sham total-energy components, integrated on the shared grid from the predicted fields.
data/
The Materials Project downloader, format detection across layouts, and mixed-source dataset assembly.
calculator.py
The ASE calculator wrapping the whole chain: geometry in, energy and fields out.
vis/
Figures and the automatic PDF report for every trained model.

The external potential is computed, never imported: the same code path builds it for training and for inference, so the model never sees at inference a field it did not also see during training.

Keep reading

The Hohenberg–Kohn map, the missing kinetic functional, and the physics constraints layered on top.

One Fourier layer for every grid shape, device- and precision-explicit, spectral resampling that preserves the electron count.

Held-out accuracy, the comparison against analytic functionals, how to get a trained model, and the roadmap beyond one element.

Try it on your own structure

Install it, or read the manual first.