Install
Get Poraquê 26.9.15
MIT licensed, pure Python. pip install poraque gets a released version; from source is what you want if you intend to change anything.
Requirements
- Python 3.11 or newer
- PyTorch 2.0+ (CUDA, Apple Metal or CPU)
- NumPy, SciPy, ASE, Matplotlib, PyYAML
- A C compiler, optional — for the faster CPU inference kernel
Install, train, predict, calibrate
Pick pip or an editable checkout — not both — then train, predict, and calibrate on labelled data before spending a DFT budget on unlabelled structures.
-
1. Install with pip
$ pip install poraque # symbolic distillation is an optional extra $ pip install "poraque[symbolic]" -
2. Install from source, editable
# if you intend to change anything: $ git clone https://github.com/seixas-research/poraque.git $ cd poraque && pip install -e . # symbolic distillation is an optional extra $ pip install -e ".[symbolic]" -
3. Train one ext2chg and one chg2tau model
$ poraque-train --config configs/train.yaml -
4. Predict a structure that has never been computed
$ poraque-inference new_structure/ \ --models models/au_w16_m8_l3/au_w16_m8_l3.pfno \ --output predictions/new_structure -
5. Calibrate, then select — in that order
# on LABELLED data first: costs nothing, and says whether # step 5 means anything at all $ poraque-committee --models "models/committee_*" --task ext2chg \ --cache data/cache/res32_potcar \ --against models/<name>/log/<name>.json # only once step 4 has passed: this one spends a DFT budget $ poraque-active-learning --models "models/committee_*" \ --task ext2chg --pool data/pool --select 5
Installing registers five console commands — poraque-train, poraque-inference, poraque-committee, poraque-active-learning and poraque-mp — which run from any directory once the environment is active. Each is the main() of the script of the same name under scripts/, so python scripts/poraque_train.py works identically and needs nothing installed.
Faster CPU inference (optional)
A small C kernel for the spectral contraction — the one part of a Fourier layer PyTorch runs poorly at batch 1, which is the shape every single-structure prediction has. It compiles itself on first use and needs no configuration.
$ python -m poraque.ml.backend --benchmark
2–3.4× on whole-model CPU inference at 24–32³ (less at 96³, where the FFTs dominate). Without a C compiler everything still works, falling back to torch.einsum; training is unaffected either way, and PORAQUE_C_BACKEND=0 turns it off explicitly.
Training on the Materials Project
poraque-mp turns a chemical space — a set of elements — into a local dataset of charge densities. Size it first: the estimate is exact, because charge densities are objects in a public S3 bucket and their sizes are read with HEAD requests that transfer no payload.
# a pure dry run: prints to the console, writes nothing at all
$ poraque-mp --elements Ag Au Pt --estimate
# download into ./data/MP, skipping anything over 20 MB
$ poraque-mp --elements Ag Au Pt --output data/MP --max-size-mb 20
chg2tau is not trainable on an MP archive — MP publishes no kinetic energy density. Set potcar_dir to the POTCAR library that generated the data; without it, the Gaussian pseudo-ion model stands in, and on the Ag–Au–Pt set the two differ by 0.38 relative L² — different fields, not different roundings of one.
Stuck on something?
The User Guide has a dedicated troubleshooting section — that is where installation and configuration problems usually live.
Documentation