Skip to content

Install

Get Poraquê 26.9.15

MIT licensed, pure Python. pip install poraque gets a released version; from source is what you want if you intend to change anything.

Requirements

  • Python 3.11 or newer
  • PyTorch 2.0+ (CUDA, Apple Metal or CPU)
  • NumPy, SciPy, ASE, Matplotlib, PyYAML
  • A C compiler, optional — for the faster CPU inference kernel

Install, train, predict, calibrate

Pick pip or an editable checkout — not both — then train, predict, and calibrate on labelled data before spending a DFT budget on unlabelled structures.

  1. 1. Install with pip

    $ pip install poraque # symbolic distillation is an optional extra $ pip install "poraque[symbolic]"
  2. 2. Install from source, editable

    # if you intend to change anything: $ git clone https://github.com/seixas-research/poraque.git $ cd poraque && pip install -e . # symbolic distillation is an optional extra $ pip install -e ".[symbolic]"
  3. 3. Train one ext2chg and one chg2tau model

    $ poraque-train --config configs/train.yaml
  4. 4. Predict a structure that has never been computed

    $ poraque-inference new_structure/ \ --models models/au_w16_m8_l3/au_w16_m8_l3.pfno \ --output predictions/new_structure
  5. 5. Calibrate, then select — in that order

    # on LABELLED data first: costs nothing, and says whether # step 5 means anything at all $ poraque-committee --models "models/committee_*" --task ext2chg \ --cache data/cache/res32_potcar \ --against models/<name>/log/<name>.json # only once step 4 has passed: this one spends a DFT budget $ poraque-active-learning --models "models/committee_*" \ --task ext2chg --pool data/pool --select 5

Installing registers five console commands — poraque-train, poraque-inference, poraque-committee, poraque-active-learning and poraque-mp — which run from any directory once the environment is active. Each is the main() of the script of the same name under scripts/, so python scripts/poraque_train.py works identically and needs nothing installed.

Faster CPU inference (optional)

A small C kernel for the spectral contraction — the one part of a Fourier layer PyTorch runs poorly at batch 1, which is the shape every single-structure prediction has. It compiles itself on first use and needs no configuration.

$ python -m poraque.ml.backend --benchmark

2–3.4× on whole-model CPU inference at 24–32³ (less at 96³, where the FFTs dominate). Without a C compiler everything still works, falling back to torch.einsum; training is unaffected either way, and PORAQUE_C_BACKEND=0 turns it off explicitly.

Training on the Materials Project

poraque-mp turns a chemical space — a set of elements — into a local dataset of charge densities. Size it first: the estimate is exact, because charge densities are objects in a public S3 bucket and their sizes are read with HEAD requests that transfer no payload.

# a pure dry run: prints to the console, writes nothing at all $ poraque-mp --elements Ag Au Pt --estimate # download into ./data/MP, skipping anything over 20 MB $ poraque-mp --elements Ag Au Pt --output data/MP --max-size-mb 20

chg2tau is not trainable on an MP archive — MP publishes no kinetic energy density. Set potcar_dir to the POTCAR library that generated the data; without it, the Gaussian pseudo-ion model stands in, and on the Ag–Au–Pt set the two differ by 0.38 relative L² — different fields, not different roundings of one.

Stuck on something?

The User Guide has a dedicated troubleshooting section — that is where installation and configuration problems usually live.

Documentation