Benchmarks
Caffeine
24 atoms · energy + forces · Apple M4 Max CPU
| Calculation | Time | Peak RSS |
|---|---|---|
| PBE / def2-SVPqc.rs · 8 CPU threads | 9.1× faster 1.70 sAtomli 15.5 sPySCF | 90% less memory 394 MiBAtomli 3.87 GiBPySCF |
| GFN2-xTBctb.rs · 8 CPU threads | 3× faster 3.89 msAtomli 11.7 mstblite | 86.6 MiBAtomli 85.2 MiBtblite |
| Nequix MP-1 PFTmlip.rs · 1 CPU thread | 1.7× faster 7.78 msAtomli 12.8 msNequix / JAX | 78% less memory 172 MiBAtomli 793 MiBNequix / JAX |
| Equiformer V3mlip.rs · 1 CPU thread | 2.1× faster 54.5 msAtomli 115 msPyTorch | 72% less memory 307 MiBAtomli 1.08 GiBPyTorch |
NVIDIA A100
| Calculation | Time | Peak RSS |
|---|---|---|
| SKALA 1.1 · 8 host CPU threads | 1.7× faster 1.56 sqc.rs 2.69 sSKALA / GPU4PySCF | 71% less memory 666 MiBqc.rs 2.22 GiBSKALA / GPU4PySCF |
| Nequix PFT · 1 host CPU thread | 2.7× faster 3.28 msmlip.rs 9.01 msNequix / JAX | 77% less memory 398 MiBmlip.rs 1.69 GiBNequix / JAX |
CPU DFT and MLIP: Atomli 0.1.4 · xTB: Atomli 0.1.6 · development build, 2026-09-20 · qc.rs A100: Atomli 0.1.5 · development build
qc.rs ↗
qc.rs / PySCF / GPU4PySCF / CP2KPBE · r²SCAN · SKALA 1.1 · molecular and periodic AO scalingxTB ↗
ctb.rs / tbliteGFN2-xTB · g-xTBMLIP ↗
mlip.rs / JAX / PyTorchNequix PFT · Equiformer V3 · CPU and GPUMachines and methods
| Machine | CPU | Memory | GPU |
|---|---|---|---|
| M3 Ultra | Apple M3 Ultra, 32 cores | 256 GiB unified | Metal |
| EPYC | AMD EPYC 7702P, 64 cores | 503 GiB | NVIDIA A100 |
| M4 Max | Apple M4 Max, 16 cores | 128 GiB unified | Metal |
QC and xTB use eight CPU threads; MLIP uses one. Each comparison uses the same CPU or physical GPU. Times are medians of three warm evaluations after construction and warm-up. QC scaling uses energy-only calls divided by SCF iterations. Complete calculations include forces and periodic stress.