Microsoft releases Skala 1.1, expands DFT model to five chemistry codes

Microsoft Research has released Skala-1.1, an updated version of its deep-learning exchange-correlation functional for density functional theory (DFT), and announced that Skala is now available inside the CP2K electronic-structure package, with integrations into Psi4, FHI-aims, ORCA and VASP in progress.
Skala-1.1 was trained on 2.5 times more data than the first public version of Skala. On GMTKN55, a widely used benchmark suite spanning 55 categories of chemistry including thermochemistry, reaction barriers and noncovalent interactions, the new model achieves a weighted average error of 2.8 kcal/mol. Microsoft says this accuracy surpasses today's leading global (range-separated) hybrid functionals while keeping the computational cost of a semi-local functional. Beyond energies, Skala-1.1 also produces accurate electron densities, dipole moments and molecular geometries. The gains come from an expanded Microsoft Research Accurate Chemistry Collection (MSR-ACC), the company's large-scale set of high-accuracy quantum-chemistry reference data generated with expensive wavefunction methods; for this release Microsoft added new categories, including electron affinities and noncovalent clusters, to increase both the size and the diversity of the training data.
Skala was designed as a continuously improving functional rather than a one-off release: each new version is meant to supersede the previous one at the same computational cost, instead of adding to what Microsoft calls the traditional "functional zoo" of accumulating, non-superseding functionals.
On distribution, Skala was first made available through an open-source community release built on (GPU4)PySCF and integrated with ASE. Microsoft has now worked with the team of Prof. Thomas D. Kühne at the Center for Advanced Systems Understanding (CASUS) to integrate Skala natively into CP2K, a package with more than 25 years of development that is widely used for large-scale DFT simulations and long-timescale molecular dynamics. The two teams built a suite of integration tests to confirm Skala produces numerically correct results inside CP2K, and describe the implementation and validation in a joint paper, "Molecular Implementation of the Machine-Learned Skala Exchange-Correlation Functional in CP2K through GauXC." Integration into Psi4, an open-source quantum chemistry platform, is underway with its developer community, which Microsoft says will bring Skala to three widely used open-source packages once complete; Microsoft is also working with the developers of FHI-aims, ORCA and VASP toward broader availability, without giving a completion date for any of these integrations.
Microsoft is also publishing a benchmarking harness and a living performance report that will track Skala's computational performance across successive releases, software implementations and hardware platforms, and that lets other package developers benchmark and validate their own Skala implementations. On performance today, Microsoft says Skala can match semi-local meta-GGA functionals on both CPUs and GPUs, with a CPU overhead that disappears for molecules larger than 20 to 30 atoms; no specific speedup or wall-clock figures are given.
Key facts
- Skala-1.1 was trained on 2.5x more data than the first public Skala and reaches a weighted average error of 2.8 kcal/mol on the 55-category GMTKN55 benchmark, which Microsoft says beats leading global hybrid functionals at semi-local cost.
- Skala is now available natively in CP2K, built with Prof. Thomas D. Kühne's team at CASUS and validated with a joint integration-test suite and paper.
- Integration into Psi4, FHI-aims, ORCA and VASP is underway but Microsoft gives no completion dates for any of them.
- Microsoft is publishing a living benchmark and performance report to track Skala's computational performance across future releases, codes and hardware.
- Skala matches semi-local meta-GGA performance on CPU and GPU, with the CPU overhead disappearing for molecules above roughly 20 to 30 atoms.
Why it matters
DFT underpins simulations across chemistry, materials science, catalysis, energy technologies and drug discovery, but its accuracy has historically traded off against computational cost. Skala's pitch is a deep-learning functional that pushes accuracy past today's hybrid functionals while keeping semi-local efficiency, and a release model where each version supersedes the last rather than adding to an ever-growing set of competing functionals. Landing inside CP2K, and heading into Psi4, FHI-aims, ORCA and VASP, moves that accuracy from a standalone tool into the codes labs already run.
Who it affects
Researchers and engineers who run DFT calculations in computational chemistry and materials science, in both academia and industry, are the direct audience. CP2K users get the integration first; users of Psi4, FHI-aims, ORCA and VASP are next once those integrations land. The CASUS team led by Prof. Thomas D. Kühne co-built the CP2K integration and its validation suite.
How to use it
Skala is already usable through an open-source community release built on (GPU4)PySCF and integrated with ASE, and now ships natively inside CP2K. Support for Psi4, FHI-aims, ORCA and VASP is being built with each package's developers but is not yet available, and Microsoft has not given a timeline for when any of these will be complete. No pricing or licensing terms are mentioned; the packages involved are established open-source and commercial electronic-structure codes.
How solid is it
The accuracy claim rests on GMTKN55, a standard 55-category chemistry benchmark suite, where Skala-1.1 posts a 2.8 kcal/mol weighted average error. The CP2K integration was checked against a dedicated suite of integration tests built jointly with the CASUS team and documented in a joint paper, which is a more concrete validation step than a bare accuracy number. The performance claims and the new living benchmark are self-published by Microsoft rather than independently audited.
Risks and caveats
Microsoft gives no timeline for the Psi4, FHI-aims, ORCA or VASP integrations, and no quantitative speedup or wall-clock numbers beyond the qualitative comparison to semi-local meta-GGAs and the 20-to-30-atom threshold. The update cadence for the living benchmark and performance report is also unstated. Because Skala is designed to have each release supersede the last, users adopting it should expect the functional itself to keep changing under them.