Microsoft's Skala 1.1 improves DFT accuracy, expands to more chemistry software

Microsoft's Skala 1.1 improves DFT accuracy, expands to more chemistry software

Microsoft Research has released Skala 1.1, the latest version of its deep-learning exchange-correlation functional for density functional theory (DFT), the computational method behind most chemistry and materials-science simulations. Trained on 2.5 times more data than the first public version of Skala, the new model delivers substantially higher accuracy across core molecular-simulation tasks: main-group thermochemistry, reaction kinetics and molecular structure prediction. On GMTKN55, a widely used benchmark suite spanning 55 categories of chemistry, including thermochemistry, reaction barriers and noncovalent interactions, Skala 1.1 reaches a weighted average error of 2.8 kcal/mol. Microsoft says this level of accuracy now surpasses today's leading global (range-separated) hybrid functionals while keeping the computational cost of a semi-local functional, and that the model also produces more accurate electron densities, dipole moments and molecular geometries. The gain traces back to an expansion of the Microsoft Research Accurate Chemistry Collection (MSR-ACC), the company's large-scale set of high-accuracy quantum-chemistry reference data generated with expensive wavefunction methods: for Skala 1.1, Microsoft added new categories, including electron affinities and noncovalent clusters, increasing both the size and the diversity of the training data.

Alongside the accuracy work, Microsoft is expanding where Skala can be used. The functional first shipped as an open-source community release built on (GPU4)PySCF and integrated with ASE, letting researchers evaluate and apply it with minimal effort while benefiting from optimized CPU and GPU performance. Microsoft now says Skala is available in CP2K, an open-source DFT package with more than 25 years of development that is widely used for large-scale systems and long-timescale molecular dynamics. That integration was built together with the team of Prof. Thomas D. Kühne at the Center for Advanced Systems Understanding (CASUS); the CASUS team's expertise in numerical verification helped design and validate a dedicated suite of integration tests, and the work is written up in a joint paper, 'Molecular Implementation of the Machine-Learned Skala Exchange-Correlation Functional in CP2K through GauXC.' Beyond CP2K, Microsoft says it is actively integrating Skala into the open-source Psi4 package together with Psi4's developer community, an effort it describes as under way rather than finished. Separately, Microsoft says it is working with the developers of FHI-aims, ORCA and VASP with the goal of making Skala broadly accessible on those platforms too, describing this as a collaboration pursuing that goal rather than integration work already in progress. Microsoft gives one reason for spreading Skala this widely: no single software package can meet the needs of every application or research community, so reaching scientists means reaching the packages they already use.

On raw performance, Microsoft says Skala already delivers speeds comparable to semi-local meta-GGA functionals on both CPUs and GPUs; on CPUs, an overhead relative to those functionals disappears once a molecule has more than 20 to 30 atoms. To track performance as new Skala releases, library updates such as to GauXC, and hardware-specific optimizations arrive, Microsoft is publishing a benchmarking harness together with a 'living performance report' that will keep being updated, covering a range of tasks and hardware platforms; the same harness lets developers of each software package benchmark, validate and improve their own Skala implementations. Microsoft describes the overall approach as a break from the traditional 'functional zoo,' where new DFT functionals accumulate without replacing older ones: each Skala release, including 1.1, is meant to supersede the one before it as new data, model architectures and training strategies become available, while keeping the same practical computational cost.

Key facts

  • Skala 1.1, trained on 2.5 times more data than the first public version of Skala, reaches a weighted average error of 2.8 kcal/mol on GMTKN55, a 55-category chemistry benchmark suite.
  • Skala is now available in the CP2K package, built together with Prof. Thomas D. Kühne's team at the Center for Advanced Systems Understanding (CASUS), and Microsoft is actively integrating it into Psi4.
  • Microsoft is also working with the developers of FHI-aims, ORCA and VASP toward bringing Skala to those platforms, a collaboration it describes as pursuing that goal rather than integration already under way.
  • Microsoft says Skala's accuracy now surpasses today's leading global range-separated hybrid functionals while keeping the computational cost of a semi-local functional.
  • Microsoft is publishing a benchmarking harness and a 'living performance report' that will keep tracking Skala's computational performance across hardware platforms as new releases and optimizations arrive.

Why it matters

DFT underpins a huge share of real-world computational chemistry and materials-science work, from catalysis and energy technologies to drug discovery, so a more accurate functional has value across all of it. Microsoft frames Skala as a deliberate break from the field's usual pattern, where new functionals pile up alongside older ones instead of replacing them, a situation the post calls a 'functional zoo.' Skala instead follows one continuously improving line: each release, including 1.1, is built to supersede the last as new data, architectures and training strategies arrive, without raising the computational cost. Skala 1.1 is Microsoft's first public demonstration that this improvement cycle works in practice, and pairing it with integration into widely used packages plus a public, ongoing performance benchmark is meant to make the gains checkable and usable outside Microsoft's own tooling, not just claimed in a blog post.

Who it affects

Researchers and engineers running DFT calculations in computational chemistry, materials science, catalysis, energy technology and drug discovery are the direct audience. Users of Microsoft's original open-source Skala Community Edition, built on (GPU4)PySCF and ASE, get the improved model automatically. CP2K users can use Skala today, following the integration built together with Prof. Thomas D. Kühne's team at CASUS. Psi4 users will get it once that in-progress integration, done together with Psi4's developer community, is complete. Users of FHI-aims, ORCA or VASP may eventually get it too, since Microsoft says it is working with those packages' developers toward that goal, though no integration timeline is given for them.

How to use it

Skala is available now in two ways, according to Microsoft. First, as an open-source community release built on (GPU4)PySCF and integrated with ASE, letting researchers apply it directly with optimized CPU and GPU performance. Second, inside the open-source CP2K package, following the completed integration with the CASUS team, which Microsoft describes as validated through a dedicated suite of integration tests and documented in a joint paper on the implementation. Integration into the open-source Psi4 package is under way but not yet described as finished. Microsoft also says it is working with the developers of FHI-aims, ORCA and VASP with the goal of bringing Skala to those platforms, without stating when or whether that will land. Anyone wanting to check Skala's speed on their own hardware can use the newly published benchmarking harness, which package developers can also use to benchmark, validate and improve their own Skala implementations; the accompanying 'living performance report' will be updated as new optimizations arrive. The source gives no price or licence terms beyond describing these releases as open source.

How solid is it

This account comes directly from Microsoft Research, the team building Skala, so the accuracy and performance figures are self-reported rather than independently verified in this post. Two things support the numbers even so: GMTKN55, the benchmark behind the 2.8 kcal/mol figure, is described as a widely used external suite rather than one Microsoft built itself, and the CP2K integration specifically was checked with a dedicated integration-test suite built together with CASUS, a team credited with deep expertise in numerical verification, and written up in a joint paper. That paper covers CP2K only. For Psi4, FHI-aims, ORCA and VASP, the post gives no benchmark numbers and no validation paper, only a stated stage of progress for each: Psi4 in progress, the other three at the collaboration-goal stage. Those parts of the story remain to be confirmed once the work is further along. The post itself carries no named author or byline; it speaks throughout in Microsoft Research's collective voice.

Risks and caveats

The post's own summary bullet blurs a distinction its body draws carefully: it groups Psi4, FHI-aims, ORCA and VASP together as all 'being integrated,' but the body says only CP2K integration is complete, Psi4 integration is in progress, and FHI-aims, ORCA and VASP are described only as a collaboration pursuing a goal, not integration work already under way. No numeric benchmark results are given for Skala's performance inside Psi4, FHI-aims, ORCA or VASP, and the new 'living performance report' is not summarized with any figures in this post. The CPU overhead that 'disappears' above roughly 20 to 30 atoms is stated only for CPUs; nothing is said one way or the other about an equivalent overhead on GPUs. The post gives no specific calendar date for the announcement beyond the word 'today,' and no individual author is named.