2026-09-09 | | Total: 35
Conventional semi-local density functional approximations for the exchange-correlation energy systematically underestimate fundamental band gaps. Advanced meta-generalized gradient approximations (meta-GGAs) can partially overcome this limitation through their orbital dependence via the kinetic energy density. In principle, the flexible meta-GGA form should allow for balanced performance across diverse properties. However, atomization energies, reaction barrier heights, and lattice constants, as well as band gaps, exhibit different sensitivities to the large parameter space of meta-GGA functionals. In this work, we explore the space of appropriate norms that guide the r$^2$SCAN-type functional construction through the interpolation function connecting two fundamental paradigms: the one-electron limit and the uniform electron gas limit. We introduce local modifications to a strongly-smoothed interpolation function and investigate their impact on functional performance. To further probe the limits of the meta-GGA framework, we include spin-unpolarized bonded systems among the set of norms. The resulting bonded-norm bn-r$^2$SCAN meta-GGA achieves balance: band-gap predictions approaching the accuracy of the ultra-nonlocal LAK meta-GGA and lattice constants of nearly r$^2$SCAN accuracy.
Model Hamiltonians (e.g., the Hubbard or Heisenberg models) provide a simple yet physically meaningful description of complex electronic phenomena through a small number of parameters, such as electron hopping integrals or magnetic exchange couplings. Their construction typically exploits the local nature of electron correlation effects and therefore relies on localized orbitals. Although many localization schemes are available, the resulting localized orbitals depend not only on the chosen localization functional but also on the orbital space to which the localization procedure is applied. These choices are often treated as technical details and are rarely discussed explicitly. Here, we demonstrate that different, chemically reasonable choices of localized orbitals can lead to substantially different model Hamiltonian parameters, even when they produce nearly identical electronic energies. Using Density Matrix Downfolding, we systematically investigate how the localization functional and, more importantly, the choice of model orbital space affect effective Hamiltonians derived from ab initio calculations for representative $π$-conjugated systems and transition metal complexes.
We present the cluster mean field(cMF) and its extensions into perturbation theory(cPT) and coupled cluster(cCC) methods for fermionic systems. We discuss how the commutation rules for fermionic systems modifies the method as well as other computational details. We then perform calculations on the Hubbard model and a few select conjugated molecules to demonstrate the efficacy of the method. We will also use these examples to compare cluster methods to CAS-SCF.
Aqueous redox-flow batteries based on TEMPO derivatives are promising for large-scale energy storage, but their practical use is limited by the chemical instability of the oxidized N -oxoammonium state. In this work, we investigate the degradation of five TEMPO derivatives using ab initio molecular dynamics combined with enhanced sampling. Two proposed degradation mechanisms, ring opening and Hofmann elimination, are examined and their corresponding activation free energies are compared. For all derivatives considered, ring opening exhibits a lower activation free energy than Hofmann elimination, identifying it as the kinetically preferred degradation pathway. The magnitude of the ring-opening barrier, however, varies significantly between molecules, showing that different functionalizations strongly influence its stability toward degradation. The predicted preference for ring opening is consistent with available experimental studies, which have identified or inferred ring-opening degradation for several TEMPO-based catholytes. These results provide an atomistic picture of degradation pathways that are difficult to resolve experimentally and highlight the importance of molecular structure in controlling the kinetic stability of TEMPO derivatives in aqueous electrolytes.
Understanding a chemical reaction requires mapping a symbolic reactant-product description onto the three-dimensional pathway through which atoms rearrange, yet chemistry benchmarks for large language models (LLMs) largely probe factual knowledge and text-based reasoning. Here we introduce TSBench, a benchmark in which an LLM agent uses structure-editing tools to construct three-dimensional transition-state (TS) guesses verified by an automated quantum-chemical pipeline, yielding a physics-grounded pass/fail verdict. Across 546 evaluations of seven frontier LLMs on 78 elementary reactions, the aggregate success rate rose from 50.4% to 66.8% under diagnosis-driven revision; the best models approached 90% on the simplest reactions, yet performance dropped sharply with mechanistic complexity. Most failed attempts produced locally plausible saddle points whose reaction paths led to the wrong reactant-product pair, revealing local geometric intuition without a robust grasp of the global reaction coordinate. TSBench establishes a mechanism-level yardstick for LLM agents in mechanism-sensitive tasks such as synthesis planning and autonomous experimentation.
In molecular discovery, molecule size is coupled to composition, structure, and other target properties. Yet most 3D generators require molecule size to be specified before generation. Here, we introduce Equivariant-Free Transformer-Autoencoded Latent Flow Matching, a two-stage generative framework that relies entirely on a single fixed-dimensional molecule-level latent representation to generate variable-size molecules. The second-stage flow matching model samples this latent vector, and an autoregressive Transformer decoder then determines molecule size while generating atom types, coordinates, and chemically informative states. Canonical atom ordering and rigid-pose alignment enable standard Transformers without equivariant layers, while joint decoding of molecular geometry and an enriched chemical state enables reliable, deterministic, chemistry-guided graph recovery without requiring a learned dense pairwise bond decoder. The same fixed-dimensional latent supports unconditional and property-conditioned flow matching, while optional property supervision adds an internal ranking readout, with no separate predictor or reference calculations. On PCQM4Mv2, EF-TALFM achieves the highest fraction of molecules that are unique, training-set novel, pass sanitization and PoseBusters sanity checks, 89.4\%, compared with 75.6\% for UAE-3D and 69.8\% for FlowMol. EF-TALFM also achieves higher measured computational throughput for training and sampling. Across ten target HOMO--LUMO gaps, internal ranking doubles the density functional theory (DFT)-verified hit rate within $0.1\,\mathrm{eV}$, while preserving 97\% novelty among unique verified hits. These results demonstrate that fixed-dimensional molecule-level generation followed by symmetry-resolved autoregressive realization provides a practical architecture for open-ended and property-directed 3D molecular design.
Projective transcorrelation recovers the short-range electron correlation by a similarity transformation with a geminal function $f(r_{12})$, at the cost of an effective Hamiltonian containing a 3-body term. We replace the term by an effective operator of rank at most two, obtained from the two-body cumulant (2C) approximation to the three-particle reduced density matrix. The 2C 3-body energy is written without cumulants, and the effective one- and two-body interactions are derived as its partial derivatives with respect to the reduced density matrices. The truncation is assessed on the atomization and reaction energies of the HEAT set with CCSD(T), and on the CAS-pTC model, whose downfolded Hamiltonian is held on a qubit register with the number of Pauli strings reduced from the sixth power of the orbital count to the fourth.
The Hemoglycin polymer of glycine derived from space is terminated by iron and silicon to form a three-dimensional regular lattice with a high density of silicon tetrahedral vertices. Divalent cation bridges bypass vertices to facilitate electron conduction axially from layer to layer of the lattice. In principle, the exact distribution of Mg or Ca ions around each vertex can encode a very large amount of information. In space where the polymer forms in a molecular cloud the current derived from photon input will fade but in-fall hemoglycin in life forms on habitable planets like Earth, could be contained behind lipid electronic barriers of the nervous system in neurons or glia. Energy minimization of the permutations for magnesium or calcium occupancy showed deeper binding for calcium-dominated configurations. A low energy state with two independent magnesium to calcium substitutions appears to be key to the maintenance of long-term chemical modification. This state could be read via a pulse train that induced the matching spatial distribution of calcium to render adjacent insulating lattice regions conductive and provide a completed conduction path, signaling recognition. Writing of information may involve spatial modulation of the Na/Ca exchanger within the lipid membrane.
Accurate prediction of high-performance liquid chromatography (HPLC) retention times (RTs) across diverse molecules and chromatographic methods remains challenging because experimental training data cover only a limited region of chemical and method spaces. Here, we develop FUSE-RT (Foundation model Unifying Simulation and Experimental supervision for Retention Time), a multitask foundation model that integrates RT data from 179 chromatographic methods and adapts to unseen molecules and methods using limited target-domain data. To extend transferability beyond experimental coverage, we introduce simulation-to-real (Sim2Real) transfer learning, in which molecular representations learned from large-scale computational data are transferred to experimental RT prediction. Specifically, we use PolyOmics, comprising 39 properties for approximately 21,400 molecules generated by molecular dynamics and density-functional theory calculations, as auxiliary supervision. We evaluate generalization under molecular, method, and joint molecular--method distribution shifts. Simulation-derived supervision substantially improves transfer beyond the experimental molecular domain, particularly under pronounced coverage gaps and few-shot adaptation. Moreover, RT-prediction error decreases systematically with increasing simulation-data size, following a significant power-law relationship. These results establish Sim2Real transfer as a scalable strategy for extending RT prediction beyond the finite coverage of experimental chromatographic data.
In this study, we implement a low-rank approximation of Moment Tensor Potential (MTP) based on the tensor train (TT) decomposition. The implemented tensor-factorized MTP (TFMTP) and the original MTP model are actively trained via a MaxVol-based algorithm during molecular dynamics simulations of a four-component molten salt mixture, LiF-NaF-KF (FLiNaK), and geometry optimizations of a five-component equiatomic MoNbTaWV random alloy. We demonstrate that under a 1.5-fold compression, TFMTP requires two times fewer configurations for fitting than the original MTP model, while maintaining an indistinguishable level of accuracy. These actively trained MTP and TFMTP models are further used to evaluate the density and viscosity of FLiNaK at temperatures ranging from 600 to 1200 K, as well as the elastic constants and bulk modulus of the MoNbTaWV alloy at zero temperature. For both atomic systems, the differences in physical properties predicted by MTP and TFMTP are negligible.
This study investigates the formation of the hydroxyl anion (OH$^-$) via dissociative electron attachment in 2-propanol as a model system for studying both organic and inorganic molecules. Using high-level CAP-EOM-EA-CCSD calculations and advanced ToF mass spectrometry, we demonstrate that OH$^-$ formation at electron energies between 7 and 11 eV is dominated by two-particle-one-hole (2p-1h) Feshbach resonances. The potential-energy curves reveal a dense manifold of anionic states coupled through numerous avoided crossings, facilitating nonadiabatic population transfer during C-OH bond dissociation. Survival-probability analysis identifies a subset of six long-lived resonances that persist long enough to drive fragmentation, with states 25 and 28 acting as primary drivers by funneling the attached electron into the localized $σ^*(\mathrm{C{-}OH})$ antibonding orbital. These theoretical predictions are confirmed by experimental observations of a prominent OH$^-$ yield peaking at 8.6 eV, supporting a site-specific fragmentation mechanism that generalizes to the broader class of molecules.
We consider a mixed quantum-classical formulation for open nonequilibrium molecular systems. We employ the pseudoparticle nonequilibrium Green's function (PP-NEGF) method as a fully quantum description and a starting point for introducing classical nuclear dynamics for an open system. We use this formulation to derive a density-matrix equation of motion describing the quantum-classical evolution. We compare our results with previous formulations and highlight important differences related to the nonequilibrium and open character of the molecular system. We also analyze the approximations needed to reduce full quantum-classical dynamics to the fewest-switches surface hopping (FSSH) method and show that these approximations are unreasonable when adiabatic surfaces come close to each other.
Degenerate electronic states govern the photophysics of excimers and the molecular dynamics near conical intersections. In regions where potential energy surfaces are degenerate or nearly degenerate, intermolecular interactions mix the corresponding electronic configurations. We introduce a first-order degenerate formulation of symmetry-adapted perturbation theory (dSAPT) with weak symmetry forcing that quantifies the mixing in terms of electrostatic and exchange (Pauli repulsion) contributions. We compare two alternative ways of resolving the degeneracy, in which antisymmetry is enforced either after or before diagonalization of the first-order eigenproblem. Using CASSCF monomer wave functions, we apply dSAPT to two model spatially degenerate systems, H$_2$--F(${}^2P$) and H$_2$--NO(${}^2Π$), and show that exchange effects contribute significantly to configuration mixing already at intermediate intermolecular separations. For the H$_2$--NO(${}^2Π$) complex, we further demonstrate that second-order dispersion energy is essential for qualitatively reproducing the angular dependence of the diabatic mixing angle. Finally, using water and benzene excimers with monomers described either by CASSCF or CIS wave functions, we show that Pauli repulsion controls delocalization of the excitation at short intermolecular separations.
Machine-learning-based inverse design can accelerate the discovery of materials with targeted properties, but conventional structure--property models often require large training datasets and generalize poorly beyond their training distribution. Here, we present a differentiable inverse design framework based on a graph neural network molecular dynamics simulator. By combining a short dynamical initialization with physics-based refinement during simulation, the framework enables optimization well beyond the conditions represented in the training data. Using disordered elastic networks, we show that a simulator trained only on non-auxetic systems with Poisson's ratios between 0.1 and 0.4 can design strongly auxetic networks with values as low as -0.3. The framework also produces stress--strain responses outside the range observed during training, creates localized mechanical defects absent from the training data, and generalizes across system size, enabling optimization of networks tested up to 5000 nodes despite being trained only on networks with fewer than 200 nodes.
We present an approach to analyze long-timescale kinetics of molecular dynamics simulations - meta-stable states and transition timescales - by a tensor-based approximation of the infinitesimal generator. We start from the variational approximation of low-lying generator eigenvalues, and discretize the associated eigenvalue problem using a tensor-product basis. We derive analytical expressions for the resulting tensor operators in tensor train (TT) format, and provide efficient algorithms to solve the corresponding linear problems. We also treat the case of a state-dependent diffusion field, which is relevant when working in reaction coordinates. We show how the number of reaction coordinates can be gradually increased without having to solve the variational problem from scratch. Using MD simulations of fast-folding proteins and TICA coordinates as reaction coordinates, we demonstrate the effectiveness of the proposed method at identifying long-timescale transitions and meta-stable states.
Delocalisation error has been argued to be the greatest outstanding challenge in density-functional theory (DFT). One of the most promising routes to minimise this error is development of local hybrid functionals, in which the fraction of exact-exchange mixing is position dependent. However, existing local hybrids capable of good thermochemical accuracy are often highly empirical, and tend to have complicated functional forms that involve some combination of range separation, calibration functions, power-series expansions, or even neural networks. In this work, we explore the limits of what can be achieved with a minimally empirical local hybrid functional form that avoids such complexities. The ``LHnz'' functional is proposed, which uses dispersionless exchange and a local mixing fraction dependent on the correlation length and effective exchange--correlation hole normalisations. With only three empirical parameters, LHnz is shown to outperform all existing global hybrid functionals for the GMTKN55 molecular-thermochemistry benchmark with no large outliers.
The Smooth Overlap of Atomic Positions (SOAP) descriptor is one of the most widely adopted representations of atomic environments in molecular machine learning, but the cost of evaluating it and its derivatives remains a principal bottleneck of SOAP-based interatomic potentials, particularly at the large radial and angular basis sizes demanded by complex, multi-species condensed-phase environments. We present cuSOAP, a GPU-accelerated generator of atom-wise SOAP vectors and their analytic derivatives. Built on PyTorch, cuSOAP evaluates closed-form expressions or quadratures for the projection coefficients and their Cartesian gradients for Gaussian-type-orbital and polynomial radial bases, through fused CUDA and Triton kernels that eliminate the multi-gigabyte intermediates of a naive tensor formulation. The package is a drop-in replacement for the CPU-based reference DScribe, reproducing its constructor signature, feature ordering, and output to within ${\sim}10^{-6}$, and accepts structures directly as Atomistic Simulation Environment (ASE) Atoms objects. On an NVIDIA Grace--Blackwell (GB200) node, a single Blackwell GPU generates the full descriptor-plus-Jacobian workload for a 1000-molecule water cluster up to two orders of magnitude faster than DScribe, i.e., from $13\times$ to $227\times$ across the entire $(n_{\max}, l_{\max})$ hyperparameter map, with the speedup growing with the angular band limit $l_{\max}$. Descriptor-only generation scales as $t \propto n^{1.34}$, close to linear, up to a million-atom water cluster, which is processed in 19.5s on one GPU and, through a shared-memory multiprocessing driver, in 5.7s on the four GPUs of the node, demonstrating $\sim$90% parallel efficiency with no sign of saturation as devices are added. These results bring on-the-fly SOAP evaluation for large-scale condensed-phase simulation within reach.
Tailoring Li+ solvation structures has emerged as a promising design principle across multiple Li metal battery (LMB) electrolytes, with localized high concentration electrolytes (LHCE) emerging as a leading candidate. However, rational design of these electrolytes remains limited by a poor quantitative understanding of how molecular properties govern Li+ solvation shell composition. Here we introduce a mean field modeling framework, built on an Ising model, that predicts solvation shell composition from donor number (DN), acceptor number (AN), molar ratio, and molecular size. This framework is end-to-end differentiable, which enables parameterization directly from molecular dynamics (MD) solvation structures via gradient-based optimization. Focusing on LHCE, the model achieves 10.7% RMSE and R2=0.87 for shell composition, 2.3% RMSE on free solvent ratio, and further reproduces solvation trends in real LHCE electrolytes, at a fraction of MD's cost. Our model also shows good generalizability on other electrolyte systems in addition to LHCE. Analysis of its interaction terms shows that DN dominates Li+ solvation energetics, providing a thermodynamic basis for empirical DN-based design rules. We further demonstrate the model's design utility on the LiTFSI/tetraglyme (G4) system absent from training. The model identifies fluorobenzene (FB) as a promising diluent and predicts a salt-concentration window with anion-rich solvation shells and low free solvent, favorable for stable anion-derived SEIs and high oxidative stability, respectively. This prediction is validated by higher-fidelity MD. This work establishes an interpretable differentiable framework for predicting Li+ solvation shell composition, with potential to extend to other electrolyte classes.
We prove that, for any given sufficiently regular one-particle density, there exists a unique external potential for which the solution to the many-body Schrödinger equation has this prescribed density. This implies in particular that one can exactly reproduce the time-dependent density of an interacting system using non-interacting electrons, which is at the heart of time-dependent density functional theory. Our result covers the Coulomb interaction and densities arising from extended nuclei.
We provide the first fully rigorous justification of time-dependent density functional theory in the continuum: Given a time-dependent density, we prove that there exists an external potential, unique up to a time-dependent constant, such that the corresponding Schrödinger equation reproduces the given density. This potential can be obtained with an iteration scheme. Our argument covers Coulomb interactions and does not need any Taylor expansion in time. It instead requires analyticity in space for the density and potential, which is compatible with extended nuclei.
Molecular diffusion in fluctuating amorphous and macromolecular media governs key transport processes across soft-matter physics, energy storage, and biological membranes. Extracting localized trapping states from single-particle tracking trajectories remains a fundamental challenge; because thermal structural breathing continuously reconfigures pore boundaries, conventional geometric algorithms suffer from severe systematic biases, erroneously merging distinct localized states during cyclic molecular returns. Here, we address this deadlock by shifting the paradigm from local geometric recurrence to a structure-informed Bayesian regularization. Leveraging discrete Morse theory, we extract the time-invariant topological skeleton of the fluctuating host matrix to construct robust, gas-specific physical priors that account for individual molecular dimensions. Trajectory steps are sequentially partitioned via a two-stage probabilistic refinement that dynamically adapts to the transport landscape. Benchmarked against a rigorous environment where synthetic particles explore the actual interconnected matrix graph, our approach eliminates systemic biases, restricting macroscopic trapping parameter deviations to just a few percent under optimal linear $O(N)$ computational scaling. Applied to hydrogen and methane transport within a type-I kerogen matrix, serving as a prototype for highly tortuous, flexible macromolecular networks, the method successfully decodes the hidden microscopic mechanisms of confined diffusion. To ensure immediate broad impact, the documented open-source code and data are made publicly available, offering an accessible strategy readily adaptable to a broad spectrum of tracking phenomena, from ion transport in battery polymers to protein trafficking within cellular environments.
In previous work, it was shown that using the cluster mean field(cMF) wavefunction as the zeroth order wavefunction for perturbation theory and coupled cluster produces a reasonable approximation for the J 1-J 2 Heisenberg model in the thermodynamic limit(TDL). However, this work was limited to cPT4 and cCCSD. The central claim of cMF is that it takes strongly correlated systems and makes them weakly correlated so that PT weakly correlated methods can be used. To confirm that this claim is true, it is important to show that the PT and CC series converge. In this paper, we extend the previous calculations to cPT7 and cCCSDTQ5 where we show that the energy is converging. We also discuss how to perform cluster based calculations efficiently in the thermodynamic limit, which is crucial for these more expensive calculations.
Predicting how molecular weight distribution and monomer sequence evolve during polymerization is central to polymer science, yet classical approaches face a trade-off between molecular resolution and computational cost: for copolymers, the number of distinguishable species grows exponentially with chain length. Quantum computing offers a potential alternative, provided the non-unitary rate matrices governing the kinetics can be embedded into unitary quantum circuits, a task known as block encoding. Here we construct explicit block-encoding circuits for two kinetic models of living polymerization: Model A, single-monomer polymerization, whose lower-bidiagonal rate matrix is encoded via a sparse-oracle construction and a two-term linear combination of unitaries (LCU) decomposition; and Model B, two-monomer copolymerization, where a bijective labeling of polymer species by an integer index (the m-index) yields a structured sparse matrix encoded via either a five-term LCU or a sparse-oracle construction. Numerical simulations with the sparse-oracle encodings reproduce the classical time evolution for reactivity ratios drawn from reported olefin copolymerization systems spanning near-random ($r_1 r_2 \simeq 1$) and blocky ($r_1 r_2 > 1$) microstructures, and the LCU encodings are verified by explicit reconstruction of the encoded matrix block. Resource estimation shows that both implementations require only $O(\log N)$ qubits in the matrix dimension $N$ (an exponential memory saving over the classical state space), with gate counts growing gradually, reaching $10^4$ to $10^5$ gates at $10^3$ system qubits. These results establish a concrete quantum circuit foundation for simulating polymerization kinetics on fault-tolerant quantum hardware, and a first step toward exploiting exponential state-space compression for high-dimensional polymer reaction networks.
In their landmark paper introducing the atomic force microscope [Phys. Rev. Lett. \textbf{56}, 930 (1986)], Binnig, Quate, and Gerber presciently anticipated that the technique would ultimately be capable of probing interactions running the gamut from weak van der Waals interactions to strong covalent bonding. They also highlighted that the tip-sample forces central to AFM are present, and often highly influential, in scanning tunnelling microscopy; indeed, this realisation directly inspired the invention of the force microscope. In this perspective for the \textit{Forty Years of AFM} special issue, we review selected aspects of two decades of work from our group at the University of Nottingham that span the force range highlighted by BQG and are united by a common, central theme: the probe as active participant rather than passive observer. Our selection of results also covers length- and correlation-scales from the microscopic right down to the single chemical bond limit, tracking a spectrum of interactions from van der Waals/Hamaker forces, through hydrogen bonding, to covalent bonds and, finally, atom-by-atom assembly of metal clusters via vertical tip-sample transfer. Echoing BQG's own observations on the prevalence of probe-sample forces in STM, we also discuss recent evidence that tip-induced heterogeneity underpins first-passage dynamics in molecular diffusion and highlight the challenges in acquiring non-invasive measurements of diffusion barriers for adsorbed molecules that are readily perturbed by the probe. We close with a perspective on machine learning's growing role in automating tip-driven atomic and molecular manipulation.
Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. Benchmarks enable reproducible evaluation but may not expose failures outside their predefined scope. Here, we show that physics-informed search can complement benchmark-based evaluation by uncovering hidden failure modes. We introduce MLIP Detective, an agentic framework for active failure mode discovery. Starting from benchmark evidence, MLIP Detective generates falsifiable, physics-informed failure hypotheses, screens them with inexpensive simulations, and escalates only the most suspicious cases to human experts together with proposed verification protocols. Without issue-specific prompting, MLIP Detective identified and characterized a systematic anomaly in MACE-MPA-0: the model predicted some relaxed adsorbate-surface systems involving O- or F-containing adsorbates to be higher in energy than their corresponding separated fragments. Using cross-model comparisons, MLIP Detective further inferred a likely training-data origin for the anomaly, consistent with recent reports.