2026-09-07 | | Total: 4
Dynamical symbolic regression methods identify governing differential equations from noisy data, balancing interpretability and predictive accuracy. However, standard methods often produce expressions that violate known physical laws. To address this, we propose FluxDisco, a physics-informed framework tailored for flux-based, stoichiometric ODE systems. By leveraging a known stoichiometry, we reduce the expression search space and ensure physical adherence. Our framework adapts the Monte Carlo Graph Search algorithm for the unique challenges associated with joint flux discovery of stoichiometric systems. We evaluate our method across a range of physical and biological systems, demonstrating its ability to accurately recover governing dynamics through interpretable equations.
The intrinsic dimension of a dataset is the number of independent directions needed to describe the space occupied by its data. Estimators based on nearest neighbors infer this number from how the probability to find a neighbor point grows around each sampled point. Because the distances $r$ between neighbor points decrease as the sample grows, the estimated dimension can depend strongly on the number of available points. Here, we derive the large-sample behavior of the TWO-NN estimator for data drawn from a smooth $d$-dimensional space. The typical nearest-neighbor distance scales as $N^{-1/d}$, and smooth deviations from a locally uniform distribution produce successive corrections proportional to $N^{-2/d}$. We test this result using the trajectories coming from ten independent $100~μ$s simulations of alanine dipeptide. Configurations are represented by all pairwise distances among the ten heavy atoms. This representation has a known geometric dimension of $3n_{\mathrm{at}}-6=24$. Over the investigated range, the TWO-NN estimate shows no systematic dependence on the temporal spacing between configurations, but increases from approximately $7.5$ to $15.6$ as the sample size grows from $10^2$ to $2\times10^5$. Extrapolations that retain corrections through $r^2$, $r^4$, and $r^6$ give limiting dimensions of $25.23$, $22.89$, and $27.00$, respectively. All three estimates lie close to the known dimension and collectively bracket it, supporting the proposed scaling. Their spread provides a direct estimate of the systematic uncertainty associated with the truncation. The derived scaling therefore explains the strong sample-size dependence of TWO-NN and provides a practical route from finite sample estimates to the underlying geometric dimension.
Cities must allocate limited resources to maintain mobility, with uncertainties about the resulting state of the system. Analyzing roughly 3,000 bus routes with more than 4 billion yearly riders across 19 metropolitan areas worldwide, we uncover a robust scaling law of the form $f \sim (d/t)^α$ with exponent $α\in [1/2,\,2/3]$, linking the service frequency $f$ to passenger demand $d$ and route duration $t$. We show that this scaling emerges from a simple optimization principle: cities implicitly minimize total passenger waiting time under a fixed operational budget when both schedule frequency and crowding are taken into account. This mechanism produces two universal regimes: a frequency-dominated regime with $α= 1/2$ when crowding is negligible, and a capacity-dominated regime with $α= 2/3$ when most routes are overloaded. Intermediate exponents arise when only part of the network operates near capacity. Furthermore, we find that the benefits of additional investment are highly uneven across systems. For instance, our model suggests that a $20\%$ budget increase yields nearly a 5-minute reduction in daily waiting time per passenger in Boston, compared to only about 1 minute in Paris. These findings place urban transit within a broader class of constrained capacity-allocation problems, while highlighting a distinct regime in which prescribed route demands shape the allocation of limited service resources. The resulting scaling laws show how simple optimization principles can generate systematic exponents in complex transport systems, beyond the dissipation-based frameworks usually considered in physical and biological flow networks.
The long-time dynamics of complex molecular systems often involves rare transitions across networks of metastable states. Building on transition-path theory, which provides a rigorous framework for describing rare transitions between two metastable states, we introduce the variational multistate committor network (VMCN), a neural framework that learns the probabilities of reaching each metastable state directly from molecular simulation data. From this representation, VMCN identifies state-specific commitment, candidate transition regions and committor-consistent pathways between state pairs, and an effective kinetic network characterized by transition rates. The model is trained using finite time-lag trajectory data together with boundary conditions defined on conservative state cores. Applications to a triple-well potential, trialanine isomerization, and the $c$--ring rotation in the V$_{\rm o}$ domain of a vacuolar ATPase show that VMCN recovers metastable organization, provides committor-consistent descriptions of transition mechanisms, and estimates state-to-state kinetics. VMCN further provides diagnostics for incomplete state decompositions and enables adaptive exploration of candidate metastable states and their connecting regions. By integrating VMCN with generative committor-guided path sampling (Gen-COMPAS) for chignolin, we start from two end point structures, identify a misfolded state and a candidate folding intermediate, and we direct subsequent sampling toward the resulting multistate transition network.