Quantitative Biology

2026-09-02 | | Total: 15

#1 PopPert: Population-level Joint-Distribution Modeling for Single-Cell Perturbation Prediction [PDF] [Copy] [Kimi] [REL]

Authors: Handong Wang, Jiaxin Qi, Haochen Feng, Baisheng Lai

Predicting transcriptional responses to specific perturbations is critical for understanding cellular regulatory mechanisms and accelerating drug discovery. Single-cell RNA sequencing destroys each measured cell, yielding only unpaired populations of control and perturbed cells. However, existing methods typically model perturbation prediction at the single-cell level and assume cell-to-cell correspondence, which conflicts with the unpaired nature of the observed data. To address this challenge, we propose PopPert, a framework that explicitly parameterizes population-level joint gene expression distributions for collective transcriptional state modeling. Given a control population distribution and a perturbation condition, PopPert predicts perturbation-induced changes in distribution parameters, eliminating the need for cell-level correspondence and reducing sensitivity to single-cell noise. To effectively capture gene co-expression patterns, PopPert leverages a low-rank Gaussian Copula to model cross-gene statistical dependencies and construct the joint gene expression distribution, additionally allowing sampling of synthetic perturbed single-cell profiles. Across multiple single-cell benchmarks spanning both genetic and chemical perturbations, PopPert achieves superior overall performance in differential expression recovery, perturbation effect estimation, and population-level distribution matching. These results establish population-level joint distribution learning as an effective paradigm for predicting transcriptional responses from unpaired single-cell populations. Code for PopPert is publicly available at https://github.com/whd1125/PopPert.

Subjects: Genomics , Artificial Intelligence

Publish: 2026-09-01 14:59:03 UTC


#2 On the interpretation of the kinetics of ligand-receptor binding [PDF] [Copy] [Kimi] [REL]

Authors: David Colquhoun, James P Higham

When the rates of ligand binding are measured by methods such as surface plasmon resonance, it is common practice to use the observed rate constants for the onset and offset of binding to estimate an equilibrium constant for ligand binding. If this agrees with the equilibrium constant found as the EC50 for binding at equilibrium, this is taken as validation of the measured rates. This is correct only when binding produces no conformation change in the receptor, and ligand binding follows a single exponential time course. Here, we investigate a simple 3 state model in which binding is followed by a conformation change in the receptor. Three special cases of this model in which the time course of onset and offset of ligand binding are close to being single exponentials are analysed. These cases are (1) when binding is much faster than the conformation change, (2) when the conformation change is much faster than binding, and (3) when the rates of ligand dissociation and receptor activation are both fast. It is concluded that the measured rates will often yield an estimate of the equilibrium constant for ligand binding that is close to the effective, or macroscopic, equilibrium constant, the EC50 found by measuring binding at equilibrium, which depends on both of the underlying microscopic equilibrium constants describing ligand binding and the conformation change. The exception to this conclusion is the case when binding is much faster than the subsequent conformation change, though the estimate of the equilibrium constant for ligand binding still depends on both of the underlying microscopic equilibrium constants.

Subject: Quantitative Methods

Publish: 2026-09-01 13:29:15 UTC


#3 From static structures to dynamic landscapes: cryo-EM redefines RNA biology [PDF] [Copy] [Kimi] [REL]

Authors: Shekhar Jadhav, Spandan Saha, Qingbin Shang, Marco Marcia

RNA molecules perform diverse biological functions by dynamically exploring multiple conformational states rather than adopting a single static structure. Capturing these ensembles is a challenge in molecular biology. Recent advances in cryoEM are now transforming this landscape by enabling the visualization of RNA molecules across a broad spectrum of functionally relevant conformations at near atomic resolution. Here, we examine how cryoEM is reshaping RNA structural biology changing focus from the analysis of static structures to dynamic conformational landscapes. Through 8 representative case studies we illustrate how cryoEM has revealed previously inaccessible mechanisms of RNA motion, including folding processes, ligand-dependent switching, and cooperative assembly. We specifically discuss emerging experimental and computational approaches that address and overcome the challenges associated with studying dynamic RNAs, particularly in construct design, sample preparation, vitrification, and data analysis. These novel methods resolve conformational variability and enable the reconstruction of discrete and continuous RNA conformational landscapes from cryoEM data, highlighting how structural heterogeneity can be harnessed to extract functional insights. Looking forward, the integration of cryoEM with complementary biophysical techniques and time resolved methodologies promises to bridge structural and temporal resolution, to routinely derive experimental molecular movies of RNA in action. These advances will not only deepen our understanding of RNA biology but also provide new opportunities for RNA targeted therapeutics and the rational design of dynamic RNA-based nanodevices. By connecting structural snapshots into coherent dynamic models, cryoEM is establishing a framework for quantitative descriptions of RNA energy landscapes and their functional roles.

Subject: Biomolecules

Publish: 2026-09-01 12:08:09 UTC


#4 Active Visual Semantics: A large-scale MEG and eye-tracking dataset for understanding visual intelligence in action [PDF] [Copy] [Kimi] [REL]

Authors: Philip Sulewski, Carmen Amme, Peter König, Martin N. Hebart, Tim C. Kietzmann

Here we present the Active Visual Semantics (AVS) dataset, a large-scale collection of magnetoencephalography (MEG) and eye-tracking data recorded while five participants freely explored 4,080 natural scenes over 10 sessions each, yielding more than 200,000 fixation epochs in total. Unlike existing neuroimaging datasets that rely on passive viewing with enforced central fixation, AVS captures brain activity during active scene exploration, including self-generated saccades and fixations. A semantic captioning task on 25% of the trials provides behavioural measures linking gaze to scene understanding and memory. In addition to neural and behavioural data, AVS includes per-fixation object category labels, human-rated annotations of the appearance of fixation targets in the scene captioning task and pupil dynamics. Individual head stabilisation casts were used during MEG data collection, which alongside with structural MRI scans, enabled precise cross-session source reconstruction. Using artificial neural network (ANN) encoding models we demonstrate that individual fixation-aligned MEG epochs hold visual content-specific signal, despite the challenges that active scene viewing poses for MEG signal quality. Further, we use fixation-aligned representation similarity analysis (RSA) and demonstrate that we can derive fixation object category averages that yield representational geometries which are highly reliable across participants. Both in MEG sensor and source space this structure is validated by its robust alignment with ANN object-level representational geometry. Taken together, AVS provides a rich resource for investigating a large variety of questions regarding the neural mechanisms of active vision, object recognition and scene captioning during natural viewing, and the relationship between gaze behaviour and memory encoding.

Subject: Neurons and Cognition

Publish: 2026-09-01 10:51:05 UTC


#5 Temporally constraining source imaging estimates in an underdetermined neural system with eigenmodes of cortical geometry [PDF] [Copy] [Kimi] [REL]

Authors: Pok Him Siu, Philippa J. Karoly, Artemio Soto-Breceda, Mark J. Cook, David B. Grayden

Geometric eigenmodes provide a compact and biologically grounded representation of large-scale neural activity. Previous work demonstrated that they can mitigate the underdetermined nature of electroencephalographic (EEG) and magnetoencephalographic (MEG) source localisation, an ill-posed inverse problem in which neural activity is reconstructed from non-invasive recordings. Beyond their spatial structure, neural field theory predicts the temporal evolution of eigenmodes through analytically derived transfer functions. Motivated by this framework, the present work investigates whether these transfer functions can be used to introduce temporal constraints into EEG source imaging. The approach is evaluated using simulated seizure dynamics generated by coupled Epileptor neural mass models. Transfer functions derived directly from neural field theory were found to be generally ineffective as temporal constraints for source localisation, primarily because they neglect cross-eigenmode coupling. Incorporating empirically estimated coupling terms substantially improves localisation performance, particularly in noisy conditions. Although estimating these eigenmode coupling interactions from experimental data remains challenging, the findings motivate dynamical source-imaging approaches that combine spatial eigenmode structure with empirically informed cross-modal dynamics.

Subjects: Neurons and Cognition , Biological Physics

Publish: 2026-09-01 07:04:50 UTC


#6 Efficiently classifying shocks in complex systems requires dormant reporters [PDF] [Copy] [Kimi] [REL]

Authors: David A. Brewster, Philippe Cluzel

Many natural and engineered systems are large complex networks of interacting components, and external perturbations drive them along different dynamical paths. Identifying which perturbation occurred matters for diagnosis, control, and prediction. Yet often times only a few components can be jointly monitored. Which components should be monitored? Experimental practice usually favors placing reporters at the most sensitive sites, where perturbations produce the largest effects. Using a simple dynamical model for complex systems with heterogeneous connectivity, we ask how sparse reporter panels should be chosen to classify shocks from partial trajectories when repeated trials only approximately reproduce an ideal initial condition. Once that reproduction is imperfect, sensitivity ranked panels fall far short of optimal, and the shortfall grows with the noise. We find that the best panels mix two kinds of reporters. A promiscuous reporter responds to most shocks, so it separates them mainly by degree, and degree fluctuates from trial to trial. A dormant reporter responds to only a few shocks but strongly, and its answers do not scatter as much between trials. As noise grows, the cost of losing a dormant reporter rises to meet the cost of losing a promiscuous one. Panels of either kind alone classify worse than the mixture, and no property of the members collected individually explains the ordering. Most of all, we find that only a minuscule number of reporters are needed on a panel to accurately identify which shock hit the system. We implement an efficient algorithm to assemble such a panel. Together these results provide a low cost practical design principle for monitoring large complex dynamical systems.

Subjects: Quantitative Methods , Disordered Systems and Neural Networks , Adaptation and Self-Organizing Systems , Physics and Society , Molecular Networks

Publish: 2026-09-01 05:00:05 UTC


#7 Operationalizing open-ended biological discovery across single-cell representations [PDF] [Copy] [Kimi] [REL]

Authors: Ningxuan Zhang, Ziwei Wang, Ning Xie, Na Liu

Single-cell studies are typically initiated from predefined research questions, leaving much of the biological information encoded within existing data unexplored. We formalize open-ended discovery as an analytical paradigm, in which data-derived signals are identified before biological context is interrogated and subsequently evaluated according to their potential to justify prospective experimental investment. Here we develop PROSPECTor, an end-to-end framework that searches for reproducible biological structures across conventional expression representations and diverse foundation-model embeddings, translating robust signals into quantitatively testable candidate hypotheses. Projection into unseen datasets then evaluates their generalizability and phenotype association, providing a scalable screen for candidates that warrant prospective validation. Supported signals emerged from different representation spaces and search strategies. PROSPECTor-nominated hypotheses were then examined in independent biological settings: fibroblast extracellular-matrix programmes demonstrated transferability to an independent mouse cohort with an intervention context, while a patient-resolved gastric-cancer T-cell programme recurred across single-cell, bulk and spatial cohorts. PROSPECTor establishes an auditable framework for systematically revisiting single-cell datasets across expanding representation spaces, turning retrospective collections into prospective resources for biological discovery that can motivate new research questions.

Subject: Quantitative Methods

Publish: 2026-09-01 03:58:21 UTC


#8 BME-like Quartet Weights for Phylogenetic Trees [PDF] [Copy] [Kimi] [REL]

Author: Peter J. Waddell

Like pairwise distances, quartets can be highly redundant and correlated on a phylogenetic tree, and their number grows on the order of n^4 rather than n^2. I explore BME-like weights for reweighting quartet scores before summing them to score a full tree. Three weights are considered on an unrooted binary tree: w_ext(q)=2^(-I_ext(q)), w_int(q)=2^(-I_int(q)), and w_tot(q)=2^(-I_tot(q))=w_ext(q)w_int(q), where the exponents count specified internal nodes in the minimal connecting subtree of a quartet. Exact tree-shape counts, total quartet-weight sums, and internal-edge crossing sums are calculated for all unlabeled unrooted binary tree shapes on 6-10 taxa. For w_ext, the total quartet weight is tree-shape-invariant and the edge-crossing sum depends only on split size. For any n-leaf tree, we prove sum_q w_ext(q)=(n-2)(n-3)/8, and the sum over quartets crossing an internal edge with split a|b equals (a-1)(b-1)/4. Exact tree-shape-specific normalizers are also derived for w_int and w_tot. A degree-corrected hard-polytomy extension is given for multifurcating trees, and a conditional consistency result shows that these positive weights preserve consistency when the underlying quartet estimates are themselves consistent for the true induced quartet states. These results provide a mathematical foundation for evaluating and applying BME-like quartet weights to reduce redundancy with the particular aim of improving statistical efficiency with finite data.

Subject: Populations and Evolution

Publish: 2026-09-01 02:58:19 UTC


#9 A distributed-delay Wilson-Cowan model of sleep-related rhythms in the corticothalamic system [PDF] [Copy] [Kimi] [REL]

Authors: Eva Kaslik, Anca Radulescu, Anca Stanoev

The corticothalamic circuit supports rhythms with timescales that differ by orders of magnitude: sleep spindles, the sigma-band events of non-rapid-eye-movement (NREM) sleep, and infra-slow fluctuations near 0.02Hz that organize when spindles occur. Because the anatomy is the same in both cases, architecture alone cannot determine which rhythm the circuit expresses. We ask whether the temporal structure of the circuit's own feedback can. In a four-population Wilson--Cowan model comprising cortical excitatory and inhibitory populations, thalamic relay cells, and the thalamic reticular nucleus (TRN), we first establish how connectivity controls access to oscillatory behavior, and then introduce temporal coupling as either a weak Gamma distributed delay or a discrete delay. We investigate three distinct connectivity levels: recurrent cortical excitation gates whether the circuit can oscillate at all, the reciprocal relay-TRN pair determines where the oscillation lies and how it is configured, sustained, and terminated, and reticular self-inhibition limits its extent. We then examine how these connectivity-dependent regimes are affected by delayed coupling. Although delay does not change the equilibria themselves, it can substantially alter their stability and the organization of the resulting oscillatory dynamics. Under weak Gamma integration, short delays support spindle-compatible oscillations in the sigma band, while longer delays give rise to a much slower regime near 0.02Hz. The discrete-delay formulation produces a qualitatively different and more complex bifurcation structure. Together, these results show that the dynamics of the corticothalamic circuit depend not only on its connectivity, but also on the temporal organization of interactions within the circuit.

Subject: Neurons and Cognition

Publish: 2026-09-01 00:39:55 UTC


#10 One Faithful Pass Over the Cuckoo's Nest [PDF] [Copy] [Kimi] [REL]

Author: Kristina Šekrst

Narrative theories of consciousness hold that conscious experience is (at least partly) constituted by a partially opaque inner narrative that does not perfectly track the underlying computation it narrates. I argue that the safety goal of making chain-of-thought (CoT) reasoning faithful and transparent is structurally incompatible with the conditions under which CoT could, even in principle, count as conscious narration. Recent empirical work suggests that CoT in large language models is largely post hoc, causally bypassed, and unreliable as a window onto internal computation. The opacity that makes CoT unreliable for alignment is exactly what narrative theories of consciousness identify as consciousness-constitutive. Alignment interventions aimed at producing faithful CoT and consciousness-detection frameworks grounded in narrative theory are therefore pulling the same architectural variable in opposite directions. I draw out the methodological consequence: alignment interventions alter the very features that consciousness-detection frameworks would need to measure, a point neither field has addressed directly.

Subject: Neurons and Cognition

Publish: 2026-07-06 13:06:10 UTC


#11 Score-Based Generative Data Assimilation for Integrating Aggregated Surveillance Data into Agent-Based Models in Epidemic Tracking [PDF] [Copy] [Kimi] [REL]

Authors: Siming Liang, Jacob Hauck, Minglei Yang, Adam Spannaus, Heidi Hanson, Guannan Zhang

Reliable epidemic monitoring often requires inferring regional infection burden and transmission heterogeneity from noisy, spatially aggregated, and potentially sparse surveillance data. Agent-based models (ABMs) are attractive for this task because they represent individual behavior, contact heterogeneity, and localized interventions, but these same features make them difficult to calibrate online. We develop a generative AI-based data-assimilation (GenDA) framework for partially observed epidemic ABMs that estimates both the epidemic state and a heterogeneous parameter field while respecting the gap between observable macrostates and latent agent-level microstates. GenDA combines a training-free, score-based generative update for macrostate correction with a direct parameter update based on macrostate discrepancies, followed by a macro-micro reassignment step that restores consistency with the ABM. In controlled and geographically explicit synthetic experiments, the framework recovers regional epidemic burden, dominant hotspot structures, and effective transmission heterogeneity from aggregated observations, while improving post-assimilation forecasts relative to state-only assimilation.

Subjects: Numerical Analysis , Quantitative Methods

Publish: 2026-09-01 15:43:20 UTC


#12 FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation [PDF] [Copy] [Kimi] [REL]

Authors: Kewei Li, Rongying Zhang, Xueli Wang, Xiwen Gong, Zhongjian Wang, Qiuchen Zhao, Lan Huang, Ruochi Zhang, Fengfeng Zhou

Token aggregation converts token-level representations into fixed-dimensional sample representations, but most pooling methods operate only in the original token space. We introduce Frequency-Domain Latent-attention Gated Pooling (FLaG), a plug-in aggregation module that re-expresses encoder outputs in the Fourier domain before final pooling. FLaG represents the nonredundant rFFT spectrum through concatenated real and imaginary components, summarizes spectral tokens with learnable latent queries, derives a sample-conditioned channel gate, and reconstructs modulated token representations for downstream aggregation. We evaluate the same architecture across ESM2-based antimicrobial peptide (AMP) activity prediction, ResNet18 image classification on CIFAR-10 and CIFAR-100, and three RoBERTa-based language tasks. FLaG achieves the best macro-averaged Spearman correlation coefficient, RMSE, and Recall@50 across four AMP backbone-species settings and the highest top-1 accuracy on CIFAR 10. It also achieves the best mean results on five of seven language metrics, although mean pooling remains strongest on STSBenchmark. AMP-side mechanistic analyses reveal low-frequency prediction sensitivity across most encoder layers, with increased relative high-frequency sensitivity in the final layer, and pronounced peptide-specific positional responses. The residual gate broadly amplifies spectral channels while preserving the low-frequency-dominated energy profile, whereas latent cross-attention exhibits sample- and species-specific spectral allocation. Overall, FLaG provides a transferable frequency-domain aggregation bias across protein, visual, and textual representations, with benefits that depend on the backbone and downstream task. Supplementary materials, source code, and data are available at https://www.healthinformaticslab.org/supp/ and https://github.com/Kewei2023/AMPCliff/tree/FLaG.

Subjects: Artificial Intelligence , Biomolecules

Publish: 2026-09-01 07:32:44 UTC


#13 Mudskippers use tail thrusting to help crutching to move on mud of various wetness [PDF] [Copy] [Kimi] [REL]

Authors: Divya Ramesh, Gargi Sadalgekar, Jiangqi Tan, Chen Li

At the water-land interface, amphibious fishes encounter wet flowable substrates made of granular solid-water mixtures, which can stay solid or flow like a fluid. As these substrates become wetter or drier, their yield strength (at which solid-fluid transition occurs) and cohesion (how sticky they are) both change, challenging locomotion. Despite substantial understanding of tetrapod locomotion on flowable substrates (mostly dry sand), we know little about how amphibious fishes cope with wet flowable substrates of various wetness. Here, we studied mudskippers on clay mud of controlled, variable wetness over the range where solid-fluid transition occurs. As mud became wetter, its strength decreased by 100-fold, leading the animal to sink deeper, with larger areas of body and fins contacting mud. By contrast, mud stuck most easily at intermediate wetness. The increased sinkage and contact and stickiness change caused more mud to stick to and pull against the animal on wetter mud. We also tested dry mud, which stuck to animal fins as its mucus dried. Despite these challenges, the mudskipper predominately used a conserved crutching gait on all except the wettest mud tested, with a modest performance reduction. When normal crutching became less effective, the animal assisted it with tail thrusting, by bending and straightening it to push downward and backward to generate additional thrust and lift, or even thrusting the tail to jump. These observations suggest that mudskipper's crutching motor program is well adapted to its native muddy substrates but inflexible, with most novelty in tail use.

Subjects: Biological Physics , Soft Condensed Matter , Robotics , Systems and Control , Quantitative Methods

Publish: 2026-09-01 01:57:28 UTC


#14 Importance and methods to control, vary, and characterize mud strength for studying locomotion [PDF] [Copy] [Kimi] [REL]

Authors: Divya Ramesh, Gargi Sadalgekar, Qiyuan Fu, Zachary Souders, Jack Rao, Chen Li

Animals and robots encounter mud at the water-land interface. Like sand, mud can stay solid or flow like a fluid. Unlike sand, the yield strength of mud at which solid-fluid transitions occur depends on not only the amount of solid relative to fluid (water in mud, air in dry sand), but also how much coarse grains and fine clay are within the solid. Despite understanding of locomotion on/within dry sand dominated by coarse grains with repulsive normal forces and friction, little is known for mud dominated by fine clay with strong cohesion. Here, we developed methods to prepare uniform mud of controlled, variable yield strength and characterize and track its drift from water evaporation. Compared to other flowable substrates, mud strength measured by upward force during penetration is weaker and can vary more, and mud sticks more during extraction to pull downward, making it more challenging for locomotion.

Subjects: Biological Physics , Soft Condensed Matter , Robotics , Systems and Control , Other Quantitative Biology

Publish: 2026-09-01 01:56:47 UTC


#15 Learning Task-Specific Antibody Representations via Function-Aware Masking [PDF] [Copy] [Kimi] [REL]

Authors: Ayan Goel, Thomas A. Walton, Amirali Aghazadeh

Antibody-specific language models pretrained via masked language modeling (MLM) learn representations that are critical for downstream sequence design and property prediction tasks. Yet, the corruption process itself is rarely leveraged as a source of inductive bias during pretraining. While preferentially masking complementarity-determining regions (CDRs) improves binding-related predictions, antibodies possess diverse biological priors over a variety of functions. Herein, we introduce function-aware masking, a family of pretraining algorithms that align mask placement with specific functional priors (e.g., from IMGT annotations or structure predictions) to shape the learned representation space. We show that these specialist masking strategies significantly improve performance on their respective objectives, yielding up to a 14% gain on structure-related tasks and up to a 5.9x improvement on CDR-related tasks. To further improve performance across multiple functional axes, we develop hybrid masking strategies that integrate multiple priors, balancing reconstruction over binding, structural, and biophysical objectives. Our results demonstrate that informed mask placement provides a parameter-free mechanism for imposing functional inductive biases in antibody language model training.

Subjects: Machine Learning , Biomolecules

Publish: 2026-09-01 00:37:56 UTC