Quantitative Biology

2026-09-22 | | Total: 29

#1 Phylogenetic Inference and the Stickiness of Fréchet Means, via Precise Asymptotics of an Embedded Random Walk [PDF] [Copy] [Kimi] [REL]

Author: Adam Quinn Jaffe

A well-known phenomenon in statistical analyses of populations of phylogenetic trees in the Billera-Holmes-Vogtmann space is that the topology of the Fréchet mean tree can contain multifurcations (i.e., internal nodes with more than two children), which raises the practical question of whether this reflects a population-level branching structure (hard polytomy) or merely sampling variability in the data (soft polytomy). This is an instance of the more general phenomenon of "stickiness" in non-Euclidean statistics, whereby the sample Fréchet mean in certain non-positively curved stratified spaces becomes permanently trapped in a lower-dimensional stratum. In this work, we identify a particular multidimensional random walk embedded within the Fréchet mean process, and we show that the time at which stickiness occurs is determined by the largest last-passage time above zero of the coordinates of this random walk. Using this representation, we develop a fully nonparametric procedure for estimating the probability that trifurcations in a sample Fréchet mean tree will bifurcate at some future time if more observations are collected. Lastly, we apply our methodology to a problem in phylogenetics where we consider whether an observed trifurcation in the species tree of primates, glires, and tree shrews is genuinely trifurcated at the population level.

Subjects: Populations and Evolution , Statistics Theory , Methodology

Publish: 2026-09-21 16:06:58 UTC


#2 Adapting Boltz-2 with limited experimental activity data improves early enrichment in virtual screening [PDF] [Copy] [Kimi] [REL]

Authors: Kairi Furui, Masahito Ohue

Virtual screening aims to prioritize active compounds from large chemical libraries within a limited experimental budget. When applying Boltz-2 to virtual screening, a key challenge is how to use limited experimental data from the target assay to improve the prioritization of active compounds. We investigated whether fine-tuning the Boltz-2 affinity heads with a small number of binary activity labels could improve early enrichment of active compounds in hit discovery. We compared fine-tuning with 40-300 labels in a retrospective evaluation on eight MF-PCBA targets. With 300 activity measurements, fine-tuning increased the number of actives in the top 1% by a geometric mean of 1.77-fold across the eight targets and improved average precision (AP) by 2.14-fold relative to the control without fine-tuning. We also investigated whether rescoring a subset of candidates could retain the improvement in hit recovery by reranking only the top-ranked Boltz-2 candidates with the fine-tuned head. Restricting rescoring to approximately 10% of the evaluation set retained hit recovery comparable to full rescoring. These findings show that affinity-head fine-tuning with limited activity labels improves early enrichment with Boltz-2 and that this benefit can be retained when rescoring a restricted set of candidates.

Subjects: Biomolecules , Artificial Intelligence , Machine Learning , Quantitative Methods

Publish: 2026-09-21 09:02:04 UTC


#3 Decoding enzyme-substrate interaction topology reveals principles underlying catalytic efficiency and mutational outcomes [PDF] [Copy] [Kimi] [REL]

Authors: Weiren Zhao, Takeyuki Tamura

The enzyme turnover number (kcat) defines catalytic efficiency and constrains quantitative models of metabolism, yet the molecular determinants governing kcat and its response to mutation remain poorly understood. Measurements are sparse and labor-intensive, and most computational approaches provide numerical predictions without explaining how enzyme-substrate interactions shape catalytic outcomes. A central challenge is therefore to identify the topological principles that determine where mutations act and how their functional outcomes are encoded within the enzyme-substrate interaction network. Here, we show that catalytic efficiency and mutational effects can be interpreted through enzyme-substrate interaction topology. We developed Interkcat, an interpretable bidirectional cross-attention framework that captures reciprocal coordination between protein residues and substrate atoms. Optimized on a unified benchmark, Interkcat achieves state-of-the-art predictive performance (R2 = 0.701). From its learned representations, we derive an Interaction Topology Score (ITS) that identifies sequence regions statistically enriched for mutation-sensitive sites without explicit structural inputs. We further demonstrate that higher-order topological features distinguish opposing mutational outcomes: lethal mutations disrupt coordinated networks, whereas activity-preserving or enhancing mutations retain sparse, globally organized coupling. These findings establish interaction topology as a unifying principle linking enzyme sequence, catalytic efficiency, and evolutionary perturbation.

Subject: Biomolecules

Publish: 2026-09-21 06:18:12 UTC


#4 Retracing the Process of Translation: Proteome-wide mapping of stable transcriptomic predictors of protein abundance in cancer cell lines [PDF] [Copy] [Kimi] [REL]

Authors: Johannes Schlüter, Alexander Schönhuth

Understanding the relationship between gene expression and protein abundance is central to molecular and systems biology. While gene expression reflects transcriptional activity, proteins are the functional molecules that determine cellular phenotypes. However, numerous post-transcriptional and translational regulatory layers complicate this relationship, and prior studies have reported only weak to moderate correlations between RNA and protein levels. Predicting protein abundance from transcriptomic data remains challenging, but it is a valuable goal for biological insight, especially when proteomic data is limited or unavailable. In this study, we applied a large-scale, Ridge regression-based feature selection strategy to identify predictive gene expression features for each of 8,423 proteins across 940 cancer cell lines. To our knowledge, this is the first work to perform such comprehensive protein-wise feature selection at this scale. Our analysis revealed both globally predictive and context-specific gene features. These included biologically meaningful modules such as immune-related genes, HOX transcription factor targets, and cytoskeletal components. The models identified stable candidate gene-protein associations that remained interpretable at the level of individual proteins and recurrent transcriptomic predictor patterns. Our approach enables interpretable modeling of protein expression from transcriptomic data and provides insight into transcriptomic features associated with protein abundance. This framework may support hypothesis generation, protein imputation in incomplete datasets, and deeper understanding of post-transcriptional regulation in cancer biology.

Subject: Quantitative Methods

Publish: 2026-09-21 02:09:37 UTC


#5 Binding-Motivated Contextuality: A Cross-Domain Cyclic Test in Perception and Judgment [PDF] [Copy] [Kimi] [REL]

Author: Adam Y. Shavit

Perceptual binding and the contextuality of judgment are studied apart, in psychophysics and decision research. We argue they share one obstruction: a nonzero class in $H^1$ of a presheaf with no global section -- though only contextuality is tested, since binding's obstruction vanishes. We build on sheaf formulations of predictive coding (Seely 2025) and contextuality (Abramsky & Brandenburger 2011): a cyclic set of pairwise judgments admits a global (noncontextual) explanation exactly when the cyclic (Suppes-Zanotti / n-cycle) inequalities hold. Its obstruction is measured by the complete contextual fraction CF (Abramsky, Barbosa & Mansfield 2017), not the weaker Cech invariant, which certifies contextuality but can miss it (Caru 2017). Penrose (1992) classifies the continuous tribar by $H^1$ over the multiplicative group $\mathbb{R}^+$ of depths -- but the $\mathbb{Z}_2$ case is what we test. We build the perceptual arena from two binary judgments per cyclic-dominance pairing, realizing the same frustrated-cycle obstruction as the survey. The central test is cross-domain: one cohort performs both; shared mechanism predicts a correlation between the arenas' signed obstruction margins -- CF before clamping -- which neither field has measured. Both are inconsistency scores, so a bare correlation is confounded by general response consistency. The prediction is therefore confound-residualized: the correlation must survive partialling out a general-consistency factor from a variance- and reliability-matched control. A sufficiently precise null would count against that account, provided the preregistered reliability checks pass and the control does not load on the obstruction. All three tests are designed but unrun.

Subject: Neurons and Cognition

Publish: 2026-09-21 01:19:48 UTC


#6 From Biological Precursors to Artificial Cognition: Consciousness, Embodiment, and the MEM Architecture [PDF] [Copy] [Kimi] [REL]

Authors: Janusz A. Starzyk, Wiesław L. Galus

This article asks under what conditions artificial intelligence could warrant a rational attribution of consciousness. Linguistic ability, multimodality, memory, planning, action control, and humanoid embodiment are not sufficient evidence of phenomenal experience. Biological precursors such as excitability, homeostasis, neural networks, and hierarchical representation instead identify functions whose counterparts may be engineered. The paper compares conventional LLMs, hybrid h-LLMs, vision-language-action systems, and embodied agents with the Motivated Emotional Mind (MEM) architecture. MEM adds receptor-grounded representations, regulatory self-monitoring and interoception, affect, semblion-based associative memory, and re-entry into lower sensory and interoceptive maps. A related formulation defines a phenomenal state as a dynamic, integrated state of the whole embodied system in which sensor-grounded modal content is recurrently stabilized and coupled to a suprathreshold interoceptive representation of the system's regulatory state. The mathematical model of MEM published on arXiv specifies implementable relations among these components and makes the proposal more precise and testable. Comparative and causal tests in humans, invertebrates, and artificial systems could strengthen or weaken the MEM capabilities. MEM is therefore presented as a falsifiable research program requiring empirical validation and ethical caution.

Subject: Neurons and Cognition

Publish: 2026-09-20 19:24:53 UTC


#7 Demographic inference of pathogen-infected populations from partially observed transmission forests [PDF] [Copy] [Kimi] [REL]

Author: Matthew Hall

Genomic epidemiology has several methods for identifying which sampled individuals are linked by close proximity in the transmission chain, whether by grouping them into clusters under a genetic distance threshold or by identifying probable direct transmission pairs. Individuals found to have no sampled neighbours---singletons---are usually set aside. We argue that they are informative. Whether any two sampled individuals prove to be linked depends on the size of the sampling frame and the number of independent lineage introductions into it, and thus the balance of linked and unlinked individuals is itself data about these quantities. We formalise this by treating the intersection of a transmission tree with a sampling frame as a labelled rooted forest, from which a fixed number of nodes are sampled uniformly at random. Using the all-minors matrix-tree theorem, we derive a closed-form expression for the number of labelled rooted $k$-forests on $N$ nodes in which a specified set of nodes is independent. From this we obtain, again in closed form, the probability that a sample of a given size contains no linked individuals at all, and the likelihood of an arbitrary observed configuration of clusters, both when the transmission structure within each is known, and when only the cluster sizes are. The approach is closer to a survey, or to mark-recapture, than to conventional phylodynamic model fitting: the information comes from the linkage structure of a single sample alone with no assumptions regarding pathogen dynamics. We outline three applications: power calculations for prospective studies, a test of the uniform sampling assumption that underlies the reading of clusters as transmission hotspots, and demographic inference itself. We finally set out the assumptions that a more flexible implementation would need to relax.

Subject: Populations and Evolution

Publish: 2026-09-20 13:15:52 UTC


#8 Minimality in Reflexive and Stoichiometric Autocatalysis [PDF] [Copy] [Kimi] [REL]

Authors: Richard Golnik, Thomas Gatter, Wim Hordijk, Peter F. Stadler, Nicola Vassena

Autocatalysis, the ability of a chemical subsystem to sustain its own constituents when supplied with sufficient food molecules, has been closely related to the origin of life on Earth. Emerging from Wilhelm Ostwald's considerations about an explicit autocatalytic reaction, different notions of autocatalysis have been developed over the years. The two most prominent are reflexively autocatalytic F-generated sets (RAFs) and stoichiometric autocatalysis. After having shown that each RAFs is, under reasonable conditions, in general stoichiometrically autocatalytic, we examine here the relationship between the two notions of minimality: irreducible RAFs and autocatalytic cores. To this end, we overcome the obstacle that RAFs and stoichiometric autocatalysis have been formalized in distinct systems of chemical reactions, i.e., catalytic reaction systems (CRS) and chemical reaction networks (CRNs), respectively. We show that reactions in a CRS constitute equivalence classes of reactions of the corresponding CRNs w.r.t. their specific catalyzations. Using the fact that CRN and CRS can be canonically identified whenever each CRS reaction is associated with a single catalyzation, we demonstrate that the Kőnig graph of a monocatalyzed, irreducible RAF is composed of strong blocks devoid of food and waste species that are separated by reaction vertices, each of which contains an autocatalytic core. In fact, a single irreducible RAF can, in general, contain exponentially many autocatalytic cores.

Subjects: Molecular Networks , Combinatorics

Publish: 2026-09-20 12:43:31 UTC


#9 Chaotic Dynamics-Regulated Topological Learning for Patient-Specific Preictal State Identification [PDF] [Copy] [Kimi] [REL]

Authors: Zihan Wang, Daixin Li, Guilin Wang, Mushal Zia, Xiaoqi Wei, Xiang Xiang Wang, Jian Jiang

Epileptic seizures arise from complex, nonlinear interactions within brain networks, yet reliable electroencephalographic (EEG) prediction remains challenging due to the nonstationary and heterogeneous nature of neural dynamics. Existing methods typically analyze EEG data as static or weakly time-dependent snapshots, overlooking the intrinsic dynamics and lacking the geometric sensitivity to capture the hierarchical, localized evolution of the epileptogenic zone. To address these limitations, we propose an offline, patient-specific evaluation of chaotic dynamics-regulated topological learning (CDRTL) for distinguishing preictal from interictal EEG states. This framework unifies chaotic dynamics, multiscale algebraic topology, and local network differentiation. Specifically, we partition EEG signals into discrete functional subnets based on correlation strengths, capturing the multi-scale connectivity of the brain. By modeling each node as a Lorenz oscillator, we embed the underlying chaotic dynamics into the network architecture. We then apply the persistent Laplacian to simultaneously extract topological invariants and geometric shape evolution through harmonic and non-harmonic spectral analysis. Additionally, a node-removal topological differentiation strategy isolates localized neural contributions. Our framework was evaluated on the CHB-MIT database using balanced preictal and interictal labels and stratified channel-level cross-validation within each patient. The results support offline discrimination of preictal and interictal channel-level nodes within fixed patient-specific networks. Because representations are constructed from the complete network, including held-out unlabeled nodes, before cross-validation, the reported performance is specific to this transductive setting and does not establish generalization to unseen EEG windows, seizures, or patients.

Subject: Neurons and Cognition

Publish: 2026-09-20 03:08:25 UTC


#10 From daylight to darkness: a nonlocal model of circadian activity cycles [PDF] [Copy] [Kimi] [REL]

Authors: Dan Bell, Timothé Boyer, Hao Wang

An animal's ability to perceive its surroundings, as well as the ecological interactions it experiences, can be strongly influenced by the sleep patterns of both itself and surrounding species. Animals with limited sensory capabilities may struggle to navigate their habitat at night, while in a predator-prey system, it may be detrimental for prey to be inactive while predators are awake and hunting. Here, we develop a general mathematical framework to study these interactions using partial differential equations, incorporating activity cycles and daylight-dependent perception into existing models of animal movement. We show that both the dominant sensory mode and the length of daylight are key factors determining the strength of aggregation in social species. Extending this framework to a two-species predator-prey system, we use game-theoretic approaches to show that the equilibrium sleep strategy of each species is shaped not only by its own sensory capabilities, but also by those of its ecological opponent. Depending on the sensory capabilities of the two species, equilibrium strategies may consist of a single sleep pattern or a combination of multiple sleep patterns.

Subject: Populations and Evolution

Publish: 2026-09-19 02:52:16 UTC


#11 Foundation-model-based multi-label phenotyping of combined hyperkinetic movement disorders [PDF] [Copy] [Kimi] [REL]

Authors: Laura Cif, Zohra Souei, Diane Demailly, Mayte Castro-Jimenez, Juan Dario Ortigoza-Escobar, Muhammad Mushhood Ur Rehman, Morgan Dornadic, Sophie Huby, Gun-Marie Hariz, Cecile Hubsch, Nathalie Dorison, Eduardo M. Moraud, Jocelyne Bloch, Gabriella Horvath, Olivier Oullier, Xavier Vasques

Movement disorders (MDs) frequently co-occur, yet phenomenological and severity assessment shows substantial inter-rater variability. Markerless video could improve reproducibility, but prior work is largely single-symptom, depends on standardized acquisition, and lacks validation and transfer across ages and sites. We combined two foundation models into one frozen backbone: Segment Anything Model 3 (SAM 3) for dense, per-frame markerless segmentation summarized into geometric, contour and grid kinematic signals, and TabICLv2, a tabular foundation model, for in-context multi-label classification of eight hyperkinetic MD phenomenologies. Trained on standardized recordings of 21 adults and 4 controls, it transferred unchanged to two independent datasets, pediatric (n=12) and tremor-dominant adult (n=20), assessed with the CODY-SAMP scale; only the patient-level decision step was recalibrated per site. Under clinician consensus labels, false positives fell to zero in both datasets. Dystonia recovered perfectly (7/7 pediatric; 15/15 adult held-out), chorea fully in children (3/3), and tremor was recovered in adults (11/15) once a tremor-rich cohort made it evaluable, through recalibration alone. Per-region effect-size analysis gave clinically coherent, phenomenology-specific signals and identified myoclonus as the principal failure. Against YOLOv8 sparse keypoints, the dense representation matched under clinician permissive labels (Jaccard 0.63 vs 0.63) and was markedly more robust under clinician-label consensus (0.93 vs 0.76). This frozen foundation-model backbone with light per-site calibration yields transferable, interpretable, conservative multi-label phenotyping of co-occurring hyperkinetic MDs across ages and from standardized to routine video, adding robustness on high-confidence, clinician-agreed labels. Prospective multi-centre validation is required before clinical use.

Subjects: Quantitative Methods , Neurons and Cognition

Publish: 2026-09-17 15:50:35 UTC


#12 2D reaction-diffusion model-based biopsy simulation for dynamic tumor growth parameter estimation [PDF] [Copy] [Kimi] [REL]

Authors: Veronika Hofmann, Pirmin Schlicke, Jan Zawallich, Heiko Enderling, Christina Kuttler

Once diagnosed, cancer requires a fast, reliable and preferably cost-efficient assessment of the current state and potential progression of the disease. A new method for estimating tumor cell diffusivity $D$ and proliferation rate $γ$ in the context of the mechanistic reaction-diffusion equation from single-point-in-time routine biopsies aims to deliver just that, and quantities computed from the parameter estimates have recently been tested as a new biomarkers for risk-stratification in radiotherapy. Here, we extend the findings of this previous work by providing a first theoretical validation. The method is applied to in-silico biopsies which are generated by solving the two-dimensional reaction-diffusion equation for different growth terms (exponential and logistic) with a Dirac-Delta initial condition, and transforming the continuous results into spatial point patterns via a form of reverse coarse-graining. If no information about tumor age is used, in short-term experiments the original dispersion length $\sqrt{D/γ}$ could be retrieved with a relative root mean squared error (RRMSE) of around 8% and an $\text{R}^2$-value of 0.97. In long-term experiments, the RRMSEs ranged from 8 to 14% and the $\text{R}^2$-values from 0.75 to 0.98. The scaled front velocity $\sqrt{D \cdotγ}$, which can only be estimated if information about tumor age is available, was retrieved with an RRMSE of 7% and an $\text{R}^2$ of 0.98 in both, the short-term and the long-term experiments.

Subject: Other Quantitative Biology

Publish: 2026-09-17 13:01:12 UTC


#13 Exploring the robustness of permutation entropy analysis to differentiate between closed-eyes and open-eyes resting states [PDF] [Copy] [Kimi] [REL]

Authors: Juan Gancio, Natalia López López, Antonio J. Pons, Giulio Tirabassi, Cristina Masoller

Electroencephalography (EEG) is a noninvasive technology that is widely used to monitor brain states, and many efforts are focused on developing reliable and efficient data analysis methods for EEG recordings. Here, we apply ordinal analysis to the EEG recording of the resting state of 109 healthy subjects measured in two different conditions: with eyes closed (EC state) or eyes open (EO state). We study the robustness of the temporal permutation entropy ($PE$) and the spatial permutation entropy ($SPE$) with respect to the presence of blinking artifacts in the EO recordings, the duration of the recordings, and the number of EEG channels analyzed. We perform a paired statistical test to assess whether these quantities differ significantly in the EO and EC recordings of the same subject. We find that $PE$ and $SPE$ perform surprisingly well, as they detect significant differences when calculated from the raw signals, even within a short time interval ($PE$) or from a reduced number of electrodes ($SPE$).

Subjects: Neurons and Cognition , Data Analysis, Statistics and Probability

Publish: 2026-09-07 16:11:36 UTC


#14 Nucleosome simulations suggest mechanisms of electrostatically-driven mesoscale chromatin evolution [PDF] [Copy] [Kimi] [REL]

Author: Siddhartha G. Jena

Nucleosomes are structures made up of proteins called histones that bind and compact DNA, driving the mesoscale organization of the chromatin polymer. Although histone proteins have diversified over evolutionary time, their contributions to the corresponding diversification of chromatin structure are poorly understood. Here, we mine protein databases for histones and create \emph{in silico} nucleosomes for 3241 organisms spanning $>$1.5B years of evolution. Using a combination of electrostatic calculations and coarse-grained molecular dynamics simulations, we reveal extensive biophysical diversification of the nucleosome unit. Finally, we perform coarse-grained oligonucleosomal simulations on a subset of evolutionarily and biophysically divergent nucleosomes, demonstrating dramatic differences in bulk phase behavior of chromatin. Taken together, our results suggest a paradigm in which histones may have evolved to facilitate particular types of mesoscale chromatin behavior.

Subjects: Soft Condensed Matter , Biomolecules

Publish: 2026-09-21 17:10:23 UTC


#15 Towards Accurate Prediction of Mutation-Induced Changes in Protein Structure [PDF] [Copy] [Kimi] [REL]

Authors: Zhuoyi Liu, Alex Calabrese, Corey S. O'Hern

Proteins can possess numerous mutations relative to their wild-type amino acid sequences with minimal impact to their structure and function. However, in other cases, even a single amino acid mutation relative to the wild-type sequence can lead to a large change in structure or even a disease phenotype. While the accuracy of wild-type protein structure prediction has improved significantly in recent years, it remains difficult to accurately predict the structure of mutant proteins. Here, we characterize the local mutation-induced structural changes in proteins for a dataset of wildtype and the corresponding single-amino acid mutant x-ray crystal structures from the Protein Data Bank (PDB). We find that mutation-induced structural changes in these proteins are localized at the site of the mutation, decaying rapidly with increasing spatial distance from the mutation site. In addition, we evaluate how well AlphaFold3 can recapitulate the observed mutation-induced structural deformations in the x-ray crystal structures. We find that the accuracy of the AlphaFold3 predictions decreases strongly with increasing mutation-induced deformation. In contrast to the results for AlphaFold3, the Pearson correlation between a single physical feature, i.e. the change in solvent accessibility, and the mutation-induced deformation does not depend on the magnitude of the deformation. Our results and analyses provide a framework for further studies aimed at predicting the structural changes in proteins caused by single amino acid mutations.

Subjects: Biological Physics , Biomolecules

Publish: 2026-09-21 16:22:36 UTC


#16 Experimental evidence of generalization in quantum machine learning in small-data regime [PDF] [Copy] [Kimi] [REL]

Authors: Leena Anthony, Artemiy Burov, Nicolas Piro, Matteo Dal Peraro, Clément Javerzac

Quantum machine learning is a promising paradigm for learning from limited data, a central bottleneck in domains such as medical imaging, clinical trials, and rare diseases. Quantum convolutional neural networks (QCNNs) are particularly attractive in this setting, combining a hierarchical architecture with strong inductive bias and a parameter count that grows only logarithmically with system size. Their appeal rests on the generalization bounds of Caro et al. (2022), which show that the generalization error of a quantum model scales with the number of trainable parameters rather than with the Hilbert-space dimension, placing QCNNs in a potentially sample-efficient regime. We develop a hardware-compatible QCNN with mid-circuit measurement and classical feed-forward, and show on a binary handwritten-digit task that strong test performance is achievable from as few as 10 training samples, with the generalization error decreasing as the training set grows. At a matched 45-parameter budget the QCNN learns where an equally small classical convolutional network stays at chance, although an unconstrained classical baseline with roughly 25,000 parameters remains strongest when data are plentiful. Transpiling amplitude and angle encoded circuits across image resolutions from 2x2 to 512x512 pixels then exposes the dominant scaling bottleneck: amplitude encoding stays qubit-efficient but grows extremely deep, whereas angle encoding stays shallow but becomes qubit-prohibitive. On the medically motivated BreastMNIST benchmark the QCNN does not surpass the unconstrained classical network, yet it learns consistently above chance using orders of magnitude fewer parameters. Our results indicate that for QCNNs, learning from few samples is attainable in practice, whereas scaling to realistic image data is constrained less by optimization than by data encoding and hardware execution.

Subjects: Quantum Physics , Quantitative Methods

Publish: 2026-09-21 14:30:06 UTC


#17 A Proof of the Global Attractor Conjecture in a Special Case [PDF] [Copy] [Kimi] [REL]

Author: Carsten Wiuf

We prove the Global Attractor Conjecture for complex balanced mass-action reaction networks whose reachable siphons satisfy two structural conditions, allowing multiple linkage classes. The proof proceeds in two steps. A structural condition relating the stoichiometric space to the reactions active within a boundary face ensures that a stoichiometric compatibility class contains at most one boundary equilibrium with any prescribed zero set. Finiteness of the possible zero sets and connectedness of the $ω$-limit set then imply that any boundary limit set consists of a single equilibrium. To exclude convergence to such an equilibrium, we consider the embedded reaction network obtained by projecting onto the vanishing species and freezing the concentrations of the surviving species at a positive limit. The structural condition guarantees complex balance of this embedded network, while a further condition on its minimal active linkage classes ensures that an explicit Chetaev function is strictly increasing near the boundary. This excludes the boundary point as an accumulation point whenever the surviving concentrations converge. Consequently, every positive trajectory converges to the unique positive equilibrium in its stoichiometric compatibility class. Examples illustrate the hypotheses and their relation to strong endotacticity.

Subjects: Dynamical Systems , Molecular Networks

Publish: 2026-09-21 13:19:04 UTC


#18 QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation [PDF] [Copy] [Kimi] [REL]

Authors: Demian Pavlyshenko, Bohdan Pavlyshenko

Functional annotation of newly sequenced proteins remains a bottleneck in molecular biology: the number of sequences in public repositories grows far faster than the capacity for manual curation. Most computational approaches consider annotation as multi-label classification over a fixed ontology, which constrains predictions to a predefined label set. In this work we study the the protein annotation as a sequence-to-text generation problem. We fine-tune the 3B-parameter Ministral 3 base model with QLoRA (4-bit NF4 quantization with low-rank adapters) on sequence annotation pairs. We assess predictions with an LLM-as-expert protocol: a GPT model prompted as a senior molecular-biology curator scores organism identification as binary and function annotation quality. We conclude that QLoRA-fine-tuned compact LLMs can generate curator-style annotations with genuine biological value for a substantial subset of proteins. We also discuss future directions in data quality, model scaling, and evidence grounding that are needed to make the approach sufficiently reliable for practical use.

Subjects: Computation and Language , Artificial Intelligence , Neural and Evolutionary Computing , Quantitative Methods

Publish: 2026-09-21 13:09:59 UTC


#19 Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea [PDF] [Copy] [Kimi] [REL]

Authors: Reza Saputra, Diah Harnoni Apriyanti, André Schuiteman, Kurt Metzger, Ashley Field, Katharina Nargar, William Edwards

New Guinea is the world's richest island flora (~2,856 orchid species), yet most species are represented by only a handful of photographs, far fewer than direct species-level classification requires. Methods for fine-grained identification in such species-rich, data-poor floras are needed, and it remains unclear which backbone architecture and pretraining strategy best support them. We built a two-stage system that first predicts the genus of a query photograph, then retrieves visually similar reference images of candidate species using FAISS. We compared four pretrained backbones -- two Vision Transformers (ViTs; DINOv2, BioCLIP 2) and two CNNs (ConvNeXt V2-L, EfficientNetV2-L) -- fine-tuned under an identical protocol on a fixed, species-stratified partition of 16,701 photographs spanning 120 genera and 1,350 species, assessing accuracy, calibration, error structure, species retrieval, and open-set detection of novel genera. DINOv2 attained the best genus performance (macro top-1 66.9%, 95% CI 63.7-70.6; global top-1 88.9%); both ViTs outranked both CNNs, and general-purpose self-supervised pretraining (DINOv2) outperformed domain-matched biological pretraining (BioCLIP 2) by 7.1 points of macro top-1. Errors concentrated on two abundant genera acting as error attractors. DINOv2 embeddings achieved species Recall@5 of 86.6% and genus Recall@5 of 98.7%; temperature scaling reduced every backbone's Expected Calibration Error to about 0.03; and a distance-based open-set gate flagged unseen genera (mean AUROC 0.958). A self-supervised Vision-Transformer backbone combined with embedding retrieval is an effective, deployable strategy for fine-grained identification in species-rich, data-poor floras. The system is released as an open web application (the New Guinea Orchid Identifier), offering a practical template for other hyperdiverse, under-documented taxa.

Subjects: Computer Vision and Pattern Recognition , Machine Learning , Quantitative Methods

Publish: 2026-09-21 03:39:10 UTC


#20 Matched-Input Estimates Differ in Sign Across Architectures: Auditing EEG Foundation Models on Motor Imagery [PDF] [Copy] [Kimi] [REL]

Authors: Kevin Zhou, Sparsh Roy

Pretrained EEG foundation models are increasingly proposed as general-purpose encoders for brain-computer interfaces, yet recent benchmarks disagree about when their representations transfer to downstream tasks. We audit LaBraM and CBraMod on motor imagery under a validation-locked protocol in which preprocessing, architecture, optimization, freeze depth, checkpoint, temperature, and method selection are determined using training-session data only. On four-class BCI Competition IV-2a, every supervised comparator evaluated here outperforms every foundation-model configuration, including validation-selected fine-tuning. We then examine a key confound: foundation models and task-specific decoders are normally evaluated using different input pipelines. Retraining three supervised architectures on the broadband arrays consumed by the foundation models produces matched-input accuracy differences of opposite sign across architectures: broadband input improves ATCNet by 0.078 accuracy while reducing EEG Conformer accuracy by 0.088. None of the three individual matched-input terms is significant after multiple-comparison correction at n = 9, so we treat the sign variation descriptively rather than as a formal architecture-by-pipeline interaction. These observed sign differences suggest that a single comparator may not provide an architecture-invariant decomposition of a pretrained-versus-supervised performance gap. The four-class deficit also does not reproduce uniformly across motor-imagery datasets: on two-class BNCI2014-004 we cannot detect the same separation between fine-tuned CBraMod and the supervised comparators. Finally, validation-fitted temperature scaling returns foundation-model calibration error to the supervised range despite substantially lower four-class accuracy.

Subjects: Machine Learning , Neurons and Cognition

Publish: 2026-09-20 22:56:32 UTC


#21 A discrete generative model of neuronal spiking activity on microelectrode arrays [PDF] [Copy] [Kimi] [REL]

Authors: Md Sayed Tanveer, Mohammed A. Mostajo-Radji, Ge Wang

Generative models of neural activity could help characterize tissue dynamics, compare experimental conditions, and simulate population activity for applications ranging from disease and drug-response studies to closed-loop experimentation. Existing approaches, however, typically assume a fixed set of sorted neurons, whereas high-density microelectrode arrays produce extremely sparse, array-wide binary spike volumes in which the observed subset of electrodes varies across assays. We introduce a discrete generative model that represents this activity using a shared vocabulary of spatiotemporal motifs. A residual vector-quantized autoencoder learns the motif vocabulary, while a factorized masked transformer predicts where activity occurs and which motif appears at each active location. We evaluate the model on 31 assays spanning human brain organoids and acute \emph{ex vivo} human hippocampal tissue. The learned motifs are broadly reused: assay identity explains only $9%$ of the entropy in motif use, and motif overlap across tissue types is comparable to overlap within them. When representation quality is evaluated independently of the generative prior, our approach achieves $5.2\times$ the voxel-level reconstruction average precision of a matched flat tokenizer. For masked completion and free generation, the full model achieves $1.4$--$2.6\times$ the site-level average precision of the matched generative baseline and outperforms it across all four families of generation metrics. These results establish a compact, reusable representation for array-wide spiking activity without learned assay-specific parameters, providing a scalable foundation for generative modeling across diverse neural preparations.

Subjects: Machine Learning , Signal Processing , Neurons and Cognition

Publish: 2026-09-20 22:31:29 UTC


#22 THz Spectroscopy of Urine Vapors from Patients with Prostate Cancer and Benign Prostatic Hyperplasia: A Pilot Analysis [PDF] [Copy] [Kimi] [REL]

Authors: V. A. Atduev, A. V. Maslennikova, V. L. Vaks, E. G. Domracheva, M. B. Chernyaeva, V. A. Anfertev, K. A. Atduev, Y. Tawalbeh, M. F. Pereira

Approximately 13% of men will be diagnosed with prostate cancer (PC) during their lifetime. Serum prostate-specific antigen (PSA) is widely used for screening and risk assessment; however, PSA elevations are not cancer-specific and may also occur in benign conditions such as prostatitis and benign prostatic hyperplasia (BPH). Additional non-invasive approaches capable of provid- ing complementary molecular information are therefore needed. Here, we present an exploratory pilot study using high-resolution terahertz (THz) spectroscopy to examine urine-derived volatile and thermal-decomposition products. Analysis of urine samples from 24 patients with PC and 14 patients with BPH identified differences in the reported molecular-assignment patterns and candi- date spectral features for further evaluation. The study was designed for candidate identification and feasibility assessment and did not evaluate diagnostic accuracy or superiority to PSA. These findings support further investigation of THz spectroscopy as a potential source of complementary molecular information alongside PSA and other established clinical assessments. The study also outlines the technical standardization and clinical-validation requirements that must be addressed before this approach can be considered for routine clinical use.

Subjects: Other Condensed Matter , Quantitative Methods

Publish: 2026-09-20 19:09:36 UTC


#23 Reconstructed holograms and explanation-aware evaluation for low-cost computational pollen analysis in veterinary cytology [PDF] [Copy] [Kimi] [REL]

Authors: Swarn Warshaneyan, Joial Danyal, Blaž Cugmas, Mindaugas Tamošiūnas, Edgars Kviesis-Kipge, Kirishanth Manivannan, Roberts Kadiķis

Automated pollen analysis supports veterinary cytology, but brightfield microscopy is costlier and more complex than lens-less digital in-line holographic microscopy. We evaluate whether reconstructed holograms can narrow this gap and whether model explanations remain reliable under modality change. Six pollen species were imaged by brightfield and holographic microscopy. Raw, single back-propagation and iterative phase retrieval holograms were evaluated with YOLOv26s detection and MobileNetV4 classification after anchor-based annotation transfer. Six attribution methods were assessed for spatial grounding and faithfulness with the Attribution Health Inspection and Repair (AHIR) protocol, which tests model brittleness under weak noise and corrects attribution-map granularity when needed. Brightfield achieved 0.6890 mAP50-95 (0.8865 mAP50) for detection and 0.9687 macro-F1 (0.9705 accuracy) for classification. Reconstructed holograms narrowed the gap with a task-dependent split: p-type was strongest for detection at 0.5324 mAP50-95 (0.8229 mAP50), while r-type was strongest for classification at 0.7695 macro-F1 (0.7866 accuracy), both far above raw-hologram baselines. Activation-based explanations localized strongly on grains, and region-based methods retained ~60 to ~80% of faithfulness under holography. The holographic detector was highly brittle to weak perturbations, saturating deletion-based evaluation while insertion remained informative. Pixel-level gradient explanations approached random floor, yet spatial smoothing restored p-type gradient faithfulness from 0.05 to 0.51. For holographic classification, perturbation-based explanations remained faithful while gradient-based methods fell below random floor. Reconstruction improves low-cost holographic pollen analysis, while AHIR distinguishes genuine attribution failure from artifacts caused by model brittleness and map granularity.

Subjects: Computer Vision and Pattern Recognition , Machine Learning , Quantitative Methods

Publish: 2026-09-19 13:35:07 UTC


#24 Experimental Design for Controller Selection in Synthetic Biology [PDF] [Copy] [Kimi] [REL]

Authors: Eric Palanques-Tost, Ron Weiss, Calin Belta

Synthetic biology enables the design of genetic circuits that act as feedback controllers. These controllers are typically designed using computational models, but mismatch between model and real dynamics can lead to controllers that fail in practice. While methods to address this issue exist, synthetic biology introduces additional structural constraints. Genetic circuits are often highly constrained by experimental limitations, reducing controller design to selection among a limited set of implementable circuits rather than an optimization over a continuous space. As a result, multiple system hypotheses may lead to the same optimal controller within the implementable set. Reducing model uncertainty may therefore be irrelevant when the models lead to the same optimal controller. In this paper, we exploit this structure to develop an algorithm for controller selection in synthetic biology, formulating the problem as a decision-oriented experimental design problem over a finite controller set. We represent plant uncertainty using a set of hypotheses and select experiments to minimize the posterior controller selection risk, rather than global model uncertainty. Across three mechanistic case studies, our method reaches the stopping criterion in fewer experimental rounds than model uncertainty and random experiment selection policies, while maintaining a comparable success rate.

Subjects: Systems and Control , Quantitative Methods

Publish: 2026-09-19 05:12:39 UTC


#25 Asymptotics and finite sample bounds for prediction and smoothing in Wright-Fisher hidden Markov models [PDF] [Copy] [Kimi] [REL]

Authors: Luigi M. Malgieri, Filippo Ascolani, Matteo Giordano, Matteo Ruggiero

We study prediction and smoothing in hidden Markov models with a latent signal given by a multi-type Wright-Fisher diffusion and discrete-time categorical observations, motivated by repeated-sampling time-series settings, including temporally binned ancient-DNA data, in which noisy frequency counts are recorded at finitely many times. Our focus is on the exact Bayesian predictive and smoothing distributions available under parent-independent mutation, in relation to their large-sample targets under repeated within-time sampling. For a fixed collection-time grid and diverging within-time sample sizes, we show that the exact Wright-Fisher predictive and smoothing distributions converge in total variation to the corresponding population transition and bridge laws at the limiting neighboring frequencies. We then derive explicit finite-sample control for the predictive law and a corresponding finite-sample bound for the marginal smoother. Finally, at the inspection times, we show that the joint conditional law concentrates at the target frequencies and that its active coordinates are asymptotically Gaussian, while coordinates with zero true frequencies converge to Gamma limits at faster rates. Our regime imposes no restriction on dependence across inspection times beyond within-time sampling. The analysis rests on a fixed-interval tail bound for Kingman's coalescent block-counting process, which is of independent interest.

Subjects: Statistics Theory , Probability , Populations and Evolution , Methodology

Publish: 2026-09-18 18:01:54 UTC