2024-10-29 | | Total: 4
Single-cell spatial transcriptomics (ST) offers a unique approach to measuring gene expression profiles and spatial cell locations simultaneously. However, most existing ST methods assume that cells in closer spatial proximity exhibit more similar gene expression patterns. Such assumption typically results in graph structures that prioritize local spatial information while overlooking global patterns, limiting the ability to fully capture the broader structural features of biological tissues. To overcome this limitation, we propose GATES (Graph Attention neTwork with global Expression fuSion), a novel model designed to capture structural details in spatial transcriptomics data. GATES first constructs an expression graph that integrates local and global information by leveraging both spatial proximity and gene expression similarity. The model then employs an autoencoder with adaptive attention to assign proper weights for neighboring nodes, enhancing its capability of feature extraction. By fusing features of both the spatial and expression graphs, GATES effectively balances spatial context with gene expression data. Experimental results across multiple datasets demonstrate that GATES significantly outperforms existing methods in identifying spatial domains, highlighting its potential for analyzing complex biological tissues. Our code can be accessed on GitHub at https://github.com/xiaoxiongtao/GATES.
In this paper, we introduce flubbles, a new definition of "bubbles" corresponding to variants in a (pan)genome graph $G$. We then show a characterization for flubbles in terms of equivalence classes regarding cycles in an intermediate data structure we built from the spanning tree of the $G$, which leads us to a linear time and space solution for finding all flubbles. Furthermore, we show how a related characterization also allows us to efficiently detect what we define as hairpin inversions: a cycle preceded and followed by the same path in the graph; being the latter necessarily traversed both ways, this structure corresponds to inversions. Finally, Inspired by the concept of Program Structure Tree introduced fifty years ago to represent the hierarchy of the control structure of a program, we define a tree representing the structure of G in terms of flubbles, the flubble tree, which we also find in linear time. The hierarchy of variants introduced by the flubble tree paves the way for new investigations of (pan)genomic structures and their decomposition for practical analyses. We have implemented our methods into a prototype tool named povu which we tested on human and yeast data. We show that povu can find flubbles and also output the flubble tree while being as fast (or faster than) well established tools that find bubbles, such as vg and BubbleGun. Moreover, we show how, within the same time, povu can find hairpin inversions that, to the best of our knowledge, no other tool is able to find. Our tool is freely available at https://github.com/urbanslug/povu/ under the MIT License.
Nanopore sequencing, a next-generation sequencing technology, holds the potential to revolutionize multiple facets of life sciences, forensics, and healthcare. While previous research has focused on its technical intricacies and biomedical applications, this paper offers a unique perspective by scrutinizing the societal dimensions (ethical, legal, and social implications) of nanopore sequencing. Employing the lenses of Diffusion and Action Network Theory, we examine the dissemination of nanopore sequencing in society as a potential consumer product, contributing to the field of the sociology of technology. We investigate the possibility of interactions between human and nonhuman actors in developing nanopore technology to analyse how various stakeholders, such as companies, regulators, and researchers, shape the trajectory of the growth of nanopore sequencing. This work offers insights into the social construction of nanopore sequencing, shedding light on the actors, power dynamics, and socio-technical networks that shape its adoption and societal impact. Understanding the sociological dimensions of this transformative technology is vital for responsible development, equitable distribution, and inclusive integration into diverse societal contexts.
This study introduces a compositional autoencoder (CAE) framework designed to disentangle the complex interplay between genotypic and environmental factors in high-dimensional phenotype data to improve trait prediction in plant breeding and genetics programs. Traditional predictive methods, which use compact representations of high-dimensional data through handcrafted features or latent features like PCA or more recently autoencoders, do not separate genotype-specific and environment-specific factors. We hypothesize that disentangling these features into genotype-specific and environment-specific components can enhance predictive models. To test this, we developed a compositional autoencoder (CAE) that decomposes high-dimensional data into distinct genotype-specific and environment-specific latent features. Our CAE framework employs a hierarchical architecture within an autoencoder to effectively separate these entangled latent features. Applied to a maize diversity panel dataset, the CAE demonstrates superior modeling of environmental influences and 5-10 times improved predictive performance for key traits like Days to Pollen and Yield, compared to the traditional methods, including standard autoencoders, PCA with regression, and Partial Least Squares Regression (PLSR). By disentangling latent features, the CAE provides powerful tool for precision breeding and genetic research. This work significantly enhances trait prediction models, advancing agricultural and biological sciences.