CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction

#1 CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction [PDF] [Copy] [Kimi] [REL]

Authors: Jaewan Lee, Changyoung Park, Hongjun Yang, Sungbin Lim, Woohyung Lim, Sehui Han

Recent advancements in graph neural networks (GNNs) have significantly enhanced the prediction of material properties by modeling crystal structures as graphs. However, GNNs often struggle to capture global structural characteristics, such as crystal systems, limiting their predictive performance. To overcome this issue, we propose CAST, a cross-attention-based multimodal model that integrates graph representations with textual descriptions of materials, effectively preserving critical structural and compositional information. Unlike previous approaches, such as CrysMMNet and MultiMat, which rely on aggregated material-level embeddings, CAST leverages cross-attention mechanisms to combine fine-grained graph node-level and text token-level features. Additionally, we introduce a masked node prediction pretraining strategy that further enhances the alignment between node and text embeddings. Our experimental results demonstrate that CAST outperforms existing baseline models across four key material properties-formation energy, band gap, bulk modulus, and shear modulus-with average relative MAE improvements ranging from 10.2% to 35.7%. Analysis of attention maps confirms the importance of pretraining in effectively aligning multimodal representations. This study underscores the potential of multimodal learning frameworks for developing more accurate and globally informed predictive models in materials science.

Subjects: Machine Learning , Materials Science , Artificial Intelligence

Publish: 2025-02-06 02:29:39 UTC

2502.06836

#1 CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction [PDF] [Copy] [Kimi] [REL]