ST-ConMa: a multimodal foundation framework for spatial transcriptomics via image-gene contrastive and matching learning.
ST-ConMa, a multimodal foundation model for spatial transcriptomics, outperforms existing methods by learning high-quality image-gene representations across diverse data.
- Why it matters: Current models are limited by task-specific design or low-resolution data, restricting their ability to generalize across different spatial transcriptomics platforms and applications.
- What they did: The authors developed ST-ConMa by pretraining on large-scale tissue images and gene profiles using contrastive and matching learning, employing a multimodal encoder with cross-attention.
- The result: ST-ConMa demonstrates superior performance on histopathology classification, gene expression prediction, and spatial clustering, establishing a versatile backbone for broad spatial transcriptomics analysis.