Interpreting embeddings from genome and protein language models.
Embeddings from genome and protein language models enable high-resolution, alignment-free insights into genomics and proteomics, surpassing traditional methods.
- Why it matters: Understanding the functional and structural aspects of DNA and proteins is crucial for advances in genomics, but current methods often rely on alignment-based techniques that can be limited in resolution and scalability.
- What they did: Researchers analyzed embeddings from large transformer-based language models trained on DNA and protein sequences, applying them to predict variant severity, splicing boundaries, homology, structure, and function across thousands of sequences.
- The result: The findings show that these embeddings offer a powerful, high-resolution framework for biological interpretation, facilitating enzyme engineering and functional genomics without the need for sequence alignment.