Mechanistic Interpretability of Fine-Tuned Protein Language Models for Nanobody Thermostability Prediction.
Fine-tuned protein language models can be interpreted to reveal biophysical principles of nanobody thermostability, with sparse autoencoders extracting meaningful features from embeddings.
- Why it matters: Understanding the physical basis of model predictions is crucial for advancing protein engineering and discovering new stability determinants, addressing the opacity of deep learning models.
- What they did: Researchers fine-tuned the ESM-2 model on nanobody thermostability data and used sparse autoencoders to decompose embeddings into interpretable features, analyzing residue-level patterns and known stabilizing elements.
- The result: This approach identified structural and residue-specific insights, including stabilizing residues and disulfide bonds, enabling the generation of testable hypotheses for rational protein design.