Automated generation of a gene perturbation transcriptomic atlas using large language models
- Open access
Automated pipeline using large language models identified 6,802 gene perturbation signatures from 4,453 GEO experiments, covering 2,907 genes with high accuracy.
- Why it matters: Structured, comprehensive perturbation data is essential for understanding gene functions, but current repositories lack organization, limiting reuse and analysis.
- What they did: The approach involved developing an automated method to find single-gene perturbation experiments, reconstruct sample groupings, and normalize metadata, supported by manual curation of 3,300 experiments.
- The result: The pipeline achieved high precision (0.925) and recall (0.836), enabling creation of an accessible atlas and a tool, perturbMatch, for exploring and querying gene expression signatures.