2024/06/10 by Yuta Nagano, Nagano, Yuta, Andrew Pyo +11 · 1 citation
Biochemistry, Genetics and Molecular Biology · Immunology and Microbiology · Medicine · #Artificial Intelligence (cs.AI) #Biomolecules (q-bio.BM) #FOS: Biological sciences #FOS: Computer and information sciences #I.2.7 #J.3 #Machine Learning (cs.LG) #Monoclonal and Polyclonal Antibodies Research #T-cell and B-cell Immunology #vaccines and immunoinformatics approaches
paper · pdf · doi:10.48550/arxiv.2406.06397
openalex publication_date 2024/06/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Computational prediction of the interaction of T cell receptors (TCRs) and their ligands is a grand challenge in immunology. Despite advances in high-throughput assays, specificity-labelled TCR data remains sparse. In other domains, the pre-training of language models on unlabelled data has been successfully used to address data bottlenecks. However, it is unclear how to best pre-train protein language models for TCR specificity prediction. Here we introduce a TCR language model called SCEPTR (Simple Contrastive Embedding of the Primary sequence of T cell Receptors), capable of data-efficient transfer learning. Through our model, we introduce a novel pre-training strategy combining autocontrastive learning and masked-language modelling, which enables SCEPTR to achieve its state-of-the-art performance. In contrast, existing protein language models and a variant of SCEPTR pre-trained without autocontrastive learning are outperformed by sequence alignment-based methods. We anticipate that contrastive learning will be a useful paradigm to decode the rules of TCR specificity.