2026/04/30 by Hannah Kockelbergh, Shelley Evans, Liam Brierley +4 · 1 voice
Biochemistry, Genetics and Molecular Biology · #Machine Learning in Bioinformatics #RNA and protein synthesis mechanisms #Metabolomics and Mass Spectrometry Studies
paper · doi:10.1371/journal.pcbi.1014211
Abstract Insights gained through interpretation of models trained on the T-cell receptor (TCR) repertoire contribute to advances in understanding of immune-mediated disease. This has the potential to improve diagnostic tests and treatments, particularly for autoimmune diseases. However, TCR repertoire datasets with samples from donors of known autoimmune disease status generally include orders of magnitude fewer samples than TCR sequences. Promising TCR repertoire classification approaches consider relationships between non-identical TCR sequences. In particular, kmer methods demonstrate strong and stable performance for small datasets. We propose a TCR repertoire representation that considers the relationships between amino acids within kmers flexibly and efficiently, which makes exploration of a wide range of TCR sequence features feasible. XGBoost models are trained and tested on kmer representations of TCR repertoire datasets including samples from patients with coeliac disease as well as donors with previous cytomegalovirus infection. We show that kmers that use small representative alphabets of amino acids are capable of training models that perform similarly or better than kmers based on all 20 amino acids. We find that, for cytomegalovirus infection status classification, defining amino acid relationships using BLOSUM62 can lead to a model with stronger performance as compared to an Atchley factor definition. Finally, we detail kmers or motifs which are important in each classification model and highlight the challenge of training truly interpretable TCR repertoire classification models which, if overcome, could lead to biomarker discovery for autoimmune diseases. Author summary TCR repertoire classification models can provide valuable understanding of autoimmune diseases if they can accurately infer autoimmune disease status and are biologically interpretable. Based on a kmer representation of the TCR repertoire, which has been shown to be most appropriate to train classification models on smaller datasets, we develop a computationally efficient method of grouping amino acid sequences to add knowledge to immune status classification model inputs, and consider its effect on interpretability. We find that most of the 4mer-based feature types we tested perform well in combination with an XGBoost model, where some benefit may be gained by applying a greatly-reduced alphabet of amino acids based on BLOSUM62 for cytomegalovirus serostatus classification. Our proposed reduced alphabet methodology is an alternative to kmer clustering which allows more efficient exploration of amino acid relationships and results in a more interpretable feature space.