2026/07/26 by Or Malca, Alona Zilberberg, Sol Efroni
Immunology and Microbiology · Biochemistry, Genetics and Molecular Biology · #T-cell and B-cell Immunology #vaccines and immunoinformatics approaches #Immune Cell Function and Interaction
paper · doi:10.1111/imr.70145
ABSTRACT The T‐cell receptor (TCR) repertoire records an individual's immunological history, but most unique CDR3 sequences in any one person are private and uninformative about anyone else. A small subset, however, recurs predictably across unrelated donors. These public TCRβ sequences are independently generated by multiple mechanisms: enrichment by thymic positive selection on a largely shared self‐peptide–MHC ligandome, further amplified by selection on common foreign antigens and convergent recombination, together, they are what makes a personal repertoire computationally legible: they provide the shared coordinate system on which otherwise incommensurable repertoires can be aligned and compared. This review takes public TCR sequences as its protagonist. We trace the biology of TCR publicity through new measurements on a 1.5‐billion‐sequence meta‐repertoire (7943 samples, 41 studies) that quantify five interrelated properties: the universe is finite, with Chao2‐bounded ceilings (i.e., a lower estimate) of ≈1.97 × 10 9 amino acid and ≈8.25 × 10 9 nucleotide CDR3β sequences of which ≈22.5% and ≈12% have already been observed; the rarefaction curve already bends within [0, 7943] by a factor of ×3.3 (amino acid) and ×1.6 (nucleotide) relative to a linear‐at‐initial‐rate extrapolation; publicness is a cohort‐scale statistic, with the heavy tail of the publicness distribution growing predictably from N = 200 to N = 7943; super‐public sequences are ≈0.014% of unique amino acid CDR3s but carry 8.45% of the total observation mass; and the recapture rate of a CDR3 already observed in another donor reaches 73.3% (amino acid) and 36.2% (nucleotide) at N ≈ 7900. We then trace the family of computational methods that exploit public sequences as features: frequency vectors, sequence‐similarity networks, self‐supervised transformer embeddings (CVC, Beaker, TCR‐BERT, SCEPTR), and our own anchor‐based graph neural networks (GraffiTee). We summarize their performance across cancer detection, autoimmunity, infectious disease, immunotherapy monitoring, and immunological aging, alongside the structural confounders (HLA, age, sex, sequencing platform) that bound generalizability. We connect repertoire‐level classification to TCR–pMHC binding prediction and antigen identity, and ask what stands between current capabilities and a universal TCR‐based diagnostic.