Matryoshka Representation Learning
2022/05/26 by Aditya Kusupati, Gantavya Bhatt, Kusupati, Aditya +20 · 4 voices · 89 citations
Computer Science · Medicine · #Artificial intelligence #COVID-19 diagnosis using AI #Computer science #Context (archaeology) #Domain Adaptation and Few-Shot Learning #Downstream (manufacturing) #Embedding #Flexibility (engineering) #Inference #Machine learning #Multimodal Machine Learning Applications #Representation (politics) #Software deployment #Software engineering #Task (project management) #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2205.13147
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/05/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Learned representations are a central component in modern ML systems, serving a multitude of downstream tasks. When training such representations, it is often the case that computational and statistical constraints for each downstream task are unknown. In this context rigid, fixed capacity representations can be either over or under-accommodating to the task at hand. This leads us to ask: can we design a flexible representation that can adapt to multiple downstream tasks with varying computational resources? Our main contribution is Matryoshka Representation Learning (MRL) which encodes information at different granularities and allows a single embedding to adapt to the computational constraints of downstream tasks. MRL minimally modifies existing representation learning pipelines and imposes no additional cost during inference and deployment. MRL learns coarse-to-fine representations that are at least as accurate and rich as independently trained low-dimensional representations. The flexibility within the learned Matryoshka Representations offer: (a) up to 14x smaller embedding size for ImageNet-1K classification at the same level of accuracy; (b) up to 14x real-world speed-ups for large-scale retrieval on ImageNet-1K and 4K; and (c) up to 2% accuracy improvements for long-tail few-shot classification, all while being as robust as the original representations. Finally, we show that MRL extends seamlessly to web-scale datasets (ImageNet, JFT) across various modalities -- vision (ViT, ResNet), vision + language (ALIGN) and language (BERT). MRL code and pretrained models are open-sourced at https://github.com/RAIVNLab/MRL.
Cited by
- Ordered Action Tokens for Visuomotor Policy Learning
- Eddy-VL 1.9B: Structural Pruning and Layered Distillation for Edge-Deployable Multimodal Embedding
- Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
- On the Theoretical Limitations of Embedding-Based Retrieval
- Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Matryoshka Quantization
- Theoretical Foundations of Scaling Law in Familial Models
- The Hitchhiker's Guide to Monoculture
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- Controlling Embedding Spaces with Text-Conditioned Transformations
- Self-attention vector output similarities reveal how machines pay attention
- C2LLM Technical Report: A New Frontier in Code Retrieval via Adaptive Cross-Attention Pooling
- No Data? No Problem: Robust Vision-Tabular Learning with Missing Values
- Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- M3DR: Towards Universal Multilingual Multimodal Document Retrieval
- Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings
- Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency
- Text-Guided Semantic Image Encoder
- Towards Hyper-Efficient RAG Systems in VecDBs: Distributed Parallel Multi-Resolution Vector Search
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
- Route Experts by Sequence, not by Token
- Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
- Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
- Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
- Eigenfunction Extraction for Ordered Representation Learning
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- No Mean Feat: Simple, Strong Baselines for Context Compression
- CoRECT: A Framework for Evaluating Embedding Compression Techniques at Scale
- Elastic ViTs from Pretrained Models without Retraining
- Dimension Mask Layer: Optimizing Embedding Efficiency for Scalable ID-based Models
- When Embedding Models Meet: Procrustes Bounds and Applications
- ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Compressed Concatenation of Small Embedding Models
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
- Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
- RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
- Semantic Compression via Multimodal Representation Learning
- MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
- EmbeddingGemma: Powerful and Lightweight Text Representations
- Kairos: Numerically Robust News Recommendation under Item Cold-Start via Cholesky-based LinUCB
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- DS@GT AnimalCLEF: Triplet Learning over ViT Manifolds with Nearest Neighbor Classification for Animal Re-identification
- FinGEAR: Financial Mapping-Guided Enhanced Answer Retrieval
- Ontology-Aligned Embeddings for Data-Driven Labour Market Analytics
- Post-training Large Language Models for Diverse High-Quality Responses
- Towards Open World Detection: A Survey
- Efficient Code Embeddings from Code Generation Models
- DiskJoin: Large-scale Vector Similarity Join with SSD
- Hierarchical Adaptive networks with Task vectors for Test-Time Adaptation
- Uncertainty-driven Embedding Convolution
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
- Single-pass Adaptive Image Tokenization for Minimum Program Search
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model
- NEAR2: A Nested Embedding Approach to Efficient Product Retrieval and Ranking
- jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
- HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition
- SweRank: Software Issue Localization with Code Ranking
- FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
- Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
- Is Hyperbolic Space All You Need for Medical Anomaly Detection?
- Residual Diffusion Models for Variable-Rate Joint Source Channel Coding of MIMO CSI
- DB-KSVD: Scalable Alternating Optimization for Disentangling High-Dimensional Embedding Spaces
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
- Exploring The Visual Feature Space for Multimodal Neural Decoding
- Table Foundation Models: on knowledge pre-training for tabular learning
- Shielding Latent Face Representations From Privacy Attacks
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
- Chain-of-Model Learning for Language Model
- Telco-oRAG: Optimizing Retrieval-augmented Generation for Telecom Queries via Hybrid Retrieval and Neural Routing
- Scalable Unit Harmonization in Medical Informatics via Bayesian-Optimized Retrieval and Transformer-Based Re-ranking
- Reverse Distillation: Consistently Scaling Protein Language Model Representations
- Learning Retrieval Models with Sparse Autoencoders
- Optimization of embeddings storage for RAG systems using quantization and dimensionality reduction techniques
- Locating acts of mechanistic reasoning in student team conversations with mechanistic machine learning
- LCSHBench: A Multilingual, Consensus-Grounded Benchmark for Library of Congress Subject Heading Assignment
- flexvec: SQL Vector Retrieval with Programmatic Embedding Modulation
- AdaVid: Adaptive Video-Language Pretraining
- One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding
- TESSERA v2: Scaling Pixel-wise Earth Foundation Models
- Hypergraph Vision Transformers: Images are More than Nodes, More than Edges
Discussions
Related