DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
2019/10/02 by Victor Sanh, Sanh, Victor, Lysandre Debut +5 · 1 voice · 708 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
paper · pdf · doi:10.48550/arxiv.1910.01108
February 2020 - Revision: fix bug in evaluation metrics, updated metrics, argumentation unchanged. 5 pages, 1 figure, 4 tables. Accepted at the 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS 2019
arxiv published 2019/10/02 · arxiv created 2020/03/01 · arxiv updated 2020/03/03
Abstract
As Transfer Learning from large-scale pre-trained models becomes more prevalent in Natural Language Processing (NLP), operating these large models in on-the-edge and/or under constrained computational training or inference budgets remains challenging. In this work, we propose a method to pre-train a smaller general-purpose language representation model, called DistilBERT, which can then be fine-tuned with good performances on a wide range of tasks like its larger counterparts. While most prior work investigated the use of distillation for building task-specific models, we leverage knowledge distillation during the pre-training phase and show that it is possible to reduce the size of a BERT model by 40%, while retaining 97% of its language understanding capabilities and being 60% faster. To leverage the inductive biases learned by larger models during pre-training, we introduce a triple loss combining language modeling, distillation and cosine-distance losses. Our smaller, faster and lighter model is cheaper to pre-train and we demonstrate its capabilities for on-device computations in a proof-of-concept experiment and a comparative on-device study.
Cited by
- Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects?
- Latent Multi-Head Attention for Small Language Models
- GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- Masked Distillation: Internalizing the Chain-of-Thought in Language Models
- TimeBill: Time-Budgeted Inference for Large Language Models
- Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
- SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
- Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
- Agentic Multi-Persona Framework for Evidence-Aware Fake News Detection
- Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
- USE: A Unified Model for Universal Sound Separation and Extraction
- Block-Recurrent Dynamics in Vision Transformers
- LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning
- Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering
- Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
- LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
- Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs
- Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
- DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- Examining the Utility of Self-disclosure Types for Modeling Annotators of Social Norms
- TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- Improving Recursive Transformers with Mixture of LoRAs
- Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- Task-Specific Sparse Feature Masks for Molecular Toxicity Prediction with Chemical Language Models
- Parallax: Runtime Parallelization for Operator Fallbacks in Heterogeneous Edge Systems
- CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation
- Provably Learning from Modern Language Models via Low Logit Rank
- SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
- Interpreto: An Explainability Library for Transformers
- Text2Graph: Combining Lightweight LLMs and GNNs for Efficient Text Classification in Label-Scarce Scenarios
- Belief Is All You Need: Modeling Narrative Archetypes in Conspiratorial Discourse
- Language-Conditioned Safe Trajectory Generation for Spacecraft Rendezvous
- Financial News Summarization: Can extractive methods still offer a true alternative to LLMs?
- Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
- MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
- A Patient-Doctor-NLP-System to contest inequality for less privileged
- Unveiling Hedge Funds: Topic Modeling and Sentiment Correlation with Fund Performance
- Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
- Optimizing LLMs Using Quantization for Mobile Execution
- Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- Fine-Tuning BERT for Domain-Specific Question Answering: Toward Educational NLP Resources at University Scale
- LLMs Know More Than Words: A Genre Study with Syntax, Metaphor & Phonetics
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- Fine-grained Narrative Classification in Biased News Articles
- Network of Theseus (like the ship)
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Bangla Hate Speech Classification with Fine-tuned Transformer Models
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- RoleMotion: A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- No Trust Issues Here: A Technical Report on the Winning Solutions for the Rayan AI Contest
- DyFuLM: An Advanced Multimodal Framework for Sentiment Analysis
- Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- Tree Matching Networks for Natural Language Inference: Parameter-Efficient Semantic Understanding via Dependency Parse Trees
- Retrieval-Augmented Few-Shot Prompting Versus Fine-Tuning for Code Vulnerability Detection
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Context-Aware Detection and Victim-Centered Response Generation for Online Harassment in Private Messaging
- A Trainable Centrality Framework for Modern Data
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- Using Text-Based Life Trajectories from Swedish Register Data to Predict Residential Mobility with Pretrained Transformers
- MorphingDB: A Task-Centric AI-Native DBMS for Model Management and Inference
- Enhancing Burmese News Classification with Kolmogorov-Arnold Network Head Fine-tuning
- Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
- BRIC: Bridging Kinematic Plans and Physical Control at Test Time
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- Deterministic Continuous Replacement: Fast and Stable Module Replacement in Pretrained Transformers
- Towards Characterizing Knowledge Distillation of PPG Heart Rate Estimation Models
- A symbolic Perl algorithm for the unification of Nahuatl word spellings
- A Systematic Study of Compression Ordering for Large Language Models
- Efficient Covariance Estimation for Sparsified Functional Data
- MTikGuard System: A Transformer-Based Multimodal System for Child-Safe Content Moderation on TikTok
- When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
- NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
- GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs
- RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue
- Gradient-descent methods for scalable quantum detector tomography
- Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
- Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system
- Decoupled Action Expert: Confining Task Knowledge to the Conditioning Pathway
- From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
- Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
- BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
- BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives
- CC30k: A Citation Contexts Dataset for Reproducibility-Oriented Sentiment Analysis
- Graph Representation-based Model Poisoning on the Heterogeneous Internet of Agents
- Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
- Beyond Uniform Deletion: A Data Value-Weighted Framework for Certified Machine Unlearning
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
- One-Shot Knowledge Transfer for Scalable Person Re-Identification
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- Quantifying the Climate Risk of Generative AI: Region-Aware Carbon Accounting with G-TRACE and the AI Sustainability Pyramid
- Advancing Equitable AI: Evaluating Cultural Expressiveness in LLMs for Latin American Contexts
- SOLVE-Med: Specialized Orchestration for Leading Vertical Experts across Medical Specialties
- Laugh, Relate, Engage: Stylized Comment Generation for Short Videos
- The Curved Spacetime of Transformer Architectures
- Targeted Error Correction in Knowledge Distillation: Small Language Models Surpass GPT
- DEEP: A Discourse Evolution Engine for Predictions about Social Movements
- Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
- Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction
- Sherlock: Reliable and Efficient Agentic Workflow Execution
- Exploring the Utilities of the Rationales from Large Language Models to Enhance Automated Essay Scoring
- Elastic Architecture Search for Efficient Language Models
- Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
- FakeZero: Real-Time, Privacy-Preserving Misinformation Detection for Facebook and X
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text
- Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case
- How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Exploring the Intersection of AI, Language, and Law: A Bibliometric Analysis
- DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
- Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
- Stay Tuned: Improving Sentiment Analysis and Stance Detection Using Large Language Models
- A machine learning-based evidence map of ocean-related options for climate change mitigation and adaptation
- MemEIC: A Step Toward Continual and Compositional Knowledge Editing
- TOPol: Capturing and Explaining Multidimensional Semantic Polarity Fields and Vectors
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Code Contribution and Credit in Science
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
- In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit Discussions
- FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
- The Lossy Horizon: Error-Bounded Predictive Coding for Lossy Text Compression (Episode I)
- Power to the Clients: Federated Learning in a Dictatorship Setting
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- Doubly-Regressing Approach for Subgroup Fairness
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- GRATING: Low-Latency and Memory-Efficient Semantic Selection on Device
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
- On Optimal Hyperparameters for Differentially Private Deep Transfer Learning
- Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent
- Calibration and Discrimination Optimization Using Clusters of Learned Representation
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder
- Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
- Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
- Sync or Sink: Bounds on Algorithmic Collective Action with Noise and Multiple Groups
- Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
- Quantifying Climate Policy Action and Its Links to Development Outcomes: A Cross-National Data-Driven Analysis
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Refugees of the Digital Space: Platform Migration from TikTok to RedNote
- MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
- Generalizing WiFi Gesture Recognition via Large-Model-Aware Semantic Distillation and Alignment
- ProtoTopic: Prototypical Network for Few-Shot Medical Topic Modeling
- Tandem Training for Language Models
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- Efficient Adaptive Transformer: An Empirical Study and Reproducible Framework
- Traveling Salesman-Based Token Ordering Improves Stability in Homomorphically Encrypted Language Models
- FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning
- A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
- Fairness Metric Design Exploration in Multi-Domain Moral Sentiment Classification using Transformer-Based Models
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Test-Time Adaptation by Causal Trimming
- From Reasoning LLMs to BERT: A Two-Stage Distillation Framework for Search Relevance
- Therapeutic AI and the Hidden Risks of Over-Disclosure: An Embedded AI-Literacy Framework for Mental Health Privacy
- Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default
- Serialized EHR make for good text representations
- Inverse Language Modeling towards Robust and Grounded LLMs
- On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
- Framing Unionization on Facebook: Communication around Representation Elections in the United States
- Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- Theoretical guarantees for change localization using conformal p-values
- STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
- AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
- Causality Guided Representation Learning for Cross-Style Hate Speech Detection
- Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
- Investigating Counterclaims in Causality Extraction from Text
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- Reasoning for Hierarchical Text Classification: The Case of Patents
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- GUIDE: Guided Initialization and Distillation of Embeddings
- LLM Bias Detection and Mitigation through the Lens of Desired Distributions
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
- Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
- Annotate Rhetorical Relations with INCEpTION: A Comparison with Automatic Approaches
- Consistent Kernel Change-Point Detection under m-Dependence for Text Segmentation
- MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
- Neural Correlates of Language Models Are Specific to Human Language
- Dissecting Transformers: A CLEAR Perspective towards Green AI
- Syntax-Guided Diffusion Language Models with User-Integrated Personalization
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
- Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
- Fine-tuning with RAG for Improving LLM Learning of New Skills
- Robust Federated Inference
- Quantifying Semantic Shift in Financial NLP: Robust Metrics for Market Prediction Stability
- Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
- Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests
- Vocabulary Customization for Efficient Domain-Specific LLM Deployment
- FITS: Towards an AI-Driven Fashion Information Tool for Sustainability
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- Can Large Language Models Express Uncertainty Like Human?
- The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact
- Why Alignment Must Precede Distillation: A Minimal Working Explanation
- Temporal Generalization: A Reality Check
- Knowledge distillation through geometry-aware representational alignment
- Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Text2Move: Text-to-moving sound generation via trajectory prediction and temporal alignment
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- An Improved Quantum Software Challenges Classification Approach using Transfer Learning and Explainable AI
- A short survey on almost orthogonal vectors in a few specific large dimensions
- CTI Dataset Construction from Telegram
- PseudoBridge: Pseudo Code as the Bridge for Better Semantic and Logic Alignment in Code Retrieval
- Overcoming Black-box Attack Inefficiency with Hybrid and Dynamic Select Algorithms
- RedHerring Attack: Testing the Reliability of Attack Detection
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI
- Confidence Calibration in Large Language Model-Based Entity Matching
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
- AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
- MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
- NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
- LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation
- A transformer-based multi-feature fusion method for detecting traffic events using Twitter data
- Otters: An Energy-Efficient SpikingTransformer via Optical Time-to-First-Spike Encoding
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- Semantic Search for Information Retrieval
- Enhancing Transformer-Based Rerankers with Synthetic Data and LLM-Based Supervision
- AIRwaves at CheckThat! 2025: Retrieving Scientific Sources for Implicit Claims on Social Media with Dual Encoders and Neural Re-Ranking
- Identifying Constructive Conflict in Online Discussions through Controversial yet Toxicity Resilient Posts
- Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
- Fine-Grained Detection of AI-Generated Text Using Sentence-Level Segmentation
- Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
- Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
- DRES: Fake news detection by dynamic representation and ensemble selection
- Random Direct Preference Optimization for Radiography Report Generation
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- Assessing Classical Machine Learning and Transformer-based Approaches for Detecting AI-Generated Research Text
- CausalSent: Interpretable Sentiment Classification with RieszNet
- RMT-KD: Random Matrix Theoretic Causal Knowledge Distillation
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- Bridging Graph and State-Space Modeling for Intensive Care Unit Length of Stay Prediction
- Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- SINAI at eRisk@CLEF 2022: Approaching Early Detection of Gambling and Eating Disorders with Natural Language Processing
- Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification
- Who Wins the Race? (R Vs Python) - An Exploratory Study on Energy Consumption of Machine Learning Algorithms
- LLM on a Budget: Active Knowledge Distillation for Efficient Classification of Large Text Corpora
- The Few-shot Dilemma: Over-prompting Large Language Models
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Exploring Training Data Attribution under Limited Access Constraints
- Spatio-Temporal Pruning for Compressed Spiking Large Language Models
- Image-Seeking Intent Prediction for Cross-Device Product Search
- User eXperience Perception Insights Dataset (UXPID): Synthetic User Feedback from Public Industrial Forums
- Cross-Modal Retrieval with Cauchy-Schwarz Divergence
- Dynamic Span Interaction and Graph-Aware Memory for Entity-Level Sentiment Classification
- LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries
- A Transformer-Based Cross-Platform Analysis of Public Discourse on the 15-Minute City Paradigm
- Efficient Hate Speech Detection: Evaluating 38 Models from Traditional Methods to Transformers
- Quantifier Scope Interpretation in Language Learners and LLMs
- Remotely Seeing Is Believing: How Trust in Cyber-Physical Systems Evolves Through Virtual Observation
- VARCO-VISION-2.0 Technical Report
- SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- Green Federated Learning via Carbon-Aware Client and Time Slot Scheduling
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
- D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Few-Shot Query Intent Detection via Relation-Aware Prompt Learning
- Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
- RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models
- Anti-establishment sentiment on TikTok: Implications for understanding influence(rs) and expertise on social media
- LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
- NoteBar: An AI-Assisted Note-Taking System for Personal Knowledge Management
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- From Confidence to Collapse in LLM Factual Robustness
- StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
- Hierarchical Motion Captioning Utilizing External Text Data Source
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA
- FLEET: A Federated Learning Emulation and Evaluation Testbed for Holistic Research
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
- Text-Driven 3D Hand Motion Generation from Sign Language Data
- MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- Speech Emotion Recognition via Entropy-Aware Score Selection
- MathBuddy: A Multimodal System for Affective Math Tutoring
- AI-Powered Detection of Inappropriate Language in Medical School Curricula
- QuesGenie: Intelligent Multimodal Question Generation
- ALSA: Anchors in Logit Space for Out-of-Distribution Accuracy Estimation
- Skill-based Explanations for Serendipitous Course Recommendation
- SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- Compressed Models are NOT Trust-equivalent to Their Large Counterparts
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- A Risk Manager for Intrusion Tolerant Systems: Enhancing HAL 9000 with New Scoring and Data Sources
- REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks
- Checkmate: interpretable and explainable RSVQA is the endgame
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- A Global Dataset of Location Data Integrity-Assessed Reforestation Efforts
- Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- MCP-Orchestrated Multi-Agent System for Automated Disinformation Detection
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
- Bhav-Net: Knowledge Transfer for Cross-Lingual Antonym vs Synonym Distinction via Dual-Space Graph Transformers
- Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
- Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
- LLMs for Law: Evaluating Legal-Specific LLMs on Contract Understanding
- Modeling and Detecting Company Risks from News: A Case Study in Bloomberg News
- TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
- LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
- Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
- Embedding Alignment in Code Generation for Audio
- Task complexity shapes internal representations and robustness in neural networks
- PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
- Leveraging large language models for SQL behavior-based database intrusion detection
- Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
- WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval
- HiTeC: Hierarchical Contrastive Learning on Text-Attributed Hypergraph with Semantic-Aware Augmentation
- SLIM-LLMs: Modeling of Style-Sensory Language RelationshipsThrough Low-Dimensional Representations
- Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
- "Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth
- Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
- Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
- Distillation-Enhanced Clustering Acceleration for Encrypted Traffic Classification
- Contextual Graph Transformer: A Small Language Model for Enhanced Engineering Document Information Extraction
- EHSAN: Leveraging ChatGPT in a Hybrid Framework for Arabic Aspect-Based Sentiment Analysis in Healthcare
- ReMoMask: Retrieval-Augmented Masked Motion Generation
- Remembering without (representational) memory: a neuro-computational study on regaining categoricity and compositionality from minimal traces
- Empowering Tabular Data Preparation with Language Models: Why and How?
- R2-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
- DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
- Classification of Psychiatry Clinical Notes by Diagnosis: A Deep Learning and Machine Learning Approach
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
- Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
- Operationalizing AI for Good: Spotlight on Deployment and Integration of AI Models in Humanitarian Work
- Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- Real-time News Story Identification
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- LLM-based Content Classification Approach for GitHub Repositories by the README Files
- Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions
- Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
- Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
- Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
- Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
- A Similarity Measure for Comparing Conversational Dynamics
- Efficient Agents: Building Effective Agents While Reducing Cost
- What does the public want their local government to hear? A data-driven case study of public comments across the state of Michigan
- OPRD: On-Policy Representation Distillation
- VMask: Tunable Label Privacy Protection for Vertical Federated Learning via Layer Masking
- VTarbel: Targeted Label Attack with Minimal Knowledge on Detector-enhanced Vertical Federated Learning
- Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
- GhostUMAP2: Measuring and Analyzing (r,d)-Stability of UMAP
- Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media
- A Survey of Deep Learning for Geometry Problem Solving
- OrdShap: Feature Position Importance for Sequential Black-Box Models
- Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses
- Cross-lingual Few-shot Learning for Persian Sentiment Analysis with Incremental Adaptation
- Evaluating gender bias in large language models in long-term care
- FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
- PlantDeBERTa: An Open Source Language Model for Plant Science
- MD-ViSCo: A Unified Model for Multi-Directional Vital Sign Waveform Conversion
- Seq vs Seq: An Open Suite of Paired Encoders and Decoders
- What Should Feature Distillation Transfer in LLMs? A Task-Tangent Geometry View
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
- Holistix: A Dataset for Holistic Wellness Dimensions Analysis in Mental Health Narratives
- Structure-Augmented Reasoning Generation
- L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training
- Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
- GNN-CNN: An Efficient Hybrid Model of Convolutional and Graph Neural Networks for Text Representation
- Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
- Improving Clustering on Occupational Text Data through Dimensionality Reduction
- Agentic-R1: Distilled Dual-Strategy Reasoning
- Taming Data Challenges in ML-based Security Tasks Using Generative AI
- Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation
- Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework
- Anomalous Decision Discovery using Inverse Reinforcement Learning
- GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
- Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM Inference
- NDAI-NeuroMAP: A Neuroscience-Specific Embedding Model for Domain-Specific Retrieval
- Assessing Small Language Models for Code Generation: An Empirical Study with Benchmarks
- QFFN-BERT: An Empirical Study of Depth, Performance, and Data Efficiency in Hybrid Quantum-Classical Transformers
- The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models
- POST: Photonic Swin Transformer for Automated and Efficient Prediction of PCSEL
- MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
- Efficient Out-of-Scope Detection in Dialogue Systems via Uncertainty-Driven LLM Routing
- Is Visual in-Context Learning for Compositional Medical Tasks within Reach?
- Matterhorn: Masked Time-to-First-Spike Encoding by Reassigning the Silent State for Sparse and Energy-Efficient Spiking Transformers
- Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
- MobileRAG: A Fast, Memory-Efficient, and Energy-Efficient Method for On-Device RAG
- MotionGPT3: Human Motion as a Second Modality
- Efficient Interleaved Speech Modeling through Knowledge Distillation
- NEU-ESC: A Comprehensive Vietnamese dataset for Educational Sentiment analysis and topic Classification toward multitask learning
- SoftStep: Learning Sparse Similarity Powers Deep Neighbor-Based Regression
- From Release to Adoption: Challenges in Reusing Pre-trained AI Models for Downstream Developers
- P2U: Progressive Precision Update For Efficient Model Distribution
- Assessing the feasibility of Large Language Models for detecting micro-behaviors in team interactions during space missions
- Detection of Personal Data in Structured Datasets Using a Large Language Model
- A Survey of LLM Inference Systems
- Early Stopping Tabular In-Context Learning
- skLEP: A Slovak General Language Understanding Benchmark
- Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
- Social Hatred: Efficient Multimodal Detection of Hatemongers
- Measuring and Guiding Monosemanticity
- Health Sentinel: An AI Pipeline For Real-time Disease Outbreak Detection
- PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications
- Can Argus Judge Them All? Comparing VLMs Across Domains
- Multimodal Medical Image Binding via Shared Text Embeddings
- Actionable Interpretability via Causal Hypergraphs: Unravelling Batch Size Effects in Deep Learning
- From Tiny Machine Learning to Tiny Deep Learning: A Survey
- LLM-driven Medical Report Generation via Communication-efficient Heterogeneous Federated Learning
- Optimal Depth of Neural Networks
- Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs
- A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
- Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues
- From Teacher to Student: Tracking Memorization Through Model Distillation
- I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
- Knowledge Distillation Framework for Accelerating High-Accuracy Neural Network-Based Molecular Dynamics Simulations
- HOIDiNi: Human-Object Interaction through Diffusion Noise Optimization
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
- Advancing Question Generation with Joint Narrative and Difficulty Control
- EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
- Theoretically Unmasking Inference Attacks Against LDP-Protected Clients in Federated Vision Models
- Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs
- Evaluating Generalization and Representation Stability in Small LMs via Prompting, Fine-Tuning and Out-of-Distribution Prompts
- TensorSLM: Energy-efficient Embedding Compression of Sub-billion Parameter Language Models on Low-end Devices
- SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
- TagRouter: Learning Route to LLMs through Tags for Open-Domain Text Generation Tasks
- Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth
- A Pluggable Multi-Task Learning Framework for Sentiment-Aware Financial Relation Extraction
- Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
- Curriculum-Guided Layer Scaling for Language Model Pretraining
- Efficient LLM Collaboration via Planning
- A Cramér-von Mises Approach to Incentivizing Truthful Data Sharing
- A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis
- LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
- RETUYT-INCO at BEA 2025 Shared Task: How Far Can Lightweight Models Go in AI-powered Tutor Evaluation?
- Causal Graph based Event Reasoning using Semantic Relation Experts
- Detecting High-Stakes Interactions with Activation Probes
- Complexity of normalized stochastic first-order methods with momentum under heavy-tailed noise
- Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
- FASCIST-O-METER: Classifier for Neo-fascist Discourse Online
- Reinforcement learning fine-tuning of language model for instruction following and math reasoning
- Masked Language Models are Good Heterogeneous Graph Generalizers
- A Culturally Rich Romanian NLP Dataset from 'Who Wants to Be a Millionaire?' Videos
- Mitigating Confounding in Speech-Based Dementia Detection through Weight Masking
- Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
- Building a Few-Shot Cross-Domain Multilingual NLU Model for Customer Care
- A Large Language Model for Feasible and Diverse Population Synthesis
- TransClean: Finding False Positives in Multi-Source Entity Matching under Real-World Conditions via Transitive Consistency
- Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
- You Only Train Once
- Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning
- Elevating Cyber Threat Intelligence against Disinformation Campaigns with LLM-based Concept Extraction and the FakeCTI Dataset
- In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration
- Stochastic Momentum Methods for Non-smooth Non-Convex Finite-Sum Coupled Compositional Optimization
- BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
- Abstract Counterfactuals for Language Model Agents
- daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
- Trajectory Prediction Meets Large Language Models: A Survey
- CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection
- KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
- Quantifying Misattribution Unfairness in Authorship Attribution
- SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
- Multi-Modal Dataset Distillation in the Wild
- The State of Large Language Models for African Languages: Progress and Challenges
- "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance Drift
- ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances
- Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
- GPR: Empowering Generation with Graph-Pretrained Retriever
- Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
- On Fairness of Task Arithmetic: The Role of Task Vectors
- ERU-KG: Efficient Reference-aligned Unsupervised Keyphrase Generation
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought
- Knowledge Distillation for Reservoir-based Classifier: Human Activity Recognition
- ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
- Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings
- Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments
- Self-supervised Learning Method Using Transformer for Multi-dimensional Sensor Data Processing
- Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs
- Research on Driving Scenario Technology Based on Multimodal Large Lauguage Model Optimization
- InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing
- Optimizing Data Augmentation through Bayesian Model Selection
- Knowledge Guided Encoder-Decoder Framework: Integrating Multiple Physical Models for Agricultural Ecosystem Modeling
- AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
- LLMPR: A Novel LLM-Driven Transfer Learning based Petition Ranking Model
- Explaining Large Language Models with gSMILE
- ReSCORE: Label-free Iterative Retriever Training for Multi-hop Question Answering with Relevance-Consistency Supervision
- MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
- Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
- Tensorization is a powerful but underexplored tool for compression and interpretability of neural networks
- SEMFED: Semantic-Aware Resource-Efficient Federated Learning for Heterogeneous NLP Tasks
- Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
- Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection
- Automatic Annotation of Ancient Greek Vowel Length
- Rethinking the Understanding Ability across LLMs through Mutual Information
- AI for Regulatory Affairs: Balancing Accuracy, Interpretability, and Computational Cost in Medical Device Classification
- Language Model Distillation: A Temporal Difference Imitation Learning Perspective
- CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning
- Locality-Sensitive Hashing for Efficient Hard Negative Sampling in Contrastive Learning
- VIBE: Vector Index Benchmark for Embeddings
- ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
- Discretization-free Multicalibration through Loss Minimization over Tree Ensembles
- FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
- On Multilingual Encoder Language Model Compression for Low-Resource Languages
- Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search
- Recursive Offloading for LLM Serving in Multi-tier Networks
- On the creation of narrow AI: hierarchy and nonlocality of neural network skills
- AdUE: Improving uncertainty estimation head for LoRA adapters in LLMs
- Visual Question Answering on Multiple Remote Sensing Image Modalities
- Know When to Abstain: Optimal Selective Classification with Likelihood Ratios
- Lethe: How Hard Is It to Forget? A Benchmark for Federated Unlearning in Medical Imaging
- Structured Agent Distillation for Large Language Model
- Saten: Sparse Augmented Tensor Networks for Post-Training Compression of Large Language Models
- InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models
- KO: Kinetics-inspired Neural Optimizer with PDE Simulation Approaches
- Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
- SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
- Latent Flow Transformer
- Deterministic Bounds and Random Estimates of Metric Tensors on Neuromanifolds
- DynaNoise: Dynamic Probabilistic Noise Injection for Defending Against Membership Inference Attacks
- HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
- SAFE: Improving LLM Systems using Sentence-Level In-generation Attribution
- Duluth at SemEval-2025 Task 7: TF-IDF with Optimized Vector Dimensions for Multilingual Fact-Checked Claim Retrieval
- Quantum Knowledge Distillation for Large Language Models
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- SecEmb: Sparsity-Aware Secure Federated Learning of On-Device Recommender System with Large Embedding
- Learning to Play Like Humans: A Framework for LLM Adaptation in Interactive Fiction Games
- MoL for LLMs: Dual-Loss Optimization to Enhance Domain Expertise While Preserving General Capabilities
- Equally Critical: Samples, Targets, and Their Mappings in Datasets
- On Membership Inference Attacks in Knowledge Distillation
- Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks
- Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
- Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation
- On the Interconnections of Calibration, Quantification, and Classifier Accuracy Prediction under Dataset Shift
- Can Large Language Models Correctly Interpret Equations with Errors?
- VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
- AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
- Demystifying AI Agents: The Final Generation of Intelligence
- Hierarchical Surgical Robot Transformer (SRT-H): Imitation Learning for Autonomous Surgery
- ADALog: Adaptive Unsupervised Anomaly detection in Logs with Self-attention Masked Language Model
- Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation
- VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing Prediction
- Ornithologist: Towards Trustworthy "Reasoning" about Central Bank Communications
- Progressive2: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression
- Text-driven Motion Generation: Overview, Challenges and Directions
- Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts
- Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
- Privacy-Preserving Real-Time Vietnamese-English Translation on iOS using Edge AI
- AI-Enabled Accurate Non-Invasive Assessment of Pulmonary Hypertension Progression via Multi-Modal Echocardiography
- KDH-MLTC: Knowledge Distillation for Healthcare Multi-Label Text Classification
- KDC-Diff: A Latent-Aware Diffusion Model with Knowledge Retention for Memory-Efficient Image Generation
- TokenProber: Jailbreaking Text-to-image Models via Fine-grained Word Impact Analysis
- Evaluating Financial Sentiment Analysis with Annotators Instruction Assisted Prompting: Enhancing Contextual Interpretation and Stock Prediction Accuracy
- Crowding Out The Noise: Algorithmic Collective Action Under Differential Privacy
- Cape: Context-Aware Prompt Perturbation Mechanism with Differential Privacy
- A Comprehensive Analysis of Adversarial Attacks against Spam Filters
- Identifying Root Cause of bugs by Capturing Changed Code Lines with Relational Graph Neural Networks
- Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- Language-Based Digital Twins for Elderly Cognitive Assistance
- CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
- Evaluating Pluralism in LLMs through Latent Perspectives
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
- DOPD: Dual On-policy Distillation
- Self-Ablating Transformers: More Interpretability, Less Sparsity
- Computational Identification of Regulatory Statements in EU Legislation
- Toward Calibrated Mixture-of-Experts Under Distribution Shift
- Fusing Semantic, Lexical, and Domain Perspectives for Recipe Similarity Estimation
- Using large language models to extract plant functional traits from unstructured text
- Offline changepoint localization using a matrix of conformal p-values
- Financial Bond Similarity Search Using Representation Learning
- Causal Motion Diffusion Models for Autoregressive Motion Generation
- Fast Transformer Inference on ARM-Based HMPSoCs
- Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
- MatMMFuse: Multi-Modal Fusion model for Material Property Prediction
- TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived Semantics
- A Survey on Parameter-Efficient Fine-Tuning for Foundation Models in Federated Learning
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
- BrightCookies at SemEval-2025 Task 9: Exploring Data Augmentation for Food Hazard Classification
- Kernel Affine Hull Machines as Compute-Efficient Encoders for Frozen Semantic Spaces
- Legilimens: Performant Video Analytics on the System-on-Chip Edge
- Fine-Grained Classification: Connecting Metadata via Cross-Contrastive Pre-Training
- ImproBR: Bug Report Improver Using LLMs
- Drift Happens: An Empirical Study of Neural Architecture Robustness to Temporal Distribution Shift
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift
- MATCHA: Matching Text via Contrastive Semantic Alignment
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
- Automated Generation of Precedence Graphs in Digital Value Chains for Automotive Production
- Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI
- ClimaEmpact: Domain-Aligned Small Language Models and Datasets for Extreme Weather Analytics
- Toward Inclusive Low-Code Development: Detecting Accessibility Issues in User Reviews
- Algorithmic Collective Action with Two Collectives
- AIBuildAI: An AI Agent for Automatically Building AI Models
- The Influence of Text Variation on User Engagement in Cross-Platform Content Sharing
- Sentiment and Social Signals in the Climate Crisis: A Survey on Analyzing Social Media Responses to Extreme Weather Events
- Leveraging Decoder Architectures for Learned Sparse Retrieval
- NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
- A Baseline Multimodal Approach to Emotion Recognition in Conversations
- Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
- Skill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation
- An evaluation of LLMs for political bias in Western media: Israel-Hamas and Ukraine-Russia wars
- Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
- ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts
- LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
- HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
- Rubrics as Privileged Information for Open-Ended Generation
- Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions
- Checked-In Secret Detection: Strings Are All You Need
- V-CEM: Bridging Performance and Intervenability in Concept-based Models
- Private Federated Learning using Preference-Optimized Synthetic Data
- Automated Extraction and Analysis of Developer's Rationale in Open Source Software
- Language Models to Support Multi-Label Classification of Industrial Data
- ViSMaP: Unsupervised Hour-long Video Summarisation by Meta-Prompting
- Saliency-driven Dynamic Token Pruning for Large Language Models
- Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
- Automated Bug Report Prioritization in Large Open-Source Projects
- NLCTables: A Dataset for Marrying Natural Language Conditions with Table Discovery
- LLM-based Semantic Augmentation for Harmful Content Detection
- W-PCA Based Gradient-Free Proxy for Efficient Search of Lightweight Language Models
- DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models
- Zero Day Malware Detection with Alpha: Fast DBI with Transformer Models for Real World Application
- Visualizing Public Opinion on X: A Real-Time Sentiment Dashboard Using VADER and DistilBERT
- EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
- Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
- Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
- Enhancing Multilingual Sentiment Analysis with Explainability for Sinhala, English, and Code-Mixed Content
- From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
- Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
- Towards Lossless Token Pruning in Late-Interaction Retrieval Models
- You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
- AdaVid: Adaptive Video-Language Pretraining
- Measuring and Detecting Harmful AI Sycophancy
- Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- A Dual-Space Framework for General Knowledge Distillation of Large Language Models
- LANGTRAJ: Diffusion Model and Dataset for Language-Conditioned Trajectory Simulation
- TD-Suite: All Batteries Included Framework for Technical Debt Classification
- QualiTagger: Automating software quality detection in issue trackers
- Improving Multimodal Hateful Meme Detection Exploiting LMM-Generated Knowledge
- RiskRAG: A Data-Driven Solution for Improved AI Model Risk Reporting
- RAG-VR: Leveraging Retrieval-Augmented Generation for 3D Question Answering in VR Environments
- MMLTC : A Novel Tolerance‐Based Clustering Framework for Multimodal Sentiment and Harmful Meme Classification in Multilingual Settings
- Semantically Encoding Activity Labels for Context-Aware Human Activity Recognition
- Exploring the Effectiveness and Interpretability of Texts in LLM-based Time Series Models
- NLP Security and Ethics, in the Wild
- Mosaic: Composite Projection Pruning for Resource-efficient LLMs
- Analyzing Examinee Comments using DistilBERT and Machine Learning to Ensure Quality Control in Exam Content
- Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
- Multi-Sense Embeddings for Language Models and Knowledge Distillation
- Mixture-of-Personas Language Models for Population Simulation
- Few Dimensions are Enough: Fine-tuning BERT with Selected Dimensions Revealed Its Redundant Nature
- SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
- Generative AI Enhanced Financial Risk Management Information Retrieval
- Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
- Prompt Optimization with Logged Bandit Data
- UNDO: Understanding Distillation as Optimization
Discussions
Related