DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
2019/10/02 by Sanh, Victor, Debut, Lysandre, Chaumond, Julien +1 · 392 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.1910.01108
Abstract
As Transfer Learning from large-scale pre-trained models becomes more prevalent in Natural Language Processing (NLP), operating these large models in on-the-edge and/or under constrained computational training or inference budgets remains challenging. In this work, we propose a method to pre-train a smaller general-purpose language representation model, called DistilBERT, which can then be fine-tuned with good performances on a wide range of tasks like its larger counterparts. While most prior work investigated the use of distillation for building task-specific models, we leverage knowledge distillation during the pre-training phase and show that it is possible to reduce the size of a BERT model by 40%, while retaining 97% of its language understanding capabilities and being 60% faster. To leverage the inductive biases learned by larger models during pre-training, we introduce a triple loss combining language modeling, distillation and cosine-distance losses. Our smaller, faster and lighter model is cheaper to pre-train and we demonstrate its capabilities for on-device computations in a proof-of-concept experiment and a comparative on-device study.
Cited by
- Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects?
- GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- Masked Distillation: Internalizing the Chain-of-Thought in Language Models
- TimeBill: Time-Budgeted Inference for Large Language Models
- Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
- SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
- Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
- Agentic Multi-Persona Framework for Evidence-Aware Fake News Detection
- Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
- USE: A Unified Model for Universal Sound Separation and Extraction
- Block-Recurrent Dynamics in Vision Transformers
- LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning
- Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering
- Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
- LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
- Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs
- Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
- DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- Examining the Utility of Self-disclosure Types for Modeling Annotators of Social Norms
- TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- Improving Recursive Transformers with Mixture of LoRAs
- Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- Task-Specific Sparse Feature Masks for Molecular Toxicity Prediction with Chemical Language Models
- Parallax: Runtime Parallelization for Operator Fallbacks in Heterogeneous Edge Systems
- CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation
- Provably Learning from Modern Language Models via Low Logit Rank
- SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
- Interpreto: An Explainability Library for Transformers
- Text2Graph: Combining Lightweight LLMs and GNNs for Efficient Text Classification in Label-Scarce Scenarios
- Belief Is All You Need: Modeling Narrative Archetypes in Conspiratorial Discourse
- Language-Conditioned Safe Trajectory Generation for Spacecraft Rendezvous
- Financial News Summarization: Can extractive methods still offer a true alternative to LLMs?
- Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
- MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
- A Patient-Doctor-NLP-System to contest inequality for less privileged
- Unveiling Hedge Funds: Topic Modeling and Sentiment Correlation with Fund Performance
- Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
- Optimizing LLMs Using Quantization for Mobile Execution
- Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- Fine-Tuning BERT for Domain-Specific Question Answering: Toward Educational NLP Resources at University Scale
- LLMs Know More Than Words: A Genre Study with Syntax, Metaphor & Phonetics
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- Fine-grained Narrative Classification in Biased News Articles
- Network of Theseus (like the ship)
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Bangla Hate Speech Classification with Fine-tuned Transformer Models
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- RoleMotion: A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- No Trust Issues Here: A Technical Report on the Winning Solutions for the Rayan AI Contest
- DyFuLM: An Advanced Multimodal Framework for Sentiment Analysis
- Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- Tree Matching Networks for Natural Language Inference: Parameter-Efficient Semantic Understanding via Dependency Parse Trees
- Retrieval-Augmented Few-Shot Prompting Versus Fine-Tuning for Code Vulnerability Detection
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Effectively Detecting and Responding to Online Harassment with Large Language Models
- A Trainable Centrality Framework for Modern Data
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- Using Text-Based Life Trajectories from Swedish Register Data to Predict Residential Mobility with Pretrained Transformers
- MorphingDB: A Task-Centric AI-Native DBMS for Model Management and Inference
- Enhancing Burmese News Classification with Kolmogorov-Arnold Network Head Fine-tuning
- Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
- BRIC: Bridging Kinematic Plans and Physical Control at Test Time
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- Deterministic Continuous Replacement: Fast and Stable Module Replacement in Pretrained Transformers
- Towards Characterizing Knowledge Distillation of PPG Heart Rate Estimation Models
- A symbolic Perl algorithm for the unification of Nahuatl word spellings
- A Systematic Study of Compression Ordering for Large Language Models
- Efficient Covariance Estimation for Sparsified Functional Data
- MTikGuard System: A Transformer-Based Multimodal System for Child-Safe Content Moderation on TikTok
- When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
- NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
- GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs
- RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue
- Gradient-descent methods for quantum detector tomography
- Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
- Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system
- Decoupled Action Head: Confining Task Knowledge to Conditioning Layers
- From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
- Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
- BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
- BiCA: Effective Biomedical Dense Retrieval with Citation-Aware Hard Negatives
- CC30k: A Citation Contexts Dataset for Reproducibility-Oriented Sentiment Analysis
- Graph Representation-based Model Poisoning on the Heterogeneous Internet of Agents
- Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
- Beyond Uniform Deletion: A Data Value-Weighted Framework for Certified Machine Unlearning
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
- One-Shot Knowledge Transfer for Scalable Person Re-Identification
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- Quantifying the Climate Risk of Generative AI: Region-Aware Carbon Accounting with G-TRACE and the AI Sustainability Pyramid
- Advancing Equitable AI: Evaluating Cultural Expressiveness in LLMs for Latin American Contexts
- SOLVE-Med: Specialized Orchestration for Leading Vertical Experts across Medical Specialties
- Laugh, Relate, Engage: Stylized Comment Generation for Short Videos
- The Curved Spacetime of Transformer Architectures
- Targeted Error Correction in Knowledge Distillation: Small Language Models Surpass GPT
- DEEP: A Discourse Evolution Engine for Predictions about Social Movements
- Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
- Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction
- Sherlock: Reliable and Efficient Agentic Workflow Execution
- Exploring the Utilities of the Rationales from Large Language Models to Enhance Automated Essay Scoring
- Elastic Architecture Search for Efficient Language Models
- Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
- FakeZero: Real-Time, Privacy-Preserving Misinformation Detection for Facebook and X
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text
- Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case
- How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Exploring the Intersection of AI, Language, and Law: A Bibliometric Analysis
- DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
- Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
- Stay Tuned: Improving Sentiment Analysis and Stance Detection Using Large Language Models
- A machine learning-based evidence map of ocean-related options for climate change mitigation and adaptation
- MemEIC: A Step Toward Continual and Compositional Knowledge Editing
- TOPol: Capturing and Explaining Multidimensional Semantic Polarity Fields and Vectors
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Code Contribution and Credit in Science
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
- In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit Discussions
- FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
- The Lossy Horizon: Error-Bounded Predictive Coding for Lossy Text Compression (Episode I)
- Power to the Clients: Federated Learning in a Dictatorship Setting
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- Doubly-Regressing Approach for Subgroup Fairness
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- GRATING: Low-Latency and Memory-Efficient Semantic Selection on Device
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
- On Optimal Hyperparameters for Differentially Private Deep Transfer Learning
- Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent
- Calibration and Discrimination Optimization Using Clusters of Learned Representation
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder
- Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
- Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
- Sync or Sink: Bounds on Algorithmic Collective Action with Noise and Multiple Groups
- Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
- Quantifying Climate Policy Action and Its Links to Development Outcomes: A Cross-National Data-Driven Analysis
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Refugees of the Digital Space: Platform Migration from TikTok to RedNote
- MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
- Generalizing WiFi Gesture Recognition via Large-Model-Aware Semantic Distillation and Alignment
- ProtoTopic: Prototypical Network for Few-Shot Medical Topic Modeling
- Tandem Training for Language Models
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- Efficient Adaptive Transformer: An Empirical Study and Reproducible Framework
- Traveling Salesman-Based Token Ordering Improves Stability in Homomorphically Encrypted Language Models
- FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning
- A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
- Fairness Metric Design Exploration in Multi-Domain Moral Sentiment Classification using Transformer-Based Models
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Test-Time Adaptation by Causal Trimming
- From Reasoning LLMs to BERT: A Two-Stage Distillation Framework for Search Relevance
- Therapeutic AI and the Hidden Risks of Over-Disclosure: An Embedded AI-Literacy Framework for Mental Health Privacy
- Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default
- Serialized EHR make for good text representations
- Inverse Language Modeling towards Robust and Grounded LLMs
- On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
- Framing Unionization on Facebook: Communication around Representation Elections in the United States
- Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- Theoretical guarantees for change localization using conformal p-values
- STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
- AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
- Causality Guided Representation Learning for Cross-Style Hate Speech Detection
- Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
- Investigating Counterclaims in Causality Extraction from Text
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- Reasoning for Hierarchical Text Classification: The Case of Patents
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- GUIDE: Guided Initialization and Distillation of Embeddings
- LLM Bias Detection and Mitigation through the Lens of Desired Distributions
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
- Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Learning from All: Concept Alignment for Autonomous Distillation from Multiple Drifting MLLMs
- Annotate Rhetorical Relations with INCEpTION: A Comparison with Automatic Approaches
- Consistent Kernel Change-Point Detection under m-Dependence for Text Segmentation
- MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
- Neural Correlates of Language Models Are Specific to Human Language
- Dissecting Transformers: A CLEAR Perspective towards Green AI
- Syntax-Guided Diffusion Language Models with User-Integrated Personalization
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
- Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
- Fine-tuning with RAG for Improving LLM Learning of New Skills
- Robust Federated Inference
- Quantifying Semantic Shift in Financial NLP: Robust Metrics for Market Prediction Stability
- Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
- Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests
- Vocabulary Customization for Efficient Domain-Specific LLM Deployment
- FITS: Towards an AI-Driven Fashion Information Tool for Sustainability
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- Can Large Language Models Express Uncertainty Like Human?
- The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact
- Why Alignment Must Precede Distillation: A Minimal Working Explanation
- Temporal Generalization: A Reality Check
- Knowledge distillation through geometry-aware representational alignment
- Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Text2Move: Text-to-moving sound generation via trajectory prediction and temporal alignment
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- An Improved Quantum Software Challenges Classification Approach using Transfer Learning and Explainable AI
- A short survey on almost orthogonal vectors in a few specific large dimensions
- CTI Dataset Construction from Telegram
- PseudoBridge: Pseudo Code as the Bridge for Better Semantic and Logic Alignment in Code Retrieval
- Overcoming Black-box Attack Inefficiency with Hybrid and Dynamic Select Algorithms
- RedHerring Attack: Testing the Reliability of Attack Detection
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI
- Confidence Calibration in Large Language Model-Based Entity Matching
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
- AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
- MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
- NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
- LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation
- A transformer-based multi-feature fusion method for detecting traffic events using Twitter data
- Otters: An Energy-Efficient SpikingTransformer via Optical Time-to-First-Spike Encoding
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- Semantic Search for Information Retrieval
- Enhancing Transformer-Based Rerankers with Synthetic Data and LLM-Based Supervision
- AIRwaves at CheckThat! 2025: Retrieving Scientific Sources for Implicit Claims on Social Media with Dual Encoders and Neural Re-Ranking
- Identifying Constructive Conflict in Online Discussions through Controversial yet Toxicity Resilient Posts
- Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
- Fine-Grained Detection of AI-Generated Text Using Sentence-Level Segmentation
- Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
- Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
- DRES: Fake news detection by dynamic representation and ensemble selection
- Random Direct Preference Optimization for Radiography Report Generation
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- Assessing Classical Machine Learning and Transformer-based Approaches for Detecting AI-Generated Research Text
- CausalSent: Interpretable Sentiment Classification with RieszNet
- RMT-KD: Random Matrix Theoretic Causal Knowledge Distillation
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- Bridging Graph and State-Space Modeling for Intensive Care Unit Length of Stay Prediction
- Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- SINAI at eRisk@CLEF 2022: Approaching Early Detection of Gambling and Eating Disorders with Natural Language Processing
- Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification
- Who Wins the Race? (R Vs Python) - An Exploratory Study on Energy Consumption of Machine Learning Algorithms
- LLM on a Budget: Active Knowledge Distillation for Efficient Classification of Large Text Corpora
- The Few-shot Dilemma: Over-prompting Large Language Models
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Exploring Training Data Attribution under Limited Access Constraints
- Spatio-Temporal Pruning for Compressed Spiking Large Language Models
- Image-Seeking Intent Prediction for Cross-Device Product Search
- User eXperience Perception Insights Dataset (UXPID): Synthetic User Feedback from Public Industrial Forums
- Cross-Modal Retrieval with Cauchy-Schwarz Divergence
- Dynamic Span Interaction and Graph-Aware Memory for Entity-Level Sentiment Classification
- LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries
- A Transformer-Based Cross-Platform Analysis of Public Discourse on the 15-Minute City Paradigm
- Efficient Hate Speech Detection: Evaluating 38 Models from Traditional Methods to Transformers
- Quantifier Scope Interpretation in Language Learners and LLMs
- Remotely Seeing Is Believing: How Trust in Cyber-Physical Systems Evolves Through Virtual Observation
- VARCO-VISION-2.0 Technical Report
- SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- Green Federated Learning via Carbon-Aware Client and Time Slot Scheduling
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
- D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Few-Shot Query Intent Detection via Relation-Aware Prompt Learning
- Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
- RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models
- Anti-establishment sentiment on TikTok: Implications for understanding influence(rs) and expertise on social media
- LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
- NoteBar: An AI-Assisted Note-Taking System for Personal Knowledge Management
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- From Confidence to Collapse in LLM Factual Robustness
- StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
- Hierarchical Motion Captioning Utilizing External Text Data Source
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA
- FLEET: A Federated Learning Emulation and Evaluation Testbed for Holistic Research
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
- Text-Driven 3D Hand Motion Generation from Sign Language Data
- MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- Speech Emotion Recognition via Entropy-Aware Score Selection
- MathBuddy: A Multimodal System for Affective Math Tutoring
- AI-Powered Detection of Inappropriate Language in Medical School Curricula
- QuesGenie: Intelligent Multimodal Question Generation
- ALSA: Anchors in Logit Space for Out-of-Distribution Accuracy Estimation
- Skill-based Explanations for Serendipitous Course Recommendation
- SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- Compressed Models are NOT Trust-equivalent to Their Large Counterparts
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- A Risk Manager for Intrusion Tolerant Systems: Enhancing HAL 9000 with New Scoring and Data Sources
- REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks
- Checkmate: interpretable and explainable RSVQA is the endgame
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- A Global Dataset of Location Data Integrity-Assessed Reforestation Efforts
- Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- MCP-Orchestrated Multi-Agent System for Automated Disinformation Detection
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
- Bhav-Net: Knowledge Transfer for Cross-Lingual Antonym vs Synonym Distinction via Dual-Space Graph Transformers
- Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
- Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
- LLMs for Law: Evaluating Legal-Specific LLMs on Contract Understanding
- Modeling and Detecting Company Risks from News: A Case Study in Bloomberg News
- TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
- LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
- Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
- Embedding Alignment in Code Generation for Audio
- Task complexity shapes internal representations and robustness in neural networks
- PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
- CoMAD: A Multiple-Teacher Self-Supervised Distillation Framework
- GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
- Leveraging large language models for SQL behavior-based database intrusion detection
- Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
- WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval
- HiTeC: Hierarchical Contrastive Learning on Text-Attributed Hypergraph with Semantic-Aware Augmentation
- SLIM-LLMs: Modeling of Style-Sensory Language RelationshipsThrough Low-Dimensional Representations
- Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
- "Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth
- Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
- Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
- Distillation-Enhanced Clustering Acceleration for Encrypted Traffic Classification
- Contextual Graph Transformer: A Small Language Model for Enhanced Engineering Document Information Extraction
- EHSAN: Leveraging ChatGPT in a Hybrid Framework for Arabic Aspect-Based Sentiment Analysis in Healthcare
- ReMoMask: Retrieval-Augmented Masked Motion Generation
- Empowering Tabular Data Preparation with Language Models: Why and How?
- R2-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
- DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
- Classification of Psychiatry Clinical Notes by Diagnosis: A Deep Learning and Machine Learning Approach
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
- Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
- Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- Real-time News Story Identification
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- LLM-based Content Classification Approach for GitHub Repositories by the README Files
- Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions
- Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
- Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
- Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
- Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
- A Similarity Measure for Comparing Conversational Dynamics
- Efficient Agents: Building Effective Agents While Reducing Cost
- What does the public want their local government to hear? A data-driven case study of public comments across the state of Michigan
- OPRD: On-Policy Representation Distillation
Related