A Comparative Study of Clinical ModernBERT and BioMedical ModernBERT on the DDXPlus Dataset
2024/12/18 by Benjamin Warner, Antoine Chaffin, Warner, Benjamin +25 · 8 voices · 171 citations
Computer Science · #Advanced Data Compression Techniques #Human Pose and Action Recognition #Video Analysis and Summarization #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2412.13663
openalex publication_date 2024/12/18 · openalex created_date 2024/12/21 · openalex updated_date 2026/07/28
Abstract
Encoder-only transformer models such as BERT offer a great performance-size tradeoff for retrieval and classification tasks with respect to larger decoder-only models. Despite being the workhorse of numerous production pipelines, there have been limited Pareto improvements to BERT since its release. In this paper, we introduce ModernBERT, bringing modern model optimizations to encoder-only models and representing a major Pareto improvement over older encoders. Trained on 2 trillion tokens with a native 8192 sequence length, ModernBERT models exhibit state-of-the-art results on a large pool of evaluations encompassing diverse classification tasks and both single and multi-vector retrieval on different domains (including code). In addition to strong downstream performance, ModernBERT is also the most speed and memory efficient encoder and is designed for inference on common GPUs.
Cited by
- A Unified Moral-Value Dataset for Instruction Tuning
- Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection
- Transition-Aware Backend Dispatch for Edge LLM Inference
- Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
- Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
- Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers
- LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models
- Latent Trajectory Discrimination for AI-Generated Text Detection
- Large Language Models as Unified Multimodal Learners for Clinical Prediction
- Beyond content: behavioral policies reveal actors in information operations
- StoryScope: Investigating idiosyncrasies in AI fiction
- Strategies for Span Labeling with Large Language Models
- On the Theoretical Limitations of Embedding-Based Retrieval
- Measuring Scalar Constructs in Social Science with LLMs
- JavelinGuard: Low-Cost Transformer Architectures for LLM Security
- Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
- ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
- Model in Distress: Sentiment Analysis on French Synthetic Social Media
- Argus: Token Aware Distributed LLM Inference Optimization
- Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- Learning Dynamics of Strategic Publishers in Generative AI Ecosystems
- Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
- Mining Legal Arguments to Study Judicial Formalism
- Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
- OnCoCo 1.0: A Public Dataset for Fine-Grained Message Classification in Online Counseling Conversations
- PolyLingua: Margin-based Inter-class Transformer for Robust Cross-domain Language Detection
- Open Polymer Challenge: Post-Competition Report
- A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data
- Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
- HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
- Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion
- BERnaT: Basque Encoders for Representing Natural Textual Diversity
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
- Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- WET -- Weighted Ensemble Transformer for Identifying Psychiatric Stressors Related to Suicide on X (formerly Twitter)
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- A Fast and Efficient Modern BERT based Text-Conditioned Diffusion Model for Medical Image Segmentation
- A Machine Learning Approach for Detection of Mental Health Conditions and Cyberbullying from Social Media
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- Predicting one-year clinical instability and mortality in heart failure patients using sequence modeling
- Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
- Classification of worldwide news articles by perceived quality, 2018-2024
- Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping
- Rethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languages
- ManufactuBERT: Efficient Continual Pretraining for Manufacturing
- Enhancing composition-based materials property prediction by cross-modal knowledge transfer
- CARMA: Comprehensive Automatically-annotated Reddit Mental Health Dataset for Arabic
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
- LIR: The First Workshop on Late Interaction and Multi Vector Retrieval @ ECIR 2026
- Supporting Workflow Reproducibility by Linking Bioinformatics Tools across Papers and Executable Code
- Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval
- Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
- Text2Score: Generating Sheet Music From Textual Prompts
- Text Simplification with Sentence Embeddings
- COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
- AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
- In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit Discussions
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- BoundRL: Efficient Structured Text Segmentation through Reinforced Boundary Generation
- NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew
- Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- BeLLMan: Controlling LLM Congestion
- Fantastic (small) Retrievers and How to Train Them: mxbai-edge-colbert-v0 Tech Report
- Embedding-Based Context-Aware Reranker
- Chinese ModernBERT with Whole-Word Masking
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion
- Simple Projection Variants Improve ColBERT Performance
- Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
- FactAppeal: Identifying Epistemic Factual Appeals in News Media
- Brick: Spatial Capability Routing for the Mixture-of-Models (MoM) Paradigm
- iBERT: Interpretable Style Embeddings via Sense Decomposition
- Stronger Re-identification Attacks through Reasoning and Aggregation
- FrameEOL: Semantic Frame Induction using Causal Language Models
- When to Reason: Semantic Router for vLLM
- EDUMATH: Generating Standards-aligned Educational Math Word Problems
- BOTANIC-0: a series of foundation models for plant genomic data
- GeneCAD: Plant Genome Annotation with a DNA Foundation Model
- LANTERN: Scalable Distillation of Large Language Models for Job-Person Fit and Explanation
- ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
- When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
- ModernVBERT: Towards Smaller Visual Document Retrievers
- Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors
- A-VERT: Agnostic Verification with Embedding Ranking Targets
- SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- ResFormer: All-Time Reservoir Memory for Long Sequence Classification
- CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
- JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
- What Is The Political Content in LLMs' Pre- and Post-Training Data?
- Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
- Peeling metric spaces of strict negative type
- AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
- The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials
- NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
- Reading the unreadable: creating a dataset of 19th century English newspapers using image-to-text language models
- Fine-Grained Detection of AI-Generated Text Using Sentence-Level Segmentation
- Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
- Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
- TraceHiding: Scalable Machine Unlearning for Mobility Data
- Multi-task Pretraining for Enhancing Interpretable L2 Pronunciation Assessment
- Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
- SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
- Patent Language Model Pretraining with ModernBERT
- Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- Op-Fed: Opinion, Stance, and Monetary Policy Annotations on FOMC Transcripts Using Active Learning
- From Embeddings to Equations: Genetic-Programming Surrogates for Interpretable Transformer Classification
- GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
- Long Context Automated Essay Scoring with Language Models
- Multimodal LLMs See Sentiment
- Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text
- Customizing the Inductive Biases of Softmax Attention using Structured Matrices
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Augmenting Human-Centered Racial Covenant Detection and Georeferencing with Plug-and-Play NLP Pipelines
- BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
- Sample-efficient Integration of New Modalities into Large Language Models
- IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
- Text Reinforcement for Multimodal Time Series Forecasting
- Neural Models and Language Model Prompting for the Multidimensional Evaluation of Open-Ended Conversations
- SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Hermes 4 Technical Report
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Transplant Then Regenerate: A New Paradigm for Text Data Augmentation
- Gener anno : A Genomic Foundation Model for Metagenomic Annotation
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable Representations
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
- Improving Document Retrieval Coherence for Semantically Equivalent Queries
- BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- CALE : Concept-Aligned Embeddings for Both Within-Lemma and Inter-Lemma Sense Differentiation
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations
- UPLME: Uncertainty-Aware Probabilistic Language Modelling for Robust Empathy Regression
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation
- TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
- Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
- GovRelBench:A Benchmark for Government Domain Relevance
- Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- NUTMEG: Separating Signal From Noise in Annotator Disagreement
- Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
- Shaping capabilities with token-level data filtering
- A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
- Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
- Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
- Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
- Seq vs Seq: An Open Suite of Paired Encoders and Decoders
- AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles
- Language Models for Adult Service Website Text Analysis
- Tiny Reward Models
Droid: A Resource Suite for AI-Generated Code Detection- Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
- Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
Discussions
- So how does it work? ModernBERT brings modern engineering ideas from LLMs over to encoder models. We did so in three core ways: 1) a modernized transformer architecture; 2) particular attention to eff [bsky, 29 points, 2 comments]
- I was super excited to read the ModernBERT paper! Love this interest in creating a better encoder model. "ModernBERT-base is the first encoder to beat DeBERTaV3-base since its release in 2021" 🤯- ar [bsky, 16 points, 0 comments]
- I'm a little behind but I finally got some time to read the ModernBERT paper. While I hate the name (never put time relative terms in names!), the paper is fabulous! Basically the entire paper can be [bsky, 6 points, 2 comments]
- For all the model design, training, and evaluation details, check out our Arxiv preprint: arxiv.org/abs/2412.13663 [bsky, 4 points, 1 comments]
- ModernBERT [hn, 3 points, 1 comments]
- It is also worth mentioning ModernBERT (arxiv.org/abs/2412.13663), which is another great parallel effort to modernize BERT but with some differences, as highlighted in our paper! (6/n) [bsky, 2 points, 1 comments]
- BERT🤗を現代化😊 arxiv.org/abs/2412.13663 [bsky, 1 points, 0 comments]
- Impressive accuracy and speed improvements over the current BERT models and a native sequence length over 8k! arxiv.org/abs/2412.13663 [bsky, 0 points, 0 comments]
Related