A Primer in BERTology: What we know about how BERT works
2020/02/27 by Anna Rogers, Olga Kovaleva, Rogers, Anna +3 · 5 voices · 120 citations
Computer Science · #cs.CL
paper · pdf · doi:10.48550/arxiv.2002.12327
Accepted to TACL. Please note that the multilingual BERT section is only available in version 1
arxiv created 2020/11/09 · arxiv updated 2020/11/10
Abstract
Transformer-based models have pushed state of the art in many areas of NLP, but our understanding of what is behind their success is still limited. This paper is the first survey of over 150 studies of the popular BERT model. We review the current state of knowledge about how BERT works, what kind of information it learns and how it is represented, common modifications to its training objectives and architecture, the overparameterization issue and approaches to compression. We then outline directions for future research.
Cited by
- surprisal is Not a Theory
- Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
- Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference
- Parallel by design? Meaning and grammar in single-stream and dual-stream neural network architectures
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- When Models Manipulate Manifolds: The Geometry of a Counting Task
- The Dead Salmons of AI Interpretability
- Research Community Perspectives on "Intelligence" and Large Language Models
- LLMs as a synthesis between symbolic and distributed approaches to language
- Beyond Context: Large Language Models' Failure to Grasp Users' Intent
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- Characterizing Mamba's Selective Memory using Auto-Encoders
- Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
- Scaling Bidirectional Spans and Span Violations in Attention Mechanism
- Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- Layer Probing Improves Kinase Functional Prediction with Protein Language Models
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Generation, Evaluation, and Explanation of Novelists' Styles with Single-Token Prompts
- A Hybrid Classical-Quantum Fine Tuned BERT for Text Classification
- N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator
- SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
- SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
- Catching Contamination Before Generation: Spectral Kill Switches for Agents
- Quantitative Bounds for Length Generalization in Transformers
- Enhancing Sentiment Classification with Machine Learning and Combinatorial Fusion
- Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
- Relation Geometry in Semantic Space of Language Models
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
- Training-Free Spectral Fingerprints of Voice Processing in Transformers
- That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
- Attention Is All You Need for KV Cache in Diffusion LLMs
- CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
- Ethic-BERT: An Enhanced Deep Learning Model for Ethical and Non-Ethical Content Classification
- Fairness Metric Design Exploration in Multi-Domain Moral Sentiment Classification using Transformer-Based Models
- Isotropy and Geometry of Pretrained Protein LMs
- Mapping Semantic & Syntactic Relationships with Geometric Rotation
- Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
- Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- Mechanistic Interpretability of Socio-Political Frames in Language Models
- Allocation of Parameters in Transformers
- Investigating Multi-layer Representations for Dense Passage Retrieval
- A short survey on almost orthogonal vectors in a few specific large dimensions
- Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
- MUGEN: A Unified Framework for Efficient Motion Understanding and Generation
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
- Documents Are People and Words Are Items: A Psychometric Approach to Textual Data with Contextual Embeddings
- Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
- Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
- Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
- Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
- Transplant Then Regenerate: A New Paradigm for Text Data Augmentation
- Semantic Anchoring in Agentic Memory: Leveraging Linguistic Structures for Persistent Conversational Context
- Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
- Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support
- Streamlining Admission with LOR Insights: AI-Based Leadership Assessment in Online Master's Program
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- Length Representations in Large Language Models
- Explainable Mapper: Charting LLM Embedding Spaces Using Perturbation-Based Explanation and Verification Agents
- Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
- Could the Road to Grounded, Neuro-symbolic AI be Paved with Words-as-Classifiers?
- On the Semantics of Large Language Models
- Identification of Potentially Misclassified Crash Narratives using Machine Learning (ML) and Deep Learning (DL)
- Discourse Heuristics For Paradoxically Moral Self-Correction
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
- Can structural correspondences ground real world representational content in Large Language Models?
- A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
- Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning
- Towards an Explainable Comparison and Alignment of Feature Embeddings
- Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
- Adaptive Task Vectors for Large Language Models
- Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
- Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
- Probing Neural Topology of Large Language Models
- Domain Pre-training Impact on Representations
- Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments
- Queueing for Civility: User Perspectives on Regulating Emotions in Online Conversations
- Token Distillation: Attention-aware Input Embeddings For New Tokens
- Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
- LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
- Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework
- The Discovery Engine: A Framework for AI-Driven Synthesis and Navigation of Scientific Knowledge Landscapes
- MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation
- Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
- Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks
- Where did the ambiguity go? Examining how multimodal models interpret polysemous words
- Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
- Designing and Contextualising Probes for African Languages
- Interpretable Risk Mitigation in LLM Agent Systems
- Chronocept: Instilling a Sense of Time in Machines
- Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging
- BOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models
- Perturbation: A simple and efficient adversarial tracer for representation learning in language models
- Adaptive Loops and Memory in Transformers: Think Harder or Know More?
- Reverse Distillation: Consistently Scaling Protein Language Model Representations
- Jekyll-and-Hyde Tipping Point in an AI's Behavior
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- The Truthfulness Spectrum Hypothesis
- Developing Students’ Statistical Expertise Through Writing in the Age of AI
- Sensitivity-Positional Co-Localization in GQA Transformers
- HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
- Can large language models recognize complex language errors such as zeugma?
- Deep Learning with Pretrained 'Internal World' Layers: A Gemma 3-Based Modular Architecture for Wildfire Prediction
- Probing then Editing Response Personality of Large Language Models
- The Joint Distance Measure: A Measure of Similarity Accounting for Spatial and Angular Distances
- LayerFlow: Layer-wise Exploration of LLM Embeddings using Uncertainty-aware Interlinked Projections
- Linguistic Interpretability of Transformer-based Language Models: a systematic review
- Few Dimensions are Enough: Fine-tuning BERT with Selected Dimensions Revealed Its Redundant Nature
- Large language model [wikipedia]
- Foundation model [wikipedia]
- BERT (language model) [wikipedia]
Discussions
- A Primer in BERTology: What We Know About How Bert Works [hn, 81 points, 16 comments]
- A Primer in BERTology: What we know about how BERT works [hn, 2 points, 0 comments]
- the "lesioning" is sometimes used in "BERTology" as well, where BeRt ology is the stdy of BERT models (which were a popular ealry pre chatGPT LLM used widely) see 6.3 for studies that kind of do lessi [bsky, 1 points, 1 comments]
- section 3.1 of arxiv.org/abs/2002.12327 mentions a number of other studies probing for linguistic structure learned by what's now a pretty simple model, let alone anything more recent/capable [bsky, 0 points, 0 comments]
- may I have a counter-argument: 1. Per your work, language models benefit from the human knowledge, but they don't fully rely on it: arxiv.org/abs/2002.12327 (my fav NLP paper) which agrees with the bi [bsky, 0 points, 2 comments]
Related