XLNet: Generalized Autoregressive Pretraining for Language Understanding
2019/06/19 by Zhilin Yang, Zihang Dai, Yang, Zhilin +9 · 1 voice · 179 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.1906.08237
Abstract
With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy. In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. Furthermore, XLNet integrates ideas from Transformer-XL, the state-of-the-art autoregressive model, into pretraining. Empirically, under comparable experiment settings, XLNet outperforms BERT on 20 tasks, often by a large margin, including question answering, natural language inference, sentiment analysis, and document ranking.
Citations
Cited by
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
- Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation
- Dice Loss for Data-imbalanced NLP Tasks
- Candidate Attended Dialogue State Tracking Using BERT
- JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
- Mind captioning: Evolving descriptive text of mental content from human brain activity
- A Survey on Diffusion Language Models
- PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- Enhancing Code Understanding for Impact Analysis by Combining Transformers and Program Dependence Graphs
- GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
- Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
- Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning
- SELECT: Detecting Label Errors in Real-world Scene Text Data
- Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats
- When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing
- Cross-platform Product Matching Based on Entity Alignment of Knowledge Graph with RAEA model
- Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
- Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- Efficient Covariance Estimation for Sparsified Functional Data
- Spanning Tree Autoregressive Visual Generation
- Zero-Shot Grammar Competency Estimation Using Large Language Model Generated Pseudo Labels
- How to Select One Among All? An Extensive Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding
- Evaluation of Sentence Representations in Polish
- Taming Pretrained Transformers for Extreme Multi-label Text Classification
- Parallel Sampling via Autospeculation
- Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
- Learning Norms from Stories: A Prior for Value Aligned Agents
- Comparing Reconstruction Attacks on Pretrained Versus Full Fine-tuned Large Language Model Embeddings on Homo Sapiens Splice Sites Genomic Data
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge
- Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction
- Reversal Invariance in Autoregressive Language Models
- Enhancing Sentiment Classification with Machine Learning and Combinatorial Fusion
- Disentangling Adaptive Gradient Methods from Learning Rates
- ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
- Portuguese Named Entity Recognition using BERT-CRF
- Efficient Attention: Attention with Linear Complexities
- Position Masking for Language Models
- Factorized Multimodal Transformer for Multimodal Sequential Learning
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- One-shot Key Information Extraction from Document with Deep Partial Graph Matching
- IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
- Enabling Language Models to Fill in the Blanks
- Multiple Structural Priors Guided Self Attention Network for Language Understanding
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- A review on the applications of Transformer-based language models for nucleotide sequence analysis
- Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
- Symmetric Regularization based BERT for Pair-wise Semantic Reasoning
- What makes us curious? analysis of a corpus of open-domain questions
- The Brownian motion in the transformer model
- ActBERT: Learning Global-Local Video-Text Representations
- Q8BERT: Quantized 8Bit BERT
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Overview of the TREC 2022 deep learning track
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
- TVDIM: Enhancing Image Self-Supervised Pretraining via Noisy Text Data
- RL makes MLLMs see better than SFT
- VL-BERT: Pre-training of Generic Visual-Linguistic Representations
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
- Large Language Models, Agency, and Why Speech Acts are Beyond Them (For Now) – A Kantian-Cum-Pragmatist Case
- Transfer Learning for Multi-lingual Tasks -- a Survey
- A Comprehensive Dataset for Human vs. AI Generated Text Detection
- SALSA: Single-pass Autoregressive LLM Structured Classification
- Robustness Verification for Transformers
- Complex Transformer: A Framework for Modeling Complex-Valued Sequence
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- IMB: An Italian Medical Benchmark for Question Answering
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG
- ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- DocVQA: A Dataset for VQA on Document Images
- Effective Unsupervised Domain Adaptation with Adversarially Trained\n Language Models
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- Isotropy and Geometry of Pretrained Protein LMs
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- Multimodal Foundation Models for Early Disease Detection
- PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
- SenWave: A Fine-Grained Multi-Language Sentiment Analysis Dataset Sourced from COVID-19 Tweets
- Using the Hammer Only on Nails: A Hybrid Method for Evidence Retrieval for Question Answering
- Reasoning for Hierarchical Text Classification: The Case of Patents
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Language models for longitudinal analysis of abusive content in Billboard Music Charts
- 12-in-1: Multi-Task Vision and Language Representation Learning
- Allocation of Parameters in Transformers
- MathBERT: A Pre-Trained Model for Mathematical Formula Understanding
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- Self-Speculative Masked Diffusions
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- <scp>AI</scp> Methods for Antimicrobial Peptides: Progress and Challenges
- Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructions
- Text Adversarial Attacks with Dynamic Outputs
- LV-BERT: Exploiting Layer Variety for BERT
- Understanding and Enhancing Mask-Based Pretraining towards Universal Representations
- Performance Consistency of Learning Methods for Information Retrieval Tasks
- Confidence Calibration in Large Language Model-Based Entity Matching
- Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
- THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
- Learning to Emphasize: Dataset and Shared Task Models for Selecting Emphasis in Presentation Slides
- Literature review on vulnerability detection using NLP technology
- Transformers and genome language models
- Evolving Character-level Convolutional Neural Networks for Text Classification
- TRANS-BLSTM: Transformer with Bidirectional LSTM for Language Understanding
- A study of word embedding models for measuring topic coherence
- Character-level Representations Improve DRS-based Semantic Parsing Even\n in the Age of BERT
- Optimizing Informer with Whale Optimization Algorithm for Enhanced Ship Trajectory Prediction
- Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
- DRES: Fake news detection by dynamic representation and ensemble selection
- A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse
- Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
- Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
- A Comparative Study of Transformer-Based Language Models on Extractive Question Answering
- Improve Transformer Models with Better Relative Position Embeddings
- Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications
- A Survey of Deep Learning for Scientific Discovery
- On Linear Identifiability of Learned Representations
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- Explaining Question Answering Models through Text Generation
- Efficient Softmax Approximation for Deep Neural Networks with Attention Mechanism
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
- IITkgp at FinCausal 2020, Shared Task 1: Causality Detection using Sentence Embeddings in Financial Reports
- A study of Turkish emotion classification with pretrained language models
- Long Context Automated Essay Scoring with Language Models
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- Beyond Token Limits: Assessing Language Model Performance on Long Text Classification
- Training Multilingual Pre-trained Language Model with Byte-level Subwords
- Hierarchical Bracketing Encodings Work for Dependency Graphs
- Diverse Image Inpainting with Bidirectional and Autoregressive Transformers
- Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- Are Transformers universal approximators of sequence-to-sequence functions?
- FastMoE: A Fast Mixture-of-Expert Training System
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
- Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply
- Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
- CoTexT: Multi-task Learning with Code-Text Transformer
- Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR
- E-Stitchup: Data Augmentation for Pre-Trained Embeddings
- Rethinking Positional Encoding
- Predicting the Order of Upcoming Tokens Improves Language Modeling
- SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- End-to-end Named Entity Recognition and Relation Extraction using Pre-trained Language Models
- Reinforced Context Order Recovery for Adaptive Reasoning and Planning
- MVP-BERT: Redesigning Vocabularies for Chinese BERT and Multi-Vocab Pretraining
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Technical report on Conversational Question Answering
- Cross-language Information Retrieval
- Structured Pruning of a BERT-based Question Answering Model
- Enhancing Rumor Detection Methods with Propagation Structure Infused Language Model
- Privacy-Preserving Tabular Synthetic Data Generation Using TabularARGN
- LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
- One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
- JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation
- User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- Domain-Specific Fine-Tuning and Prompt-Based Learning: A Comparative Study for developing Natural Language-Based BIM Information Retrieval Systems
- The Bidirectional Process Reward Model
- Beyond English-Only Reading Comprehension: Experiments in Zero-Shot\n Multilingual Transfer for Bulgarian
- Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
- Color as the Impetus: Transforming Few-Shot Learner
- List of large language models [wikipedia]
- Transformer (deep learning) [wikipedia]
- XLNet [wikipedia]
Discussions
Related