XLNet: Generalized Autoregressive Pretraining for Language Understanding
2019/06/19 by Zhilin Yang, Zihang Dai, Yang, Zhilin +9 · 1 voice · 1,857 citations
Computer Science · Mathematics · #Artificial intelligence #Autoregressive model #Computer science #Econometrics #Inference #Language model #Machine learning #Margin (machine learning) #Mathematics #Natural Language Processing Techniques #Natural language processing #Ranking (information retrieval) #Speech Recognition and Synthesis #Speech recognition #Topic Modeling #Transformer #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1906.08237
published in arXiv (Cornell University) (Cornell University) · Pretrained models and code are available at https://github.com/zihangdai/xlnet
openalex publication_date 2019/06/19 · arxiv created 2020/01/02 · arxiv updated 2020/01/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy. In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. Furthermore, XLNet integrates ideas from Transformer-XL, the state-of-the-art autoregressive model, into pretraining. Empirically, under comparable experiment settings, XLNet outperforms BERT on 20 tasks, often by a large margin, including question answering, natural language inference, sentiment analysis, and document ranking.
Citations
Cited by
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
- Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation
- Dice Loss for Data-imbalanced NLP Tasks
- Candidate Attended Dialogue State Tracking Using BERT
- JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
- Mind captioning: Evolving descriptive text of mental content from human brain activity
- A Survey on Diffusion Language Models
- PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- Enhancing Code Understanding for Impact Analysis by Combining Transformers and Program Dependence Graphs
- GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
- Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
- Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning
- TreeBERT: A Tree-Based Pre-Trained Model for Programming Language
- SELECT: Detecting Label Errors in Real-world Scene Text Data
- Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats
- When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing
- Cross-platform Product Matching Based on Entity Alignment of Knowledge Graph with RAEA model
- Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
- Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- Efficient Covariance Estimation for Sparsified Functional Data
- Spanning Tree Autoregressive Visual Generation
- Zero-Shot Grammar Competency Estimation Using Large Language Model Generated Pseudo Labels
- How to Select One Among All? An Extensive Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding
- Evaluation of Sentence Representations in Polish
- Taming Pretrained Transformers for Extreme Multi-label Text Classification
- Parallel Sampling via Autospeculation
- Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
- Learning Norms from Stories: A Prior for Value Aligned Agents
- Comparing Reconstruction Attacks on Pretrained Versus Full Fine-tuned Large Language Model Embeddings on Homo Sapiens Splice Sites Genomic Data
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge
- Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction
- Reversal Invariance in Autoregressive Language Models
- Enhancing Sentiment Classification with Machine Learning and Combinatorial Fusion
- Disentangling Adaptive Gradient Methods from Learning Rates
- ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
- Portuguese Named Entity Recognition using BERT-CRF
- Efficient Attention: Attention with Linear Complexities
- Position Masking for Language Models
- Factorized Multimodal Transformer for Multimodal Sequential Learning
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- One-shot Key Information Extraction from Document with Deep Partial Graph Matching
- IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
- Enabling Language Models to Fill in the Blanks
- Multiple Structural Priors Guided Self Attention Network for Language Understanding
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- A review on the applications of Transformer-based language models for nucleotide sequence analysis
- Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
- Symmetric Regularization based BERT for Pair-wise Semantic Reasoning
- What makes us curious? analysis of a corpus of open-domain questions
- The Brownian motion in the transformer model
- ActBERT: Learning Global-Local Video-Text Representations
- Q8BERT: Quantized 8Bit BERT
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Overview of the TREC 2022 deep learning track
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
- TVDIM: Enhancing Image Self-Supervised Pretraining via Noisy Text Data
- RL makes MLLMs see better than SFT
- VL-BERT: Pre-training of Generic Visual-Linguistic Representations
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
- Large Language Models, Agency, and Why Speech Acts are Beyond Them (For Now) – A Kantian-Cum-Pragmatist Case
- Transfer Learning for Multi-lingual Tasks -- a Survey
- A Comprehensive Dataset for Human vs. AI Generated Text Detection
- SALSA: Single-pass Autoregressive LLM Structured Classification
- Robustness Verification for Transformers
- Complex Transformer: A Framework for Modeling Complex-Valued Sequence
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- IMB: An Italian Medical Benchmark for Question Answering
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG
- ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- DocVQA: A Dataset for VQA on Document Images
- Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- Isotropy and Geometry of Pretrained Protein LMs
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- Multimodal Foundation Models for Early Disease Detection
- PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
- SenWave: A Fine-Grained Multi-Language Sentiment Analysis Dataset Sourced from COVID-19 Tweets
- Using the Hammer Only on Nails: A Hybrid Method for Evidence Retrieval for Question Answering
- Reasoning for Hierarchical Text Classification: The Case of Patents
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Language models for longitudinal analysis of abusive content in Billboard Music Charts
- 12-in-1: Multi-Task Vision and Language Representation Learning
- Allocation of Parameters in Transformers
- MathBERT: A Pre-Trained Model for Mathematical Formula Understanding
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- Self-Speculative Masked Diffusions
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- AI Methods for Antimicrobial Peptides: Progress and Challenges
- Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructions
- Text Adversarial Attacks with Dynamic Outputs
- LV-BERT: Exploiting Layer Variety for BERT
- Understanding and Enhancing Mask-Based Pretraining towards Universal Representations
- Performance Consistency of Learning Methods for Information Retrieval Tasks
- Confidence Calibration in Large Language Model-Based Entity Matching
- Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
- THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
- Learning to Emphasize: Dataset and Shared Task Models for Selecting Emphasis in Presentation Slides
- Literature review on vulnerability detection using NLP technology
- Transformers and genome language models
- Evolving Character-level Convolutional Neural Networks for Text Classification
- TRANS-BLSTM: Transformer with Bidirectional LSTM for Language Understanding
- A study of word embedding models for measuring topic coherence
- Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT
- Optimizing Informer with Whale Optimization Algorithm for Enhanced Ship Trajectory Prediction
- Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
- DRES: Fake news detection by dynamic representation and ensemble selection
- A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse
- Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
- Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
- A Comparative Study of Transformer-Based Language Models on Extractive Question Answering
- Improve Transformer Models with Better Relative Position Embeddings
- Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications
- A Survey of Deep Learning for Scientific Discovery
- On Linear Identifiability of Learned Representations
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- Explaining Question Answering Models through Text Generation
- Efficient Softmax Approximation for Deep Neural Networks with Attention Mechanism
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
- IITkgp at FinCausal 2020, Shared Task 1: Causality Detection using Sentence Embeddings in Financial Reports
- A study of Turkish emotion classification with pretrained language models
- Long Context Automated Essay Scoring with Language Models
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- Beyond Token Limits: Assessing Language Model Performance on Long Text Classification
- Training Multilingual Pre-trained Language Model with Byte-level Subwords
- Hierarchical Bracketing Encodings Work for Dependency Graphs
- Diverse Image Inpainting with Bidirectional and Autoregressive Transformers
- Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- Are Transformers universal approximators of sequence-to-sequence functions?
- FastMoE: A Fast Mixture-of-Expert Training System
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
- Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply
- Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- CoTexT: Multi-task Learning with Code-Text Transformer
- Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR
- E-Stitchup: Data Augmentation for Pre-Trained Embeddings
- Rethinking Positional Encoding
- Predicting the Order of Upcoming Tokens Improves Language Modeling
- SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- End-to-end Named Entity Recognition and Relation Extraction using Pre-trained Language Models
- Reinforced Context Order Recovery for Adaptive Reasoning and Planning
- MVP-BERT: Redesigning Vocabularies for Chinese BERT and Multi-Vocab Pretraining
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Technical report on Conversational Question Answering
- Cross-language Information Retrieval
- PatentTransformer-2: Controlling Patent Text Generation by Structural Metadata
- Structured Pruning of a BERT-based Question Answering Model
- Enhancing Rumor Detection Methods with Propagation Structure Infused Language Model
- Privacy-Preserving Tabular Synthetic Data Generation Using TabularARGN
- Enhancing second language speaking assessment: Integrating large language models for Finnish and Finland Swedish proficiency scoring
- LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
- One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
- Pre-Trained and Attention-Based Neural Networks for Building Noetic Task-Oriented Dialogue Systems
- JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation
- User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- Domain-Specific Fine-Tuning and Prompt-Based Learning: A Comparative Study for developing Natural Language-Based BIM Information Retrieval Systems
- The Bidirectional Process Reward Model
- Beyond English-Only Reading Comprehension: Experiments in Zero-Shot Multilingual Transfer for Bulgarian
- Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
- Color as the Impetus: Transforming Few-Shot Learner
- Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
- Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations
- A novel language model for predicting serious adverse event results in clinical trials from their prospective registrations
- Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
- Explicit Pairwise Word Interaction Modeling Improves Pretrained Transformers for English Semantic Similarity Tasks
- BERTScore: Evaluating Text Generation with BERT
- How Can BERT Help Lexical Semantics Tasks?
- SpecBPP: A Self-Supervised Learning Approach for Hyperspectral Representation and Soil Organic Carbon Estimation
- Small-Bench NLP: Benchmark for small single GPU trained models in Natural Language Processing
- Sketch-BERT: Learning Sketch Bidirectional Encoder Representation from Transformers by Self-supervised Learning of Sketch Gestalt
- Reducing Transformer Depth on Demand with Structured Dropout
- Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection
- OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
- Musical Speech: A Transformer-based Composition Tool
- memeBot: Towards Automatic Image Meme Generation
- Incorporating BERT into Neural Machine Translation
- A Survey of Quantum Theory Inspired Approaches to Information Retrieval
- ProTo: Program-Guided Transformer for Program-Guided Tasks
- Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
- DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
- Natural Language Inference in Context -- Investigating Contextual Reasoning over Long Texts
- Combining Language and Topic Models for Hierarchical Text Classification
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses
- Unsupervised Keyphrase Extraction by Jointly Modeling Local and Global Context
- A Survey on Text Classification: From Shallow to Deep Learning
- A neural document language modeling framework for spoken document retrieval
- Pre-Trained Models: Past, Present and Future
- Multimodal Representation Alignment for Cross-modal Information Retrieval
- Neural, Symbolic and Neural-Symbolic Reasoning on Knowledge Graphs
- Billion-scale Pre-trained E-commerce Product Knowledge Graph Model
- The Eighth Dialog System Technology Challenge
- Multi-label Classification for Automatic Tag Prediction in the Context of Programming Challenges
- From BERT to Qwen: Hate Detection across architectures
- Holistix: A Dataset for Holistic Wellness Dimensions Analysis in Mental Health Narratives
- Fine-tuning Pre-trained Contextual Embeddings for Citation Content Analysis in Scholarly Publication
- Optimizing Deeper Transformers on Small Datasets
- Emergent Natural Language with Communication Games for Improving Image Captioning Capabilities without Additional Data
- A document is worth a structured record: Principled inductive bias design for document recognition
- Mitigating Shortcut Learning with InterpoLated Learning
- A Comparison of LSTM and BERT for Small Corpus
- Multi-node Bert-pretraining: Cost-efficient Approach
- OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
- Graph Collaborative Attention Network for Link Prediction in Knowledge Graphs
- Occlusion Aware Kernel Correlation Filter Tracker using RGB-D
- A Short Survey of Pre-trained Language Models for Conversational AI-A NewAge in NLP
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning
- LEDOM: Reverse Language Model
- Why Multi-Interest Fairness Matters: Hypergraph Contrastive Multi-Interest Learning for Fair Conversational Recommender System
- Edit Flows: Flow Matching with Edit Operations
- Applications and Modeling of Keystroke Logs in Writing Assessments
- CORD19STS: COVID-19 Semantic Textual Similarity Dataset
- An Experimental Evaluation of Transformer-based Language Models in the Biomedical Domain
- Normalization of Input-output Shared Embeddings in Text Generation Models
- Offensive Language Detection on Social Media Using XLNet
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
- Health App Reviews for Privacy & Trust (HARPT): A Corpus for Analyzing Patient Privacy Concerns, Trust in Providers and Trust in Applications
- Health Sentinel: An AI Pipeline For Real-time Disease Outbreak Detection
- Focus Your Attention: Towards Data-Intuitive Lightweight Vision Transformers
- Efficient Nearest Neighbor Language Models
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
- Enhancement Report Approval Prediction: A Comparative Study of Large Language Models
- Unsupervised Domain Adaptation of Language Models for Reading Comprehension
- Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis
- CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model
- Assessing the Limits of In-Context Learning beyond Functions using Partially Ordered Relation
- Temporal FiLM: Capturing Long-Range Sequence Dependencies with Feature-Wise Modulations
- INTERPOS: Interaction Rhythm Guided Positional Morphing for Mobile App Recommender Systems
- Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering
- Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
- Corrector Sampling in Language Models
- Invenio: Discovering Hidden Relationships Between Tasks/Domains Using Structured Meta Learning
- RecGPT: A Foundation Model for Sequential Recommendation
- A Two-Sample Test of Text Generation Similarity
- SoK: Are Watermarks in LLMs Ready for Deployment?
- CoFrNets: Interpretable Neural Architecture Inspired by Continued Fractions
- A MISMATCHED Benchmark for Scientific Natural Language Inference
- GlobalMind: Global multi-head interactive self-attention network for hyperspectral change detection
- A Deep Recurrent Survival Model for Unbiased Ranking
- Pre-trained Language Model Based Active Learning for Sentence Matching
- BERT Embeddings Can Track Context in Conversational Search
- Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots
- DFraud3- Multi-Component Fraud Detection freeof Cold-start
- ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
- HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models
- Zero-Resource Knowledge-Grounded Dialogue Generation
- Towards Domain Adaptation from Limited Data for Question Answering Using Deep Neural Networks
- An Empirical Investigation of Contextualized Number Prediction
- PipeMare: Asynchronous Pipeline Parallel DNN Training
- A Survey on Long-Tailed Visual Recognition
- Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs
- Xinyu AI Search: Enhanced Relevance and Comprehensive Results with Rich Answer Presentations
- Sentence Analogies: Exploring Linguistic Relationships and Regularities in Sentence Embeddings
- Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
- On VLMs for Diverse Tasks in Multimodal Meme Classification
- Exploring Fluent Query Reformulations with Text-to-Text Transformers and Reinforcement Learning
- MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection
- Real-Time Execution of Large-scale Language Models on Mobile
- Towards Learning Cross-Modal Perception-Trace Models
- Masked ELMo: An evolution of ELMo towards fully contextual RNN language models
- Adaptive Data-Resilient Multi-Modal Hierarchical Multi-Label Book Genre Identification
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and Languages
- A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports
- Partition Generative Modeling: Masked Modeling Without Masks
- Multi-Party Conversational Agents: A Survey
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision
- Attention Mechanisms in Computer Vision: A Survey
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- A Position Paper on the Automatic Generation of Machine Learning Leaderboards
- XGPT: Cross-modal Generative Pre-Training for Image Captioning
- Why Can You Lay Off Heads? Investigating How BERT Heads Transfer
- Evaluating Document Coherence Modelling
- Adapt or Get Left Behind: Domain Adaptation through BERT Language Model Finetuning for Aspect-Target Sentiment Classification
- KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding
- Bi-directional Cognitive Thinking Network for Machine Reading Comprehension
- Entailed Opinion Matters: Improving the Fact-Checking Performance of Language Models by Relying on their Entailment Ability
- Improving Diversity of Neural Text Generation via Inverse Probability Weighting
- Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review
- Bilingual Language Modeling, A transfer learning technique for Roman Urdu
- Quellcodekritik. Zur Philologie von Algorithmen
- Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
- MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark
- Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks
- Foundation Models for AI-Enabled Biological Design
- Hierarchical Bracketing Encodings for Dependency Parsing as Tagging
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- Examining the rhetorical capacities of neural language models
- Low-Shot Classification: A Comparison of Classical and Deep Transfer Machine Learning Approaches
- Larger-Context Tagging: When and Why Does It Work?
- Improving speech recognition models with small samples for air traffic control systems
- Multi-Token Prediction Needs Registers
- An Effective Contextual Language Modeling Framework for Speech Summarization with Augmented Features
- Current Limitations of Language Models: What You Need is Retrieval
- An empirical study of task and feature correlations in the reuse of pre-trained models
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishing
- Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies
- Structural-Temporal Coupling Anomaly Detection with Dynamic Graph Transformer
- Memeify: A Large-Scale Meme Generation System
- Generative Deep Learning Techniques for Password Generation
- I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model Inference
- Boosting Neural Language Inference via Cascaded Interactive Reasoning
- Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions
- Alignment Attention by Matching Key and Query Distributions
- BOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models
- Hide-and-Seek: A Template for Explainable AI
- TTTTTackling WinoGrande Schemas
- A Character-based Diffusion Embedding Algorithm for Enhancing the Generation Quality of Generative Linguistic Steganographic Texts
- Cracking the Contextual Commonsense Code: Understanding Commonsense Reasoning Aptitude of Deep Contextual Representations
- Pretrained AI Models: Performativity, Mobility, and Change
- Mischief: A Simple Black-Box Attack Against Transformer Architectures
- Reweighted Proximal Pruning for Large-Scale Language Representation
- Generalized Discrete Diffusion from Snapshots
- Multi-Stage Conversational Passage Retrieval: An Approach to Fusing Term Importance Estimation and Neural Query Rewriting
- Language Model for Large-Text Transmission in Noisy Quantum Communications
- Enhancing Vulnerability Reports with Automated and Augmented Description Summarization
- Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
- The State-Prediction Separation Hypothesis
- baller2vec++: A Look-Ahead Multi-Entity Transformer For Modeling Coordinated Agents
- Syntax-driven Iterative Expansion Language Models for Controllable Text Generation
- Tree-structured Attention with Hierarchical Accumulation
- Heads-up! Unsupervised Constituency Parsing via Self-Attention Heads
- Bridging Cognition and Emotion: Empathy-Driven Multimodal Misinformation Detection
- RAGAT-Mind: A Multi-Granular Modeling Approach for Rumor Detection Based on MindSpore
- HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
- Low-Resource Knowledge-Grounded Dialogue Generation
- The Ultimate Cookbook for Invisible Poison: Crafting Subtle Clean-Label Text Backdoors with Style Attributes
- A Pairwise Probe for Understanding BERT Fine-Tuning on Machine Reading Comprehension
- TathyaNyaya and FactLegalLlama: Advancing Factual Judgment Prediction and Explanation in the Indian Legal Context
- Distilling Specialized Orders for Visual Generation
- NUIG-Shubhanker@Dravidian-CodeMix-FIRE2020: Sentiment Analysis of Code-Mixed Dravidian text using XLNet
- Sentiment Analysis in Software Engineering: Evaluating Generative Pre-trained Transformers
- VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
- Enhancing Review Comprehension with Domain-Specific Commonsense
- Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
- Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
- Generating Accurate Assert Statements for Unit Test Cases using Pretrained Transformers
- You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
- IMMENSE: Inductive Multi-perspective User Classification in Social Networks
- Looking beyond the next token
- C-MTCSD: A Chinese Multi-Turn Conversational Stance Detection Dataset
- Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
- Leveraging Linguistic Coordination in Reranking N-Best Candidates For End-to-End Response Selection Using BERT
- SapiensID: Foundation for Human Recognition
- Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
- List of large language models [wikipedia]
- Transformer (deep learning) [wikipedia]
- XLNet [wikipedia]
Discussions
Related