Transformers: State-of-the-Art Natural Language Processing
2020/01/01 by Thomas Wolf, Lysandre Debut, Victor Sanh +19 · 52 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
paper · pdf · doi:10.18653/v1/2020.emnlp-demos.6
openalex publication_date 2020/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, Alexander Rush. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 2020.
Cited by
- Immersive exposure to simulated visual hallucinations modulates high-level human cognition
- Gender biases and hate speech: Promoters and targets in the Argentinean political context
- SciMON: Scientific Inspiration Machines Optimized for Novelty
- Sensor movement drives emergent attention and scalability in active neural cellular automata
- Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study
- Fine-tuning Small Language Models (SLMs) for autonomous web-based geographical information systems (AWebGIS)
- Rethinking Data Use in Large Language Models
- AgentTOD: A Task-Oriented Dialogue Agent with a Flexible and Adaptive API Calling Paradigm
- GPT-GNN: Generative Pre-Training of Graph Neural Networks
- Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
- Out of Context: How important is Local Context in Neural Program Repair?
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- Linguistic Characterization of Divisive Topics Online: Case Studies on Contentiousness in Abortion, Climate Change, and Gun Control
- Tokenization Changes Meaning in Large Language Models: Evidence from Chinese
- The impact of tokenizer selection in genomic language models
- Assessing and Understanding Creativity in Large Language Models
- Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
- LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions
- Social Processes in the Intensification of Online Hate: The Effects of Verbal Replies to Anti-Muslim and Anti-Jewish Posts Following 7 October 2023
- Algorithmic reproduction of social inequality: language attitude bias in Chinese and English pre-trained language models
- Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior
- Voice user interfaces for effortless navigation in medical virtual reality environments
- Transformers enable accurate prediction of acute and chronic chemical toxicity in aquatic organisms
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- A short trajectory is all you need: A transformer-based model for long-time dissipative quantum dynamics
- Multiple agroecological practices use and climate change mitigation. A review
- Investigating Idiomaticity in Word Representations
- Literature review on vulnerability detection using NLP technology
- RetroSynFormer: Planning multi-step chemical synthesis routes via a Decision Transformer
- Rice Yield Prediction and Model Interpretation Based on Satellite and Climatic Indicators Using a Transformer Method
- Assessing the reproducibility of a bioimage analysis workflow characterising tissue flow in <i>Drosophila</i>
- Adversarial Sequence Mutations in AlphaFold and ESMFold Reveal Nonphysical Structural Invariance, Confidence Failures, and Concerns for Protein Design
- Towards Effective and Efficient Sparse Neural Information Retrieval
- Divergences between Language Models and Human Brains
- REFINER: Reasoning Feedback on Intermediate Representations
- Behavioral Homophily in Social Media via Inverse Reinforcement Learning: A Reddit Case Study
- Investigating Transfer Learning Capabilities of Vision Transformers and CNNs by Fine-Tuning a Single Trainable Block
- Generating language assessment content free from representational harms
- Enhancing second language speaking assessment: Integrating large language models for Finnish and Finland Swedish proficiency scoring
- Jointly Extracting Interventions, Outcomes, and Findings from RCT Reports with LLMs
- Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models
- Experimental Standards for Deep Learning in Natural Language Processing Research
- Longformer for MS MARCO Document Re-ranking Task
- Improving Fine-Grained Emotion Detection in Text with BERT and GoEmotions
- Improving Fine-Grained Emotion Detection in Text with BERT and GoEmotions: An Experimental Study
- Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys
- Bayesian Surprise Predicts Human Event Segmentation in Story Listening
- Modal and scenario noise reduction for multimodal sarcasm detection
- Hippocampo-neocortical interaction as compressive retrieval-augmented generation
- Generative Models for Security: Attacks, Defenses, and Opportunities
- Turn-level Dialog Evaluation with Dialog-level Weak Signals for Bot-Human Hybrid Customer Service Systems
- Decoding customer satisfaction in Michelin Green Star restaurants: BERTopic modelling and sentiment analysis
Related