Distributed Representations of Sentences and Documents
2014/05/16 by Quoc V. Le, Tomas Mikolov, Tomáš Mikolov +2 · 4 voices · 149 citations
Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.1405.4053
Abstract
Many machine learning algorithms require the input to be represented as a fixed-length feature vector. When it comes to texts, one of the most common fixed-length features is bag-of-words. Despite their popularity, bag-of-words features have two major weaknesses: they lose the ordering of the words and they also ignore semantics of the words. For example, "powerful," "strong" and "Paris" are equally distant. In this paper, we propose Paragraph Vector, an unsupervised algorithm that learns fixed-length feature representations from variable-length pieces of texts, such as sentences, paragraphs, and documents. Our algorithm represents each document by a dense vector which is trained to predict words in the document. Its construction gives our algorithm the potential to overcome the weaknesses of bag-of-words models. Empirical results show that Paragraph Vectors outperform bag-of-words models as well as other techniques for text representations. Finally, we achieve new state-of-the-art results on several text classification and sentiment analysis tasks.
Citations
Cited by
- Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
- Enhancing Code Understanding for Impact Analysis by Combining Transformers and Program Dependence Graphs
- Citation importance-aware document representation learning for large-scale science mapping
- Data-driven inverse uncertainty quantification: application to the Chemical Vapor Deposition Reactor Modeling
- COIVis: Eye-tracking-based Visual Exploration of Concept Learning in MOOC Videos
- ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding
- Estimating Early Fundraising Performance of Innovations via Graph-based Market Environment Model
- Retrieving Semantically Similar Decisions under Noisy Institutional Labels: Robust Comparison of Embedding Methods
- Neural Attentive Bag-of-Entities Model for Text Classification
- Learning Transferable Visual Models From Natural Language Supervision
- Provably Secure Generative Linguistic Steganography
- deepFEPS: Deep Learning-Oriented Feature Extraction for Biological Sequences
- Bot-Match: Social Bot Detection with Recursive Nearest Neighbors Search
- Wikibook-Bot - Automatic Generation of a Wikipedia Book
- Feature extraction using Latent Dirichlet Allocation and Neural\n Networks: A case study on movie synopses
- Visual Framing of Science Conspiracy Videos: Integrating Machine Learning with Communication Theories to Study the Use of Color and Brightness
- American social media users have ideological differences of opinion about the War in Ukraine
- Structural-Aware Sentence Similarity with Recursive Optimal Transport
- Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space
- Two-path Deep Semi-supervised Learning for Timely Fake News Detection
- Parameter-Efficient Transfer Learning with Diff Pruning
- Multimodal Prediction based on Graph Representations
- Streaming Social Event Detection and Evolution Discovery in Heterogeneous Information Networks
- Inclusion of Role into Named Entity Recognition and Ranking
- Prototype Selection Using Topological Data Analysis
- KScaNN: Scalable Approximate Nearest Neighbor Search on Kunpeng
- IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
- LAGOS-AND: A Large Gold Standard Dataset for Scholarly Author Name Disambiguation
- From Word Embeddings to Item Recommendation
- Harnessing Deep Neural Networks with Logic Rules
- Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations
- REMOD: Relation Extraction for Modeling Online Discourse
- Exploiting Behavioral Consistence for Universal User Representation
- baller2vec: A Multi-Entity Transformer For Multi-Agent Spatiotemporal Modeling
- Corruption Is Not All Bad: Incorporating Discourse Structure into Pre-training via Corruption for Essay Scoring
- RiTUAL-UH at TRAC 2018 Shared Task: Aggression Identification
- Latent Variable Modeling with Diversity-Inducing Mutual Angular Regularization
- Image Captioning and Visual Question Answering Based on Attributes and External Knowledge
- You Shall Know a User by the Company It Keeps: Dynamic Representations\n for Social Media Users in NLP
- JTAV: Jointly Learning Social Media Content Representation by Fusing\n Textual, Acoustic, and Visual Features
- Selfie: Self-supervised Pretraining for Image Embedding
- Deep Learning in Protein Structural Modeling and Design
- Review Regularized Neural Collaborative Filtering
- Learning Blended, Precise Semantic Program Embeddings
- Graph Convolutional Network for Swahili News Classification
- CoAID: COVID-19 Healthcare Misinformation Dataset
- Real-Time Steganalysis for Stream Media Based on Multi-channel Convolutional Sliding Windows
- Document Embedding for Scientific Articles: Efficacy of Word Embeddings vs TFIDF
- Phishing Detection through Email Embeddings
- Job relatedness, local skill coherence and economic performance: a job postings approach
- Towards Making the Most of BERT in Neural Machine Translation
- Investigating Meta-Learning Algorithms for Low-Resource Natural Language Understanding Tasks
- A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair
- Skip-Thought Vectors
- Offensive Language and Hate Speech Detection for Danish
- Wasserstein-Fisher-Rao Document Distance
- Fine-grained Event Categorization with Heterogeneous Graph Convolutional Networks
- Learning Action Models from Disordered and Noisy Plan Traces
- Deep Graph Similarity Learning: A Survey
- From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction
- Fake Reviews Detection through Ensemble Learning
- Unsupervised Graph Representation by Periphery and Hierarchical Information Maximization
- SAFE: Self-Attentive Function Embeddings for Binary Similarity
- SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning
- Log-based Anomaly Detection Without Log Parsing
- Audio-Linguistic Embeddings for Spoken Sentences
- Evaluating Bayesian Deep Learning Methods for Semantic Segmentation
- A Bug or a Suggestion? An Automatic Way to Label Issues
- Natural language analysis of the structure of altered states of consciousness
- A Robust Classification Method using Hybrid Word Embedding for Early Diagnosis of Alzheimer's Disease
- What is missing from this picture? Persistent homology and mixup barcodes as a means of investigating negative embedding space
- DeepHunter: A Graph Neural Network Based Approach for Robust Cyber Threat Hunting
- MIARec: Mutual-influence-aware Heterogeneous Network Embedding for Scientific Paper Recommendation
- Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood Embedding
- RelEmb: A relevance-based application embedding for Mobile App retrieval\n and categorization
- GrASP: A Generalizable Address-based Semantic Prefetcher for Scalable Transactional and Analytical Workloads
- Dialogue Act Classification with Context-Aware Self-Attention
- Utilizing FastText for Venue Recommendation
- Analysis of Moral Judgement on Reddit
- Towards a Flexible Embedding Learning Framework
- Supervised and Semi-Supervised Text Categorization using LSTM for Region Embeddings
- Equation Embeddings
- GT-SEER: Geo-Temporal SEquential Embedding Rank for Point-of-interest Recommendation
- Not All Bugs Are the Same: Understanding, Characterizing, and Classifying the Root Cause of Bugs
- Multi-Category Materials Information Extraction (Composition, Processing, Microstructure, Properties)
- PEARL: Performance-Enhanced Aggregated Representation Learning
- A multi-label classification method using a hierarchical and transparent representation for paper-reviewer recommendation
- Self-supervised Learning on Graphs: Deep Insights and New Direction
- An Improved Framework for Scaling Party Positions from Texts with Transformer
- Sequential Learning of Convolutional Features for Effective Text\n Classification
- Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
- Speeding up Word Mover's Distance and its variants via properties of distances between embeddings
- Parameter-Efficient Transfer Learning for NLP
- An end-to-end Neural Network Framework for Text Clustering
- Automated Discovery and Classification of Training Videos for Career Progression
- miGAP: miRNA–Gene Association Prediction Method Based on Deep Learning Model
- Towards Successful Social Media Advertising: Predicting the Influence of Commercial Tweets
- Improving Clinical Outcome Predictions Using Convolution over Medical Entities with Multimodal Learning
- Emerging App Issue Identification via Online Joint Sentiment-Topic Tracing
- Multi-view and Multi-source Transfers in Neural Topic Modeling with Pretrained Topic and Word Embeddings
- Generating Realistic Sequences of Customer-level Transactions for Retail Datasets
- Predicting the descent into extremism and terrorism
- Correlation Coefficients and Semantic Textual Similarity
- MIC: Model-agnostic Integrated Cross-channel Recommenders
- An Unsupervised Domain-Independent Framework for Automated Detection of Persuasion Tactics in Text
- Normalisation of SWIFT Message Counterparties with Feature Extraction and Clustering
- Adversarial Training Based Multi-Source Unsupervised Domain Adaptation for Sentiment Analysis
- An Intelligent CNN-VAE Text Representation Technology Based on Text Semantics for Comprehensive Big Data
- Twitter Sentiment on Affordable Care Act using Score Embedding
- Drug Package Recommendation via Interaction-aware Graph Induction
- Causal evidence of racial and institutional biases in accessing paywalled articles and scientific data
- Dividing and Conquering Cross-Modal Recipe Retrieval: from Nearest Neighbours Baselines to SoTA
- An Edit-centric Approach for Wikipedia Article Quality Assessment
- Towards Reliable Online Clickbait Video Detection: A Content-Agnostic Approach
- Similarity-Based Supervised User Session Segmentation Method for Behavior Logs
- "A Passage to India": Pre-trained Word Embeddings for Indian Languages
- Detect All Abuse! Toward Universal Abusive Language Detection Models
- Semantic Classification of Tabular Datasets via Character-Level Convolutional Neural Networks
- P-SIF: Document Embeddings Using Partition Averaging
- Deep Learning Based Concurrency Bug Detection and Localization
- Text mining policy: Classifying forest and landscape restoration policy agenda with neural information retrieval
- CUSATNLP@HASOC-Dravidian-CodeMix-FIRE2020:Identifying Offensive Language from ManglishTweets
- A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMs
- Inductive Subgraph Embedding for Link Prediction
- End-to-End Resume Parsing and Finding Candidates for a Job Description using BERT
- DEFENDCLI: Command-Line Driven Attack Provenance Examination
- Towards the Next-generation Bayesian Network Classifiers
- Fedora And Debian Software Package Dependency Networks Along With Description Text Associated With Nodes
- SoulMate: Short-text author linking through Multi-aspect temporal-textual embedding
- Learning Dense Representations of Phrases at Scale
- A New Approach for Topic Detection using Adaptive Neural Networks
- Mind the Nuisance: Gaussian Process Classification using Privileged\n Noise
- Representation Learning for Words and Entities
- Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
- Compressed Deep Networks: Goodbye SVD, Hello Robust Low-Rank Approximation
- WikiContradiction: Detecting Self-Contradiction Articles on Wikipedia
- Graph Classification Based on Skeleton and Component Features
- Using Paragraph Vectors to improve our existing code review assisting\n tool-CRUSO
- Comparison of Information Retrieval Techniques Applied to IT Support Tickets
- PROVCREATOR: Synthesizing Complex Heterogenous Graphs with Node and Edge Attributes
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Beyond Social Fragmentation: Coexistence of Cultural Diversity and Structural Connectivity Is Possible with Social Constituent Diversity
- An empirical comparison of deep-neural-network architectures for next activity prediction using context-enriched process event logs
- Types of artificial neural networks [wikipedia]
- Word2vec [wikipedia]
- Feature learning [wikipedia]
- Entity linking [wikipedia]
- Quoc V. Le [wikipedia]
Discussions
Related