A vector space model for automatic indexing
1975/11/01 by G. Salton, Gerard Salton, A. Wong +3 · 7,464 citations
Computer Science · Mathematics · #Algorithm #Computation #Computer science #Data Management and Algorithms #Data Mining Algorithms and Applications #Data mining #Function (biology) #Image Retrieval and Classification Techniques #Information retrieval #Matching (statistics) #Mathematics #Property (philosophy) #Search engine indexing #Space (punctuation) #Vector space model
paper · pdf · doi:10.1145/361219.361220
published in Communications of the ACM 18(11), 613-620 (Association for Computing Machinery)
openalex publication_date 1975/11/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Abstract
In a document retrieval, or other pattern matching environment where stored entities (documents) are compared with each other or with incoming patterns (search requests), it appears that the best indexing (property) space is one where each entity lies as far away from the others as possible; in these circumstances the value of an indexing system may be expressible as a function of the density of the object space; in particular, retrieval performance may correlate inversely with space density. An approach based on space density computations is used to choose an optimum indexing vocabulary for a collection of documents. Typical evaluation results are shown, demonstating the usefulness of the model.
Cited by
- Large‐Sample Evidence on Firms’ Year‐over‐Year MD&A Modifications
- Algorithmic Trading and Forward‐Looking MD&A Disclosures
- Word2Vec vs DBnary: Augmenting METEOR using Vector Representations or Lexical Resources?
- Research Progress of News Recommendation Methods
- Beyond Cosine Similarity
- FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
- A Connection-Centric Survey of Recommender Systems Research
- Statistical Automatic Summarization in Organic Chemistry
- GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Neural Embeddings of Scholarly Periodicals Reveal Complex Disciplinary Organizations
- amc: The Automated Mission Classifier for Telescope Bibliographies
- Network-Clustered Multi-Modal Bug Localization
- How Far Are We from Genuinely Useful Deep Research Agents?
- A Survey on Malware Detection Using Data Mining Techniques
- MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
- Feature extraction using Latent Dirichlet Allocation and Neural Networks: A case study on movie synopses
- A Survey On Neural Word Embeddings
- Leveraging Deep Neural Networks and Knowledge Graphs for Entity Disambiguation
- Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
- Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
- IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
- AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding
- On the effectiveness of feature set augmentation using clusters of word embeddings
- Exploring term-document matrices from matrix models in text mining
- GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
- Discovering Political Topics in Facebook Discussion threads with Graph Contextualization
- An alternative text representation to TF-IDF and Bag-of-Words
- WeSSQoS: A Configurable SOA System for Quality-aware Web Service Selection
- Automatic Code Summarization: A Systematic Literature Review
- Evaluating Methods to Rediscover Missing Web Pages from the Web Infrastructure
- A Survey on Metric Learning for Feature Vectors and Structured Data
- Coevolution of Network Structure and Content
- Semantically enhanced pseudo relevance feedback for Arabic information retrieval
- Making Sense of Word Embeddings
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- RMD: Robust Modal Decomposition with Constrained Bandwidth
- On Strategyproof Conference Peer Review
- Sherlock Your Queries: Learning to Ask the Right Questions for Dialogue-Based Retrieval
- Dense Text Retrieval Based on Pretrained Language Models: A Survey
- Generative AI and College Students: Use and Perceptions
- Distributional Semantics: Meaning Through Culture and Interaction
- The effects of multiple query evidences on social image retrieval
- The Impact of Classifier Configuration and Classifier Combination on Bug Localization
- RelEmb: A relevance-based application embedding for Mobile App retrieval\n and categorization
- Detecting Hallucinations in Authentic LLM-Human Interactions
- Ornitología Virtual: Caracterizando a #Chile en Twitter
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- Recurrent Binary Embedding for GPU-Enabled Exhaustive Retrieval from Billion-Scale Semantic Vectors
- From keywords to semantics: Perceptions of large language models in data discovery
- ALARB: An Arabic Legal Argument Reasoning Benchmark
- ACT: Agentic Classification Tree
- Neural Language Priors
- Policy Influence and Private Returns from Lobbying in the Energy Sector
- Unsupervised record matching with noisy and incomplete data
- Modeling Document Interactions for Learning to Rank with Regularized Self-Attention
- Enhancing LLM-based Fault Localization with a Functionality-Aware Retrieval-Augmented Generation Framework
- Related Fact Checks: a tool for combating fake news
- From Recommendation Systems to Facility Location Games
- Semantic Search by Latent Ontological Features
- How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
- Similarity Field Theory: A Mathematical Framework for Intelligence
- Calculating Semantic Similarity between Academic Articles using Topic Event and Ontology
- CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completion
- Large-Scale Cover Song Detection in Digital Music Libraries Using Metadata, Lyrics and Audio Features
- Compositional Concept Generalization with Variational Quantum Circuits
- LLM Architecture, Scaling Laws, and Economics: A Quick Summary
- Towards computational fluorescence microscopy: Machine learning-based integrated prediction of morphological and molecular tumor profiles
- Search-Based Software Bugs Localization, Triage and Prioritization
- A Smoothed Dual Approach for Variational Wasserstein Problems
- An Enhanced Model-based Approach for Short Text Clustering
- Improving Code Summarization with Block-wise Abstract Syntax Tree Splitting
- Explainable Graph Spectral Clustering For GloVe-like Text Embeddings
- Analyzing User Activities Using Vector Space Model in Online Social Networks
- Multiple Metric Learning for Structured Data
- Scaling Personality Control in LLMs with Big Five Scaler Prompts
- Improving Persian Document Classification Using Semantic Relations between Words
- Evaluating, Synthesizing, and Enhancing for Customer Support Conversation
- Exploring the GIS Knowledge Domain Using CiteSpace
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- Text Categorization via Similarity Search: An Efficient and Effective Novel Algorithm
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Semi-orthogonal Non-negative Matrix Factorization with an Application in Text Mining
- Semantic-Sensitive Web Information Retrieval Model for HTML Documents
- On the Estimation and Use of Statistical Modelling in Information Retrieval
- LLM-based Embedders for Prior Case Retrieval
- A Brief Survey of Text Mining: Classification, Clustering and Extraction Techniques
- Incorporating Semantic Knowledge into Latent Matching Model in Search
- A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
- LScDC-new large scientific dictionary
- Indexing by Latent Dirichlet Allocation and Ensemble Model
- Down the Rabbit Hole: Robust Proximity Search and Density Estimation in Sublinear Space
- Vector representations of text data in deep learning
- A Density-Based Approach to the Retrieval of Top-K Spatial Textual Clusters
- Music Information Retrieval: Recent Developments and Applications
- MedSTS: A Resource for Clinical Semantic Textual Similarity
- Non-distributive logics: from semantics to meaning
- A neural document language modeling framework for spoken document retrieval
- Learning Semantic Sentence Embeddings using Sequential Pair-wise Discriminator
- An Information Retrieval Approach to Short Text Conversation
- InsurTech innovation using natural language processing
- Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval
- SemRAG: Semantic Knowledge-Augmented RAG for Improved Question-Answering
- Poly-encoders: Transformer Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring
- Large Language Model for Extracting Complex Contract Information in Industrial Scenes
- Semantic Analysis for Automated Evaluation of the Potential Impact of Research Articles
- Keyphrase Annotation with Graph Co-Ranking
- MyTrackingChoices: Pacifying the Ad-Block War by Enforcing User Privacy Preferences
- Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
- Predicting missing links via correlation between nodes
- Curating art exhibitions using machine learning
- Team LA at SCIDOCA shared task 2025: Citation Discovery via relation-based zero-shot retrieval
- Harnessing the power of Social Bookmarking for improving tag-based Recommendations
- Neural Prioritisation for Web Crawling
- The model of information retrieval based on the theory of hypercomplex numerical systems
- Enhancing software requirements classification: a comparative study of deep learning model integration with embedding techniques
- A Practical Guide for Evaluating LLMs and LLM-Reliant Systems
- Quizbowl: The Case for Incremental Question Answering
- TF-IDFC-RF: A Novel Supervised Term Weighting Scheme
- Sentiment Analysis for Education with R: packages, methods and practical applications
- Infusing Collaborative Recommenders with Distributed Representations
- Multi-owner Secure Encrypted Search Using Searching Adversarial Networks
- 2P-Med: Building a Personalization Platform for Mediation Systems
- Semantic, Efficient, and Secure Search over Encrypted Cloud Data
- A Deep Look into Neural Ranking Models for Information Retrieval
- A Review of Behavioral Closed-Loop Paradigm from Sensing to Intervention for Ingestion Health
- C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models
- Automating App Review Response Generation
- A Review on Human-Computer Interaction and Intelligent Robots
- Prioritizing documentation effort: Can we do better?
- Measures of Cluster Informativeness for Medical Evidence Aggregation and Dissemination
- TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
- Exploring Semantic Incrementality with Dynamic Syntax and Vector Space Semantics
- Should Semantic Vector Composition be Explicit? Can it be Linear?
- Connecting Independently Trained Modes via Layer-Wise Connectivity
- A Slicing-Based Approach for Detecting and Patching Vulnerable Code Clones
- Refining Neural Activation Patterns for Layer-Level Concept Discovery in Neural Network-Based Receivers
- An Unsupervised Language-Independent Entity Disambiguation Method and its Evaluation on the English and Persian Languages
- Model Compression with Multi-Task Knowledge Distillation for Web-scale Question Answering System
- CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
- MAT: A simple yet strong baseline for identifying self-admitted technical debt
- Batched Self-Consistency Improves LLM Relevance Assessment and Ranking
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
- MyAdChoices: Bringing Transparency and Control to Online Advertising
- The Counting Power of Transformers
- Visualizing knowledge domains
- Linguistic Matrix Theory
- Guiding Data Collection via Factored Scaling Curves
- Latent Dirichlet Allocation in R
- Generative Interest Estimation for Document Recommendations
- Large scale biomedical texts classification: a kNN and an ESA-based approaches
- Code-Enhanced Cross-Perspective Bug Question Retrieval
- From Frequency to Meaning: Vector Space Models of Semantics
- G-Bean: an ontology-graph based web tool for biomedical literature retrieval
- A case for applying an abstracted quantum formalism to cognition
- Combinatorial Spaces And Order Topologies
- Topic Modeling the Reading and Writing Behavior of Information Foragers
- The History of Information Retrieval Research
- Bias-aware news analysis using matrix-based news aggregation
- ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections
- ImproBR: Bug Report Improver Using LLMs
- Multi-sense embeddings through a word sense disambiguation process
- Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
- Punjabi Documents Clustering System
- Exploring Diversity, Novelty, and Popularity Bias in ChatGPT's Recommendations
- Inverted files for text search engines
- High-dimensional distributed semantic spaces for utterances
- Dependency-Based Construction of Semantic Space Models
- Chinese-Portuguese Machine Translation: A Study on Building Parallel Corpora from Comparable Texts
- Improving Update Summarization by Revisiting the MMR Criterion
- Using Large Language Models to Support Automation of Failure Management in CI/CD Pipelines: A Case Study in SAP HANA
- Patent Analytics Based on Feature Vector Space Model: A Case of IoT
- An extensive study on smell-aware bug localization
- Machine learning in automated text categorization
- The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
- Cognitive Database: A Step towards Endowing Relational Databases with Artificial Intelligence Capabilities
- An Efficient Approach to Learning Chinese Judgment Document Similarity Based on Knowledge Summarization
- Supervised Metric Learning with Generalization Guarantees
- Exploiting a comparability mapping to improve bi-lingual data categorization: a three-mode data analysis perspective
- Comparing heterogeneous entities using artificial neural networks of trainable weighted structural components and machine-learned activation functions
- "What is relevant in a text document?": An interpretable machine learning approach
- Multiclass Classification for Self-Admitted Technical Debt Based on XGBoost
- On Linear Representations and Pretraining Data Frequency in Language Models
- Privacy-Preserving Distributed Link Predictions Among Peers in Online Classrooms Using Federated Learning
- TACOA: taxonomic classification of environmental genomic fragments using a kernelized nearest neighbor approach. [europepmc]
- Ranked retrieval of Computational Biology models. [europepmc]
- Development of a classification scheme for disease-related enzyme information. [europepmc]
- Combining position weight matrices and document-term matrix for efficient extraction of associations of methylated genes and diseases from free text. [europepmc]
- An effective method of large scale ontology matching. [europepmc]
- Learning Effective Connectivity Network Structure from fMRI Data Based on Artificial Immune Algorithm. [europepmc]
- A novel procedure on next generation sequencing data analysis using text mining algorithm. [europepmc]
- Semantic Health Knowledge Graph: Semantic Integration of Heterogeneous Medical Knowledge and Services. [europepmc]
- NoGOA: predicting noisy GO annotations using evidences and sparse representation. [europepmc]
- Protein Function Prediction Using Deep Restricted Boltzmann Machines. [europepmc]
- "What is relevant in a text document?": An interpretable machine learning approach. [europepmc]
- InfAcrOnt: calculating cross-ontology term similarities using information flow by a random walk. [europepmc]
- Automatic extraction of informal topics from online suicidal ideation. [europepmc]
- MBOSS: A Symbolic Representation of Human Activity Recognition Using Mobile Sensors. [europepmc]
- Short-term plasticity at cerebellar granule cell to molecular layer interneuron synapses expands information processing. [europepmc]
- Detection of medical text semantic similarity based on convolutional neural network. [europepmc]
- Hate speech detection: Challenges and solutions. [europepmc]
- Efficient Reuse of Natural Language Processing Models for Phenotype-Mention Identification in Free-text Electronic Medical Records: A Phenotype Embedding Approach. [europepmc]
- Hyperalignment: Modeling shared information encoded in idiosyncratic cortical topographies. [europepmc]
- Hard for humans, hard for machines: predicting readmission after psychiatric hospitalization using narrative notes. [europepmc]
- Disease Concept-Embedding Based on the Self-Supervised Method for Medical Information Extraction from Electronic Health Records and Disease Retrieval: Algorithm Development and Validation Study. [europepmc]
- Multi-Similarities Bilinear Matrix Factorization-Based Method for Predicting Human Microbe-Disease Associations. [europepmc]
- Impact of word embedding models on text analytics in deep learning environment: a review. [europepmc]
- Enabling Early Health Care Intervention by Detecting Depression in Users of Web-Based Forums using Language Models: Longitudinal Analysis and Evaluation. [europepmc]
- Classification of tumor types using XGBoost machine learning model: a vector space transformation of genomic alterations. [europepmc]
Related