The Rise of Language Models in Mining Software Repositories: A Survey
2026/04/30 by Miguel Romero-Arjona, Saman Barakat, Ana B. Sánchez +1
Computer Science · #cs.SE
paper · pdf · doi:10.48550/arxiv.2604.00787
arxiv created 2026/08/03 · arxiv updated 2026/08/04
Abstract
The Mining Software Repositories (MSR) field focuses on analyzing the rich data contained in software repositories to derive actionable insights into software processes and products. Mining repositories at scale requires techniques capable of handling large volumes of heterogeneous data, a challenge for which language models (LMs) are increasingly well-suited. Since the advent of Transformer-based architectures, LMs have been rapidly adopted across a wide range of MSR tasks. This article presents a comprehensive survey of the use of LMs in MSR, based on an analysis of 177 papers. We examine how LMs are applied, the types of artifacts analyzed, which models are used, how their adoption has evolved over time, and the availability of supplementary materials and tools supporting reproducibility and reuse. Building on this analysis, we propose a taxonomy of LM applications in MSR, identify key trends shaping the field, and highlight open challenges alongside actionable directions for future research.
Citations
- Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies
- What Types of Code Review Comments Do Developers Most Frequently Resolve?
- Reverse Engineering User Stories from Code using Large Language Models
- Are Prompts All You Need? Evaluating Prompt-Based Large Language Models (LLM)s for Software Requirements Classification
- Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
- Does AI Code Review Lead to Code Changes? A Case Study of GitHub Actions
- On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository
- Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
- A Methodological Framework for LLM-Based Mining of Software Repositories
- CASPER: Contrastive Approach for Smart Ponzi Scheme Detecter with More Negative Samples
- SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
- SAGE: A Context-Aware Approach for Mining Privacy Requirements Relevant Reviews from Mental Health Apps
Droid: A Resource Suite for AI-Generated Code Detection- CMER: A Context-Aware Approach for Mining Ethical Concern-related App Reviews
- Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
- Towards Trustworthy Sentiment Analysis in Software Engineering: Dataset Characteristics and Tool Selection
- From Release to Adoption: Challenges in Reusing Pre-trained AI Models for Downstream Developers
- What Characteristics Make ChatGPT Effective for Software Issue Resolution? An Empirical Study of Task, Project, and Conversational Signals in GitHub Issues
- Social Media Reactions to Open Source Promotions: AI-Powered GitHub Projects on Hacker News
- How Far Have LLMs Come Toward Automated SATD Taxonomy Construction?
- What About Emotions? Guiding Fine-Grained Emotion Extraction from Mobile App Reviews
- Rethinking Code Review Workflows with LLM Assistance: An Empirical Study
- Towards an Interpretable Analysis for Estimating the Resolution Time of Software Issues
- Applying Large Language Models to Issue Classification: Revisiting with Extended Data and New Models
- Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
- Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code
- Do Developers Depend on Deprecated Library Versions? A Mining Study of Log4j
- Unveiling Ruby: Insights from Stack Overflow and Developer Survey
- Exploring the Role of Women in Hugging Face Organizations
- A Bot-based Approach to Manage Codes of Conduct in Open-Source Projects
- Towards Refining Developer Questions using LLM-Based Named Entity Recognition for Developer Chatroom Conversations
- FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs
- LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights
- Combining Large Language Models with Static Analyzers for Code Review Generation
- Are the Majority of Public Computational Notebooks Pathologically Non-Executable?
- Should Code Models Learn Pedagogically? A Preliminary Evaluation of Curriculum Learning for Real-World Software Engineering Tasks
- Harnessing Large Language Models for Curated Code Reviews
- Too Noisy To Learn: Enhancing Data Quality for Code Review Comment Generation
- SkillScope: A Tool to Predict Fine-Grained Skills Needed to Solve Issues on GitHub
- Towards Detecting Prompt Knowledge Gaps for Improved LLM-guided Issue Resolution
- Analyzing the Evolution and Maintenance of Quantum Software Repositories
- Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories
- Why Do Developers Engage with ChatGPT in Issue-Tracker? Investigating Usage and Reliance on ChatGPT-Generated Code
- Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
- Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
- Generative Language Models Potential for Requirement Engineering Applications: Insights into Current Strengths and Limitations
- Do Developers Adopt Green Architectural Tactics for ML-Enabled Systems? A Mining Software Repository Study
- A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages
- CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
- What do we know about Hugging Face? A systematic literature review and quantitative validation of qualitative claims
- A Survey on Large Language Models for Code Generation
- A Systematic Literature Review on Large Language Models for Automated Program Repair
- Lessons Learned from Mining the Hugging Face Repository
- How to Refactor this Code? An Exploratory Study on Developer-ChatGPT Refactoring Conversations
- Incivility in Open Source Projects: A Comprehensive Annotated Dataset of Locked GitHub Issue Threads
- Towards Summarizing Code Snippets Using Pre-Trained Transformers
- Shedding Light on Software Engineering-specific Metaphors and Idioms
- Analyzing the Evolution and Maintenance of ML Models on Hugging Face
- Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models
- "I see models being a whole other thing": An Empirical Study of Pre-Trained Model Naming Conventions and A Tool for Enhancing Naming Consistency
- Can GitHub Issues Help in App Review Classifications?
- Software Testing with Large Language Models: Survey, Landscape, and Vision
- Software Testing With Large Language Models: Survey, Landscape, and Vision
- Large Language Models for Software Engineering: Survey and Open Problems
- CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
- PICASO: Enhancing API Recommendations with Relevant Stack Overflow Posts
- A Study of Gender Discussions in Mobile Apps
- LLMSecEval: A Dataset of Natural Language Prompts for Security Evaluations
- GPT-4 Technical Report
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Automated Detection, Categorisation and Developers' Experience with the Violations of Honesty in Mobile Apps
- AUGER: Automatically Generating Review Comments with Pre-training Models
- Data Augmentation for Improving Emotion Recognition in Software Engineering Communication
- Automatic Pull Request Title Generation
- Using Transfer Learning for Code-Related Tasks
- Can language models learn from explanations in context?
- Supporting Developers in Addressing Human-centric Issues in Mobile Apps
- Automating Code Review Activities by Large-Scale Pre-training
- On the Violation of Honesty in Mobile Apps: Automated Detection and Categories
- Efficient Search of Live-Coding Screencasts from Online Videos
- Using Pre-Trained Models to Boost Code Review Automation
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Automatically Matching Bug Reports With Related App Reviews
- Predicting the Objective and Priority of Issue Reports in Software Repositories
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Language Models are Few-Shot Learners
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language\n Generation, Translation, and Comprehension
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Boosting Automatic Commit Classification Into Maintenance Activities By Utilizing Source Code Changes
- Attention Is All You Need
- A Survey on Metamorphic Testing
- The Kappa Statistic in Reliability Studies: Use, Interpretation, and Sample Size Requirements
- Analyzing the Past to Prepare for the Future: Writing a Literature Review
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding