Mistral 7B
2023/10/10 by Albert Q. Jiang, Alexandre Sablayrolles, Jiang, Albert Q. +35 · 5 voices · 265 citations
Computer Science · #Machine Learning and Data Classification #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2310.06825
openalex publication_date 2023/10/10 · openalex created_date 2023/10/12 · openalex updated_date 2026/07/29
Abstract
We introduce Mistral 7B v0.1, a 7-billion-parameter language model engineered for superior performance and efficiency. Mistral 7B outperforms Llama 2 13B across all evaluated benchmarks, and Llama 1 34B in reasoning, mathematics, and code generation. Our model leverages grouped-query attention (GQA) for faster inference, coupled with sliding window attention (SWA) to effectively handle sequences of arbitrary length with a reduced inference cost. We also provide a model fine-tuned to follow instructions, Mistral 7B -- Instruct, that surpasses the Llama 2 13B -- Chat model both on human and automated benchmarks. Our models are released under the Apache 2.0 license.
Cited by
- Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- Statistical Inference for Generative Model Comparison
- TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
- Multi2: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments
- CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
- MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference
- jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation
- Octopus v4: Graph of language models
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- Regularize or Localize: When Training-Time KV-Cache Geometry Pays Under Quantization
- TopoTuner: Topological Finetuning of Large Language Models
- SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning
- Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study
- Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
- Intent2QoS: Language Model-Driven Automation of Traffic Shaping Configurations
- Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
- DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
- Text-to-LoRA: Instant Transformer Adaption
- Do Chinese models speak Chinese languages?
- Transformers without Normalization
- Large Language Diffusion Models
- Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
- On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- A scaling law of contextual persistence in human language
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
- OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
- Large Language Models as Discounted Bayesian Filters
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications
- PPSEBM: An Energy-Based Model with Progressive Parameter Selection for Continual Learning
- Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
- Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis
- REMODEL-LLM: Transforming C code to Java using LLMs
- Watermarks for Language Models via Probabilistic Automata
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- Large Language Model-Based Generation of Discharge Summaries
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- Large Language Models as Generalist Policies for Network Optimization
- VACoT: Rethinking Visual Data Augmentation with VLMs
- Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
- Towards Active Synthetic Data Generation for Finetuning Language Models
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights
- CacheTrap: Injecting Trojans in LLMs without Leaving any Traces in Inputs or Weights
- Generative models for crystalline materials
- FlockVote: LLM-Empowered Agent-Based Modeling for Simulating U.S. Presidential Elections
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Skypilot: Fine-Tuning LLM with Physical Grounding for AAV Coverage Search
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- PersonaAgent with GraphRAG: Community-Aware Knowledge Graphs for Personalized LLM
- Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
- Contrastive vision-language learning with paraphrasing and negation
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
- Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning
- A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
- CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- The Open Syndrome Definition as a Machine-Readable Standard for Public Health: Design and Implementation Study
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Adaptive Testing for Segmenting Watermarked Texts From Language Models
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- Logit-Entropy Adaptive Stopping Heuristic for Efficient Chain-of-Thought Reasoning
- Direct Semantic Communication Between Large Language Models via Vector Translation
- OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
- Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
- AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
- Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- RDQ: Residual Distribution Quantization for Large Language Models
- SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
- RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
- Multimodal learning enables chat-based exploration of single-cell data
- Semantic Similarity in Radiology Reports via LLMs and NER
- How do large-language models respond to moral dilemmas? Insights from the defining issues test
- Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
- ProofSketch: Efficient Verified Reasoning for Large Language Models
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- Probing Knowledge Holes in Unlearned LLMs
- DETECT: Determining Ease and Textual Clarity of German Text Simplifications
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- Personalized Chain-of-Thought Summarization of Financial News for Investor Decision Support
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
- A Benchmark Dataset And LLMs Comparison For NFR Classification With Explainable AI
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- An Evaluation of LLMs Inference on Popular Single-board Computers
- Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- End-to-End Multi-Modal Diffusion Mamba
- Stable LLM Ensemble: Interaction between Example Representativeness and Diversity
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Taming the Fragility of KV Cache Eviction in LLM Inference
- Beyond Imitation: Recovering Dense Rewards from Demonstrations
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- PromptLocate: Localizing Prompt Injection Attacks
- Representation-Based Exploration for Language Models: From Test-Time to Post-Training
- F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
- SoundReactor: Frame-level Online Video-to-Audio Generation
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
- OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
- LLM Based Long Code Translation using Identifier Replacement
- RustAssure: Differential Symbolic Testing for LLM-Transpiled C-to-Rust Code
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Models
- ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- Measuring and Mitigating Identity Bias in Multi-Agent Debate via Anonymization
- From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents
- Verifying Memoryless Sequential Decision-making of Large Language Models
- Revisiting Long-context Modeling from Context Denoising Perspective
- The New Quant: A Survey of Large Language Models in Financial Prediction and Trading
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Scalable In-context Ranking with Generative Models
- Imperceptible Jailbreaking against Large Language Models
- NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- Aligning Language Models with Clinical Expertise: DPO for Heart Failure Nursing Documentation in Critical Care
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- Negative Pre-activations Differentiate Syntax
- Analyzing and Evaluating Unbiased Language Model Watermark
- An Ensemble Framework for Unbiased Language Model Watermarking
- HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based Fuzzing
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
- LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
- MMPB: It's Time for Multi-Modal Personalization
- Multidimensional Uncertainty Quantification via Optimal Transport
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- Beyond Stars: Bridging the Gap Between Ratings and Review Sentiment with LLM
- MMSE-Calibrated Few-Shot Prompting for Alzheimer's Detection
- Divergence Decoding: Training-Free Capability Fusion
- OPENXRD: a comprehensive benchmark framework for LLM/MLLM XRD question answering
- LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
- References Improve LLM Alignment in Non-Verifiable Domains
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Bounded PCTL Model Checking of Large Language Model Outputs
- Model selection meets clinical semantics: Optimizing ICD-10-CM prediction via LLM-as-Judge evaluation, redundancy-aware sampling, and section-aware fine-tuning
- Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
- MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
- Understanding Post-Training Structural Changes in Large Language Models
- Probabilistic Token Alignment for Large Language Model Fusion
- Semantic Representation Attack against Aligned Large Language Models
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- The Few-shot Dilemma: Over-prompting Large Language Models
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Selective Risk Certification for LLM Outputs via Information-Lift Statistics: PAC-Bayes, Robustness, and Skeleton Design
- Continually Adding New Languages to Multilingual Language Models
- HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- RefactorCoderQA: Benchmarking LLMs for Multi-Domain Coding Question Solutions in Cloud and Edge Deployment
- Towards Understanding Visual Grounding in Visual Language Models
- One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
- Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- Do small language models generate realistic variable-quality fake news headlines?
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- LLM-based Triplet Extraction for Automated Ontology Generation in Software Engineering Standards
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- Ensemble Debates with Local Large Language Models for AI Alignment
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
- Generics and Default Reasoning in Large Language Models
- FedSODA: Federated Fine-tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment
- Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Improving Detection of Watermarked Language Models
- ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
- Diffusion is a code repair operator and generator
- AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- SinLlama -- A Large Language Model for Sinhala
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Leveraging LLMs for Smart Cities Qualitative Data Analysis
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
- 5G Core Fault Detection and Root Cause Analysis using Machine Learning and Generative AI
- Latent Knowledge Scalpel: Precise and Massive Knowledge Editing for Large Language Models
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration
- MemoCue: Empowering LLM-Based Agents for Human Memory Recall via Strategy-Guided Querying
- KLLM: Fast LLM Inference with K-Means Quantization
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
- Strategic Deflection: Defending LLMs from Logit Manipulation
- Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
- HLSDebugger: Identification and Correction of Logic Bugs in HLS Code with LLM Solutions
- Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
- AI-Driven Generation of Old English: A Framework for Low-Resource Languages
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- The Carbon Cost of Conversation, Sustainability in the Age of Language Models
- Toward Revealing Nuanced Biases in Medical LLMs
Discussions
- Mistral 7B [hn, 267 points, 123 comments]
- Mistral 7B model [lemmy, 24 points, 8 comments]
- Here's a PDF containing the Mistral prompt: arxiv.org/pdf/2310.068... [bsky, 4 points, 1 comments]
- I started using it because it was one of the first open source models (for some contentious definition of open source) and it gave consistently good answers. Cracked team, just look at their paper. ar [bsky, 1 points, 1 comments]
- Here’s the documentation for Mistral 7B from a French LLM lab: arxiv.org/abs/2310.06825 [bsky, 0 points, 0 comments]
Related