Mistral 7B
2023/10/10 by Albert Q. Jiang, Alexandre Sablayrolles, Jiang, Albert Q. +35 · 5 voices · 549 citations
Computer Science · #Algorithm #Artificial intelligence #Code (set theory) #Computer science #Inference #License #MIT License #Machine Learning and Data Classification #Natural Language Processing Techniques #Operating system #Programming language #Sliding window protocol #Topic Modeling #Window (computing) #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2310.06825
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/10/10 · openalex created_date 2023/10/12 · openalex updated_date 2026/08/05
Abstract
We introduce Mistral 7B v0.1, a 7-billion-parameter language model engineered for superior performance and efficiency. Mistral 7B outperforms Llama 2 13B across all evaluated benchmarks, and Llama 1 34B in reasoning, mathematics, and code generation. Our model leverages grouped-query attention (GQA) for faster inference, coupled with sliding window attention (SWA) to effectively handle sequences of arbitrary length with a reduced inference cost. We also provide a model fine-tuned to follow instructions, Mistral 7B -- Instruct, that surpasses the Llama 2 13B -- Chat model both on human and automated benchmarks. Our models are released under the Apache 2.0 license.
Cited by
- Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- Statistical Inference for Generative Model Comparison
- TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
- Multi2: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments
- CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
- MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference
- jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation
- Octopus v4: Graph of language models
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- Regularize or Localize: When Training-Time KV-Cache Geometry Pays Under Quantization
- TopoTuner: Topological Finetuning of Large Language Models
- SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning
- Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study
- Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
- Intent2QoS: Language Model-Driven Automation of Traffic Shaping Configurations
- Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
- DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
- Text-to-LoRA: Instant Transformer Adaption
- Do Chinese models speak Chinese languages?
- Transformers without Normalization
- Large Language Diffusion Models
- Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
- On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- A scaling law of contextual persistence in human language
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
- OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
- Large Language Models as Discounted Bayesian Filters
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications
- PPSEBM: An Energy-Based Model with Progressive Parameter Selection for Continual Learning
- Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
- Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis
- REMODEL-LLM: Transforming C code to Java using LLMs
- Watermarks for Language Models via Probabilistic Automata
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- Large Language Model-Based Generation of Discharge Summaries
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- Large Language Models as Generalist Policies for Network Optimization
- VACoT: Rethinking Visual Data Augmentation with VLMs
- Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
- Towards Active Synthetic Data Generation for Finetuning Language Models
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights
- CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
- Generative models for crystalline materials
- FlockVote: LLM-Empowered Agent-Based Modeling for Simulating U.S. Presidential Elections
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Skypilot: Fine-Tuning LLM with Physical Grounding for AAV Coverage Search
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- PersonaAgent with GraphRAG: Community-Aware Knowledge Graphs for Personalized LLM
- Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
- Contrastive vision-language learning with paraphrasing and negation
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
- Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning
- A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
- CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- The Open Syndrome Definition as a Machine-Readable Standard for Public Health: Design and Implementation Study
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Adaptive Testing for Segmenting Watermarked Texts From Language Models
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- Logit-Entropy Adaptive Stopping Heuristic for Efficient Chain-of-Thought Reasoning
- Direct Semantic Communication Between Large Language Models via Vector Translation
- OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
- Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
- AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
- Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- RDQ: Residual Distribution Quantization for Large Language Models
- SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
- RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
- Multimodal learning enables chat-based exploration of single-cell data
- Semantic Similarity in Radiology Reports via LLMs and NER
- How do large-language models respond to moral dilemmas? Insights from the defining issues test
- Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
- ProofSketch: Efficient Verified Reasoning for Large Language Models
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- Probing Knowledge Holes in Unlearned LLMs
- DETECT: Determining Ease and Textual Clarity of German Text Simplifications
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- Personalized Chain-of-Thought Summarization of Financial News for Investor Decision Support
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
- A Benchmark Dataset And LLMs Comparison For NFR Classification With Explainable AI
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- An Evaluation of LLMs Inference on Popular Single-board Computers
- Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- End-to-End Multi-Modal Diffusion Mamba
- Stable LLM Ensemble: Interaction between Example Representativeness and Diversity
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Taming the Fragility of KV Cache Eviction in LLM Inference
- Beyond Imitation: Recovering Dense Rewards from Demonstrations
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- PromptLocate: Localizing Prompt Injection Attacks
- Representation-Based Exploration for Language Models: From Test-Time to Post-Training
- F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
- SoundReactor: Frame-level Online Video-to-Audio Generation
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
- OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
- LLM Based Long Code Translation using Identifier Replacement
- RustAssure: Differential Symbolic Testing for LLM-Transpiled C-to-Rust Code
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Models
- ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- Measuring and Mitigating Identity Bias in Multi-Agent Debate via Anonymization
- From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents
- Verifying Memoryless Sequential Decision-making of Large Language Models
- Revisiting Long-context Modeling from Context Denoising Perspective
- The New Quant: A Survey of Large Language Models in Financial Prediction and Trading
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Scalable In-context Ranking with Generative Models
- Imperceptible Jailbreaking against Large Language Models
- NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- Aligning Language Models with Clinical Expertise: DPO for Heart Failure Nursing Documentation in Critical Care
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- Negative Pre-activations Differentiate Syntax
- Analyzing and Evaluating Unbiased Language Model Watermark
- An Ensemble Framework for Unbiased Language Model Watermarking
- HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based Fuzzing
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
- LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
- MMPB: It's Time for Multi-Modal Personalization
- Multidimensional Uncertainty Quantification via Optimal Transport
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- Beyond Stars: Bridging the Gap Between Ratings and Review Sentiment with LLM
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- MMSE-Calibrated Few-Shot Prompting for Alzheimer's Detection
- Divergence Decoding: Training-Free Capability Fusion
- OPENXRD: a comprehensive benchmark framework for LLM/MLLM XRD question answering
- LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
- References Improve LLM Alignment in Non-Verifiable Domains
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Bounded PCTL Model Checking of Large Language Model Outputs
- Model selection meets clinical semantics: Optimizing ICD-10-CM prediction via LLM-as-Judge evaluation, redundancy-aware sampling, and section-aware fine-tuning
- Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
- MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
- Understanding Post-Training Structural Changes in Large Language Models
- Probabilistic Token Alignment for Large Language Model Fusion
- Semantic Representation Attack against Aligned Large Language Models
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- The Few-shot Dilemma: Over-prompting Large Language Models
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Selective Risk Certification for LLM Outputs via Information-Lift Statistics: PAC-Bayes, Robustness, and Skeleton Design
- Continually Adding New Languages to Multilingual Language Models
- HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- RefactorCoderQA: Benchmarking LLMs for Multi-Domain Coding Question Solutions in Cloud and Edge Deployment
- Towards Understanding Visual Grounding in Visual Language Models
- One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
- Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- Do small language models generate realistic variable-quality fake news headlines?
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- LLM-based Triplet Extraction for Automated Ontology Generation in Software Engineering Standards
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- Ensemble Debates with Local Large Language Models for AI Alignment
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
- Generics and Default Reasoning in Large Language Models
- FedSODA: Federated Fine-tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment
- Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Improving Detection of Watermarked Language Models
- ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
- Diffusion is a code repair operator and generator
- AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- SinLlama -- A Large Language Model for Sinhala
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Leveraging LLMs for Smart Cities Qualitative Data Analysis
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
- 5G Core Fault Detection and Root Cause Analysis using Machine Learning and Generative AI
- Latent Knowledge Scalpel: Precise and Massive Knowledge Editing for Large Language Models
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration
- MemoCue: Empowering LLM-Based Agents for Human Memory Recall via Strategy-Guided Querying
- OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
- Strategic Deflection: Defending LLMs from Logit Manipulation
- Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
- Foundation Models and Transformers for Anomaly Detection: A Survey
- HLSDebugger: Identification and Correction of Logic Bugs in HLS Code with LLM Solutions
- Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
- PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants
- AI-Driven Generation of Old English: A Framework for Low-Resource Languages
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- The Carbon Cost of Conversation, Sustainability in the Age of Language Models
- Toward Revealing Nuanced Biases in Medical LLMs
- Mellum2 Technical Report
- Security-by-Design for LLM-Based Code Generation: Leveraging Internal Representations for Concept-Driven Steering Mechanisms
- KROMA: Ontology Matching with Knowledge Retrieval and Large Language Models
- Large Language Models in the Travel Domain: An Industrial Experience
- InTraVisTo: Inside Transformer Visualisation Tool
- Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
- DesignLab: Designing Slides Through Iterative Detection and Correction
- Cooling Matters: Benchmarking Large Language Models and Vision-Language Models on Liquid-Cooled Versus Air-Cooled H100 GPU Systems
- Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
- Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
- Aligning Knowledge Graphs and Language Models for Factual Accuracy
- Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- From Matching to Generation: A Survey on Generative Information Retrieval
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
- Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks
- KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
- LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
- Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- SAGE: A Context-Aware Approach for Mining Privacy Requirements Relevant Reviews from Mental Health Apps
- ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making
- NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
- Multilingual Multimodal Software Developer for Code Generation
- Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?
- Scalable Medication Extraction and Discontinuation Identification from Electronic Health Records Using Large Language Models
- Extracting epilepsy‐related information from unstructured clinic letters using large language models
- GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
- TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
- SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression
- From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations
- Toward Better Generalisation in Uncertainty Estimators: Leveraging Data-Agnostic Features
- OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
- Transforming Calabi-Yau Constructions: Generating New Calabi-Yau Manifolds with Transformers
- ReservoirChat: Interactive Documentation Enhanced with LLM and Knowledge Graph for ReservoirPy
- TACOS: Open Tagging and Comparative Scoring for Instruction Fine-Tuning Data Selection
- Blackbox Dataset Inference for LLM
- GDC Cohort Copilot: An AI Copilot for Curating Cohorts from the Genomic Data Commons
- The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
- CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
- Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks
- OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation
- PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER)
- Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
- A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
- TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
- Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
- Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
- "I wasn't sure if this is indeed a security risk": Data-driven Understanding of Security Issue Reporting in GitHub Repositories of Open Source npm Packages
- Data Efficacy for Language Model Training
- How to Retrieve Examples in In-context Learning to Improve Conversational Emotion Recognition using Large Language Models?
- DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
- How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?
- Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
- Spark Transformer: Reactivating Sparsity in FFN and Attention
- SafeLawBench: Towards Safe Alignment of Large Language Models
- Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
- CommVQ: Commutative Vector Quantization for KV Cache Compression
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
- CS-KG 2.0: A Large-scale Knowledge Graph of Computer Science
- Object-aware Sound Source Localization via Audio-Visual Scene Understanding
- End-to-End Spoken Grammatical Error Correction
- From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium
- TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
- Moment Sampling in Video LLMs for Long-Form Video QA
- A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
- LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
- From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
- LLM-Powered Swarms: A New Frontier or a Conceptual Stretch?
- Value-Free Policy Optimization via Reward Partitioning
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations
- PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
- Explaining Recovery Trajectories of Older Adults Post Lower-Limb Fracture Using Modality-wise Multiview Clustering and Large Language Models
- Large Language Models for Detection of Life-Threatening Texts
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- TransXSSM: A Hybrid Transformer State Space Model with Unified Rotary Position Embedding
- You Only Fine-tune Once: Many-Shot In-Context Fine-Tuning for Large Language Models
- Defending against Indirect Prompt Injection by Instruction Detection
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
- Elementary Math Word Problem Generation using Large Language Models
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- Power Law Guided Dynamic Sifting for Efficient Attention
- SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
- GEM: Empowering LLM for both Embedding Generation and Language Understanding
- CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation
- Large Means Left: Political Bias in Large Language Models Increases with Their Number of Parameters
- RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
- High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
- SAGE:Specification-Aware Grammar Extraction for Automated Test Case Generation with LLMs
- A Statistical Physics of Language Model Reasoning
- Enhancing Decision-Making of Large Language Models via Actor-Critic
- Geospatial Mechanistic Interpretability of Large Language Models
- Should LLM Safety Be More Than Refusing Harmful Instructions?
- Adaptive Task Vectors for Large Language Models
- Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection
- MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching
- CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection
- StochasTok: Improving Fine-Grained Subword Understanding in LLMs
- Reasoning-Based Approach with Chain-of-Thought for Alzheimer's Detection Using Speech and Large Language Models
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- SPACE: Your Genomic Profile Predictor is a Powerful DNA Foundation Model
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
- Earley-Driven Dynamic Pruning for Efficient Structured Decoding
- Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
- A Simple Linear Patch Revives Layer-Pruned Large Language Models
- Interpretable phenotyping of Heart Failure patients with Dutch discharge letters
- Cross-Attention Speculative Decoding
- Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
- A Multi‑Region Brain Model to Elucidate the Role of Hippocampus in Spatially Embedded Decision‑Making
- Spoken Language Modeling with Duration-Penalized Self-Supervised Units
- From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
- GenIC: An LLM-Based Framework for Instance Completion in Knowledge Graphs
- Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
- Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
- Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical Reshape
- Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
- Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
- Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
- Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM
- Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning
- RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
- Emotion-aware Dual Cross-Attentive Neural Network with Label Fusion for Stance Detection in Misinformative Social Media Content
- EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
- Can Large Language Models Predict Audio Effects Parameters from Natural Language?
- Efficient Large Language Model Inference with Neural Block Linearization
- Pretrained LLMs Learn Multiple Types of Uncertainty
- Incentivizing Inclusive Contributions in Model Sharing Markets
- ResSVD: Residual Compensated SVD for Large Language Model Compression
- SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
- DoctorAgent-RL: A Multi-Agent Collaborative Reinforcement Learning System for Multi-Turn Clinical Dialogue
- The Role of Diversity in In-Context Learning for Large Language Models
- Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
- Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- FairPO: Robust Preference Optimization for Fair Multi-Label Learning
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
- EnvSDD: Benchmarking Environmental Sound Deepfake Detection
- Rethinking the Understanding Ability across LLMs through Mutual Information
- NextG-GPT: Leveraging GenAI for Advancing Wireless Networks and Communication Research
- Benchmarking and Rethinking Knowledge Editing for Large Language Models
- Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models
- ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation
- Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
- Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning
- Misaligning Reasoning with Answers -- A Framework for Assessing LLM CoT Robustness
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
- Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
- Advancing LLM Safe Alignment with Safety Representation Ranking
- Feed-Forward Steering in Transformer Residual Dynamics
- Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
- Set-LLM: A Permutation-Invariant LLM
- Position: Agentic Systems Constitute a Key Component of Next-Generation Intelligent Image Processing
- Model Merging is Secretly Certifiable: Non-Vacuous Generalisation Bounds for Low-Shot Learning
- UltraEdit: Training-, Subject-, and Memory-Free Lifelong Editing in Language Models
- Rank-K: Test-Time Reasoning for Listwise Reranking
- Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
- Fragments to Facts: Partial-Information Fragment Inference from LLMs
- Beyond Words: Multimodal LLM Knows When to Speak
- Text Generation Beyond Discrete Token Sampling
- Safety Alignment Can Be Not Superficial With Explicit Safety Signals
- RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
- Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
- Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review
- MARGE: Improving Math Reasoning for LLMs with Guided Exploration
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- Fine-grained Contrastive Learning for ECG-Report Alignment with Waveform Enhancement
- Data-driven material screening of secondary and natural cementitious precursors
- RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings
- Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
- Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMs
- Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
- ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale
- EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
- Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning
- Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
- Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
- Multimodal Cancer Modeling in the Age of Foundation Model Embeddings
- JobHop: A Large-Scale Dataset of Career Trajectories
- Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
- Reassessing Large Language Model Boolean Query Generation for Systematic Reviews
- Turning LLM Activations Quantization-Friendly
- Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
- Empowering Scientific Workflows with Federated Agents
- Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
- Geometric Analysis of Token Selection in Multi-Head Attention
- TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition
- Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
- R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
- SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
- Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency
- Advancing AI Research Assistants with Expert-Involved Learning
- Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
- Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
- On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
- LLMs can construct powerful representations and streamline sample-efficient supervised learning
- SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
- BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
- Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs
- Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models
- Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
- Small or Large? Zero-Shot or Finetuned? Guiding Language Model Choice for Specialized Applications in Healthcare
- A Survey on Parameter-Efficient Fine-Tuning for Foundation Models in Federated Learning
- DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
- LLM Enhancer: Merged Approach using Vector Embedding for Reducing Large Language Model Hallucinations with External Knowledge
- Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
- HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation
- HiFloat4 Format for Language Model Inference
- BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
- Online Safety Monitoring for LLMs
- Language-Critique Imitation Learning from Suboptimal Demonstrations
- Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
- MATCHA: Matching Text via Contrastive Semantic Alignment
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation
- Zero-shot adaptable task planning for autonomous construction robots: a comparative study of lightweight single and multi-AI agent systems
- Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
- Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints
- Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review
- AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents
- Reversing Arrows in Large Language Models
- The Role of Open-Source LLMs in Shaping the Future of GeoAI
- CoheMark: A Novel Sentence-Level Watermark for Enhanced Text Quality
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning
- Maglev: Sliding Recurrent Memory
- Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
- Language Models Are Implicitly Continuous
- Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
- SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
- Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
- Saliency-driven Dynamic Token Pruning for Large Language Models
- Vidi: Large Multimodal Models for Video Understanding and Editing
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
- Fine-Tuning a 7B Advisor on Free-Tier GPUs: An Adapter-Handoff Recipe and a Synthetic-Data Reliability Caution
- Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation
- ForgeBench: A Machine Learning Benchmark Suite and Auto-Generation Framework for Next-Generation HLS Tools
- Kuwain 1.5B: An Arabic SLM via Language Injection
- A Hierarchical Framework for Measuring Scientific Paper Innovation via Large Language Models
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection
- A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content
- Activated LoRA: Fine-tuned LLMs for Intrinsics
- Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
- The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
- Subitizing-InspiredLargeLanguageModelsforFloorplanning
- Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents
- A Dual-Space Framework for General Knowledge Distillation of Large Language Models
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
- SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
- Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
- Improving In-Context Learning with Reasoning Distillation
- Alleviating the Fear of Losing Alignment in LLM Fine-tuning
- Large Language Models Could Be Rote Learners
- Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
- Efficient Tuning of Large Language Models for Knowledge-Grounded Dialogue Generation
- LSR-MCTS: Alleviating Long Range Dependency in Code Generation
- LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
- ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
- Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization
- ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
- On the Impact of Language Nuances on Sentiment Analysis with Large Language Models: Paraphrasing, Sarcasm, and Emojis
- REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
- Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
- Generative Large Language Model usage in Smart Contract Vulnerability Detection
- Evaluating Knowledge Graph Based Retrieval Augmented Generation Methods under Knowledge Incompleteness
- How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks
- Prompt Optimization with Logged Bandit Data
- Large language model for post‐earthquake structural damage assessment of buildings
Discussions
- Mistral 7B [hn, 267 points, 123 comments]
- Mistral 7B model [lemmy, 24 points, 8 comments]
- Here's a PDF containing the Mistral prompt: arxiv.org/pdf/2310.068... [bsky, 4 points, 1 comments]
- I started using it because it was one of the first open source models (for some contentious definition of open source) and it gave consistently good answers. Cracked team, just look at their paper. ar [bsky, 1 points, 1 comments]
- Here’s the documentation for Mistral 7B from a French LLM lab: arxiv.org/abs/2310.06825 [bsky, 0 points, 0 comments]
Related