Gemma 2: Improving Open Language Models at a Practical Size
2024/07/31 by Morgane Rivière, Gemma Team, Shreya Pathak +292 · 312 citations
Computer Science · #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2408.00118
Abstract
In this work, we introduce Gemma 2, a new addition to the Gemma family of lightweight, state-of-the-art open models, ranging in scale from 2 billion to 27 billion parameters. In this new version, we apply several known technical modifications to the Transformer architecture, such as interleaving local-global attentions (Beltagy et al., 2020a) and group-query attention (Ainslie et al., 2023). We also train the 2B and 9B models with knowledge distillation (Hinton et al., 2015) instead of next token prediction. The resulting models deliver the best performance for their size, and even offer competitive alternatives to models that are 2-3 times bigger. We release all our models to the community.
Cited by
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs
- A Brain-like Synergistic Core in LLMs Drives Behaviour and Learning
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
- Where Steering Signals Come From: Activation Source Selection in Activation Steering
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- Reference Feature Atlases for Mechanistic Auditing of Language Models
- RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
- CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- Pioneering Multimodal Emotion Recognition in the Era of Large Models: From Closed Sets to Open Vocabularies
- LookPlanGraph: Embodied Instruction Following Method with VLM Graph Augmentation
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
- Measuring Mechanistic Independence: Can Bias Be Removed Without Erasing Demographics?
- Making Large Language Models Efficient Dense Retrievers
- Fine-Tuned In-Context Learners for Efficient Adaptation
- Large Language Models as Discounted Bayesian Filters
- FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content
- Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
- Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- PathFinder: MCTS and LLM Feedback-based Path Selection for Multi-Hop Question Answering
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
- IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
- Local LLM Ensembles for Zero-shot Portuguese Named Entity Recognition
- Scaling Behavior of Discrete Diffusion Language Models
- A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
- PCMind-2.1-Kaiyuan-2B Technical Report
- Uncovering Competency Gaps in Large Language Models and Their Benchmarks
- LMSpell: Neural Spell Checking for Low-Resource Languages
- Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- AITutor-EvalKit: Exploring the Capabilities of AI Tutors
- Towards Contextual Sensitive Data Detection
- EngChain: A Symbolic Benchmark for Verifiable Multi-Step Reasoning in Engineering
- Towards Active Synthetic Data Generation for Finetuning Language Models
- Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
- Can LLMs extract human-like fine-grained evidence for evidence-based fact-checking?
- Training Introspective Behavior: Fine-Tuning Induces Reliable Internal State Detection in a 7B Model
- Controlling changes to attention logits
- LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
- SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
- PixelDiT: Pixel Diffusion Transformers for Image Generation
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- Vision-Language Models for Automated 3D PET/CT Report Generation
- ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
- Comparative Analysis of LoRA-Adapted Embedding Models for Clinical Cardiology Text Representation
- Building Domain-Specific Small Language Models via Guided Data Generation
- Building Resilient Information Ecosystems: Large LLM-Generated Dataset of Persuasion Attacks
- Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
- Anatomy of an Idiom: Tracing Non-Compositionality in Language Models
- Enhancing Breast Cancer Prediction with LLM-Inferred Confounders
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
- Steganographic Backdoor Attacks in NLP: Ultra-Low Poisoning and Defense Evasion
- Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys
- KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
- Defending Unauthorized Model Merging via Dual-Stage Weight Protection
- ParliaBench: An Evaluation and Benchmarking Framework for LLM-Generated Parliamentary Speech
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Voice-Interactive Surgical Agent for Multimodal Patient Data Control
- FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
- More Agents Helps but Adversarial Robustness Gap Persists
- Beyond English: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
- Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
- GEMMA-SQL: A Novel Text-to-SQL Model Based on Large Language Models
- SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
- Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
- In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
- Dynamic Reflections: Probing Video Representations with Text Alignment
- Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
- ParaScopes: What do Language Models Activations Encode About Future Text?
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
- Angular Steering: Behavior Control via Rotation in Activation Space
- Interpreting LLMs as Credit Risk Classifiers: Do Their Feature Explanations Align with Classical ML?
- Simulating hashtag dynamics with networked groups of generative agents
- Alibaba International E-commerce Product Search Competition DcuRAGONs Team Technical Report
- Localized Adaptation Reveals Distinct Learning Signatures in Transformers
- FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA
- OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
- Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
- SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- Retrieval and Argumentation Enhanced Multi-Agent LLMs for Judgmental Forecasting
- HACK: Hallucinations Along Certainty and Knowledge Axes
- Utilising Large Language Models for Generating Effective Counter Arguments to Anti-Vaccine Tweets
- Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
- FARMER: Flow AutoRegressive Transformer over Pixels
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit Discussions
- A Comprehensive Dataset for Human vs. AI Generated Text Detection
- Evaluating LLMs' Reasoning Over Ordered Procedural Steps
- Model-Aware Tokenizer Transfer
- Large Language Models as Model Organisms for Human Associative Learning
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Designing and Evaluating Hint Generation Systems for Science Education
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Fast Inference via Hierarchical Speculative Decoding
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- ActivationReasoning: Logical Reasoning in Latent Activation Spaces
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
- ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
- Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback
- DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- Intent Clustering with Shared Pseudo-Labels
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- VaultGemma: A Differentially Private Gemma Model
- To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
- Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- Keep Calm and Avoid Harmful Content: Concept Alignment and Latent Manipulation Towards Safer Answers
- Don't Walk the Line: Boundary Guidance for Filtered Generation
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- Investigating Large Language Models' Linguistic Abilities for Text Preprocessing
- The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
- ADVICE: Answer-Dependent Verbalized Confidence Estimation
- Topological Alignment of Shared Vision-Language Embedding Space
- FactAppeal: Identifying Epistemic Factual Appeals in News Media
- Calibrating Generative Models
- Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- Augmenting Dialog with Think-Aloud Utterances for Modeling Individual Personality Traits by LLM
- Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
- Opponent Shaping in LLM Agents
- Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- EDUMATH: Generating Standards-aligned Educational Math Word Problems
- End-to-End Test-Time Training for Long Context
- Mid-Training of Large Language Models: A Survey
- TWIST: Training-free and Label-free Short Text Clustering through Iterative Vector Updating with LLMs
- Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
- VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Staircase Streaming for Low-Latency Multi-Agent Inference
- When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- Activation Steering with a Feedback Controller
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
- CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
- Reliability Crisis of Reference-free Metrics for Grammatical Error Correction
- ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Generative Value Conflicts Reveal LLM Priorities
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Training Agents Inside of Scalable World Models
- Evaluating Program Semantics Reasoning with Type Inference in System F
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
- Multiplayer Nash Preference Optimization
- Meta-Awareness Enhances Reasoning Models: Self-Alignment Reinforcement Learning
- OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- A circuit for predicting hierarchical structure in-context in Large Language Models
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- Towards Atoms of Large Language Models
- Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
- Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
- The potential and limitations of large language models for automatic classification of teachers' motivational messages in educational research
- DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
- OPENXRD: a comprehensive benchmark framework for LLM/MLLM XRD question answering
- Large Language Model Automated Extraction of Clinical Signs and Symptoms From Emergency Department Reports for Machine Learning Prediction Models: Development and Validation Study
- Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
- Benchmarking Humans and Machines on Complex Multilingual Speech Understanding Tasks
- Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
- Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark
- Variation in Verification: Understanding Verification Dynamics in Large Language Models
- Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text
- DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context
- Understanding Post-Training Structural Changes in Large Language Models
- SLICET5: Static Program Slicing using Language Models with Copy Mechanism and Constrained Decoding
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models
- DISCO: Disentangled Communication Steering for Large Language Models
- 'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
- Randomized Smoothing Meets Vision-Language Models
- Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
- Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
- Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
- Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
- Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning
- Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
- Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- Do Large Language Models Understand Word Senses?
- All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
- Automated and Context-Aware Code Documentation Leveraging Advanced LLMs
- InfoGain-RAG: Boosting Retrieval-Augmented Generation via Document Information Gain-based Reranking and Filtering
- LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Fluent but Unfeeling: The Emotional Blind Spots of Language Models
- LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
- TORSO: Template-Oriented Reasoning Towards General Tasks
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- MERLIN: Multi-Stage Curriculum Alignment for Multilingual Encoder-LLM Integration in Cross-Lingual Reasoning
- Customizing the Inductive Biases of Softmax Attention using Structured Matrices
- Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communities
- HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- Do LLMs exhibit the same commonsense capabilities across languages?
- HyFedRAG: A Federated Retrieval-Augmented Generation Framework for Heterogeneous and Privacy-Sensitive Data
- SpikingBrain: Spiking Brain-inspired Large Models
- FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
- Delta Activations: A Representation for Finetuned Large Language Models
- MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
- FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
- Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
- DriveQA: Passing the Driving Knowledge Test
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
- ArgCMV: An Argument Summarization Benchmark for the LLM-era
- The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation
- WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
- HebID: Detecting Social Identities in Hebrew-language Political Text
- Improving LLMs for Machine Translation Using Synthetic Preference Data
- Evaluating Sparse Autoencoders for Monosemantic Representation
- Generics and Default Reasoning in Large Language Models
- GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
- LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
- SEA-BED: Southeast Asia Embedding Benchmark
- Retrieval-augmented reasoning with lean language models
- Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
- Reverse Physician-AI Relationship: Full-process Clinical Diagnosis Driven by a Large Language Model
- Constrained Decoding of Diffusion LLMs with Context-Free Grammars
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- AINL-Eval 2025 Shared Task: Detection of AI-Generated Scientific Abstracts in Russian
- Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
- TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
- Large Language Models for Subjective Language Understanding: A Survey
- Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
- HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways
- EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
- Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
- Majority Bit-Aware Watermarking For Large Language Models
- RegMean++: Enhancing Effectiveness and Generalization of Regression Mean for Model Merging
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
- Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
- Is neural semantic parsing good at ellipsis resolution, or isn't it?
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
- Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
- Unveiling the Influence of Amplifying Language-Specific Neurons
- Beyond Single Labels: Improving Conversational Recommendation through LLM-Powered Data Augmentation
- Context-aware Rotary Position Embedding
- When Truthful Representations Flip Under Deceptive Instructions?
- Libra: Large Chinese-based Safeguard for AI Content
- Kimi K2: Open Agentic Intelligence
- SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding
- Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System Theory
- GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
- Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving
- Shaping capabilities with token-level data filtering
Related