InternLM2 Technical Report
2024/03/26 by Zheng Cai, Maosong Cao, Cai, Zheng +199 · 2 voices · 108 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Business #Machine Learning in Healthcare #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2403.17297
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/03/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advancements in open-source models has been challenging. This paper introduces InternLM2, an open-source LLM that outperforms its predecessors in comprehensive evaluations across 6 dimensions and 30 benchmarks, long-context modeling, and open-ended subjective evaluations through innovative pre-training and optimization techniques. The pre-training process of InternLM2 is meticulously detailed, highlighting the preparation of diverse data types including text, code, and long-context data. InternLM2 efficiently captures long-term dependencies, initially trained on 4k tokens before advancing to 32k tokens in pre-training and fine-tuning stages, exhibiting remarkable performance on the 200k ``Needle-in-a-Haystack" test. InternLM2 is further aligned using Supervised Fine-Tuning (SFT) and a novel Conditional Online Reinforcement Learning from Human Feedback (COOL RLHF) strategy that addresses conflicting human preferences and reward hacking. By releasing InternLM2 models in different training stages and model sizes, we provide the community with insights into the model's evolution.
Cited by
- Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
- VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
- CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
- Why Do Vision Language Models Struggle To Recognize Human Emotions?
- Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
- Do Chinese models speak Chinese languages?
- Transformers without Normalization
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VACoT: Rethinking Visual Data Augmentation with VLMs
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Code-driven Number Sequence Calculation: Enhancing the inductive Reasoning Abilities of Large Language Models
- Confidence as a Reward: Transforming LLMs into Reward Models
- The Harder The Better: Maintaining Supervised Fine-tuning Generalization with Less but Harder Data
- LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation
- Conjecturing: An Overlooked Step in Formal Mathematical Reasoning
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Mid-Training of Large Language Models: A Survey
- Online Rubrics Elicitation from Pairwise Comparisons
- Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial Trading
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
- Probabilistic Token Alignment for Large Language Model Fusion
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
- 3D Aware Region Prompted Vision Language Model
- ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
- 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
- SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
- Docopilot: Improving Multimodal Models for Document-Level Understanding
- The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
- FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
- MultiJustice: A Chinese Dataset for Multi-Party, Multi-Charge Legal Prediction
- Advancing Financial Engineering with Foundation Models: Progress, Applications, and Challenges
- Pre-Trained Policy Discriminators are General Reward Models
- Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
- Inference-Time Reward Hacking in Large Language Models
- Synthetic Visual Genome
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- Mixture of Small and Large Models for Chinese Spelling Check
- Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
- Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
- ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
- Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
- RM-R1: Reward Modeling as Reasoning
- CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- Knowledge Grafting of Large Language Models
- Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance
- AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear Mapping
- Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
- Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
- Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM
- Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
- UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- Large Language Models for Computer-Aided Design: A Survey
- AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
- HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models
- Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation
- From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- CPG-EVAL: A Multi-Tiered Benchmark for Evaluating the Chinese Pedagogical Grammar Competence of Large Language Models
- Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Discussions
Related