InternLM2 Technical Report
2024/03/26 by Zheng Cai, Maosong Cao, Cai, Zheng +199 · 2 voices · 57 citations
Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Topic Modeling #Machine Learning in Healthcare
paper · pdf · doi:10.48550/arxiv.2403.17297
Abstract
The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advancements in open-source models has been challenging. This paper introduces InternLM2, an open-source LLM that outperforms its predecessors in comprehensive evaluations across 6 dimensions and 30 benchmarks, long-context modeling, and open-ended subjective evaluations through innovative pre-training and optimization techniques. The pre-training process of InternLM2 is meticulously detailed, highlighting the preparation of diverse data types including text, code, and long-context data. InternLM2 efficiently captures long-term dependencies, initially trained on 4k tokens before advancing to 32k tokens in pre-training and fine-tuning stages, exhibiting remarkable performance on the 200k ``Needle-in-a-Haystack" test. InternLM2 is further aligned using Supervised Fine-Tuning (SFT) and a novel Conditional Online Reinforcement Learning from Human Feedback (COOL RLHF) strategy that addresses conflicting human preferences and reward hacking. By releasing InternLM2 models in different training stages and model sizes, we provide the community with insights into the model's evolution.
Cited by
- Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
- VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
- CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
- Why Do Vision Language Models Struggle To Recognize Human Emotions?
- Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
- Do Chinese models speak Chinese languages?
- Transformers without Normalization
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VACoT: Rethinking Visual Data Augmentation with VLMs
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Code-driven Number Sequence Calculation: Enhancing the inductive Reasoning Abilities of Large Language Models
- Confidence as a Reward: Transforming LLMs into Reward Models
- The Harder The Better: Maintaining Supervised Fine-tuning Generalization with Less but Harder Data
- LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation
- Conjecturing: An Overlooked Step in Formal Mathematical Reasoning
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Mid-Training of Large Language Models: A Survey
- Online Rubrics Elicitation from Pairwise Comparisons
- Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial Trading
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
- Probabilistic Token Alignment for Large Language Model Fusion
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
- 3D Aware Region Prompted Vision Language Model
- ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
- 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
- SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
Discussions
Related