Yi: Open Foundation Models by 01.AI
2024/03/07 by 01. AI, AI, AI, 01. +65 · 2 voices · 166 citations
Computer Science · #Computer science #Foundation (evidence) #Image Processing and 3D Reconstruction #Law #Neural Networks and Applications #Political science
paper · pdf · doi:10.48550/arxiv.2403.04652
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/03/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce the Yi model family, a series of language and multimodal models that demonstrate strong multi-dimensional capabilities. The Yi model family is based on 6B and 34B pretrained language models, then we extend them to chat models, 200K long context models, depth-upscaled models, and vision-language models. Our base models achieve strong performance on a wide range of benchmarks like MMLU, and our finetuned chat models deliver strong human preference rate on major evaluation platforms like AlpacaEval and Chatbot Arena. Building upon our scalable super-computing infrastructure and the classical transformer architecture, we attribute the performance of Yi models primarily to its data quality resulting from our data-engineering efforts. For pretraining, we construct 3.1 trillion tokens of English and Chinese corpora using a cascaded data deduplication and quality filtering pipeline. For finetuning, we polish a small scale (less than 10K) instruction dataset over multiple iterations such that every single instance has been verified directly by our machine learning engineers. For vision-language, we combine the chat language model with a vision transformer encoder and train the model to align visual representations to the semantic space of the language model. We further extend the context length to 200K through lightweight continual pretraining and demonstrate strong needle-in-a-haystack retrieval performance. We show that extending the depth of the pretrained checkpoint through continual pretraining further improves performance. We believe that given our current results, continuing to scale up model parameters using thoroughly optimized data will lead to even stronger frontier models.
Cited by
- TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
- CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
- Octopus v4: Graph of language models
- OmniSVG: A Unified Scalable Vector Graphics Generation Model
- Do Chinese models speak Chinese languages?
- ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
- An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data
- Towards Long-window Anchoring in Vision-Language Model Distillation
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
- HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
- REMODEL-LLM: Transforming C code to Java using LLMs
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- A Latent Variable Framework for Scaling Laws in Large Language Models
- Completion by Comprehension: Guiding Code Generation with Multi-Granularity Understanding
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
- CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
- ParliaBench: An Evaluation and Benchmarking Framework for LLM-Generated Parliamentary Speech
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- Towards Human-Centric Intelligent Treatment Planning for Radiation Therapy
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
- Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
- Topological Alignment of Shared Vision-Language Embedding Space
- Active Model Selection for Large Language Models
- Lossless Vocabulary Reduction for Auto-Regressive Language Models
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
- CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
- Mechanistic Interpretability of Socio-Political Frames in Language Models
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- Vision Function Layer in Multimodal LLMs
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Exploring Similarity between Neural and LLM Trajectories in Language Processing
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Mapping Overlaps in Benchmarks through Perplexity in the Wild
- Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection
- MMPB: It's Time for Multi-Modal Personalization
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responses
- BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
- RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
- LoRALib: A Standardized Benchmark for Evaluating LoRA-MoE Methods
- From Parameters to Performance: A Data-Driven Study on LLM Structure and Development
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
- Empirical Study of Code Large Language Models for Binary Security Patch Detection
- LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
- DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
- RepoMark: A Data-Usage Auditing Framework for Code Large Language Models
- Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
- FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
- ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Generics and Default Reasoning in Large Language Models
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
- PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
- Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
- MemoCue: Empowering LLM-Based Agents for Human Memory Recall via Strategy-Guided Querying
- Libra: Large Chinese-based Safeguard for AI Content
- A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat
- MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
- AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
- ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
- Psychology-Driven Enhancement of Humour Translation
- SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
- KAT-V1: Kwai-AutoThink Technical Report
- Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
- MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
- Train-before-Test Harmonizes Language Model Rankings
- Pre-Trained Policy Discriminators are General Reward Models
- Legal text summarization via judicial syllogism with large language models
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- Linearly Decoding Refused Knowledge in Aligned Language Models
- Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection
- WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- Probing the Robustness of Large Language Models Safety to Latent Perturbations
- Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language Models
- Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality
- Sampling from Your Language Model One Byte at a Time
- Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
- MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
- Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
- CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making
- Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models
- Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
- Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
- Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
- AdaptiveLLM: A Framework for Selecting Optimal Cost-Efficient LLM for Code-Generation Based on CoT Length
- Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
- LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data
- RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
- DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes
- LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
- Multiple LLM Agents Debate for Equitable Cultural Alignment
- Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
- UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
- RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
- Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation
- Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
- Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
- Efficient Multi-modal Long Context Learning for Training-free Adaptation
- 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
- A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
- Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning
- Understanding Gated Neurons in Transformers from Their Input-Output Functionality
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
- Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
- PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
- PsyMem: Fine-grained psychological alignment and Explicit Memory Control for Advanced Role-Playing LLMs
- Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
- Advancing Sequential Numerical Prediction in Autoregressive Models
- Noise Injection Systemically Degrades Large Language Model Safety Guardrails
- Improve Rule Retrieval and Reasoning with Self-Induction and Relevance ReEstimate
- Exploring the Feasibility of Multilingual Grammatical Error Correction with a Single LLM up to 9B parameters: A Comparative Study of 17 Models
- FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
- Reasoning Capabilities and Invariability of Large Language Models
- Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
- Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation
- Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
- UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
- UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
- C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
- Redefining Machine Translation on Social Network Services with Large Language Models
Discussions
Related