Phi-4 Technical Report
2024/12/12 by Marah Abdin, Jyoti Aneja, Abdin, Marah +51 · 4 voices · 175 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2412.08905
Abstract
We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go beyond distillation. Despite minimal changes to the phi-3 architecture, phi-4 achieves strong performance relative to its size -- especially on reasoning-focused benchmarks -- due to improved data, training curriculum, and innovations in the post-training scheme.
Cited by
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data
- Can abstract concepts from LLM improve SLM performance?
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- When Reasoning Meets Its Laws
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
- AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org
- REMODEL-LLM: Transforming C code to Java using LLMs
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- The Effect of Belief Boxes and Open-mindedness on Persuasion
- BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
- Menta: A Small Language Model for On-Device Mental Health Prediction
- Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System
- WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
- OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
- Progressive Code Integration for Abstractive Bug Report Summarization
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Unexplored flaws in multiple-choice VQA evaluations
- Can LLMs extract human-like fine-grained evidence for evidence-based fact-checking?
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- AppSelectBench: Application-Level Tool Selection Benchmark
- CLASH: A Benchmark for Cross-Modal Contradiction Detection
- LLMs-Powered Real-Time Fault Injection: An Approach Toward Intelligent Fault Test Cases Generation
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
- Integrating Symbolic Natural Language Understanding and Language Models for Word Sense Disambiguation
- Contrastive vision-language learning with paraphrasing and negation
- NLP Datasets for Idiom and Figurative Language Tasks
- Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
- Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Characterizing AI Manipulation Risks in Brazilian YouTube Climate Discourse
- CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
- Generating Software Architecture Description from Source Code using Reverse Engineering and Large Language Model
- Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
- In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
- LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
- DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
- Proactive DDoS Detection and Mitigation in Decentralized Software-Defined Networking via Port-Level Monitoring and Zero-Training Large Language Models
- Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
- Automated Extract Method Refactoring with Open-Source LLMs: A Comparative Study
- From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
- BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
- Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Learning to Reason Efficiently with Discounted Reinforcement Learning
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- Thought Communication in Multiagent Collaboration
- Citation Failure: Definition, Analysis and Efficient Mitigation
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
- Composition-Grounded Instruction Synthesis for Visual Reasoning
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- Toward Cybersecurity-Expert Small Language Models
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
- Representation-Based Exploration for Language Models: From Test-Time to Post-Training
- Investigating Large Language Models' Linguistic Abilities for Text Preprocessing
- LogiNumSynth: Synthesizing Joint Logical-Numerical Reasoning Problems for Language Models
- A Survey on Agentic Multimodal Large Language Models
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
- TinyGraphEstimator: Adapting Lightweight Language Models for Graph Structure Inference
- Neuron-Level Analysis of Cultural Understanding in Large Language Models
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- Mid-Training of Large Language Models: A Survey
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- ContextNav: Towards Agentic Multimodal In-Context Learning
- NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
- Contrastive Retrieval Heads Improve Attention-Based Re-Ranking
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Quantitative Certification of Agentic Tool Selection
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
- OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
- Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
- DyFlow: Dynamic Workflow Framework for Agentic Reasoning
- DescribeEarth: Describe Anything for Remote Sensing Images
- Reference-Free Rating of LLM Responses via Latent Information
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective
- Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
- LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- MalLoc: Toward Fine-grained Android Malicious Payload Localization via LLMs
- AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
- MolPILE -- large-scale, diverse dataset for molecular representation learning
- ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
- Session-Level Spoken Language Assessment with a Multimodal Foundation Model via Multi-Target Learning
- From Hype to Insight: Rethinking Large Language Model Integration in Visual Speech Recognition
- A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
- Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Biomedical Hypothesis Explainability with Graph-Based Context Retrieval
- The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis
- Measuring Epistemic Humility in Multimodal Large Language Models
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- Temporal Counterfactual Explanations of Behaviour Tree Decisions
- Large language models surpass domain-specific architectures for antepartum electronic fetal monitoring analysis
- Text-Trained LLMs Can Zero-Shot Extrapolate PDE Dynamics, Revealing a Three-Stage In-Context Learning Mechanism
- Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
- Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
- GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
- Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
- Do small language models generate realistic variable-quality fake news headlines?
- When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment
- Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
- When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
- Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Datarus-R1: An Adaptive Multi-Step Reasoning LLM for Automated Data Analysis
- ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
- Learning Facts at Scale with Active Reading
- DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
- HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways
- Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- Use of a genetic algorithm to find solutions to introductory physics problems
- Who is a Better Player: LLM against LLM
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
- Activation-Guided Local Editing for Jailbreaking Attacks
- Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- PhysicsEval: Inference-Time Techniques to Improve the Reasoning Proficiency of Large Language Models on Physics Problems
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
- GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
Discussions
Related