Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
2024/04/22 by Marah Abdin, Abdin, Marah, Jyoti Aneja +262 · 8 voices · 349 citations
Computer Science · #Computational Physics and Python Applications
paper · pdf · doi:10.48550/arxiv.2404.14219
Abstract
We introduce phi-3-mini, a 3.8 billion parameter language model trained on 3.3 trillion tokens, whose overall performance, as measured by both academic benchmarks and internal testing, rivals that of models such as Mixtral 8x7B and GPT-3.5 (e.g., phi-3-mini achieves 69% on MMLU and 8.38 on MT-bench), despite being small enough to be deployed on a phone. Our training dataset is a scaled-up version of the one used for phi-2, composed of heavily filtered publicly available web data and synthetic data. The model is also further aligned for robustness, safety, and chat format. We also provide parameter-scaling results with a 7B, 14B models trained for 4.8T tokens, called phi-3-small, phi-3-medium, both significantly more capable than phi-3-mini (e.g., respectively 75%, 78% on MMLU, and 8.7, 8.9 on MT-bench). To enhance multilingual, multimodal, and long-context capabilities, we introduce three models in the phi-3.5 series: phi-3.5-mini, phi-3.5-MoE, and phi-3.5-Vision. The phi-3.5-MoE, a 16 x 3.8B MoE model with 6.6 billion active parameters, achieves superior performance in language reasoning, math, and code tasks compared to other open-source models of similar scale, such as Llama 3.1 and the Mixtral series, and on par with Gemini-1.5-Flash and GPT-4o-mini. Meanwhile, phi-3.5-Vision, a 4.2 billion parameter model derived from phi-3.5-mini, excels in reasoning tasks and is adept at handling both single-image and text prompts, as well as multi-image and text prompts.
Cited by
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments
- QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization
- LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization
- From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
- RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
- PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image
- Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field
- Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
- Octopus v4: Graph of language models
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- In-Place Tokenizer Expansion for Pre-trained LLMs
- From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing
- CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models
- Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study
- Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating
- Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
- GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
- Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
- Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought
- Fifty Shades of Greenwashing: The Political Economy of Climate Change Advertising on Social Media
- LLMs can hide text in other text of the same length
- SAM 3D: 3Dfy Anything in Images
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
- Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
- Small Language Models are the Future of Agentic AI
- Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- Enough Coin Flips Can Make LLMs Act Bayesian
- Layers at Similar Depths Generate Similar Activations Across LLM Architectures
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
- Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
- Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
- CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
- TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
- CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generation
- Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control
- When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs
- AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- CheXPO-v2: Preference Optimization for Chest X-ray VLMs with Knowledge Graph Consistency
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
- Sigma-MoE-Tiny Technical Report
- FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- NL2SpaTiaL: Generating Geometric Spatio-Temporal Logic Specifications from Natural Language for Manipulation Tasks
- Integrating Causal Reasoning into Automated Fact-Checking
- Moment and Highlight Detection via MLLM Frame Segmentation
- Using GUI Agent for Electronic Design Automation
- Leveraging LLMs for Title and Abstract Screening for Systematic Review: A Cost-Effective Dynamic Few-Shot Learning Approach
- CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
- VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
- Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
- Local LLM Ensembles for Zero-shot Portuguese Named Entity Recognition
- Chasing Shadows: Pitfalls in LLM Security Research
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
- David vs. Goliath: Can Small Models Win Big with Agentic AI in Hardware Design?
- AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- Jina-VLM: Small Multilingual Vision Language Model
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Unexplored flaws in multiple-choice VQA evaluations
- Mirror, Mirror on the Wall -- Which is the Best Model of Them All?
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Skypilot: Fine-Tuning LLM with Physical Grounding for AAV Coverage Search
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Multimodal Evaluation of Russian-language Architectures
- Attention Grounded Enhancement for Visual Document Retrieval
- Detecting and Steering LLMs' Empathy in Action
- TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- ChexFract: From General to Specialized -- Enhancing Fracture Description Generation
- Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
- Seeing Straight: Document Orientation Detection for Efficient OCR
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- UTF-8 Plumbing: Byte-level Tokenizers Unavoidably Enable LLMs to Generate Ill-formed UTF-8
- TAUE: Training-free Noise Transplant and Cultivation Diffusion Model
- Assessing LLM Reasoning Steps via Principal Knowledge Grounding
- ShadowLogic: Backdoors in Any Whitebox LLM
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
- ChartAB: A Benchmark for Chart Grounding & Dense Alignment
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
- Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
- Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
- GReF: A Unified Generative Framework for Efficient Reranking via Ordered Multi-token Prediction
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- A Survey on Efficient Vision-Language-Action Models
- Lightweight Robust Direct Preference Optimization
- LightKGG: Simple and Efficient Knowledge Graph Generation from Textual Data
- A Survey on LLM Mid-Training
- Agentic Meta-Orchestrator for Multi-task Copilots
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- Efficient Low Rank Attention for Long-Context Inference in Large Language Models
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- A Benchmark Dataset And LLMs Comparison For NFR Classification With Explainable AI
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- An Evaluation of LLMs Inference on Popular Single-board Computers
- LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding
- Train a Unified Multimodal Data Quality Classifier with Synthetic Data
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
- Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
- Bridging Semantics & Structure for Software Vulnerability Detection using Hybrid Network Models
- The Achilles' Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
- On the Representations of Entities in Auto-regressive Large Language Models
- Zero-shot image privacy classification with Vision-Language Models
- Understanding the Effects of Domain Finetuning on LLMs
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Automatic Text Box Placement for Supporting Typographic Design
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
- Mid-Training of Large Language Models: A Survey
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- JAI-1: A Thai-Centric Large Language Model
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- H-DDx: A Hierarchical Evaluation Framework for Differential Diagnosis
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
- Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
- ModernVBERT: Towards Smaller Visual Document Retrievers
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
- Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
- From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
- Pushing LLMs to Their Logical Reasoning Bound: The Role of Data Reasoning Intensity
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Analyzing and Evaluating Unbiased Language Model Watermark
- An Ensemble Framework for Unbiased Language Model Watermarking
- PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- Evaluating Program Semantics Reasoning with Type Inference in System F
- LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
- Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
- Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
- Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
- Query-Centric Graph Retrieval Augmented Generation
- Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
- Accelerate Creation of Product Claims Using Generative AI
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
- Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks
- Polarity Detection of Sustainable Detection Goals in News Text
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- From Global to Local: Social Bias Transfer in CLIP
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
- AccessEval: Benchmarking Disability Bias in Large Language Models
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- mmExpert: Integrating Large Language Models for Comprehensive mmWave Data Synthesis and Understanding
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- AToken: A Unified Tokenizer for Vision
- TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
- Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- Automated and Context-Aware Code Documentation Leveraging Advanced LLMs
- Positional Encoding via Token-Aware Phase Attention
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- A Dynamic Knowledge Update-Driven Model with Large Language Models for Fake News Detection
- Pluralistic Off-policy Evaluation and Alignment
- Continually Adding New Languages to Multilingual Language Models
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- LoRALib: A Standardized Benchmark for Evaluating LoRA-MoE Methods
- LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
- TrueSkin: Towards Fair and Accurate Skin Tone Recognition and Generation
- Humanizing Automated Programming Feedback: Fine-Tuning Generative Models with Student-Written Feedback
- RefactorCoderQA: Benchmarking LLMs for Multi-Domain Coding Question Solutions in Cloud and Edge Deployment
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communities
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
- HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
- MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
- Self-Aligned Reward: Towards Effective and Efficient Reasoners
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
- How Small is Enough? Empirical Evidence of Quantized Small Language Models for Automated Program Repair
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- RoboBuddy in the Classroom: Exploring LLM-Powered Social Robots for Storytelling in Learning and Integration Activities
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
- Dynamic Sparse Attention on Mobile SoCs
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Kwai Keye-VL 1.5 Technical Report
- DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
- VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
- DriveQA: Passing the Driving Knowledge Test
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- Scalable Object Detection in the Car Interior With Vision Foundation Models
- Ensemble Debates with Local Large Language Models for AI Alignment
- Reflective Agreement: Combining Self-Mixture of Agents with a Sequence Tagger for Robust Event Extraction
- MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
- Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
- Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning
- PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- Prompt Orchestration Markup Language
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
- Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
- Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
- Diffusion is a code repair operator and generator
- AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
- Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
- Apriel-Nemotron-15B-Thinker
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- Learning Facts at Scale with Active Reading
- From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
- SinLlama -- A Large Language Model for Sinhala
- Classifier Language Models: Unifying Sparse Finetuning and Adaptive Tokenization for Specialized Classification Tasks
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- Sample-efficient LLM Optimization with Reset Replay
- LLM Unlearning Without an Expert Curated Dataset
- InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
- DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
- RL-MoE: An Image-Based Privacy Preserving Approach In Intelligent Transportation System
- Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- EvoGraph: Hybrid Directed Graph Evolution toward Software 3.0
- Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
- Modelling and Classifying the Components of a Literature Review
- Sustainability assessment using multimodal AI agents
- ChartCap: Mitigating Hallucination of Dense Chart Captioning
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models
- What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
- Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Enhancing Spatial Reasoning through Visual and Textual Thinking
- CodeNER: Code Prompting for Named Entity Recognition
- Infogen: Generating Complex Statistical Infographics from Documents
- How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
- True Multimodal In-Context Learning Needs Attention to the Visual Context
- The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer
- Beyond Isolated Dots: Benchmarking Structured Table Construction as Deep Knowledge Extraction
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
- Characterizing State Space Model (SSM) and SSM-Transformer Hybrid Language Model Performance with Long Context Length
- AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
- Probing for Arithmetic Errors in Language Models
- Improving Contextual ASR via Multi-grained Fusion with Large Language Models
- POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
- IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
- Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models
- Scaling Laws for Optimal Data Mixtures
- Multilingual Multimodal Software Developer for Code Generation
- Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
- Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
- Orchestration for Domain-specific Edge-Cloud Language Models
- Multi-Actor Generative Artificial Intelligence as a Game Engine
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- Towards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models
- Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses
- CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale
- Beyond Scale: Small Language Models are Comparable to GPT-4 in Mental Health Understanding
- Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
- BlueLM-2.5-3B Technical Report
- Vision-Language Models Can't See the Obvious
- A Technical Survey of Reinforcement Learning Techniques for Large Language Models
- Conformal Information Pursuit for Interactively Guiding Large Language Models
- BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
- List of large language models [wikipedia]
Discussions
- Phi-3 Technical Report [hn, 411 points, 130 comments]
- Phi-3 is out, from MSFT. 1) Mini = small enough to deploy on a phone. 2) Medium = small enough to run on a laptop, but better than GPT-3.5. 3) They claim the advance is almost entirely due to data *qu [bsky, 14 points, 4 comments]
- arxiv.org/abs/2404.14219 なんだとお?! スマホでGPT-3.5クラスやと?! [bsky, 2 points, 1 comments]
- arxiv.org/abs/2404.14219 [bsky, 1 points, 0 comments]
- Microsoft Launches Phi-3 (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- arxiv.org/abs/2404.14219 [bsky, 0 points, 1 comments]
- Phi3-small, our new open model, at 7B parameters rivals GPT3.5, proving the quality of training data matters more than pure size: https://arxiv.org/pdf/2404.14219.pdf [bsky, 0 points, 0 comments]
- Phi-3 Technical Report https://arxiv.org/abs/2404.14219 [bsky, 0 points, 0 comments]
Related