The Falcon Series of Open Language Models
2023/11/28 by Ebtesam Almazrouei, Hamza Alobeidli, Almazrouei, Ebtesam +23 · 122 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2311.16867
openalex publication_date 2023/11/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce the Falcon series: 7B, 40B, and 180B parameters causal decoder-only models trained on a diverse high-quality corpora predominantly assembled from web data. The largest model, Falcon-180B, has been trained on over 3.5 trillion tokens of text--the largest openly documented pretraining run. Falcon-180B significantly outperforms models such as PaLM or Chinchilla, and improves upon concurrently developed models such as LLaMA 2 or Inflection-1. It nears the performance of PaLM-2-Large at a reduced pretraining and inference cost, making it, to our knowledge, one of the three best language models in the world along with GPT-4 and PaLM-2-Large. We report detailed evaluations, as well as a deep dive into the methods and custom tooling employed to pretrain Falcon. Notably, we report on our custom distributed training codebase, allowing us to efficiently pretrain these models on up to 4,096 A100s on cloud AWS infrastructure with limited interconnect. We release a 600B tokens extract of our web dataset, as well as the Falcon-7/40/180B models under a permissive license to foster open-science and accelerate the development of an open ecosystem of large language models.
Cited by
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- Subjective Question Generation and Answer Evaluation using NLP
- Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- REMODEL-LLM: Transforming C code to Java using LLMs
- Interpreto: An Explainability Library for Transformers
- MindShift: Analyzing Language Models' Reactions to Psychological Prompts
- Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- Menta: A Small Language Model for On-Device Mental Health Prediction
- Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Evaluation of Large Language Models for Numeric Anomaly Detection in Power Systems
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- The PLLuM Instruction Corpus
- AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
- Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
- Secu-Table: a Comprehensive security table dataset for evaluating semantic table interpretation systems
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
- Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
- Mediocrity is the key for LLM as a Judge Anchor Selection
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- Model-Aware Tokenizer Transfer
- ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- Active Model Selection for Large Language Models
- Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Models
- Mid-Training of Large Language Models: A Survey
- On the Role of Unobserved Sequences on Sample-based Uncertainty Quantification for LLMs
- DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
- VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Exploring Similarity between Neural and LLM Trajectories in Language Processing
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- Best-of-∞ -- Asymptotic Performance of Test-Time LLM Ensembling
- Predicting LLM Reasoning Performance with Small Proxy Model
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Understanding Subword Compositionality of Large Language Models
- DRES: Fake news detection by dynamic representation and ensemble selection
- Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
- SENTRA: Selected-Next-Token Transformer for LLM Text Detection
- HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
- HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
- Bias after Prompting: Persistent Discrimination in Large Language Models
- EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics
- AnomalyExplainer Explainable AI for LLM-based anomaly detection using BERTViz and Captum
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- QU-NLP at QIAS 2025 Shared Task: A Two-Phase LLM Fine-Tuning and Retrieval-Augmented Generation Approach for Islamic Inheritance Reasoning
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- SinLlama -- A Large Language Model for Sinhala
- Artificial intelligence-driven computational methods for antibody design and optimization
- MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- Defend LLMs Through Self-Consciousness
- Authorship Attribution in Multilingual Machine-Generated Texts
- GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
- Knowledge Editing for Multi-Hop Question Answering Using Semantic Analysis
- SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding
- Cloud Native System for LLM Inference Serving
- Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
- SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
- MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning
- Stylometry recognizes human and LLM-generated texts in short samples
- Should We Still Pretrain Encoders with Masked Language Modeling?
- Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
- Spark Transformer: Reactivating Sparsity in FFN and Attention
- CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems
- A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations
- Overview of the ClinIQLink 2025 Shared Task on Medical Question-Answering
- SeqPE: Transformer with Sequential Position Encoding
- Just Go Parallel: Improving the Multilingual Capabilities of Large Language Models
- Detecting Hard-Coded Credentials in Software Repositories via LLMs
- MALM: A Multi-Information Adapter for Large Language Models to Mitigate Hallucination
- QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
- Exploring Cultural Variations in Moral Judgments with Large Language Models
- DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains
- Elementary Math Word Problem Generation using Large Language Models
- SoK: Are Watermarks in LLMs Ready for Deployment?
- Position: EU AI Act's Research Exemptions Can Break the Publication Norms of Major AI Conferences
- Beyond Text Compression: Evaluating Tokenizers Across Scales
- Comparing LLM-generated and human-authored news text using formal syntactic theory
- Leveraging Knowledge Graphs and LLMs for Structured Generation of Misinformation
- Stepsize anything: A unified learning rate schedule for budgeted-iteration training
- Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
- Speeding up Model Loading with fastsafetensors
- Self-Critique and Refinement for Faithful Natural Language Explanations
- Ratas framework: A comprehensive genai-based approach to rubric-based marking of real-world textual exams
- Explaining Large Language Models with gSMILE
- Optimization-Inspired Few-Shot Adaptation for Large Language Models
- A Position Paper on the Automatic Generation of Machine Learning Leaderboards
- Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
- Why ‘open’ AI systems are actually closed, and why this matters
- Entailed Opinion Matters: Improving the Fact-Checking Performance of Language Models by Relying on their Entailment Ability
- Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
- Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
- Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
- EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
- Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
- Towards AI-Driven Human-Machine Co-Teaming for Adaptive and Agile Cyber Security Operation Centers
- Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
- LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
- UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- Buy versus Build an LLM: A Decision Framework for Governments
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
- gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
- NNTile: a machine learning framework capable of training extremely large GPT language models on a single node
- Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
- Video Summarization with Large Language Models
- Can We Edit LLMs for Long-Tail Biomedical Knowledge?
- Can LLMs Revolutionize the Design of Explainable and Efficient TinyML Models?
- Layer-Aware Embedding Fusion for LLMs in Text Classifications
- On the Impact of Language Nuances on Sentiment Analysis with Large Language Models: Paraphrasing, Sarcasm, and Emojis
Related