Ray: A Distributed Framework for Emerging AI Applications
2017/12/16 by Philipp Moritz, Moritz, Philipp, Robert Nishihara +19 · 1 voice · 85 citations
Computer Science · #Cloud Computing and Resource Management #Parallel Computing and Optimization Techniques #Reinforcement Learning in Robotics #cs.AI #cs.DC #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1712.05889
openalex publication_date 2017/12/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The next generation of AI applications will continuously interact with the environment and learn from these interactions. These applications impose new and demanding systems requirements, both in terms of performance and flexibility. In this paper, we consider these requirements and present Ray---a distributed system to address them. Ray implements a unified interface that can express both task-parallel and actor-based computations, supported by a single dynamic execution engine. To meet the performance requirements, Ray employs a distributed scheduler and a distributed and fault-tolerant store to manage the system's control state. In our experiments, we demonstrate scaling beyond 1.8 million tasks per second and better performance than existing specialized systems for several challenging reinforcement learning applications.
Citations
Cited by
- Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
- AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning
- Retriever: Composing Closed-Loop Asynchronous Robot Programs
- KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
- JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
- A Probabilistic Framework for LLM-Based Model Discovery
- RollArt: Disaggregated Multi-Task Agentic RL Training at Scale
- Role-Based Fault Tolerance System for LLM RL Post-Training
- DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training
- BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
- Scaling Point-based Differentiable Rendering for Large-scale Reconstruction
- RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
- IslandRun: Privacy-Aware Multi-Objective Orchestration for Distributed AI Inference
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation
- Geometrically-Constrained Agent for Spatial Reasoning
- Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- Meta-Black-Box Optimization with Bi-Space Landscape Analysis and Dual-Control Mechanism for SAEA
- SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs
- FlowPath: Learning Data-Driven Manifolds with Invertible Flows for Robust Irregularly-sampled Time Series Classification
- TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
- Continuum Dropout for Neural Differential Equations
- Experiences Building Enterprise-Level Privacy-Preserving Federated Learning to Power AI for Science
- Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
- Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
- MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments
- Fix: externalizing network I/O in serverless computing
- FlowMesh: A Service Fabric for Composable LLM Workflows
- ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform
- ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
- TridentServe: A Stage-level Serving System for Diffusion Pipelines
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- Dara: Automated multiple-hypothesis phase identification and refinement from powder X-ray diffraction
- HEADER: Hierarchical Robot Exploration via Attention-Based Deep Reinforcement Learning with Expert-Guided Reward
- An Elastic Job Scheduler for HPC Applications on the Cloud
- Deducing Closed-Form Expressions for Bright-Solitons in Strongly Magnetized Plasmas with Physics Informed Symbolic Regression (PISR)
- Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
- Dynamic SBI: Round-free Sequential Simulation-Based Inference with Adaptive Datasets
- Laminar: A Scalable Asynchronous RL Post-Training Framework
- A Decentralized Microservice Scheduling Approach Using Service Mesh in Cloud-Edge Systems
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- AlphaApollo: Orchestrating Foundation Models and Professional Tools into a Self-Evolving System for Deep Agentic Reasoning
- Lattica: A Decentralized Cross-NAT Communication Framework for Scalable AI Inference and Training
- SCUBA: Salesforce Computer Use Benchmark
- HAPT: Heterogeneity-Aware Automated Parallel Training on Heterogeneous Clusters
- A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
- Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
- Experience Deploying Containerized GenAI Services at an HPC Center
- OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC
- APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification
- Learning Interpretable Differentiable Logic Networks for Time-Series Classification
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- Driver analysis of subarctic wildfire severity over a 35-year period
- Federated Learning for Deforestation Detection: A Distributed Approach with Satellite Imagery
- Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
- GRATE: a Graph transformer-based deep Reinforcement learning Approach for Time-efficient autonomous robot Exploration
- TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
- On Hyperparameters and Backdoor-Resistance in Horizontal Federated Learning
- On the Duality of Task and Actor Programming Models
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
- Batch Query Processing and Optimization for Agentic Workflows
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Nash Q-Network for Multi-Agent Cybersecurity Simulation
- HyperFlexis: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
- DEPTH: Hallucination-Free Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
- Declarative Data Pipeline for Large Scale ML Services
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- Reinforcement learning in densely recurrent biological networks
- WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
- Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment
- High-Performance Statistical Computing (HPSC): Challenges, Opportunities, and Future Directions
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- Qwen-Image Technical Report
- RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
- Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
- G-Core: A Simple, Scalable and Balanced RLHF Trainer
- Toward Trusted Onboard AI: Advancing Small Satellite Operations using Reinforcement Learning
Discussions
Related