MapReduce
2008/01/01 by Jeffrey Dean, Jay B. Dean, Sanjay Ghemawat · 304 citations
Computer Science · #Cloud Computing and Resource Management #Advanced Data Storage Technologies #Parallel Computing and Optimization Techniques
paper · pdf · doi:10.1145/1327452.1327492
Abstract
MapReduce is a programming model and an associated implementation for processing and generating large datasets that is amenable to a broad variety of real-world tasks. Users specify the computation in terms of a map and a reduce function, and the underlying runtime system automatically parallelizes the computation across large-scale clusters of machines, handles machine failures, and schedules inter-machine communication to make efficient use of the network and disks. Programmers find the system easy to use: more than ten thousand distinct MapReduce programs have been implemented internally at Google over the past four years, and an average of one hundred thousand MapReduce jobs are executed on Google's clusters every day, processing a total of more than twenty petabytes of data per day.
Cited by
- PaSh
- A Pearson’s correlation coefficient based decision tree and its parallel implementation
- MTHAEL: Cross-Architecture IoT Malware Detection Based on Neural Network Advanced Ensemble Learning
- Incremental Query Processing on Big Data Streams
- Partitioning networks into clusters of synchronized nodes via the message-passing algorithm: A scalable approach
- More Parts Than Elements: How Databases Multiply
- Pipelined Gradient Coding
- Evidence-Aware MapReduce for Forkable Compute
- Language Model Teams as Distributed Systems
- DataRater: Meta-Learned Dataset Curation
- New lower bounds for Massively Parallel Computation from query complexity
- Accurate and Fast Federated Learning via IID and Communication-Aware Grouping
- Big Data Staging with MPI-IO for Interactive X-ray Science
- Communication Steps for Parallel Query Processing
- Scalable and Efficient Construction of Suffix Array with MapReduce and In-Memory Data Store System
- Theoretical and Empirical Analysis of a Parallel Boosting Algorithm
- Partout: A Distributed Engine for Efficient RDF Processing
- On the Approximability of Related Machine Scheduling under Arbitrary Precedence
- Robust Gradient Descent via Moment Encoding with LDPC Codes
- Revisiting Degree Distribution Models for Social Graph Analysis
- Enabling Loosely-Coupled Serial Job Execution on the IBM BlueGene/P Supercomputer and the SiCortex SC5832
- Flexible Scheduling of Distributed Analytic Applications
- Masked LARk: Masked Learning, Aggregation and Reporting worKflow
- SD-CPS: Taming the Challenges of Cyber-Physical Systems with a\n Software-Defined Approach
- A New Parallelization Method for K-means
- Improved Algorithms for Distributed Entropy Monitoring
- Declarative, Secure, Convergent Edge Computation
- Parallel Knowledge Embedding with MapReduce on a Multi-core Processor
- Relationship Queries on Large graphs using Pregel
- i2MapReduce: Incremental MapReduce for Mining Evolving Big Data
- Parallel Algorithms for Summing Floating-Point Numbers
- Subsampling MCMC - An introduction for the survey statistician
- TensorFlow: A system for large-scale machine learning
- Towards Chip-on-Chip Neuroscience: Fast Mining of Frequent Episodes Using Graphics Processors
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Scaling Point-based Differentiable Rendering for Large-scale Reconstruction
- Online Self-Evolving Anomaly Detection in Cloud Computing Environments
- Joint Design of Embedded Index Coding and Beamforming for MIMO-based Distributed Computing via Multi-Agent Reinforcement Learning
- On the Design and Implementation of Structured P2P VPNs
- Simulations between Strongly Sublinear MPC and Node-Capacitated Clique
- Synchromodulametry: From Coincidence Detection to Coherent State Measurement
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- TurKPF: TurKontrol as a Particle Filter
- Scalable, Fast Cloud Computing with Execution Templates
- Managing Schema Evolution in NoSQL Data Stores
- Distributed Graphical Simulation in the Cloud
- Workflow is All You Need: Escaping the "Statistical Smoothing Trap" via High-Entropy Information Foraging and Adversarial Pacing
- Large scale citation matching using Apache Hadoop
- HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
- Actors vs Shared Memory: two models at work on Big Data application frameworks
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- Optimal Load Balancing in Bipartite Graphs
- An Empirical Study of Cross-Language Interoperability in Replicated Data Systems
- Understanding the Nature of System-Related Issues in Machine Learning Frameworks: An Exploratory Study
- StarDist: A Code Generator for Distributed Graph Algorithms
- On the Feasibility of Distributed Kernel Regression for Big Data
- Cost Analysis of Nondeterministic Probabilistic Programs
- AgentODRL: A Large Language Model-based Multi-agent System for ODRL Generation
- BoPF: Mitigating the Burstiness-Fairness Tradeoff in Multi-Resource Clusters
- Trustless Federated Learning at Edge-Scale: A Compositional Architecture for Decentralized, Verifiable, and Incentive-Aligned Coordination
- Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
- Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
- Meta-Learning surrogate models for sequential decision making
- CPL: A Core Language for Cloud Computing -- Technical Report
- Near-Optimal Massively Parallel Graph Connectivity
- Towards Federated Learning at Scale: System Design
- Chronos: A Unifying Optimization Framework for Speculative Execution of Deadline-critical MapReduce Jobs
- TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Deadline is not Enough: How to Achieve Importance-aware Server-centric Data Centers via a Cross Layer Approach
- Serverless seismic imaging in the cloud
- Efficient Approximation of Volterra Series for High-Dimensional Systems
- Yesquel: scalable SQL storage for Web applications
- A Fundamental Tradeoff between Computation and Communication in Distributed Computing
- DINGO: Distributed Newton-Type Method for Gradient-Norm Optimization
- Making problems tractable on big data via preprocessing with polylog-size output
- Sorting, Searching, and Simulation in the MapReduce Framework
- Using Span Queries to Optimize for Cache and Attention Locality
- Private Map-Secure Reduce: Infrastructure for Efficient AI Data Markets
- ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
- Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
- Graph Processing on FPGAs: Taxonomy, Survey, Challenges
- Optimal Load Allocation for Coded Distributed Computation in Heterogeneous Clusters
- Skew Handling in Aggregate Streaming Queries on GPUs
- A Data–Driven Approximation of the Koopman Operator: Extending Dynamic Mode Decomposition
- Implementation of Algorithms for Right-Sizing Data Centers
- Fusion: An Analytics Object Store Optimized for Query Pushdown
- MLitB: Machine Learning in the Browser
- Online Machine Learning in Big Data Streams
- SwitchAgg:A Further Step Towards In-Network Computation
- Survey and Taxonomy of Lossless Graph Compression and Space-Efficient Graph Representations
- Real-time semiparametric regression for distributed data sets
- Contrasting Effects of Replication in Parallel Systems: From Overload to Underload and Back
- A Survey of Blocking and Filtering Techniques for Entity Resolution
- Dynamic Memory Allocation Policies for Postings in Real-Time Twitter Search
- Optimizing MapReduce for Highly Distributed Environments
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Stochastic Primal-Dual Coordinate Method for Regularized Empirical Risk\n Minimization
- Does The Cloud Need Stabilizing?
- Non-clairvoyant Scheduling of Coflows
- On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML
- Upper and Lower Bounds on the Cost of a Map-Reduce Computation
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Assignment Problems of Different-Sized Inputs in MapReduce
- When Gaussian Process Meets Big Data: A Review of Scalable GPs
- Teaching Machine Learning to Software Engineers
- Time Warp on the Go
- Data Mining Scheme for Globally Distributed Big Data
- Big Data at HPC Wales
- Bayesian Federated Learning over Wireless Networks
- Simple and sharp analysis of k-means||
- Fast Clustering using MapReduce
- Arrows for Parallel Computation
- Distributed Data Processing Frameworks for Big Graph Data
- A Preliminary Review of Influential Works in Data-Driven Discovery
- Parallel Bayesian Additive Regression Trees
- Robust Scheduling for Flexible Processing Networks
- Coflow Scheduling in Data Centers: Routing and Bandwidth Allocation
- TGE-viz : Transition Graph Embedding for Visualization of Plan Traces and Domains
- Evidential instance selection for K-nearest neighbor classification of big data
- Forecasting the 2013–2014 Influenza Season Using Wikipedia
- On the Complexity of Processing Massive, Unordered, Distributed Data
- The OoO VLIW JIT Compiler for GPU Inference
- GLB: Lifeline-based Global Load Balancing library in X10
- Evaluating Device-First Continuum AI (DFC-AI) for Autonomous Operations in the Energy Sector
- A Unified Coding Framework for Distributed Computing with Straggling Servers
- Heterogeneous Coded Distributed Computing: Joint Design of File Allocation and Function Assignment
- Representing emotions with knowledge graphs for movie recommendations
- Scientific Computing Meets Big Data Technology: An Astronomy Use Case
- Dynamic Deferral of Workload for Capacity Provisioning in Data Centers
- Approximate Gradient Coding for Distributed Learning with Heterogeneous Stragglers
- ProGQL: A Provenance Graph Query System for Cyber Attack Investigation
- Institutional Metaphors for Designing Large-Scale Distributed AI versus AI Techniques for Running Institutions
- Embed and Conquer: Scalable Embeddings for Kernel k-Means on MapReduce
- Analysis of Input-Output Mappings in Coinjoin Transactions with Arbitrary Values
- Towards Stochastically Optimizing Data Computing Flows
- Distributed Stochastic Variance Reduced Gradient Methods and A Lower Bound for Communication Complexity
- GPU Tensor Cores for Fast Arithmetic Reductions
- Machine Learning on Graphs: A Model and Comprehensive Taxonomy
- Centrality Metric for Dynamic Networks
- A Decentralized Microservice Scheduling Approach Using Service Mesh in Cloud-Edge Systems
- Graph3S: A Simple, Speedy and Scalable Distributed Graph Processing System
- NOMAD: Non-locking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion
- Agentic Troubleshooting Guide Automation for Incident Management
- Category-Theoretic Quantitative Compositional Distributional Models of Natural Language Semantics
- Scalable Data Cube Analysis over Big Data
- Coflow Scheduling in Input-Queued Switches: Optimal Delay Scaling and Algorithms
- CloudSVM : Training an SVM Classifier in Cloud Computing Systems
- Big Data and Cross-Document Coreference Resolution: Current State and Future Opportunities
- Scalable Similarity Joins of Tokenized Strings
- Parallel Markov Chain Monte Carlo for Bayesian Hierarchical Models with Big Data, in Two Stages
- Algorithms for a Topology-aware Massively Parallel Computation Model
- Distributed Machine Learning for Wireless Communication Networks: Techniques, Architectures, and Applications
- Parallel D2-Clustering: Large-Scale Clustering of Discrete Distributions
- Distributed Machine Learning via Sufficient Factor Broadcasting
- A Combinatorial Design for Cascaded Coded Distributed Computing on General Networks
- Distributed Algorithms for Finding Local Clusters Using Heat Kernel Pagerank
- Fundamental Limits of Distributed Computing for Linearly Separable Functions
- Pilot-Data: An Abstraction for Distributed Data
- CLAMShell: Speeding up Crowds for Low-latency Data Labeling
- Mapping and Reducing the Brain on the Cloud
- Exploratory Analysis of a Terabyte Scale Web Corpus
- Integrating SysML and AUTOSAR: model transformation in automotive MBSE
- PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development
- Energy-efficient photonic neural networks for high-speed AI computation
- FlashR: R-Programmed Parallel and Scalable Machine Learning using SSDs
- UStore: A Distributed Storage With Rich Semantics
- Polypus: a Big Data Self-Deployable Architecture for Microblogging Text Extraction and Real-Time Sentiment Analysis
- Application of Deep Learning Models for Real-Time Automatic Malware Detection
- Global and Local Implications of Computational Artifacts
- Prediction of Video Popularity in the Absence of Reliable Data from\n Video Hosting Services: Utility of Traces Left by Users on the Web
- AGL: a Scalable System for Industrial-purpose Graph Machine Learning
- On new data sources for the production of official statistics
- Jug: Software for Parallel Reproducible Computation in Python
- Pretrained Transformers for Text Ranking: BERT and Beyond
- On the inequality of the 3V's of Big Data Architectural Paradigms: A case for heterogeneity
- Graph Summarization Methods and Applications: A Survey
- ElasticBroker: Combining HPC with Cloud to Provide Realtime Insights into Simulations
- Optimal bandwidth-aware VM allocation for Infrastructure-as-a-Service
- Processing Database Joins over a Shared-Nothing System of Multicore\n Machines
- Equi-depth Histogram Construction for Big Data with Quality Guarantees
- OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC
- Deep Learning At Scale and At Ease
- A Map-Reduce Parallel Approach to Automatic Synthesis of Control\n Software
- Distributed Exploration in Multi-Armed Bandits
- TF-Replicator: Distributed Machine Learning for Researchers
- ALID: Scalable Dominant Cluster Detection
- Thrill: High-Performance Algorithmic Distributed Batch Data Processing\n with C++
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- Thinking Like a Vertex: a Survey of Vertex-Centric Frameworks for Distributed Graph Processing
- Easy Acceleration with Distributed Arrays
- On the Mathematics of Data Centre Network Topologies
- Matrix Factorization at Scale: a Comparison of Scientific Data Analytics in Spark and C+MPI Using Three Case Studies
- Real-Time Machine Learning: The Missing Pieces
- Architectural Tactics for Big Data Cybersecurity Analytic Systems: A Review
- Distributed Proximal Gradient Algorithm for Partially Asynchronous\n Computer Clusters
- L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems
- FLCD: A Flexible Low Complexity Design of Coded Distributed Computing
- Comparative Performance Analysis of Intel Xeon Phi, GPU, and CPU
- Density Estimations for Approximate Query Processing on SIMD Architectures
- Alternative statistical inference for the first normalized incomplete moment
- SAF: Simulated Annealing Fair Scheduling for Hadoop Yarn Clusters
- A Uniqueness Theorem for Distributed Computation under Physical Constraint
- Parallel Approximate Undirected Shortest Paths Via Low Hop Emulators
- ATLAS: a failure-aware scheduler for Hadoop
- Characterizing Data Analysis Workloads in Data Centers
- Deep Learning with Apache SystemML
- Fast attention mechanisms: a tale of parallelism
- Hoplite: Efficient and Fault-Tolerant Collective Communication for Task-Based Distributed Systems
- Communication-Efficient Edge AI: Algorithms and Systems
- ZhiFangDanTai: Fine-tuning Graph-based Retrieval-Augmented Generation Model for Traditional Chinese Medicine Formula
- Benchmarking DataStax Enterprise/Cassandra with HiBench
- A MapReduce Approach to NoSQL RDF Databases
- Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training
- Limited Random Walk Algorithm for Big Graph Data Clustering
- The Two Quadrillionth Bit of Pi is 0! Distributed Computation of Pi with\n Apache Hadoop
- Counterfactual simulations for large scale systems with burnout variables
- From Data Fusion to Knowledge Fusion
- Merlin: A Language for Provisioning Network Resources
- Delivery, consistency, and determinism: rethinking guarantees in distributed stream processing
- Connected Components on a PRAM in Log Diameter Time
- Improving Nonpreemptive Multiserver Job Scheduling with Quickswap
- A Novel IaaS Tax Model as Leverage Towards Green Cloud Computing
- Visualizing a Million Time Series with the Density Line Chart
- HiCR, an Abstract Model for Distributed Heterogeneous Programming
- Collaborative Learning with Limited Interaction: Tight Bounds for Distributed Exploration in Multi-Armed Bandits
- Scaling-Up Reasoning and Advanced Analytics on BigData
- Parallel Programming for FPGAs
- Diversity, Productivity, and Growth of Open Source Developer Communities
- Parallelization in Scientific Workflow Management Systems
- Analysis of building maintenance requests using a text mining approach: building services evaluation
- A Generalized Performance Evaluation Framework for Parallel Systems with\n Output Synchronization
- MMLSpark: Unifying Machine Learning Ecosystems at Massive Scales
- Parallel Evaluation of Interaction Nets: Case Studies and Experiments
- Exploring heterogeneity of unreliable machines for p2p backup
- Efficient Straggler Replication in Large-scale Parallel Computing
- A Reliable Effective Terascale Linear Learning System
- Ray: A Distributed Framework for Emerging AI Applications
- DiskJoin: Large-scale Vector Similarity Join with SSD
- MapReduce Meets Fine-Grained Complexity: MapReduce Algorithms for APSP, Matrix Multiplication, 3-SUM, and Beyond
- Transduction is All You Need for Structured Data Workflows
- Behavioral Simulations in MapReduce
- Scalable Protein Sequence Similarity Search using Locality-Sensitive Hashing and MapReduce
- How to Optimally Allocate Resources for Coded Distributed Computing?
- GraphLab: A New Framework For Parallel Machine Learning
- A novel approach for fast mining frequent itemsets use N-list structure based on MapReduce
- Ripple : Simplified Large-Scale Computation on Heterogeneous Architectures with Polymorphic Data Layout
- Coded Computing for Master-Aided Distributed Computing Systems
- Cloudpress 2.0: A MapReduce Approach for News Retrieval on the Cloud
- Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
- MovePattern: Interactive Framework to Provide Scalable Visualization of Movement Patterns
- Oseba: Optimization for Selective Bulk Analysis in Big Data Processing
- A Stream Pipeline Framework for Digital Payment Programming based on Smart Contracts
- A Comparative Study of Asynchronous Many-Tasking Runtimes: Cilk, Charm++, ParalleX and AM++
- An Alternating Direction Method Approach to Cloud Traffic Management
- Evaluating Complex Task through Crowdsourcing: Multiple Views Approach
- Comparing Spark vs MPI/OpenMP On Word Count MapReduce
- Asynchronous Complex Analytics in a Distributed Dataflow Architecture
- SensorCloud: Towards the Interdisciplinary Development of a Trustworthy\n Platform for Globally Interconnected Sensors and Actuators
- Coding for Distributed Fog Computing
- Estimation of Passenger Route Choice Pattern Using Smart Card Data for Complex Metro Systems
- Online Job Scheduling with Redundancy and Opportunistic Checkpointing: A Speedup-Function-Based Analysis
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- Fork and Join Queueing Networks with Heavy Tails: Scaling Dimension and Throughput Limit
- FLOSS: Federated Learning with Opt-Out and Straggler Support
- Introduction to Rank-polymorphic Programming in Remora (Draft)
- Towards a Periodic Table of Computer System Design Principles
- The NOESIS Network-Oriented Exploration, Simulation, and Induction System
- Collaborative State Machines: A Better Programming Model for the Cloud-Edge-IoT Continuum
- Locality Optimization for Data Parallel Programs
- Query and Resource Optimizations: A Case for Breaking the Wall in Big\n Data Systems
- Whiz: A Fast and Flexible Data Analytics System
- Quantum Adiabatic Evolution for Global Optimization in Big Data
- A Comparative Taxonomy and Survey of Public Cloud Infrastructure Vendors
- Using Variational Inference and MapReduce to Scale Topic Modeling
- Friendship Prediction in Composite Social Networks
- Rapid AkNN Query Processing for Fast Classification of Multidimensional\n Data in the Cloud
- Parallel Hierarchical Affinity Propagation with MapReduce
- Fast Fourier-Based Generation of the Compression Matrix for\n Deterministic Compressed Sensing
- Filter and refine [wikipedia]
- Quake: quality-aware detection and correction of sequencing errors. [europepmc]
- SeqWare Query Engine: storing and searching sequence data in the cloud. [europepmc]
- Nipype: a flexible, lightweight and extensible neuroimaging data processing framework in python. [europepmc]
- Disk-based k-mer counting on a PC. [europepmc]
- The role and challenges of exome sequencing in studies of human diseases. [europepmc]
- Bioinformatics for precision medicine in oncology: principles and application to the SHIVA clinical trial. [europepmc]
- Applications of the MapReduce programming framework to clinical big data analysis: current landscape and future trends. [europepmc]
- aTRAM - automated target restricted assembly method: a fast method for assembling loci across divergent taxa from next-generation sequencing data. [europepmc]
- Forecasting the 2013-2014 influenza season using Wikipedia. [europepmc]
- Big Data Analytics in Healthcare. [europepmc]
- Methodological challenges and analytic opportunities for modeling and interpreting Big Healthcare Data. [europepmc]
- The real cost of sequencing: scaling computation to keep pace with data generation. [europepmc]
- Scalable metagenomics alignment research tool (SMART): a scalable, rapid, and complete search heuristic for the classification of metagenomic sequences from complex sequence populations. [europepmc]
- A novel process of viral vector barcoding and library preparation enables high-diversity library generation and recombination-free paired-end sequencing. [europepmc]
- Patient Similarity: Emerging Concepts in Systems and Precision Medicine. [europepmc]
- Cloud computing for genomic data analysis and collaboration. [europepmc]
- Opportunities and obstacles for deep learning in biology and medicine. [europepmc]
- Clustering algorithms: A comparative approach. [europepmc]
- CaImAn an open source tool for scalable calcium imaging data analysis. [europepmc]
- Machine Learning and Integrative Analysis of Biomedical Big Data. [europepmc]
- Social big data: Recent achievements and new challenges. [europepmc]
- Data-Driven Molecular Dynamics: A Multifaceted Challenge. [europepmc]
- A fast and scalable framework for large-scale and ultrahigh-dimensional sparse regression with application to the UK Biobank. [europepmc]
- mbkmeans: Fast clustering for single cell data using mini-batch k-means. [europepmc]
- Single-Cell Transcriptomics: Current Methods and Challenges in Data Acquisition and Analysis. [europepmc]
Related