MapReduce
2008/01/01 by Jeffrey Dean, Jay B. Dean, Sanjay Ghemawat · 18,575 citations
Computer Science · #Advanced Data Storage Technologies #Artificial intelligence #Big data #Cloud Computing and Resource Management #Computation #Computer science #Distributed computing #Function (biology) #Operating system #Parallel Computing and Optimization Techniques #Parallel computing #Petabyte #Programming language #Programming paradigm #Variety (cybernetics)
paper · pdf · doi:10.1145/1327452.1327492
published in Communications of the ACM 51(1), 107-113 (Association for Computing Machinery)
openalex publication_date 2008/01/01 · openalex created_date 2016/06/24 · openalex updated_date 2026/08/01
Abstract
MapReduce is a programming model and an associated implementation for processing and generating large datasets that is amenable to a broad variety of real-world tasks. Users specify the computation in terms of a map and a reduce function, and the underlying runtime system automatically parallelizes the computation across large-scale clusters of machines, handles machine failures, and schedules inter-machine communication to make efficient use of the network and disks. Programmers find the system easy to use: more than ten thousand distinct MapReduce programs have been implemented internally at Google over the past four years, and an average of one hundred thousand MapReduce jobs are executed on Google's clusters every day, processing a total of more than twenty petabytes of data per day.
Cited by
- PaSh: Light-touch Data-Parallel Shell Processing
- A Pearson’s correlation coefficient based decision tree and its parallel implementation
- MTHAEL: Cross-Architecture IoT Malware Detection Based on Neural Network Advanced Ensemble Learning
- Incremental Query Processing on Big Data Streams
- Partitioning networks into clusters of synchronized nodes via the message-passing algorithm: A scalable approach
- More Parts Than Elements: How Databases Multiply
- Pipelined Gradient Coding
- Evidence-Aware MapReduce for Forkable Compute
- Language Model Teams as Distributed Systems
- DataRater: Meta-Learned Dataset Curation
- New lower bounds for Massively Parallel Computation from query complexity
- Accurate and Fast Federated Learning via IID and Communication-Aware Grouping
- Big Data Staging with MPI-IO for Interactive X-ray Science
- Communication Steps for Parallel Query Processing
- Scalable and Efficient Construction of Suffix Array with MapReduce and In-Memory Data Store System
- Theoretical and Empirical Analysis of a Parallel Boosting Algorithm
- Partout: A Distributed Engine for Efficient RDF Processing
- On the Approximability of Related Machine Scheduling under Arbitrary Precedence
- Robust Gradient Descent via Moment Encoding with LDPC Codes
- Revisiting Degree Distribution Models for Social Graph Analysis
- Enabling Loosely-Coupled Serial Job Execution on the IBM BlueGene/P Supercomputer and the SiCortex SC5832
- Flexible Scheduling of Distributed Analytic Applications
- Masked LARk: Masked Learning, Aggregation and Reporting worKflow
- SD-CPS: Taming the Challenges of Cyber-Physical Systems with a Software-Defined Approach
- A New Parallelization Method for K-means
- Improved Algorithms for Distributed Entropy Monitoring
- Declarative, Secure, Convergent Edge Computation
- Parallel Knowledge Embedding with MapReduce on a Multi-core Processor
- Relationship Queries on Large graphs using Pregel
- i2MapReduce: Incremental MapReduce for Mining Evolving Big Data
- Parallel Algorithms for Summing Floating-Point Numbers
- Subsampling MCMC - An introduction for the survey statistician
- TensorFlow: A system for large-scale machine learning
- Towards Chip-on-Chip Neuroscience: Fast Mining of Frequent Episodes Using Graphics Processors
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Scaling Point-based Differentiable Rendering for Large-scale Reconstruction
- Online Self-Evolving Anomaly Detection in Cloud Computing Environments
- Joint Design of Embedded Index Coding and Beamforming for MIMO-based Distributed Computing via Multi-Agent Reinforcement Learning
- On the Design and Implementation of Structured P2P VPNs
- Simulations between Strongly Sublinear MPC and Node-Capacitated Clique
- Synchromodulametry: From Coincidence Detection to Coherent State Measurement
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- What Teachers Should Know About the Bootstrap: Resampling in the Undergraduate Statistics Curriculum
- TurKPF: TurKontrol as a Particle Filter
- Scalable, Fast Cloud Computing with Execution Templates
- Managing Schema Evolution in NoSQL Data Stores
- Distributed Graphical Simulation in the Cloud
- Workflow is All You Need: Escaping the "Statistical Smoothing Trap" via High-Entropy Information Foraging and Adversarial Pacing
- Large scale citation matching using Apache Hadoop
- HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
- Actors vs Shared Memory: two models at work on Big Data application frameworks
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- Optimal Load Balancing in Bipartite Graphs
- An Empirical Study of Cross-Language Interoperability in Replicated Data Systems
- Understanding the Nature of System-Related Issues in Machine Learning Frameworks: An Exploratory Study
- StarDist: A Code Generator for Distributed Graph Algorithms
- On the Feasibility of Distributed Kernel Regression for Big Data
- Cost Analysis of Nondeterministic Probabilistic Programs
- AgentODRL: A Large Language Model-based Multi-agent System for ODRL Generation
- BoPF: Mitigating the Burstiness-Fairness Tradeoff in Multi-Resource Clusters
- Trustless Federated Learning at Edge-Scale: A Compositional Architecture for Decentralized, Verifiable, and Incentive-Aligned Coordination
- Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
- Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
- Meta-Learning surrogate models for sequential decision making
- CPL: A Core Language for Cloud Computing -- Technical Report
- Near-Optimal Massively Parallel Graph Connectivity
- Towards Federated Learning at Scale: System Design
- Chronos: A Unifying Optimization Framework for Speculative Execution of Deadline-critical MapReduce Jobs
- TD-Orch: Efficient Task-Data Orchestration for Distributed Systems with Application to Graph Processing
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Deadline is not Enough: How to Achieve Importance-aware Server-centric Data Centers via a Cross Layer Approach
- Serverless seismic imaging in the cloud
- Efficient Approximation of Volterra Series for High-Dimensional Systems
- Yesquel: scalable SQL storage for Web applications
- A Fundamental Tradeoff between Computation and Communication in Distributed Computing
- DINGO: Distributed Newton-Type Method for Gradient-Norm Optimization
- Making problems tractable on big data via preprocessing with polylog-size output
- Sorting, Searching, and Simulation in the MapReduce Framework
- Using Span Queries to Optimize for Cache and Attention Locality
- Private Map-Secure Reduce: Infrastructure for Efficient AI Data Markets
- ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
- Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
- Graph Processing on FPGAs: Taxonomy, Survey, Challenges
- Optimal Load Allocation for Coded Distributed Computation in Heterogeneous Clusters
- Skew Handling in Aggregate Streaming Queries on GPUs
- A Data-Driven Approximation of the Koopman Operator: Extending Dynamic Mode Decomposition
- Implementation of Algorithms for Right-Sizing Data Centers
- Fusion: An Analytics Object Store Optimized for Query Pushdown
- MLitB: Machine Learning in the Browser
- Online Machine Learning in Big Data Streams
- SwitchAgg:A Further Step Towards In-Network Computation
- Survey and Taxonomy of Lossless Graph Compression and Space-Efficient Graph Representations
- Real-time semiparametric regression for distributed data sets
- Contrasting Effects of Replication in Parallel Systems: From Overload to Underload and Back
- A Survey of Blocking and Filtering Techniques for Entity Resolution
- Dynamic Memory Allocation Policies for Postings in Real-Time Twitter Search
- Optimizing MapReduce for Highly Distributed Environments
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Stochastic Primal-Dual Coordinate Method for Regularized Empirical Risk Minimization
- Does The Cloud Need Stabilizing?
- Non-clairvoyant Scheduling of Coflows
- On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML
- Upper and Lower Bounds on the Cost of a Map-Reduce Computation
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Assignment Problems of Different-Sized Inputs in MapReduce
- When Gaussian Process Meets Big Data: A Review of Scalable GPs
- Teaching Machine Learning to Software Engineers
- Time Warp on the Go (Updated Version)
- Data Mining Scheme for Globally Distributed Big Data
- Big Data at HPC Wales
- Bayesian Federated Learning over Wireless Networks
- Simple and sharp analysis of k-means||
- Fast Clustering using MapReduce
- Arrows for Parallel Computation
- Distributed Data Processing Frameworks for Big Graph Data
- A Preliminary Review of Influential Works in Data-Driven Discovery
- Parallel Bayesian Additive Regression Trees
- Robust Scheduling for Flexible Processing Networks
- Coflow Scheduling in Data Centers: Routing and Bandwidth Allocation
- TGE-viz : Transition Graph Embedding for Visualization of Plan Traces and Domains
- Evidential instance selection for K-nearest neighbor classification of big data
- Forecasting the 2013--2014 Influenza Season using Wikipedia
- On the Complexity of Processing Massive, Unordered, Distributed Data
- The OoO VLIW JIT Compiler for GPU Inference
- GLB: Lifeline-based Global Load Balancing library in X10
- Evaluating Device-First Continuum AI (DFC-AI) for Autonomous Operations in the Energy Sector
- A Unified Coding Framework for Distributed Computing with Straggling Servers
- Heterogeneous Coded Distributed Computing: Joint Design of File Allocation and Function Assignment
- Representing emotions with knowledge graphs for movie recommendations
- Scientific Computing Meets Big Data Technology: An Astronomy Use Case
- Dynamic Deferral of Workload for Capacity Provisioning in Data Centers
- Approximate Gradient Coding for Distributed Learning with Heterogeneous Stragglers
- ProGQL: A Provenance Graph Query System for Cyber Attack Investigation
- Institutional Metaphors for Designing Large-Scale Distributed AI versus AI Techniques for Running Institutions
- Embed and Conquer: Scalable Embeddings for Kernel k-Means on MapReduce
- Analysis of Input-Output Mappings in Coinjoin Transactions with Arbitrary Values
- Towards Stochastically Optimizing Data Computing Flows
- Distributed Stochastic Variance Reduced Gradient Methods and A Lower Bound for Communication Complexity
- Machine Learning on Graphs: A Model and Comprehensive Taxonomy
- Centrality Metric for Dynamic Networks
- A Decentralized Microservice Scheduling Approach Using Service Mesh in Cloud-Edge Systems
- Graph3S: A Simple, Speedy and Scalable Distributed Graph Processing System
- NOMAD: Non-locking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion
- Agentic Troubleshooting Guide Automation for Incident Management
- Category-Theoretic Quantitative Compositional Distributional Models of Natural Language Semantics
- Scalable Data Cube Analysis over Big Data
- Coflow Scheduling in Input-Queued Switches: Optimal Delay Scaling and Algorithms
- CloudSVM : Training an SVM Classifier in Cloud Computing Systems
- Big Data and Cross-Document Coreference Resolution: Current State and Future Opportunities
- Scalable Similarity Joins of Tokenized Strings
- Parallel Markov Chain Monte Carlo for Bayesian Hierarchical Models with Big Data, in Two Stages
- Algorithms for a Topology-aware Massively Parallel Computation Model
- Distributed Machine Learning for Wireless Communication Networks: Techniques, Architectures, and Applications
- Parallel D2-Clustering: Large-Scale Clustering of Discrete Distributions
- Distributed Machine Learning via Sufficient Factor Broadcasting
- A Combinatorial Design for Cascaded Coded Distributed Computing on General Networks
- Distributed Algorithms for Finding Local Clusters Using Heat Kernel Pagerank
- Fundamental Limits of Distributed Computing for Linearly Separable Functions
- Pilot-Data: An Abstraction for Distributed Data
- CLAMShell: Speeding up Crowds for Low-latency Data Labeling
- Mapping and Reducing the Brain on the Cloud
- Exploratory Analysis of a Terabyte Scale Web Corpus
- Integrating SysML and AUTOSAR: model transformation in automotive MBSE
- PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development
- Energy-efficient photonic neural networks for high-speed AI computation
- FlashR: R-Programmed Parallel and Scalable Machine Learning using SSDs
- UStore: A Distributed Storage With Rich Semantics
- Polypus: a Big Data Self-Deployable Architecture for Microblogging Text Extraction and Real-Time Sentiment Analysis
- Application of Deep Learning Models for Real-Time Automatic Malware Detection
- Global and Local Implications of Computational Artifacts
- Prediction of Video Popularity in the Absence of Reliable Data from Video Hosting Services: Utility of Traces Left by Users on the Web
- AGL: a Scalable System for Industrial-purpose Graph Machine Learning
- On new data sources for the production of official statistics
- Jug: Software for Parallel Reproducible Computation in Python
- Pretrained Transformers for Text Ranking: BERT and Beyond
- On the inequality of the 3V's of Big Data Architectural Paradigms: A case for heterogeneity
- Graph Summarization Methods and Applications: A Survey
- ElasticBroker: Combining HPC with Cloud to Provide Realtime Insights into Simulations
- Optimal bandwidth-aware VM allocation for Infrastructure-as-a-Service
- Processing Database Joins over a Shared-Nothing System of Multicore Machines
- Equi-depth Histogram Construction for Big Data with Quality Guarantees
- OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC
- Deep Learning At Scale and At Ease
- A Map-Reduce Parallel Approach to Automatic Synthesis of Control\n Software
- Distributed Exploration in Multi-Armed Bandits
- TF-Replicator: Distributed Machine Learning for Researchers
- ALID: Scalable Dominant Cluster Detection
- Thrill: High-Performance Algorithmic Distributed Batch Data Processing with C++
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- Thinking Like a Vertex: a Survey of Vertex-Centric Frameworks for Distributed Graph Processing
- Easy Acceleration with Distributed Arrays
- On the Mathematics of Data Centre Network Topologies
- Matrix Factorization at Scale: a Comparison of Scientific Data Analytics in Spark and C+MPI Using Three Case Studies
- Real-Time Machine Learning: The Missing Pieces
- Architectural Tactics for Big Data Cybersecurity Analytic Systems: A Review
- Distributed Proximal Gradient Algorithm for Partially Asynchronous Computer Clusters
- L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems
- FLCD: A Flexible Low Complexity Design of Coded Distributed Computing
- Comparative Performance Analysis of Intel Xeon Phi, GPU, and CPU
- Density Estimations for Approximate Query Processing on SIMD Architectures
- Alternative statistical inference for the first normalized incomplete moment
- SAF: Simulated Annealing Fair Scheduling for Hadoop Yarn Clusters
- A Uniqueness Theorem for Distributed Computation under Physical Constraint
- Parallel Approximate Undirected Shortest Paths Via Low Hop Emulators
- ATLAS: An Adaptive Failure-aware Scheduler for Hadoop
- Characterizing Data Analysis Workloads in Data Centers
- Deep Learning with Apache SystemML
- Fast attention mechanisms: a tale of parallelism
- Hoplite: Efficient and Fault-Tolerant Collective Communication for Task-Based Distributed Systems
- Communication-Efficient Edge AI: Algorithms and Systems
- ZhiFangDanTai: Fine-tuning Graph-based Retrieval-Augmented Generation Model for Traditional Chinese Medicine Formula
- Benchmarking DataStax Enterprise/Cassandra with HiBench
- A MapReduce Approach to NoSQL RDF Databases
- Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training
- Limited Random Walk Algorithm for Big Graph Data Clustering
- The Two Quadrillionth Bit of Pi is 0! Distributed Computation of Pi with Apache Hadoop
- Counterfactual simulations for large scale systems with burnout variables
- From Data Fusion to Knowledge Fusion
- Merlin: A Language for Provisioning Network Resources
- Delivery, consistency, and determinism: rethinking guarantees in distributed stream processing
- Connected Components on a PRAM in Log Diameter Time
- Improving Nonpreemptive Multiserver Job Scheduling with Quickswap
- A Novel IaaS Tax Model as Leverage Towards Green Cloud Computing
- Visualizing a Million Time Series with the Density Line Chart
- HiCR, an Abstract Model for Distributed Heterogeneous Programming
- Collaborative Learning with Limited Interaction: Tight Bounds for Distributed Exploration in Multi-Armed Bandits
- Scaling-Up Reasoning and Advanced Analytics on BigData
- Parallel Programming for FPGAs
- Diversity, Productivity, and Growth of Open Source Developer Communities
- Parallelization in Scientific Workflow Management Systems
- Analysis of building maintenance requests using a text mining approach: building services evaluation
- A Generalized Performance Evaluation Framework for Parallel Systems with Output Synchronization
- MMLSpark: Unifying Machine Learning Ecosystems at Massive Scales
- Parallel Evaluation of Interaction Nets: Case Studies and Experiments
- Exploring heterogeneity of unreliable machines for p2p backup
- Efficient Straggler Replication in Large-scale Parallel Computing
- A Reliable Effective Terascale Linear Learning System
- Ray: A Distributed Framework for Emerging AI Applications
- DiskJoin: Large-scale Vector Similarity Join with SSD
- MapReduce Meets Fine-Grained Complexity: MapReduce Algorithms for APSP, Matrix Multiplication, 3-SUM, and Beyond
- Transduction is All You Need for Structured Data Workflows
- Behavioral Simulations in MapReduce
- Scalable Protein Sequence Similarity Search using Locality-Sensitive Hashing and MapReduce
- How to Optimally Allocate Resources for Coded Distributed Computing?
- A Simple and Efficient MapReduce Algorithm for Data Cube Materialization
- GraphLab: A New Framework For Parallel Machine Learning
- A novel approach for fast mining frequent itemsets use N-list structure based on MapReduce
- Ripple : Simplified Large-Scale Computation on Heterogeneous Architectures with Polymorphic Data Layout
- CADDeLaG: Framework for distributed anomaly detection in large dense graph sequences
- Coded Computing for Master-Aided Distributed Computing Systems
- Cloudpress 2.0: A MapReduce Approach for News Retrieval on the Cloud
- Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
- MovePattern: Interactive Framework to Provide Scalable Visualization of Movement Patterns
- Oseba: Optimization for Selective Bulk Analysis in Big Data Processing
- A Stream Pipeline Framework for Digital Payment Programming based on Smart Contracts
- A Comparative Study of Asynchronous Many-Tasking Runtimes: Cilk, Charm++, ParalleX and AM++
- An Alternating Direction Method Approach to Cloud Traffic Management
- Evaluating Complex Task through Crowdsourcing: Multiple Views Approach
- Comparing Spark vs MPI/OpenMP On Word Count MapReduce
- Asynchronous Complex Analytics in a Distributed Dataflow Architecture
- Sparkle: Optimizing Spark for Large Memory Machines and Analytics
- Mobile Cloud Computing: A Comparison of Application Models
- SensorCloud: Towards the Interdisciplinary Development of a Trustworthy Platform for Globally Interconnected Sensors and Actuators
- Coding for Distributed Fog Computing
- Estimation of Passenger Route Choice Pattern Using Smart Card Data for Complex Metro Systems
- Optimization for Speculative Execution of Multiple Jobs in a MapReduce-like Cluster
- Online Job Scheduling with Redundancy and Opportunistic Checkpointing: A Speedup-Function-Based Analysis
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- Fork and Join Queueing Networks with Heavy Tails: Scaling Dimension and Throughput Limit
- Maiter: An Asynchronous Graph Processing Framework for Delta-based Accumulative Iterative Computation
- Byzantine-Resilient Distributed Computation via Task Replication and Local Computations
- FLOSS: Federated Learning with Opt-Out and Straggler Support
- Introduction to Rank-polymorphic Programming in Remora (Draft)
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- Towards a Periodic Table of Computer System Design Principles
- The NOESIS Network-Oriented Exploration, Simulation, and Induction System
- Collaborative State Machines: A Better Programming Model for the Cloud-Edge-IoT Continuum
- Locality Optimization for Data Parallel Programs
- Query and Resource Optimizations: A Case for Breaking the Wall in Big Data Systems
- Whiz: A Fast and Flexible Data Analytics System
- Quantum Adiabatic Evolution for Global Optimization in Big Data
- A Comparative Taxonomy and Survey of Public Cloud Infrastructure Vendors
- Using Variational Inference and MapReduce to Scale Topic Modeling
- Friendship Prediction in Composite Social Networks
- Rapid AkNN Query Processing for Fast Classification of Multidimensional Data in the Cloud
- Parallel Hierarchical Affinity Propagation with MapReduce
- Fast Fourier-Based Generation of the Compression Matrix for Deterministic Compressed Sensing
- Big Data Energy Systems: A Survey of Practices and Associated Challenges
- Social Influence and Radicalization: A Social Data Analytics Study
- Optimizing Redundancy Levels in Master-Worker Compute Clusters for Straggler Mitigation
- Performance Provisioning and Energy Efficiency in Cloud and Distributed Computing Systems
- Odysseus/DFS: Integration of DBMS and Distributed File System for Transaction Processing of Big Data
- Semantic Support for Log Analysis of Safety-Critical Embedded Systems
- Sparse Tensor Algebra as a Parallel Programming Model
- On the Complexity of Sorted Neighborhood
- Multinomial Loss on Held-out Data for the Sparse Non-negative Matrix Language Model
- Optimizing Prediction Serving on Low-Latency Serverless Dataflow
- Optimizing Stochastic Scheduling in Fork-Join Queueing Models: Bounds and Applications
- ClusterCluster: Parallel Markov Chain Monte Carlo for Dirichlet Process Mixtures
- clusterNOR: A NUMA-Optimized Clustering Framework
- CrowdFusion: A Crowdsourced Approach on Data Fusion Refinement
- Coded Distributed Computing over Packet Erasure Channels
- Effective Techniques for Message Reduction and Load Balancing in Distributed Graph Computation
- Meta-MapReduce: A Technique for Reducing Communication in MapReduce Computations
- The Open Cloud Testbed: A Wide Area Testbed for Cloud Computing Utilizing High Performance Network Services
- A Tale of Two Data-Intensive Paradigms: Applications, Abstractions, and Architectures
- Collage Inference: Using Coded Redundancy for Low Variance Distributed Image Classification
- An Abstract View of Big Data Processing Programs
- An efficient K-means algorithm for Massive Data
- A Low Complexity Decentralized Neural Net with Centralized Equivalence using Layer-wise Learning
- Energy-efficient Analytics for Geographically Distributed Big Data
- P4COM: In-Network Computation with Programmable Switches
- Pilot-Abstraction: A Valid Abstraction for Data-Intensive Applications on HPC, Hadoop and Cloud Infrastructures?
- Jupiter Rising
- Aethon: A Reference-Based Replication Primitive for Constant-Time Instantiation of Stateful AI Agents
- Towards Efficient Post-training Quantization of Pre-trained Language Models
- Big Data application in congestion detection and classification using Apache spark
- Two-level Data Staging ETL for Transaction Data
- Top-k Spatial-keyword Publish/Subscribe Over Sliding Window
- MapReduce for Integer Factorization
- GYM: A Multiround Join Algorithm In MapReduce
- Twitch Plays Pokemon, Machine Learns Twitch: Unsupervised Context-Aware Anomaly Detection for Identifying Trolls in Streaming Data
- Intent Models for Contextualising and Diversifying Query Suggestions
- Analyzing Web Application Log Files to Find Hit Count Through the\n Utilization of Hadoop MapReduce in Cloud Computing Environment
- Scaling Datalog for Machine Learning on Big Data
- A Bloom Filter Survey: Variants for Different Domain Applications
- Splash: User-friendly Programming Interface for Parallelizing Stochastic Algorithms
- Scalable Facility Location for Massive Graphs on Pregel-like Systems
- Embarrassingly Parallel Time Series Analysis for Large Scale Weak Memory Systems
- Distributed Parameter Map-Reduce
- A Survey of Coded Distributed Computing
- SOFA: An Extensible Logical Optimizer for UDF-heavy Dataflows
- Parallel queues with synchronization
- Blockchain Cohomology
- CIAO: An Optimization Framework for Client-Assisted Data Loading
- Cut to Fit: Tailoring the Partitioning to the Computation
- Coded Distributed Computing with Heterogeneous Function Assignments
- Scalable Inference of System-level Models from Component Logs
- Towards Automated Management and Analysis of Heterogeneous Data Within Cannabinoids Domain
- On data skewness, stragglers, and MapReduce progress indicators
- Simulation Based Formal Verification of Cyber-Physical Systems
- Automatic Performance Debugging of SPMD Parallel Programs
- Flexible Support for Fast Parallel Commutative Updates
- GPU Tensor Cores for fast Arithmetic Reductions
- Coding Method for Parallel Iterative Linear Solver
- In Search of a Fast and Efficient Serverless DAG Engine
- Falkirk Wheel: Rollback Recovery for Dataflow Systems
- Launchpad: A Programming Model for Distributed Machine Learning Research
- A study of big data processing constraints on a low-power Hadoop cluster
- Optimizing performance and power consumption for an ARM-based big data cluster
- A Survey and Taxonomy of Resource Optimisation for Executing Bag-of-Task Applications on Public Clouds
- Large-scale Artificial Neural Network: MapReduce-based Deep Learning
- Will solid-state drives accelerate your bioinformatics? In-depth profiling, performance analysis, and beyond
- STEP : A Distributed Multi-threading Framework Towards Efficient Data Analytics
- Distributed computing of Seismic Imaging Algorithms
- Real-time Data Infrastructure at Uber
- Scalable and Efficient Statistical Inference with Estimating Functions in the MapReduce Paradigm for Big Data
- A Survey on Large-scale Machine Learning
- Bayesian computation: a perspective on the current state, and sampling backwards and forwards
- Matrix Computations and Optimization in Apache Spark
- Distributed Robust Learning
- Distributed Averaging CNN-ELM for Big Data
- Improved Constructions for Secure Multi-Party Batch Matrix Multiplication
- City on the Sky: Flexible, Secure Data Sharing on the Cloud
- Analytics for the Internet of Things: A Survey
- Introducing Distributed Dynamic Data-intensive (D3) Science: Understanding Applications and Infrastructure
- Stream programs are monoid homomorphisms with state
- Incremental Techniques for Large-Scale Dynamic Query Processing
- Typing a Core Binary Field Arithmetic in a Light Logic
- Design Patterns for Securing LLM Agents against Prompt Injections
- The AnyLog Edge Data Fabric
- Exploring Block Anomaly Detection In HDFS Log Data Analysis
- Efficient Distributed Locality Sensitive Hashing
- Analysis of Research in Healthcare Data Analytics
- META-pipe - Pipeline Annotation, Analysis and Visualization of Marine Metagenomic Sequence Data
- Building your Cross-Platform Application with RHEEM
- Towards Serverless Processing of Spatiotemporal Big Data Queries
- Leveraging User Diversity to Harvest Knowledge on the Social Web
- On the Evaluation of RDF Distribution Algorithms Implemented over Apache Spark
- Caramel: Accelerating Decentralized Distributed Deep Learning with Computation Scheduling
- Streaming Algorithms for News and Scientific Literature Recommendation: Submodular Maximization with a d-Knapsack Constraint
- Parallel Matrix Factorization for Binary Response
- Studying the Impact of Power Capping on MapReduce-based, Data-intensive Mini-applications on Intel KNL and KNM Architectures
- Intermediate Data Caching Optimization for Multi-Stage and Parallel Big Data Frameworks
- Positive region preserved random sampling: an efficient feature selection method for massive data
- Federated Learning on Stochastic Neural Networks
- ArcLink: Optimization Techniques to Build and Retrieve the Temporal Web Graph
- Coded Elastic Computing
- Blaze: Simplified High Performance Cluster Computing
- WarpFlow: Exploring Petabytes of Space-Time Data
- Improved MPC Algorithms for MIS, Matching, and Coloring on Trees and Beyond
- The Scalability for Parallel Machine Learning Training Algorithm: Dataset Matters
- Tolerating Correlated Failures in Massively Parallel Stream Processing Engines
- Hardness of Virtual Network Embedding with Replica Selection
- Pregelix: Big(ger) Graph Analytics on A Dataflow Engine
- Programming and Deployment of Autonomous Swarms using Multi-Agent Reinforcement Learning
- Analyzing Big Datasets of Genomic Sequences: Fast and Scalable Collection of k-mer Statistics
- Iterative MapReduce for Large Scale Machine Learning
- Detecting Group Anomalies in Tera-Scale Multi-Aspect Data via Dense-Subtensor Mining
- Distributed k-Core Decomposition
- Mobile Edge Computing Empowers Internet of Things
- Serverless Straggler Mitigation using Local Error-Correcting Codes
- Extract ABox Modules for Efficient Ontology Querying
- A Framework for Application-aware Networking by Delegating Traffic Management of SDNs
- A Novel Approach to Finding Near-Cliques: The Triangle-Densest Subgraph Problem
- (α, k)-Minimal Sorting and Skew Join in MPI and MapReduce
- FlashGraph: Processing Billion-Node Graphs on an Array of Commodity SSDs
- Mapping the Evolution of Research Contributions using KnoVo
- CORE: Augmenting Regenerating-Coding-Based Recovery for Single and Concurrent Failures in Distributed Storage Systems
- A Scalable Framework for Wireless Distributed Computing
- Avalon: Building an Operating System for Robotcenter
- Greedy Column Subset Selection for Large-scale Data Sets
- Energy-efficient data centers
- Exploring Social Influence for Recommendation - A Probabilistic Generative Model Approach
- A parallel and distributed C4.5 algorithm in cloud computing environments
- HARMONY: A Scalable Distributed Vector Database for High-Throughput Approximate Nearest Neighbor Search
- Non-Asymptotic Delay Bounds for Multi-Server Systems with Synchronization Constraints
- Solving Large-Scale Granular Resource Allocation Problems Efficiently with POP
- Towards Cost-Optimal Policies for DAGs to Utilize IaaS Clouds with Online Learning
- GENMR: Generalized Query Processing through Map Reduce In Cloud Database Management System
- Space and Time Efficient Parallel Graph Decomposition, Clustering, and Diameter Approximation
- Coded Alternating Least Squares for Straggler Mitigation in Distributed Recommendations
- Scalable and Fault Tolerant Computation with the Sparse Grid Combination Technique
- Cognitive Internet of Things: A New Paradigm beyond Connection
- Primitives for Dynamic Big Model Parallelism
- Model-Parallel Inference for Big Topic Models
- A Survey and Evaluation of Data Center Network Topologies
- InstaCluster: Building A Big Data Cluster in Minutes
- A Study on Individual Spatiotemporal Activity Generation Method Using MCP-Enhanced Chain-of-Thought Large Language Models
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation
- A Survey on Array Storage, Query Languages, and Systems
- A Distributed Cubic-Regularized Newton Method for Smooth Convex Optimization over Networks
- Generating Long Semantic IDs in Parallel for Recommendation
- Faster MPC Algorithms for Approximate Allocation in Uniformly Sparse Graphs
- Effective Spatial Data Partitioning for Scalable Query Processing
- MACH: Fast Randomized Tensor Decompositions
- Dimension Independent Similarity Computation
- Simulating fluid vortex interactions on a superconducting quantum processor
- Evaluating Hive and Spark SQL with BigBench
- MultiHead MultiModal Deep Interest Recommendation Network
- Enumerating Subgraph Instances Using Map-Reduce
- A New Combinatorial Coded Design for Heterogeneous Distributed Computing
- A Formal, Resource Consumption-Preserving Translation of Actors to Haskell
- Storage, Computation, and Communication: A Fundamental Tradeoff in Distributed Computing
- Adaptive Partitioning for Very Large RDF Data
- Technical Report: Optimistic Execution in Key-Value Store
- Wiggins: Detecting Valuable Information in Dynamic Networks Using Limited Resources
- Ripple: A Practical Declarative Programming Framework for Serverless Compute
- Differentially Private Densest Subgraph Detection
- MIRAGE: An Iterative MapReduce based FrequentSubgraph Mining Algorithm
- First-Class Functions for First-Order Database Engines
- Approximate Computation and Implicit Regularization for Very Large-scale Data Analysis
- Unifying Data, Model and Hybrid Parallelism in Deep Learning via Tensor Tiling
- Analyzing Self-Driving Cars on Twitter
- Future Robotics Database Management System along with Cloud TPS
- Subgraph Enumeration in Massive Graphs
- Network calculus for parallel processing
- Programming Scalable Cloud Services with AEON
- How to perform research in Hadoop environment not losing mental equilibrium - case study
- GRE: A Graph Runtime Engine for Large-Scale Distributed Graph-Parallel Applications
- GraphLab: A Distributed Framework for Machine Learning in the Cloud
- A formal definition of Big Data based on its essential features
- DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index
- JUNO Conceptual Design Report
- Coded Sparse Matrix Multiplication
- Experience: Type alignment on DBpedia and Freebase
- Polygon Queries for Convex Hulls of Points
- BigDataBench: A Scalable and Unified Big Data and AI Benchmark Suite
- Evaluation of Codes with Inherent Double Replication for Hadoop
- A Query Language for Summarizing and Analyzing Business Process Data
- An Order-Aware Dataflow Model for Parallel Unix Pipelines
- Elastic Processing of Analytical Query Workloads on IaaS Clouds
- P6: A Declarative Language for Integrating Machine Learning in Visual Analytics
- Neural Article Pair Modeling for Wikipedia Sub-article Matching
- Coded Computation Against Distributed Straggling Channel Decoders in the Cloud for Gaussian Uplink Channels
- GraphMP: I/O-Efficient Big Graph Analytics on a Single Commodity Machine
- Reasoning about Block-based Cloud Storage Systems
- Learning over inherently distributed data
- Distributed Streaming Analytics on Large-scale Oceanographic Data using Apache Spark
- Distributed Variational Inference in Sparse Gaussian Process Regression and Latent Variable Models
- A fast, lock-free approach for efficient parallel counting of occurrences of k -mers
- On Efficient Data Transfers Across Geographically Dispersed Datacenters
- Energy-Efficient Mechanism for Smart Communication in Cellular Networks
- RankMap: A Platform-Aware Framework for Distributed Learning from Dense Datasets
- Patterns and Rewrite Rules for Systematic Code Generation (From High-Level Functional Patterns to High-Performance OpenCL Code)
- Implementing Randomized Matrix Algorithms in Parallel and Distributed Environments
- WTF
- Elastic and Secure Energy Forecasting in Cloud Environments
- Collective Creativity: Where we are and where we might go
- MaRe: a MapReduce-Oriented Framework for Processing Big Data with Application Containers
- Measuring the Optimality of Hadoop Optimization
- Computing the Schulze Method for Large-Scale Preference Data Sets
- Fundamental Limits of Coded Linear Transform
- On the Local Communication Complexity of Counting and Modular Arithmetic
- Big other: Surveillance Capitalism and the Prospects of an Information Civilization
- The Stratosphere platform for big data analytics
- Predicting Scheduling Failures in the Cloud
- A Fundamental Storage-Communication Tradeoff for Distributed Computing with Straggling Nodes
- Large-Scale Reasoning with OWL
- Modeling Events with Cascades of Poisson Processes
- A Review and Analysis of a Parallel Approach for Decision Tree Learning from Large Data Streams
- AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications
- Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture Slides
- Augmenting recommendation systems using a model of semantically-related terms extracted from user behavior
- Support vector regression model for BigData systems
- Optimized Composition: Generating Efficient Code for Heterogeneous Systems from Multi-Variant Components, Skeletons and Containers
- DALiuGE: A Graph Execution Framework for Harnessing the Astronomical Data Deluge
- Statistical Physics and Information Theory Perspectives on Linear Inverse Problems
- Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
- Monoidify! Monoids as a Design Principle for Efficient MapReduce Algorithms
- Enabling Operator Reordering in Data Flow Programs Through Static Code Analysis
- End-to-End Entity Resolution for Big Data: A Survey
- A Survey of Semantics-Aware Performance Optimization for Data-Intensive Computing
- ZenLDA: An Efficient and Scalable Topic Model Training System on Distributed Data-Parallel Platform
- On Optimal Batch Size in Coded Computing
- Harmonic Coding: An Optimal Linear Code for Privacy-Preserving Gradient-Type Computation
- Understanding Stragglers in Large Model Training Using What-if Analysis
- Parallel Structure from Motion from Local Increment to Global Averaging
- BigRoots: An Effective Approach for Root-cause Analysis of Stragglers in Big Data System
- Polynomial Codes: an Optimal Design for High-Dimensional Coded Matrix Multiplication
- "vcd2df" -- Leveraging Data Science Insights for Hardware Security Research
- On the Delay-Storage Trade-off in Content Download from Coded Distributed Storage Systems
- MURS: Mitigating Memory Pressure in Service-oriented Data Processing Systems
- Parallelized Kendall's Tau Coefficient Computation via SIMD Vectorized Sorting On Many-Integrated-Core Processors
- Stochastic Non-preemptive Co-flow Scheduling with Time-Indexed Relaxation
- Distributed Data Stream Processing and Edge Computing: A Survey on Resource Elasticity and Future Directions
- Cloud-based Privacy Preserving Image Storage, Sharing and Search
- Query-driven Frequent Co-occurring Term Extraction over Relational Data using MapReduce
- Identifying Duplicate and Contradictory Information in Wikipedia
- Hadoop Scheduling Base On Data Locality
- Cloud-based Manufacturing: Old Wine in New Bottles?
- G-thinker: Big Graph Mining Made Easier and Faster
- Transparently Resilient Task Parallelism for Chapel
- Sector and Sphere: the design and implementation of a high-performance data cloud
- Parallel Priority-Flood Depression Filling For Trillion Cell Digital Elevation Models On Desktops Or Clusters
- Parallel Actors and Learners: A Framework for Generating Scalable RL Implementations
- Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation
- Tupleware: Redefining Modern Analytics
- A survey of algorithmic skeleton frameworks: high‐level structured parallel programming enablers
- Big Data: A Survey
- High-Performance Computation in Residue Number System Using Floating-Point Arithmetic
- From Frequency to Meaning: Vector Space Models of Semantics
- Computing n-Gram Statistics in MapReduce
- Demystifying Parallel and Distributed Deep Learning
- Discretionary Information Flow Control for Interaction-Oriented Specifications
- MERIT: Tensor Transform for Memory-Efficient Vision Processing on Parallel Architectures
- DDSL: Efficient Subgraph Listing on Distributed and Dynamic Graphs
- A Survey of State Management in Big Data Processing Systems
- PISA
- State Space Exploration of RT Systems in the Cloud
- Parallel Computation of PDFs on Big Spatial Data Using Spark
- Mathematical and Algorithmic Analysis of Network and Biological Data
- To pipeline or not to pipeline, that is the question
- PFO: A Parallel Friendly High Performance System for Online Query and Update of Nearest Neighbors
- A Taxonomy and Survey on eScience as a Service in the Cloud
- Parallel Bayesian Additive Regression Trees
- Thriving in a crowded and changing world: C++ 2006–2020
- On Batch-Processing Based Coded Computing for Heterogeneous Distributed Computing Systems
- Column-Oriented Storage Techniques for MapReduce
- Scalable Ontological Query Processing over Semantically Integrated Life Science Datasets using MapReduce
- How Computers Work: Computational Thinking for Everyone
- Privacy-Preserving Access of Outsourced Data via Oblivious RAM Simulation
- Inspector: A Data Provenance Library for Multithreaded Programs
- FractalSortCPU: Bandwidth-Efficient Compressed Radix Sort on CPU
- APWA: A Distributed Architecture for Parallelizable Agentic Workflows
- Analyzing Large-Scale, Distributed and Uncertain Data
- Distributed aggregation for data-parallel computing
- Scheduling Data Intensive Workloads through Virtualization on MapReduce based Clouds
- Evolving a language in and for the real world
- Parallel Evaluation Of Multi-Semi-Joins
- On the Complexity of List Ranking in the Parallel External Memory Model
- Enhance parallel input/output with cross-bundle aggregation
- RecSplit: Minimal Perfect Hashing via Recursive Splitting
- InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking
- Databases
- bigMICE: Multiple Imputation of Big Data
- Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining
- Integrating Multi-Armed Bandit, Active Learning, and Distributed Computing for Scalable Optimization
- Coreset-based Strategies for Robust Center-type Problems
- Cassandra
- Dynamic Approximate Maximum Matching in the Distributed Vertex Partition Model
- MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing
- A Unifying Framework for Parallel and Distributed Processing in R using Futures
- Occam’s Razor for Big Data? On Detecting Quality in Large Unstructured Datasets
- PRIMAL: PRofIt Maximization Avatar pLacement for Mobile Edge Computing
- Forecasting: theory and practice
- Power-efficient Assignment of Virtual Machines to Physical Machines
- Performance and Fault Tolerance in the StoreTorrent Parallel Filesystem
- Cascading map-side joins over HBase for scalable join processing
- Declarative Machine Learning - A Classification of Basic Properties and Types
- GraphX: Unifying Data-Parallel and Graph-Parallel Analytics
- Fully Scalable MPC Algorithms for Euclidean k-Center
- PaPy: Parallel and Distributed Data-processing Pipelines in Python
- Transformative effects of IoT, Blockchain and Artificial Intelligence on cloud computing: Evolution, vision, trends and open challenges
- Performance optimization of MapReduce-based Apriori algorithm on Hadoop cluster
- Recent advances in feature selection and its applications
- Distributed Programming over Time-Series Graphs
- COMET: A Recipe for Learning and Using Large Ensembles on Massive Data
- Streaming Algorithm for Euler Characteristic Curves of Multidimensional Images
- Large-Scale Intelligent Microservices
- A Comparative Study of Data Storage and Processing Architectures for the Smart Grid
- ARENA: Asynchronous Reconfigurable Accelerator Ring to Enable Data-Centric Parallel Computing
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- PEGASUS: mining peta-scale graphs
- Space-round tradeoffs for MapReduce computations
- Steering Semantic Data Processing With DocWrangler
- A Distributed Approach Toward Discriminative Distance Metric Learning
- Fault Tolerance for Stream Processing Engines
- NXgraph: An efficient graph processing system on a single machine
- Towards Polyglot Data Processing in Social Networks using the Hadoop-Spark ecosystem
- A Cloud Infrastructure Service Recommendation System for Optimizing Real-time QoS Provisioning Constraints
- Efficient Task Replication for Fast Response Times in Parallel Computation
- Round Compression for Parallel Graph Algorithms in Strongly Sublinear\n Space
- Deep learning for video classification and captioning
- Trends and Advancements in Deep Neural Network Communication
- Towards enabling I/O awareness in task-based programming models
- Scheduling MapReduce Jobs under Multi-Round Precedences
- High-Performance Cloud Computing: A View of Scientific Applications
- ESTemd: A Distributed Processing Framework for Environmental Monitoring based on Apache Kafka Streaming Engine
- Distributed Deep Learning Strategies For Automatic Speech Recognition
- Improving Distributed Similarity Join in Metric Space with Error-bounded Sampling
- The core decomposition of networks: theory, algorithms and applications
- Enumerating Maximal Bicliques from a Large Graph using MapReduce
- Parallel Batch-Dynamic Trees via Change Propagation
- Parallel non-divergent flow accumulation for trillion cell digital elevation models on desktops or clusters
- Distributed ReliefF-based feature selection in Spark
- Next-generation genotype imputation service and methods
- A Review of Botnet Detection Approaches Based on DNS Traffic Analysis
- Code Reborn AI-Driven Legacy Systems Modernization from COBOL to Java
- A new thesis concerning synchronised parallel computing – simplified parallel ASM thesis
- Apache VXQuery: A Scalable XQuery Implementation
- String Problems in the Congested Clique Model
- Parallelizing Machine Learning as a service for the end-user
- Zerrow: True Zero-Copy Arrow Pipelines in Bauplan
- Filter and refine [wikipedia]
- Quake: quality-aware detection and correction of sequencing errors. [europepmc]
- SeqWare Query Engine: storing and searching sequence data in the cloud. [europepmc]
- Nipype: a flexible, lightweight and extensible neuroimaging data processing framework in python. [europepmc]
- Disk-based k-mer counting on a PC. [europepmc]
- The role and challenges of exome sequencing in studies of human diseases. [europepmc]
- Bioinformatics for precision medicine in oncology: principles and application to the SHIVA clinical trial. [europepmc]
- Applications of the MapReduce programming framework to clinical big data analysis: current landscape and future trends. [europepmc]
- aTRAM - automated target restricted assembly method: a fast method for assembling loci across divergent taxa from next-generation sequencing data. [europepmc]
- Forecasting the 2013-2014 influenza season using Wikipedia. [europepmc]
- Big Data Analytics in Healthcare. [europepmc]
- Methodological challenges and analytic opportunities for modeling and interpreting Big Healthcare Data. [europepmc]
- The real cost of sequencing: scaling computation to keep pace with data generation. [europepmc]
- Scalable metagenomics alignment research tool (SMART): a scalable, rapid, and complete search heuristic for the classification of metagenomic sequences from complex sequence populations. [europepmc]
- A novel process of viral vector barcoding and library preparation enables high-diversity library generation and recombination-free paired-end sequencing. [europepmc]
- Patient Similarity: Emerging Concepts in Systems and Precision Medicine. [europepmc]
- Cloud computing for genomic data analysis and collaboration. [europepmc]
- Opportunities and obstacles for deep learning in biology and medicine. [europepmc]
- Clustering algorithms: A comparative approach. [europepmc]
- CaImAn an open source tool for scalable calcium imaging data analysis. [europepmc]
- Machine Learning and Integrative Analysis of Biomedical Big Data. [europepmc]
- Social big data: Recent achievements and new challenges. [europepmc]
- Data-Driven Molecular Dynamics: A Multifaceted Challenge. [europepmc]
- A fast and scalable framework for large-scale and ultrahigh-dimensional sparse regression with application to the UK Biobank. [europepmc]
- mbkmeans: Fast clustering for single cell data using mini-batch k-means. [europepmc]
- Single-Cell Transcriptomics: Current Methods and Challenges in Data Acquisition and Analysis. [europepmc]
Related