2016/08/20 by Marius Poke, Torsten Hoefler, Poke, Marius +3 · 1 voice
Computer Science · #Cloud Computing and Resource Management #Distributed systems and fault tolerance #Software System Performance and Reliability #cs.DC
paper · pdf · doi:10.48550/arxiv.1608.05866
openalex publication_date 2016/08/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Many distributed systems require coordination between the components involved. With the steady growth of such systems, the probability of failures increases, which necessitates scalable fault-tolerant agreement protocols. The most common practical agreement protocol, for such scenarios, is leader-based atomic broadcast. In this work, we propose AllConcur, a distributed system that provides agreement through a leaderless concurrent atomic broadcast algorithm, thus, not suffering from the bottleneck of a central coordinator. In AllConcur, all components exchange messages concurrently through a logical overlay network that employs early termination to minimize the agreement latency. Our implementation of AllConcur supports standard sockets-based TCP as well as high-performance InfiniBand Verbs communications. AllConcur can handle up to 135 million requests per second and achieves 17x higher throughput than today's standard leader-based protocols, such as Libpaxos. Thus, AllConcur is highly competitive with regard to existing solutions and, due to its decentralized approach, enables hitherto unattainable system designs in a variety of fields.