2003/01/20 by Armando Fox, E.A. Brewer · 1 citation
Computer Science · #Distributed systems and fault tolerance #Software System Performance and Reliability #Advanced Data Storage Technologies
paper · doi:10.1109/hotos.1999.798396
openalex publication_date 2003/01/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
The cost of reconciling consistency and state management with high availability is highly magnified by the unprecedented scale and robustness requirements of today's Internet applications. We propose two strategies for improving overall availability using simple mechanisms that scale over large applications whose output behavior tolerates graceful degradation. We characterize this degradation in terms of harvest and yield, and map it directly onto engineering mechanisms that enhance availability by improving fault isolation, and in some cases also simplify programming. By collecting examples of related techniques in the literature and illustrating the surprising range of applications that can benefit from these approaches, we hope to motivate a broader research program in this area.