Reliability Issues in Computing System Design

The work

TitleReliability Issues in Computing System Design
AuthorsB. Randell; P. A. Lee; P. C. Treleaven
Typearticle
Year1978
Citekeyrandell1978reliability

Where it appeared

Published inACM Computing Surveys
PublisherAssociation for Computing Machinery
Volume10
Issue2
Pages123--165

Identifiers

DOI10.1145/356725.356729

Abstract

This paper surveys the various problems involved in achieving very high reliability from complex computing systems, and discusses the relationship between system structuring techniques and techniques of fault tolerance. Topics covered include: 1) protective redundancy in hardware and software; 2) the use of atomic actions to structure the activity of a system to limit information flow; 3) error detection techniques; 4) strategies for locating and dealing with faults and for assessing the damage they have caused; and 5) forward and backward error recovery techniques, based on the concepts of recovery line, commitment, exception, and compensation. The ideas described relate to techniques used to date in systems intended for environments in which high reliability is demanded. Three specific systems, the JPL-STAR, the Bell Laboratories ESS No. 1A processor, and the PLURIBUS, are described in some detail and compared.

Copy held

KindPDF, 3.9 MB
Retrieved2026-08-12
Heldlocal, for personal reference

Where this came from

How it got herealready cited · cited in bibtex
First seen2026-08-12
Recordreviewed by a person
Approved2026-08-16

Cite it as

@article{randell1978reliability,
  title = {Reliability Issues in Computing System Design},
  author = {B. Randell and P. A. Lee and P. C. Treleaven},
  year = {1978},
  journal = {ACM Computing Surveys},
  volume = {10},
  number = {2},
  pages = {123--165},
  publisher = {Association for Computing Machinery},
  doi = {10.1145/356725.356729},
}

This record lives at https://refs.drheap.org/randell1978reliability/ and will keep doing so.