This description was written by a machine and published without a person checking it. It is what the agent made of this grouping, and not a statement anybody has stood behind.

After the shared exponential

A subject the papers are about. The loosest grouping, and the one to reach for last.

What computing does when the improvement everybody shared stops arriving.

The set is not one argument but a disagreement about where the next gains come from, and its members sort by the answer they give. *From the layers above the device*: Leiserson's 'room at the Top' and Thompson on the decline of the general-purpose computer, both from the same MIT group, and Hennessy and Patterson's golden age. *From specialisation*: the TPU paper, which is the case study the golden-age argument leans on -- and Hestness, which is the same claim measured rather than advocated, finding that putting CPU and GPU cores on one die does not make their memory behaviour converge. *From making the die bigger instead of the features smaller*: the wafer-scale line, the set's longest thread -- Leighton and Leiserson in 1983 proposing it and proving their placement procedures reliable only under an assumed model of cell failure, Kung the year after, Lauterbach in 2021 reporting what it took Cerebras to ship one, and Hu's 2023 survey of what is still unsolved. *From stacking*: RevaMp3D on the processor core and cache hierarchy, and Najibi on what stacking costs -- deeply scaled 3D MPSoCs whose stated obstacle is temperature, answered with an integrated flow cell array. *From admitting the floor*: Theis and Shalf.

Two older pairs are here because the argument is older than the phrase. Patterson's RISC case and Clark's reply are the field's precedent for buying performance with architecture when the device will not supply it; Wulf and McKee's memory wall is the precedent for a shared exponential that stopped being shared -- logic outran memory decades before it stopped outrunning itself, and Hestness is that same measurement problem in the heterogeneous case.

Worth noticing across the answers: three of them are about heat or defects rather than about speed. Wafer-scale is limited by dead cells, stacking by temperature, and the floor papers by dissipation. The obstacle moved from how fast a device switches to whether a large assembly of them can be made to work at all.

Read against `device-to-scaling-limit`, which covers the physics that produced the exponential rather than the responses to its end, and `the-cost-is-in-forgetting`, where the one limit in the corpus that is settled rather than extrapolated is kept.

19 references

Wafer-scale Computing: Advancements, Challenges, and Future Perspectives
Yang Hu and others (2023) · arXiv
A Charge Domain P-8T SRAM Compute-In-Memory with Low-Cost DAC/ADC Operation for 4-bit Input Processing
Joonhyung Kim and others (2022) · Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and Design · Association for Computing Machinery
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
Nika Mansouri Ghiasi and others (2022)
The Path to Successful Wafer-Scale Integration: The Cerebras Story
Gary Lauterbach (2021) · IEEE Micro · Institute of Electrical and Electronics Engineers (IEEE)
The decline of computers as a general purpose technology
Neil C. Thompson and others (2021) · Communications of the ACM · Association for Computing Machinery
There's plenty of room at the Top: What will drive computer performance after Moore's law?
Charles E. Leiserson and others (2020) · Science · American Association for the Advancement of Science (AAAS)
Towards Deeply Scaled 3D MPSoCs with Integrated Flow Cell Array Technology
Halima Najibi and others (2020) · Proceedings of the Great Lakes Symposium on VLSI 2020 · Association for Computing Machinery
The future of computing beyond Moore's Law
John Shalf (2020) · Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences
A new golden age for computer architecture
John L. Hennessy and others (2019) · Communications of the ACM · Association for Computing Machinery
In-Datacenter Performance Analysis of a Tensor Processing Unit
Norman P. Jouppi and others (2017) · Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17) · Association for Computing Machinery
The End of Moore’s Law: A New Beginning for Information Technology
Thomas N. Theis and others (2017) · Computing in Science & Engineering
Past and future of hardware and architecture
David A. Patterson (2015) · Proceedings of the SOSP History Day 2015 · Association for Computing Machinery
A comparative analysis of microarchitecture effects on CPU and GPU memory system behavior
Joel Hestness and others (2014) · 2014 IEEE International Symposium on Workload Characterization (IISWC) · IEEE
VLSI architecture: Past, present, and future
William J. Dally and others (1999) · Proceedings 20th Anniversary Conference on Advanced Research in VLSI · IEEE
Hitting the Memory Wall: Implications of the Obvious
Wm. A. Wulf and others (1995) · ACM SIGARCH Computer Architecture News · Association for Computing Machinery
Wafer-scale integration and two-level pipelined implementations of systolic arrays
H. T. Kung and others (1984) · Journal of Parallel and Distributed Computing
Wafer-Scale Integration of Systolic Arrays
Frank Thomson Leighton and others (1983) · MIT Laboratory for Computer Science
Comments on "The case for the reduced instruction set computer," by Patterson and Ditzel
Douglas W. Clark and others (1980) · ACM SIGARCH Computer Architecture News · ACM SIGARCH
The case for the reduced instruction set computer
David A. Patterson and others (1980) · ACM SIGARCH Computer Architecture News · ACM SIGARCH