After the shared exponential
A subject the papers are about. The loosest grouping, and the one to reach for last.
What computing does when the improvement everybody shared stops arriving.
The set is not one argument but a disagreement about where the next gains come from, and its members sort by the answer they give. *From the layers above the device*: Leiserson's 'room at the Top' and Thompson on the decline of the general-purpose computer, both from the same MIT group, and Hennessy and Patterson's golden age. *From specialisation*: the TPU paper, which is the case study the golden-age argument leans on -- and Hestness, which is the same claim measured rather than advocated, finding that putting CPU and GPU cores on one die does not make their memory behaviour converge. *From making the die bigger instead of the features smaller*: the wafer-scale line, the set's longest thread -- Leighton and Leiserson in 1983 proposing it and proving their placement procedures reliable only under an assumed model of cell failure, Kung the year after, Lauterbach in 2021 reporting what it took Cerebras to ship one, and Hu's 2023 survey of what is still unsolved. *From stacking*: RevaMp3D on the processor core and cache hierarchy, and Najibi on what stacking costs -- deeply scaled 3D MPSoCs whose stated obstacle is temperature, answered with an integrated flow cell array. *From admitting the floor*: Theis and Shalf.
Two older pairs are here because the argument is older than the phrase. Patterson's RISC case and Clark's reply are the field's precedent for buying performance with architecture when the device will not supply it; Wulf and McKee's memory wall is the precedent for a shared exponential that stopped being shared -- logic outran memory decades before it stopped outrunning itself, and Hestness is that same measurement problem in the heterogeneous case.
Worth noticing across the answers: three of them are about heat or defects rather than about speed. Wafer-scale is limited by dead cells, stacking by temperature, and the floor papers by dissipation. The obstacle moved from how fast a device switches to whether a large assembly of them can be made to work at all.
Read against `device-to-scaling-limit`, which covers the physics that produced the exponential rather than the responses to its end, and `the-cost-is-in-forgetting`, where the one limit in the corpus that is settled rather than extrapolated is kept.