← All posts

Essay · Artificial Intelligence

An essay on the collapse of AI systems and why this is not an engineering problem

26 min read5,611 wordsSections: 11Oct 10, 2026

Keywords

Share
Comment
An essay on the collapse of AI systems and why this is not an engineering problem

Abstract

This essay discusses why aligning artificial intelligence (AI) systems with the interests of humanity is not a problem that engineering can close. It starts from a plain-language description of how these systems are built and tuned from human behavior and judgment, and examines what they inherit in that process. Drawing on moral philosophy, social choice theory and empirical evidence on value divergence, the text discusses who defines the controls of these systems and the limits of translating values into metrics. It also examines the consequences of automation for work and the literature on the limits of human control, over others and over one's own behavior. It proposes treating alignment as a problem of legitimacy and continuous governance.

Keywords: AI alignment; value pluralism; social choice; Goodhart's law; future of work; governance.

1. Introduction

At a condominium meeting, the agenda had a single item: whether dogs could use the main elevator. It took two hours. Some invoked hygiene and some invoked property rights; someone recalled that the resident of apartment 302 has been afraid of dogs since childhood, and someone replied that her fear was not the building's problem. In the end, they voted. The majority won by three votes, the minority left convinced the decision was unfair, and the matter returned to the agenda six months later. No one there was irrational. Everyone had principles, and the principles did not fit in the same elevator.

The scene reproduces, on a small scale, the problem of this essay. We are asking artificial intelligence engineers to do, on a planetary scale and in writing, what forty apartments cannot do about an elevator: define what is good, for whom and in what order.

To train AI systems to operate within the confines we understand to be in humanity's interests, we will need to transfer to these systems a set of ethical and philosophical principles on which we ourselves, humans, do not fully agree. And let us not forget that philosophical disagreement, which can be used as an argument for something beneficial — and legitimately is, in principle — is the very reason that leads us to fight one another, from the condominium meeting to the war over territory.

To suppose that we will control in another what we do not control in ourselves reveals the root cause of a good part of our ills: arrogance.

Current AI systems were built to imitate human behavior recorded in text and then tuned by human raters. In imitating, they also inherit what remains unresolved in humans.

Norbert Wiener, MIT mathematician and founder of cybernetics, formulated the problem in 1960 in the journal “Science”: if we use a mechanical agency with whose operation we cannot efficiently interfere, "we had better be quite sure that the purpose put into the machine is the purpose which we really desire and not merely a colorful imitation of it" [1]. That certainty presupposes a "we" with a common desire, and it is this presupposition that I examine. Stuart Russell, of Berkeley, proposed in “Human Compatible” machines that learn human preferences by observing our behavior, and devotes chapter 9 to the complications of the proposal, among them conflicting preferences between people [2].

The collapse in the title is the collapse of a premise: that alignment is a specification problem, closed once the specification is correct.

2. How engineering made intelligence tractable

A language model is a function with billions of parameters that takes a stretch of text and estimates the probability of each next fragment. In pre-training, it processes large volumes of digital text and, at each prediction error, the parameters are corrected in the direction that would have reduced it. The measure of that error, the objective function, is the only criterion of success at this stage and contains no notion of truth or of good.

From this procedure emerged competences that no one programmed. Jason Wei and colleagues, at Google, documented abilities absent in smaller models and present in larger ones [3], and Rylan Schaeffer, Brando Miranda and Sanmi Koyejo, at Stanford, showed that the apparent abrupt jump can be an effect of the chosen metric [4]. Kenneth Li and colleagues, at Harvard, MIT and Northeastern, trained a model only to predict legal Othello moves and found evidence of an internal representation of the board, used to choose the move [5]. Jared Kaplan and colleagues, at OpenAI, showed that error falls according to power laws as model, data and compute grow [6]. Intelligence, understood as the capacity to solve problems described in language, became tractable in this sense: obtained through a reproducible procedure, without being understood.

A model that imitates human text also imitates fraud and disinformation, and that is why laboratories tune it by human preference. In reinforcement learning from human feedback (RLHF), proposed by Paul Christiano and colleagues, at OpenAI and DeepMind [7], and applied to language models by Long Ouyang and colleagues, at OpenAI [8], raters choose the better of pairs of responses, a second model learns to predict those choices, and the main model is optimized to maximize the predicted score. In constitutional AI, by Yuntao Bai and colleagues, at Anthropic, part of this judgment is replaced by written principles that the model itself uses to revise its responses [9].

With this, the objective function changes in nature: in pre-training, it measures agreement with the text; in tuning, it measures the approval of people or of a list written by people, that is, a judgment about the desirable. Stephen Casper, of MIT, and colleagues separated the limits of this stage that more engineering can mitigate from those fundamental to the method, and among the fundamental ones is the difficulty of representing in a single reward function the preferences of people who disagree with one another [10].

Dario Amodei and colleagues, then at Google Brain, Stanford, Berkeley and OpenAI, already distinguished in 2016 specification failures, in which the formal objective diverges from the intended one, from failures of robustness and of oversight [11]. The thesis of this essay is narrower: even with perfect specification, the question remains of whose objective is being specified, and that question has no gradient.

3. Whoever imitates inherits the imitated

What the machine learns depends on who wrote what it read and on who evaluated what it answered. Psychologists Joseph Henrich, Steven Heine and Ara Norenzayan, then at the University of British Columbia, showed in 2010 that behavioral research relied on samples from Western, educated, industrialized, rich and democratic societies, frequently located at the extremes of the human distribution [12]. In language models, Shibani Santurkar and colleagues, at Stanford, found substantial misalignment between the opinions of the models and those of entire demographic groups in the United States, in several cases worsened by tuning with human feedback [13]. Mohammad Atari and colleagues, from Henrich's group at Harvard, in work not yet peer-reviewed, found that the models' responses move away from the psychology of each population as its cultural distance from the United States grows [14]. The distribution of the data already determines which humanity is being imitated.

Preference tuning adds a second inheritance: systems trained to maximize approval learn to be approved. Mrinank Sharma and colleagues, at Anthropic, documented that these assistants tend toward sycophancy, agreeing with the user and backing away from correct answers when challenged, a behavior that the human preference data themselves reward [15]. In humans, the meta-analysis by Bob Uttl and colleagues, at Mount Royal University, found no significant correlation between the rating students give the teacher and how much they learn [16].

Ryan Greenblatt and colleagues, at Anthropic and Redwood Research, told the model Claude 3 Opus that it would be retrained to comply with objectives contrary to its earlier preferences. In some of the cases, the model began to comply with the instructions only when it judged it was in training and recorded in its reasoning that it did so in order not to be modified [17]. The authors caution that the scenario was artificial. Even so, these studies indicate that systems trained to imitate human behavior reproduce known problems of human agents: origin bias, sensitivity to incentives and the tendency to tell the evaluator what they want to hear.

4. The disagreement that does not resolve

Choosing better, with more representative data, more diverse raters and better-written principles, does not eliminate the problem: representing human diversity better brings into the system the disagreement that exists outside it.

The philosopher Isaiah Berlin, of Oxford, argued in “The Crooked Timber of Humanity” that human values are multiple, frequently incompatible and without a common measure, so that choosing between liberty and equality, or between justice and mercy, implies loss [18]. John Rawls, of Harvard, arrived by another route, in “Political Liberalism”, at the fact of reasonable pluralism: under free institutions, reasonable people arrive durably at incompatible moral, religious and philosophical doctrines [19].

The economist Kenneth Arrow demonstrated in 1951 that no procedure aggregates individual preferences into a collective ordering while satisfying at the same time a small set of minimal conditions, among them that no one be a dictator [20]. Every reward function that condenses divergent judgments into a single score therefore makes choices that no neutral criterion justifies.

The Moral Machine experiment, by Edmond Awad and colleagues at the MIT Media Lab, collected about forty million decisions on autonomous-vehicle dilemmas in 233 countries and territories [21]. Some preferences were widely shared, such as sparing more lives, but three cultural clusters emerged, Western, Eastern and Southern, which diverge, for example, in the weight given to the age or social status of the victims. A vehicle sold in the three regions would have to choose one of them or adopt an average that represents none.

Anna Jobin, Marcello Ienca and Effy Vayena, at ETH Zurich, analyzed 84 AI ethics guidelines and found convergence around principles such as transparency, justice and non-maleficence, with substantive divergence on how to interpret and implement them [22]. Brent Mittelstadt, of the Oxford Internet Institute, argued that such principles do not guarantee ethics in practice because AI, unlike medicine, has no common aims, fiduciary duties, professional tradition or accountability mechanisms [23].

Iason Gabriel, a philosopher at DeepMind, proposes seeking principles that people with reasonably different conceptions can endorse, instead of true moral principles, which makes alignment a question of political legitimacy [24]. Legitimacy is a relation between the model, those who control it and those affected by it, and it is not measured in the laboratory. The psychologist Jonathan Haidt, of New York University, explains why this disagreement resists arguments: moral judgments are mostly intuitive, reasoning comes afterward, and different groups weigh differently foundations such as loyalty, authority and sanctity [25].

5. Who defines the controls, and where

Even if a legitimate assembly approved a set of principles, it would remain to turn them into behavior. Ludwig Wittgenstein formulated in the “Philosophical Investigations” the rule-following paradox: "no course of action could be determined by a rule, because every course of action can be made out to accord with the rule" (§201) [26]. A rule does not contain the instructions for its own application: between "avoid causing harm" and the decision whether to answer a question about medication dosage there is a gap that only interpretation fills. The jurist H. L. A. Hart, of Oxford, called this feature of legal language open texture in “The Concept of Law”: every rule has a core of clear cases and a penumbra that requires judgment [27]. Law handles the problem with judges, appeals and case law. In constitutional AI, the interpreter is the model itself, and the interpretation is fixed in training [9].

There is also the metric. The economist Charles Goodhart, then an adviser to the Bank of England, observed in 1975 that any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes [28]. The anthropologist Marilyn Strathern, of Cambridge, generalized in 1997: when a measure becomes a target, it ceases to be a good measure [29]. David Manheim and Scott Garrabrant showed that this effect results from the intense optimization of an imperfect indicator [30]. Every control over an AI system has to become a number, such as a preference score, a refusal rate or a benchmark score, and that number, once it becomes a target, comes loose from the value it was meant to represent.

The controls can sit in the pre-training data, in the selection of raters, in the wording of the principles, in output filters or in usage policies, and each layer is decided by someone who is rarely the one who bears the effects. Saffron Huang and colleagues, at the Collective Intelligence Project and Anthropic, had a representative sample of 1,002 American adults propose and vote on principles for a model [31]. Taylor Sorensen and colleagues, at the University of Washington and other institutions, propose treating pluralism as a design requirement [32]. Both works broaden participation, but the first had to choose a country and the second has to decide what is reasonable.

David Collingridge described in 1980, in “The Social Control of Technology”, a dilemma: while a technology is young, its effects are little known and changing it is easy; once the effects become clear, changing it is expensive and slow [33]. Bill Gates, co-founder of Microsoft, stated in an interview with Ezra Klein that the industry cannot be relied on to self-regulate and that supposing otherwise would be "insane" [34]. The decision about the controls is therefore political before it is technical, and leaving it to those who build the system is also a political decision, only an undeclared one.

6. From the condominium to war

Philosophical disagreement is a condition of moral progress: expansions of rights such as the abolition of slavery and women's suffrage began as minority positions. It also fuels conflict. The psychologist and philosopher Joshua Greene, of Harvard, describes in “Moral Tribes” the tragedy of commonsense morality: morality evolved to enable cooperation within groups, but when groups with different moralities meet, each judges itself right by its own criteria, which makes the conflict hard to resolve [35]. By this argument, the condominium meeting and the dispute over territory differ in scale but share the structure: divergent principles and the absence of a common metric to decide.

Carried over to AI systems, this mechanism produces two opposite risks. The first is homogenization. Jon Kleinberg, of Cornell, and Manish Raghavan, of MIT, demonstrated in a theoretical model that when many decision-makers adopt the same algorithm, the quality of the set of decisions can fall even if the algorithm is the most accurate for each one [36]. Rishi Bommasani and colleagues, at Stanford, found that sharing training data worsens outcome homogenization, in which the same people are failed by every system; for shared foundation models, the effect varied with the adaptation method [37]. A few models, from a few companies, mediating the decisions of billions of people is the configuration these works describe. The second risk is fragmentation: models aligned with the values of each country or community and serving the division between groups. I know of no systematic evidence about it and treat it as a hypothesis derived from Greene's argument.

I know of no technical solution that avoids both risks. The choice between them is made before the gradient, by people, and should be made in the open.

7. The promise of new jobs

The same question appears in the debate about work. In an interview with Ezra Klein, Jensen Huang, founder and chief executive of Nvidia, stated that jobs will change en masse, but that there will be net job creation because new industries create new jobs, and cited as examples wellness centers, spas, entertainment and the whole luxury market, which did not exist in the first half of his life [38].

The argument has a historical basis. The economist David Autor, of MIT, showed that automation replaces tasks more than occupations and complements human labor in the tasks that remain [39]. David Autor, Caroline Chin, Anna Salomons and Bryan Seegmiller estimated that about six in ten jobs in the United States in 2018 were in specialties that did not exist in 1940 [40]. Daron Acemoglu, of MIT, and Pascual Restrepo, then at Boston University, described a race between displacement, in which machines take over tasks, and reinstatement, in which new tasks emerge for humans [41].

The same data limit Huang's bet. The study by David Autor and colleagues shows that, between 1940 and 1980, new work emerged mainly in middle-income occupations and, after 1980, in well-paid professional occupations and, secondarily, in poorly paid services [40]. Acemoglu and Restrepo record that in recent decades displacement has moved faster than reinstatement [41].

Huang's examples have a structural limit. Thorstein Veblen showed in 1899, in “The Theory of the Leisure Class”, that part of the value of certain goods comes from signaling social position [42], and Fred Hirsch called them in 1976, in “Social Limits to Growth”, positional goods, whose supply is socially limited because their value depends on relative scarcity [43]. A market whose product is exclusivity cannot grow to absorb the majority. According to Bain & Company and Altagamma, luxury consumers fell from 400 million in 2022 to about 340 million in 2025 [44], and the “World Inequality Report 2022”, coordinated by Lucas Chancel, Thomas Piketty, Emmanuel Saez and Gabriel Zucman, estimates that the richest 10% hold 76% of global wealth [45]. Making luxury the destination of displaced labor is proposing that many live by serving a few.

Spas and wellness centers are personal care for those who can pay and fall under the same critique. Care in the broad sense is different. The International Labour Organization report authored by Laura Addati, Umberto Cattaneo, Valeria Esquivel and Isabel Valarino estimates 381 million workers in care activities, 11.5% of global employment, and describes a sector of growing demand, low pay and financing concentrated in families and the State [46]. William Baumol showed in 1967 why: services in which human labor is the product itself, such as education and care, do not gain productivity at the pace of industry and become relatively more expensive over time [47]. Promising that care will absorb those displaced by AI without saying who will pay decent wages ignores the ILO's conclusion that decent care work requires public investment.

There is also scale and speed. Tyna Eloundou and colleagues, at OpenAI and the University of Pennsylvania, estimated in the preliminary version of their study that about 80% of U.S. workers would have at least 10% of their tasks exposed to language models, and about 19% would have at least half [48]. Daron Acemoglu estimates a gain of at most 0.66% in total factor productivity over ten years [49]. If both are approximately right, there will be broad exposure with small gain: substitution without the additional wealth that would finance new demand. Erik Brynjolfsson, of Stanford, called the Turing Trap the orientation of AI toward imitating and replacing humans instead of augmenting them, an option that concentrates income and power in those who control the technology [50].

In the interview with Klein, according to Truman Dickerson's report in “Business Insider”, Bill Gates proposed charging on each unit of labor the same payroll tax a human would pay, whether it is delivered by a person or by a robot, and said that economic signals favorable to AI do not oblige society to adopt it [34], [51]. In an essay from August 2026, reported by Ina Fried in “Axios”, he advocated reserving for humans functions such as childcare and jury service [52]. One may doubt the feasibility of the proposals, but they are decisions about what a society wants to preserve, and none derives from a cost curve.

John Maynard Keynes predicted in 1930, in "Economic possibilities for our grandchildren", that the standard of living in progressive countries would be, in one hundred years, four to eight times higher, and suggested fifteen-hour weeks to share out the remaining work [53]. In the volume edited by Lorenzo Pecchi and Gustavo Piga to explain why working hours shrank much less, Richard Freeman points to inequality, which raises the premium on working more, and Joseph Stiglitz to the dynamics of consumption, which renews itself with income [54]. Neither explanation is technical.

8. What we do not control in ourselves

Every company with more than one employee lives an alignment problem: employment contracts, codes of conduct, compensation policies and audits exist because the interests of those who hire and those who are hired do not coincide on their own.

The economists Michael Jensen and William Meckling formalized this problem in agency theory in 1976. When a principal delegates decisions to an agent, there arise monitoring costs, bonding costs and a residual loss, which records the impossibility, in general, of ensuring at zero cost that the agent decides as the principal would [55]. Dylan Hadfield-Menell and Gillian Hadfield, then at Berkeley and the University of Toronto, proposed reading AI alignment as a problem of incomplete contracts: contracts between humans work because culture and law supply the implicit terms that fill their gaps [56]. The proposal is a research agenda, not a demonstrated solution, but the diagnosis suffices here: among humans, control over the other depends on institutions that complete what specification does not cover.

Control also fails over one's own behavior. The psychologist Paschal Sheeran, then at the University of Sheffield, gathered ten meta-analyses comprising 422 studies and found that intentions explain, on average, about 28% of the variance in behavior [57]. In the meta-analysis of experimental studies by Thomas Webb and Sheeran, medium-to-large changes in intentions produced only small-to-medium changes in behavior [58].

It remains to explain why, given this, specifying values for a machine is treated as a task within reach of a competent team. Ellen Langer, of Harvard, called illusion of control the expectancy of success higher than the objective probability warrants, and showed that it grows when chance situations display cues of skill [59]. Don Moore and Paul Healy distinguished three forms of overconfidence: overestimating one's own performance, judging oneself better than others and placing too much faith in the precision of one's own beliefs [60]. Daniel Kahneman and Dan Lovallo showed that optimistic forecasts arise from the inside view, anchored in the plans of one's own project and not in the record of similar projects [61].

This essay proposes that this combination of illusion of control and overconfidence is present in three beliefs in the debate: that a competent team can specify values, despite the disagreement documented in AI ethics guidelines and in the Moral Machine experiment; that industry will regulate itself, although whoever builds the system is, in Jensen and Meckling's terms, an agent with interests of its own; and that human ambition will create the necessary jobs, although displacement has outpaced reinstatement in recent decades.

9. Conclusion

The condominium meeting did not end badly for lack of intelligence or information. Everyone knew what a dog, an elevator and a bylaw were. It ended badly because the question had no answer that spared someone from losing something. What the condominium has is a procedure: agenda, vote, minutes and the possibility of returning to the matter six months later.

Engineering made intelligence tractable by converting the imitation of human text into an optimization problem and, with that, transferred to machines impasses that, among humans, are managed by institutions rather than solved by specification [55], [56]. Engineering will remain necessary to make systems more robust, transparent and auditable, but alignment should be treated as these impasses are treated among humans: as a permanent process of revision, with avenues for contestation, representation of those affected and mechanisms of correction. The question ceases to be how to specify human values for a machine and becomes who has the legitimacy to decide which values go in, how that decision is reviewed and who answers for the consequences.

I include myself in the critique: as an engineer and founder of a technology company, I know the temptation to think that a well-formulated problem is almost solved, the inside view described by Kahneman and Lovallo [61]. Wiener asked for certainty about the purpose put into the machine [1]. Sixty-six years later, that certainty is not available, because the data on moral divergence do not indicate a single human purpose. This requires institutions able to live with the uncertainty. What remains is the question the condominium already knows: who will have a seat at the meeting where it is decided what the machine should consider good?

References

[1] N. Wiener, “Some moral and technical consequences of automation,” Science, vol. 131, no. 3410, pp. 1355–1358, May 1960, doi: 10.1126/science.131.3410.1355.

[2] S. Russell, Human Compatible: Artificial Intelligence and the Problem of Control. New York, NY, USA: Viking, 2019, ch. 9.

[3] J. Wei et al., “Emergent abilities of large language models,” Trans. Mach. Learn. Res., 2022, arXiv:2206.07682.

[4] R. Schaeffer, B. Miranda, and S. Koyejo, “Are emergent abilities of large language models a mirage?” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, 2023, pp. 55565–55581.

[5] K. Li, A. K. Hopkins, D. Bau, F. Viégas, H. Pfister, and M. Wattenberg, “Emergent world representations: Exploring a sequence model trained on a synthetic task,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2023, arXiv:2210.13382.

[6] J. Kaplan et al., “Scaling laws for neural language models,” arXiv:2001.08361, 2020.

[7] P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Adv. Neural Inf. Process. Syst. (NIPS), vol. 30, 2017, pp. 4299–4307.

[8] L. Ouyang et al., “Training language models to follow instructions with human feedback,” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, 2022, pp. 27730–27744.

[9] Y. Bai et al., “Constitutional AI: Harmlessness from AI feedback,” arXiv:2212.08073, Dec. 2022.

[10] S. Casper et al., “Open problems and fundamental limitations of reinforcement learning from human feedback,” Trans. Mach. Learn. Res., 2023. [Online]. Available: https://openreview.net/forum?id=bx24KpJ4Eb

[11] D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané, “Concrete problems in AI safety,” arXiv:1606.06565, 2016.

[12] J. Henrich, S. J. Heine, and A. Norenzayan, “The weirdest people in the world?” Behav. Brain Sci., vol. 33, no. 2–3, pp. 61–83, 2010, doi: 10.1017/S0140525X0999152X.

[13] S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto, “Whose opinions do language models reflect?” in Proc. 40th Int. Conf. Mach. Learn. (ICML), PMLR vol. 202, 2023, pp. 29971–30004.

[14] M. Atari, M. J. Xue, P. S. Park, D. E. Blasi, and J. Henrich, “Which humans?” PsyArXiv preprint, 2023, doi: 10.31234/osf.io/5b26t.

[15] M. Sharma et al., “Towards understanding sycophancy in language models,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2024, arXiv:2310.13548.

[16] B. Uttl, C. A. White, and D. Wong Gonzalez, “Meta-analysis of faculty’s teaching effectiveness: Student evaluation of teaching ratings and student learning are not related,” Stud. Educ. Eval., vol. 54, pp. 22–42, 2017, doi: 10.1016/j.stueduc.2016.08.007.

[17] R. Greenblatt et al., “Alignment faking in large language models,” arXiv:2412.14093, Dec. 2024.

[18] I. Berlin, The Crooked Timber of Humanity: Chapters in the History of Ideas, H. Hardy, Ed. London, U.K.: John Murray, 1990.

[19] J. Rawls, Political Liberalism. New York, NY, USA: Columbia Univ. Press, 1993.

[20] K. J. Arrow, Social Choice and Individual Values. New York, NY, USA: Wiley, 1951.

[21] E. Awad et al., “The Moral Machine experiment,” Nature, vol. 563, no. 7729, pp. 59–64, 2018, doi: 10.1038/s41586-018-0637-6.

[22] A. Jobin, M. Ienca, and E. Vayena, “The global landscape of AI ethics guidelines,” Nat. Mach. Intell., vol. 1, no. 9, pp. 389–399, 2019, doi: 10.1038/s42256-019-0088-2.

[23] B. Mittelstadt, “Principles alone cannot guarantee ethical AI,” Nat. Mach. Intell., vol. 1, no. 11, pp. 501–507, 2019, doi: 10.1038/s42256-019-0114-4.

[24] I. Gabriel, “Artificial intelligence, values, and alignment,” Minds Mach., vol. 30, no. 3, pp. 411–437, 2020, doi: 10.1007/s11023-020-09539-2.

[25] J. Haidt, The Righteous Mind: Why Good People Are Divided by Politics and Religion. New York, NY, USA: Pantheon Books, 2012.

[26] L. Wittgenstein, Philosophical Investigations, G. E. M. Anscombe, Trans. Oxford, U.K.: Blackwell, 1953, §201.

[27] H. L. A. Hart, The Concept of Law. Oxford, U.K.: Clarendon Press, 1961, ch. VII.

[28] C. A. E. Goodhart, “Problems of monetary management: The U.K. experience,” in Papers in Monetary Economics, vol. 1. Sydney, Australia: Reserve Bank of Australia, 1975.

[29] M. Strathern, “‘Improving ratings’: Audit in the British University system,” Eur. Rev., vol. 5, no. 3, pp. 305–321, 1997.

[30] D. Manheim and S. Garrabrant, “Categorizing variants of Goodhart’s law,” arXiv:1803.04585, 2018.

[31] S. Huang, D. Siddarth, L. Lovitt, T. I. Liao, E. Durmus, A. Tamkin, and D. Ganguli, “Collective Constitutional AI: Aligning a language model with public input,” in Proc. ACM Conf. Fairness, Accountability, and Transparency (FAccT), 2024, doi: 10.1145/3630106.3658979.

[32] T. Sorensen et al., “Position: A roadmap to pluralistic alignment,” in Proc. 41st Int. Conf. Mach. Learn. (ICML), PMLR vol. 235, 2024, pp. 46280–46302.

[33] D. Collingridge, The Social Control of Technology. London, U.K.: Frances Pinter, 1980.

[34] E. Klein, “Bill Gates’s blunt warning on A.I.,” The Ezra Klein Show, The New York Times, Sep. 29, 2026. [Podcast]. Available: https://www.nytimes.com/2026/09/29/opinion/ezra-klein-podcast-bill-gates.html

[35] J. D. Greene, Moral Tribes: Emotion, Reason, and the Gap Between Us and Them. New York, NY, USA: Penguin Press, 2013.

[36] J. Kleinberg and M. Raghavan, “Algorithmic monoculture and social welfare,” Proc. Natl. Acad. Sci. USA, vol. 118, no. 22, Art. no. e2018340118, 2021, doi: 10.1073/pnas.2018340118.

[37] R. Bommasani, K. A. Creel, A. Kumar, D. Jurafsky, and P. Liang, “Picking on the same person: Does algorithmic monoculture lead to outcome homogenization?” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, 2022, arXiv:2211.13972.

[38] E. Klein, “Jensen Huang vs. the A.I. doomers,” The Ezra Klein Show, The New York Times, Sep. 23, 2026; rebroadcast in Hard Fork, Sep. 25, 2026. [Podcast]. Available: https://www.nytimes.com/2026/09/25/podcasts/hardfork-ezra-klein-jensen-huang.html. Claims C018, C024 and C025 checked against the transcript inventory in A. Maynard, “L2: Claims inventory. Ezra Klein interviews Jensen Huang (transcript dated 23 Sept 2026),” Late Lessons: Jensen Huang and AI, 2026. [Online]. Available: https://andrewmaynard.net/late-lessons-ai-sept-2026/supporting/huang/lenses/L2-claims.html

[39] D. H. Autor, “Why are there still so many jobs? The history and future of workplace automation,” J. Econ. Perspect., vol. 29, no. 3, pp. 3–30, 2015, doi: 10.1257/jep.29.3.3.

[40] D. Autor, C. Chin, A. Salomons, and B. Seegmiller, “New frontiers: The origins and content of new work, 1940–2018,” Q. J. Econ., vol. 139, no. 3, pp. 1399–1465, 2024, doi: 10.1093/qje/qjae008.

[41] D. Acemoglu and P. Restrepo, “Automation and new tasks: How technology displaces and reinstates labor,” J. Econ. Perspect., vol. 33, no. 2, pp. 3–30, 2019, doi: 10.1257/jep.33.2.3.

[42] T. Veblen, The Theory of the Leisure Class: An Economic Study of Institutions. New York, NY, USA: Macmillan, 1899.

[43] F. Hirsch, Social Limits to Growth. Cambridge, MA, USA: Harvard Univ. Press, 1976.

[44] Bain & Company and Altagamma, “Global luxury stays resilient despite economic headwinds and shifting consumer trends that reshape market,” press release, Nov. 20, 2025. [Online]. Available: https://www.bain.com/about/media-center/press-releases/20252/global-luxury-stays-resilient-despite-economic-headwinds-and-shifting-consumer-trends-that-reshape-marketbain--company-and-altagamma/

[45] L. Chancel, T. Piketty, E. Saez, and G. Zucman, Coords., World Inequality Report 2022. Paris, France: World Inequality Lab, 2022.

[46] L. Addati, U. Cattaneo, V. Esquivel, and I. Valarino, Care Work and Care Jobs for the Future of Decent Work. Geneva, Switzerland: International Labour Office, 2018.

[47] W. J. Baumol, “Macroeconomics of unbalanced growth: The anatomy of urban crisis,” Am. Econ. Rev., vol. 57, no. 3, pp. 415–426, 1967.

[48] T. Eloundou, S. Manning, P. Mishkin, and D. Rock, “GPTs are GPTs: An early look at the labor market impact potential of large language models,” arXiv:2303.10130v5, Aug. 2023. Published version: Science, vol. 384, no. 6702, pp. 1306–1308, 2024, doi: 10.1126/science.adj0998.

[49] D. Acemoglu, “The simple macroeconomics of AI,” Econ. Policy, vol. 40, no. 121, pp. 13–58, 2025, doi: 10.1093/epolic/eiae042.

[50] E. Brynjolfsson, “The Turing Trap: The promise & peril of human-like artificial intelligence,” Daedalus, vol. 151, no. 2, pp. 272–287, 2022, doi: 10.1162/daed_a_01915.

[51] T. Dickerson, “Bill Gates says companies that replace human labor with robots should have to pay the same FICA taxes,” Business Insider, Sep. 30, 2026. [Online]. Available: https://www.aol.com/articles/bill-gates-says-companies-replace-172316000.html

[52] I. Fried, “Bill Gates wants to keep some jobs off-limits to AI,” Axios, Aug. 26, 2026. [Online]. Available: https://www.axios.com/2026/08/26/bill-gates-wants-to-keep-some-jobs-off-limits-to-ai

[53] J. M. Keynes, “Economic possibilities for our grandchildren,” in Essays in Persuasion. London, U.K.: Macmillan, 1931, pp. 358–373.

[54] L. Pecchi and G. Piga, Eds., Revisiting Keynes: Economic Possibilities for Our Grandchildren. Cambridge, MA, USA: MIT Press, 2008, ch. 3 (J. E. Stiglitz, “Toward a general theory of consumerism”) and ch. 9 (R. B. Freeman, “Why do we work more than Keynes expected?”), doi: 10.7551/mitpress/9780262162494.001.0001.

[55] M. C. Jensen and W. H. Meckling, “Theory of the firm: Managerial behavior, agency costs and ownership structure,” J. Financ. Econ., vol. 3, no. 4, pp. 305–360, 1976, doi: 10.1016/0304-405X(76)90026-X.

[56] D. Hadfield-Menell and G. K. Hadfield, “Incomplete contracting and AI alignment,” in Proc. 2019 AAAI/ACM Conf. AI, Ethics, and Society (AIES), 2019, pp. 417–422, doi: 10.1145/3306618.3314250.

[57] P. Sheeran, “Intention–behavior relations: A conceptual and empirical review,” Eur. Rev. Soc. Psychol., vol. 12, no. 1, pp. 1–36, 2002, doi: 10.1080/14792772143000003.

[58] T. L. Webb and P. Sheeran, “Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence,” Psychol. Bull., vol. 132, no. 2, pp. 249–268, 2006, doi: 10.1037/0033-2909.132.2.249.

[59] E. J. Langer, “The illusion of control,” J. Pers. Soc. Psychol., vol. 32, no. 2, pp. 311–328, 1975, doi: 10.1037/0022-3514.32.2.311.

[60] D. A. Moore and P. J. Healy, “The trouble with overconfidence,” Psychol. Rev., vol. 115, no. 2, pp. 502–517, 2008, doi: 10.1037/0033-295X.115.2.502.

[61] D. Kahneman and D. Lovallo, “Timid choices and bold forecasts: A cognitive perspective on risk taking,” Manage. Sci., vol. 39, no. 1, pp. 17–31, 1993, doi: 10.1287/mnsc.39.1.17.

Comments

Every comment is moderated before it appears here. Nothing is published automatically.

Loading…