← All posts

Essay · Artificial Intelligence

The Layer That Teaches: compression, abstraction, and the formation of judgement in software engineering

65 min read14,259 wordsSections: 39Images: 5Oct 10, 2026

Keywords

Share
Comment
The Layer That Teaches: compression, abstraction, and the formation of judgement in software engineering

Abstract

We discovered fire by striking stones, and today we get it from a lighter without missing the capability we lost. This paper examines what changes when abstraction moves from task execution to the formulation of the solution, and argues that compression is abstraction and that abstraction applied to the foundation degrades learning. It first addresses formation and employability, disputing the premise that AI literacy, as currently taught, solves the entry of junior professionals into the market, and it locates the argument in the labour-process literature on deskilling, distinguishing those mechanisms from this one. The evidence shows performance gains alongside losses in comprehension, though the hypothesis that this harm concentrates among those who know least remains without the interaction term that would settle it. The conclusion about juniors does not depend on that hypothesis, because the asymmetry lies in stock position, where those already formed lose an advantage while those not yet formed never arrive. To avoid repeating a prediction the history of computing has already seen fail, the paper offers as an operational criterion of foundation the minimum layer whose absence prevents diagnosing failure at the level immediately above. It then argues that today’s senior engineer is valuable partly for having been formed without AI, that this stock does not replenish itself, and that from this follows the rule of treating AI as a tool and not as a methodology, with the caveat that delegating method is a bet on a shrinking stock, not an announced ruin.

Keywords— Engineering education, learning and automation, deskilling, tacit knowledge, AI assistance, formation of judgement.

1. Introduction

The average adult today cannot make fire. Anyone can “obtain” fire in under a second, with a lighter that costs what a coffee costs, and the difference between obtaining and making costs nothing day to day. The ability to produce fire by friction, for tens of thousands of years the boundary between surviving and not surviving, became a survivalist’s pastime without anyone mourning the loss.

The trade was a good one for two reasons that usually go unnoticed. The first is that the abstracted capability stopped being necessary in order to judge the result, since you need not know how to make fire to know that the food burned. The second is that the compression was local, because while fire became a lighter other spheres of life kept demanding first-hand learning, and the species kept learning, only other things.

The question of this paper is what happens when those two conditions stop holding at the same time, when the abstracted layer comes to be precisely the one needed to judge what the abstraction produced and compression stops being local, coming to operate across nearly every sphere at once, including those in which the species used to compensate for what it lost elsewhere.

The first condition already fails, and the argument depends on it; the second is conjecture, treated as such in Section 3.6 and kept off the critical path, because none of the conclusions about software engineering requires compression to be simultaneous. Software engineering is where the first failure becomes visible first, being the domain of fastest adoption and densest measurement, to the point of concentrating more than half of enterprise use of AI tools while other professions still register in single digits [1].

The thesis. Compression is abstraction, in the sense that every technology that compresses the time of an activity abstracts the path that produced knowledge about it. This is benign as long as the abstracted path is not the foundation on which the result is judged; when it is, compression degrades formation, and degrades it silently, because delivery indicators improve over the same period in which the capacity to judge deteriorates.

Formation and employability. Industry’s answer to the collapse in junior hiring is AI literacy, understood as operational fluency. I argue that this fluency does not replace formation in the foundation, and that early, unrestricted exposure to assistance produces performance without comprehension.

Software engineering. The senior engineer’s edge is usually described as domain knowledge, which matters but does not explain everything, because seniors in 2026 understand the layer the agent abstracts by virtue of having been formed before assistance. The current stock of seniority is, in that sense, a historical artefact rather than a renewable resource, and the practical consequence is a rule of organisational design, that of using AI as a tool and not as a methodology.

I run an engineering company, I teach at a school of computing, and I have argued elsewhere that the effort which consolidates knowledge is exactly what no tool delivers ready-made. The thesis that human formation remains the scarce resource favours both positions I hold, and since declaring that interest does not fix the problem, I applied to the text the test of looking, for each central claim, for the evidence that would cost me something. Sections 3.4, 3.5, 4.2 and 5.2 exist because of it, and each one narrows a claim I would have preferred to keep whole.

2. Method and scope

This is a work of synthesis and argument, without a primary study, whose contribution to the literature it cites lies in the articulation of one mechanism, the abstraction of the foundation, linking findings from the psychology of learning, human factors, labour economics and software engineering that today circulate separately.

Three rules were applied to the sources. The first was to trace every numerical claim to its primary source and to report the real scope of the finding, not the summarised version in circulation. The second was to identify preprints and vendor reports as such in the body, because an argument about epistemic rigour cannot be built from material treated as firmer than it is. Under the third, evidence that contradicts the thesis gets its own section rather than a footnote, and there are two such cases in the learning literature (Section 4.3) and one in the strategy literature (Section 9).

Applying the first rule showed that several of the most cited statistics in executive material on AI and engineering [2] do not survive checking, less through bad faith than through chain citation, in which each link cites the previous one and nobody returns to the source. Online Resource 1 lists the eight claims I traced, with what the primary source actually supports. The same treatment applies to the sources I depend on here, and the reference list below is complete so that the reader can apply it.

Since the central thesis speaks of a horizon of years and the available evidence measures sessions and months, it has the standing of a grounded hypothesis, with the scope limits declared in Section 12.

3. Compression is abstraction

3.1. What compression removes

I call compression any technology that reduces the time between intention and result, and abstraction what it does to the path, because the intermediate steps go on existing, only hidden. The lighter encapsulated friction in a mechanism the user need not understand, without making it unnecessary.

The claim that there is no compression algorithm for experience [3] is usually read as praise for seniority, but it says something stronger about what does not compress. Information and procedure compress; experience does not, because it is the residue of a lived process, and residue is not transmitted. When technology compresses the process, the very residue that mattered fails to form.

The psychology of learning treats that residue as a variable, along three paths. Slamecka and Graf showed that an item produced by the subject is retained better than the same item merely read [4]. Sweller located the mechanism in cognitive load: building the schemas that distinguish expert from novice consumes the same capacity that solving the problem demands, which is why solving and learning compete [5]. Bjork and Bjork showed that conditions which worsen performance during practice improve retention and transfer afterwards [6], [7]. In all three, friction appears as part of the mechanism of learning, which is uncomfortable for any productivity tool applied to a learner.

One point of vocabulary avoids confusion later, because “judgement” carries three senses in this discussion. The first sense is cognitive and designates the capacity to assess whether an output is right; that is the capacity I consider at risk. The second is economic, opposing judgement to intelligence in order to classify tasks by how codifiable they are; that is the sense used in Section 7.2. The third is organisational, naming the decision about method and architecture that Section 10 addresses. The three touch without being interchangeable, and where judgement appears alone, without qualification, the first is meant.

Education research already has both a name and an instrument tradition for that first sense. Tai and colleagues call it evaluative judgement, the capability to make decisions about the quality of work of self and others, and treat it as something a curriculum can develop deliberately instead of leaving it as a residue of experience [8]. Bearman and Ajjawi move in the direction the present case requires when they argue that, faced with systems whose internals cannot be inspected, the pedagogical answer is not to make the box transparent but to teach students to work with it, which requires exactly the standards of quality that evaluative judgement supplies [9]. The difference between their prescription and mine is where the standard comes from. Their pedagogy assumes a learner who can acquire standards while working with the box; I claim that the standards for formulation in software were acquired by having formulated, which is the acquisition the assistant now performs. Whichever is right, the construct is measurable, and that removes one excuse for not measuring it.

3.2. Why this compression differs from earlier ones

The obvious objection is the lighter. Humanity has always abstracted, there was always someone lamenting it and whoever lamented was almost always wrong, since the calculator did not destroy mathematics, nor did the compiler destroy programming or the IDE the engineer, and it remains to explain why now would be different.

The difference lies in the layer each technology abstracted. The calculator removed arithmetic execution and left the formulation of the problem intact, just as the compiler removed translation to instructions without touching algorithm design and the IDE removed syntax recall while leaving the programmer the decision of what to write. The generative assistant removes decomposition, the choice of approach, and the entire first version, returning to the human only acceptance or rejection of the result (Figure 1).

How far each compression reaches. Earlier technologies stopped below formulation. The generative assistant reaches it, and returns to the human only the layer that depends on it.
How far each compression reaches. Earlier technologies stopped below formulation. The generative assistant reaches it, and returns to the human only the layer that depends on it.

Figure 1. How far each compression reaches. Earlier technologies stopped below formulation. The generative assistant reaches it, and returns to the human only the layer that depends on it.

This is a difference of kind and not of degree, because earlier compressions abstracted execution and preserved formulation, whereas the generative assistant abstracts formulation and returns to the human only judgement, when it was formulation that produced the capacity to judge. The loop closes on itself, since judging the output well requires having walked the path the output dispenses with.

Bainbridge described this structure four decades ago as the irony of automation. The more advanced the automatic system, the more critical the human operator’s contribution in anomalies becomes, and it is precisely automation that removes the practice by which that operator would maintain the competence to intervene [10]. Parasuraman and Manzey added the attentional mechanism underlying uncritical acceptance of the machine’s recommendation [11], and Carr generalised the effect to knowledge work [12]. The difference from that picture is one of temporal order, because Bainbridge [10] described operators who already knew how to operate and were losing the practice, and the present case is that of apprentices who will never have it.

3.3. What the labour-process literature already established

The claim that technology removes the path by which skill is formed is fifty years old. Braverman argued that the separation of conception from execution was not a by-product of mechanisation but its purpose, since management captures the worker’s knowledge, encodes it into procedure and machine, and returns to the worker a task stripped of discretion [13], and the vocabulary of deskilling comes from there. The present case differs in two respects. The first is agency, because Braverman’s separation is imposed from above, against the worker’s interest, and resisted, whereas the compression described here is adopted from below, tool by tool, by the very person whose formation it interrupts, and it arrives as a favour. The second is direction, since in Braverman the worker’s output does not improve as discretion falls, while here the output improves immediately, which makes the loss hard to see and hard to argue against. Attewell’s review of the controversy corrects the thesis where it needs correcting by showing that the deskilling argument over-generalised, because the same period upgraded skill in other occupations, and that the level of analysis chosen among task, job and occupation decided the answer [14]. The offshoring episode of Section 5.2 applies that correction to my own claim.

Zuboff’s field studies [15] are the closest historical analogue. Where work became mediated by symbols on a screen, operators lost the action-centred skill that came from handling the material directly, and what had to replace it was an intellective skill built by interpreting the representation. Two of her findings carry over to the present case. The first is that the outcome was a design decision and not a property of the machine, since whether operators developed the new skill depended on being given access to the underlying data and the time to interpret it, and the safeguarded arm of Bastani et al. forces the same conclusion on this paper (Section 4.3). The second is that the operators who had run the plant by hand before the conversion read the screen in a way their successors could not, which is the stock argument of Section 8 appearing in a factory four decades earlier.

Collins supplies the epistemic half by showing that tacit knowledge is not one thing and that the part he calls collective tacit knowledge is acquired by socialisation into a practice, which instruction cannot do [16]. With Evans he separates interactional expertise, the ability to speak the language of a domain without being able to contribute to it, from contributory expertise, which requires having practised [17]. When AI literacy is taught as operational fluency, it produces interactional expertise in a practice the student has never performed, convincing to the student and to the interviewer alike, and the distinction drawn by Collins and Evans puts the objection of Section 5.3 in stronger terms than I had for it.

The contribution still available against that lineage is narrow. Braverman’s mechanism is organisational, Zuboff’s is the mediation of work by representation and Collins’s is epistemic, and none of the three describes a tool that performs the formulation step itself and returns a complete and plausible artefact, leaving the human only acceptance or rejection. Nor do they pose the question of stock and flow raised in Section 8, which asks not whether the current worker loses skill, but whether the next one can be produced at all, once the tasks by which newcomers entered are the first the tool absorbs.

3.4. The precedent this argument must face

Arguments like this one have been made before, by people with greater technical authority than mine, and they failed.

In June 1975, in manuscript EWD498, Dijkstra wrote that it is practically impossible to teach good programming to students previously exposed to BASIC, because as potential programmers they were mentally mutilated beyond hope of regeneration [18]. Backus, in his history of FORTRAN, records that in 1954 scepticism about automatic programming was widespread, and that the community did not believe a system could produce efficient code, because efficiency was human judgement and human judgement could not be automated [19]. Ensmenger documents the full cycle of these attempts and the reactions to them [20].

Those arguments had the same structure as mine, in which abstraction removes the professional from the foundation and therefore destroys the formation of judgement, and both predictions were falsifiable and failed.

My argument lacks, as Dijkstra’s did, a criterion of foundation independent of the layer in which the writer happened to be trained, and without it the thesis says only that the next generation knows less of what mine learned, which every generation says.

I offer the following criterion, verifiable:

“Foundation is the minimum layer whose absence prevents diagnosing failure at the level immediately above.”

The criterion leaves things behind without nostalgia. Microcode and manual register allocation left the foundation because almost no current failure requires descending to them, whereas concurrency semantics stayed, because a large class of production failures is diagnosable only by those who understand it, and computational cost stayed by the same test.

Applied to the present case, the criterion says where the risk lies, because the characteristic failure of a system built with assistance is one of formulation rather than of syntax or translation. It shows up as a plausible solution to the wrong problem, as a structure that works in the demonstrated case and breaks in the neighbouring one, or as an architectural decision nobody explicitly made.

The claim is verifiable, and Apiiro’s analysis of roughly 62,000 enterprise repositories offers a first measurement of it. Comparing code with and without assistance, syntax errors fall by 76% and logic bugs by 60%, while architectural design flaws rise by 153% and privilege-escalation paths by 322% [21]. In the authors’ qualitative reading, those using assistance tend toward design-level flaws and those not using it toward logic mistakes, so the defect changed layers without diminishing, and changed in the direction the criterion predicts. Being vendor analysis with only partly published methodology, it counts as a first measurement and does not yet have the standing of an established finding.

Diagnosing failures of formulation requires precisely the layer that assistance absorbs. Unlike BASIC, which abstracted a layer whose absence did not prevent diagnosing the level above, the assistant abstracts the very layer diagnosis requires.

Two properties of the criterion bound it. The first is that it is indexed to the current distribution of failure modes, and that distribution is itself partly produced by the tools in use. Microcode left the foundation because current failures rarely require descending to it, and if assistants stopped producing formulation-level failures the criterion would reclassify formulation in the same way. That is a feature for discarding nostalgia and a limitation for prediction, because the criterion says what the foundation is now without saying what it will be. The second property concerns evidence, since the support for the applied criterion is, at present, one vendor analysis with partly published methodology. An adequate test would compare defect classes across matched projects, with assistance as an assigned rather than an observed condition, and with defects classified blind to that condition. Until such a test exists, the criterion is a proposal with one measurement in its favour.

The argument may be wrong, and the way to show it is to exhibit a relevant class of failures in assisted systems that is diagnosable without the formulation layer.

3.5. Is the capacity the individual’s or the system’s?

When I say there will be nobody left to judge, I treat the capacity to judge as internalised in persons. A consolidated tradition denies that premise.

Wegner formalised the transactive memory system, in which the group’s knowledge is each member’s memory plus a directory of who knows what, so that nobody needs to know everything for the system to know everything [22]. Hutchins showed, in the field, that a ship’s position is computed by a sociotechnical system of which no single navigator executes or comprehends the whole, and that the memory of reference speeds is a property of the cockpit rather than of the pilots [23]. Hollan, Hutchins and Kirsh extended the programme to the design of interactive systems [24].

At this point the lighter from the opening works against me. Nobody knows how to make fire, and even so society’s capacity to control fire rose, institutionalised in fire services, building codes and fire engineering; the capacity migrated from the individual to the system, and the system is better than any individual was.

The right question becomes whether the capacity can be institutionalised, more than whether individual capacity falls, and here there is a difference that seems to me decisive between the generative assistant and the artefacts studied by the distributed cognition literature. The sextant, the takeoff checklist and the compiler are deterministic and auditable, and the capacity they carry can be inspected, versioned, taught and corrected by whoever operates them, even without that person reconstructing it from scratch. The generative model has none of those properties, because its output varies between runs, its justification is produced after the decision and the error is not localisable in the artefact.

The question does not close there, and Section 11.2 describes the attempt to return the agent to the category of institutionalisable artefacts, by making deterministic and auditable the layer that decides what it may do. That layer settles permission, making auditable what the agent is entitled to execute, but it does not say whether its output is right. The half of the objection concerning safety has a built answer; the half concerning judgement depends on continuous quality evaluation in production, the least mature piece of the stack. Until that piece exists with the same rigour, the sextant analogy holds only for the half of the objection that does not worry me.

3.6. The objection from scale

The second condition that made the lighter trade benign was locality: you lose the ability to make fire and gain time to learn something else, so the social sum of learning shifts rather than falls. That is the implicit premise of nearly every optimistic argument about AI and work, and it is reasonable while compression is local, but it evaporates when compression is simultaneous. If the same technology abstracts the formulation of the solution in programming, in writing and in diagnosis at once, and across every level of training, no compensating sphere remains where first-hand learning would go on happening. Displacement presupposes somewhere to displace to, and the generality of the model puts in question the availability of that somewhere, more than human adaptability.

I did not locate empirical evidence on the aggregate effect of simultaneous compression across domains, for the simple reason that the phenomenon is three years old. The hypothesis stands as conjecture, repeated as such in Section 12, and what the available evidence supports is the in-domain effect.

4. What the learning evidence supports

4.1. The pattern that repeats

Five independent studies, with different populations, tasks and methods, find the same pattern: assistance improves the product and worsens formation (Table 1).

Table 1. Experimental evidence on AI assistance and learning.

StudyPopulationFinding
[25], PNAS≈1,000 high-school students, fieldWith access during exercises, performance rises 48% (open interface) and 127% (tutor with safeguards). With access withdrawn, the open-interface group performs 17% worse than those who never had access. The safeguards eliminate almost all the harm
[26], BJET117 university students, randomised experimentThe assisted group improves its essay grade and does not improve knowledge or transfer. What changes is self-regulation, outsourced to the model, which the authors call metacognitive laziness
[27], ICERNovice programmers, lab study with eye trackingAI widens the distance between strong and weak: those who already knew what they wanted to write gained; those who did not had their metacognitive difficulties aggravated and ended with inflated self-assessment, which the authors describe as an illusion of competence
[28]Software engineers learning a new libraryConceptual comprehension, code reading and debugging ability worsen, with no significant average efficiency gain. Those who delegated entirely gained productivity at the cost of not learning the library. In a related experiment, 50% against 67% correct on a comprehension test
[29], CHIKnowledge workers, 936 real usesHigher confidence in the tool is associated with lower self-reported critical effort, and the effort that remains shifts from producing to verifying

Two of them say something beyond the number. Prather et al. [27] observe an unequal effect, with gains for those who already have a formed schema and aggravated metacognitive difficulties for those who do not; Section 4.2 returns to this, because it is a laboratory observation and the field evidence points the other way. Bastani et al. [25] bring the strongest causal evidence available, with the particularity that the harm appears only when the crutch is withdrawn; while the assistant is present, the performance indicator registers the problem with its sign inverted, which is worse than hiding it.

Kosmyna and colleagues report weaker brain connectivity and lower ownership of one’s own text in the group that wrote with an assistant [30], which suggests a mechanism without establishing one, since this is a preprint with eighteen participants in the critical session, answered by a technical comment on statistical power, reproducibility and reporting transparency [31].

4.2. The gradient the literature measures is not the gradient the thesis needs

The tempting formulation would say that the harm of assistance concentrates among those with less prior knowledge, because it would connect the learning effect directly to the collapse in junior hiring. Today that formulation is untenable, and all the evidence on heterogeneity by skill level points the other way (Table 2).

Table 2. Studies estimating effects by skill level or seniority.

StudyNDesignDirection of the effect
[32], “Management Science”4,867 devsThree field RCTsAdoption and gains higher among recent hires and junior positions; no effect among those with longer tenure
[33], “QJE”5,172Staggered difference-in-differences+34% for novice and low-skill workers; effect near zero for the experienced
[34]758Pre-registered field RCT+43% below the performance average and +17% above; and a 19 percentage point drop on a task outside the capability frontier
[35], “Science”453Pre-registered RCTInequality between workers decreases
[36]70 completersRCT, single taskLess experienced developers gain more, with marginal significance

Taken together, that is over eleven thousand subjects, and in every study the sign is the same, with a larger gain for those who know less.

Two readings can be drawn from this result. The lazy one concludes that the concern about formation is refuted, which does not follow, because these studies measure productivity while the question is about learning, and nothing prevents assistance from raising juniors’ output more while reducing their learning more. There is even a plausible mechanism for this, the same as in Section 3, whereby the greater the distance between what the person knows and what the tool delivers, the greater the immediate gain and the greater the portion of the path not walked.

The other reading acknowledges that the interaction term between prior knowledge and learning has never been estimated, which I checked in the two available meta-analyses. Maier et al. aggregate 23 studies and 27 effect sizes, and find a positive and significant effect on productivity, g = 0.33 with an interval of [0.09, 0.58], against an effect indistinguishable from zero on learning, g = 0.14 with an interval of [−0.18, 0.47] [37]. Alanazi et al. aggregate 35 controlled studies and reach the same shape, with performance at 0.86 and comprehension at 0.16, p = 0.41 [38]. Neither reports moderation by prior knowledge, and Alanazi et al. record that absence as a declared limitation.

The state of the evidence on learning is therefore one of a null average effect with high heterogeneity, not of a negative effect. That is compatible with harm concentrated in one subgroup and gain in another, as I propose, and equally compatible with no effect at all, so that deciding between the two hypotheses rests on an estimate nobody has produced.

The evidence supports that assistance dissociates performance from comprehension, and that is measured in specific populations [25], [26], [28], but it does not support that the harm is greater for those who know less, which remains a hypothesis, and the gap that would test it is identified above.

The conclusion about juniors does not depend on the harm being greater for them, only on two things, namely that assistance dissociates performance from comprehension, which holds at any level, and that the junior is the one who still has to form. The asymmetry therefore comes from stock position, since seniors who degrade lose part of an advantage they already have, whereas juniors who do not form never reach the position from which they would degrade. If the effect is uniform, the conclusion holds and only the target of the warning changes, extending to the erosion of those already formed. If the effect is progressive, as the productivity evidence suggests, juniors deliver more and still do not form, which aggravates the problem, because it makes their delivery an even worse indicator of the capacity they are building.

4.3. Two tensions the thesis must face

The first is the expertise reversal effect. Kalyuga and colleagues showed that instructional scaffolds which help the novice lose effectiveness and come to harm the advanced learner [39]. Read hastily, the finding contradicts the thesis, because if support helps the novice there would be no reason to fear assistance. The correct reading is the inverse, since the reversal effect establishes that the optimal level of support is a function of prior knowledge, and the generative assistant does not modulate, applying maximum, undifferentiated, on-demand support to someone with zero prior knowledge. Moreover, the support it offers is not a worked example, which exposes the path, but a finished solution, which dispenses with it.

The second comes from Bastani et al. themselves. The tutor with safeguards, which refuses to hand over the answer and guides the student, produces a larger performance gain than the open interface “and” preserves learning [25]. This weakens any deterministic version of the thesis, because what degrades formation is not AI, but AI designed to deliver a result rather than to guide a path. The prescription is one of design, and Sections 6 and 10 address it.

5. The market and the literacy fallacy

5.1. Two curves in opposite directions

Rather than shrinking or growing as a whole, the technology labour market is splitting by seniority (Table 3).

Table 3. Labour market indicators: the two simultaneous directions. The sources differ in rigour, and include vendor reports, a news item and an opinion piece; the scope and caveat of each row are stated in the middle column.

IndicatorScope and caveatSource
−73% in entry-level hiringEuropean tech, year on year. P1-level hiring rate falling from 35% to 8%[40]
−50% or more in new-graduate hiringLarge technology firms, against the pre-pandemic level[41]
−7.7% junior headcount at AI-adopting firms62 million résumés. The revised working paper reports ≈−9%[42]
−20% employment among developers aged 22 to 25Against the 2022 peak. The 35 to 49 brackets keep growing[43]
+1.3 million AI-related jobsGlobal, two years, LinkedIn data. An opinion piece published by the WEF, not a report[44]
67,000 open engineering roles against 52,000 layoffsSimultaneous, Q1 2026, base of 9,000 companies[45]
“AI Engineer” as fastest-growing roleList of the 25 fastest-growing US roles in 2026[46]

The period of these series coincides with the post-pandemic hiring correction and with high interest rates, and none of the studies fully isolates the AI effect from the cycle effect. Hosseini Maasoum and Lichtinger come closest, comparing adopting with non-adopting firms, and even so measure adoption by “proxy” [42]. Attributing the whole decline to AI goes beyond the data, but the shape of the split is supported, and in it the variable separating the two trajectories is seniority.

5.2. The closest control case we have

The split by seniority has a precedent in a natural experiment we already ran. Between 2000 and 2015 the West moved exactly the junior layer of technology work abroad, and the argument of the day repeated today’s almost word for word, holding that without entry work nobody learns and seniority dries up.

Tambe and Hitt documented the displacement with about half a million workers, showing that firms with their own offshore centres came to employ 8% less of their domestic workforce in tradable occupations because they stopped hiring, without resorting to layoffs [47]. Vardi, revisiting the ACM task force five years later, found no evidence of measurable impact on aggregate technology employment in developed countries [48]. The bank-teller precedent has the same shape, with fewer tellers per branch, more branches and the remaining task changing content and becoming more valuable [49].

I found no study measuring the effect of that displacement on the formation of seniors at the firms that did it. The absence of literature about a collapse is weak evidence that none occurred, but fifteen years without a visible seniority crisis is a fact that any thesis like mine has to explain.

My explanation is that, under offshoring, entry work went on being performed by human beings who formed while performing it; the pipeline moved country without stopping, and the effect on world production of seniors was geographic reallocation. For the analogy not to hold, this difference must be true, and it is falsifiable: if the entry work absorbed by agents reappears elsewhere as formative human work, we are looking at another episode of reallocation rather than an interruption.

The comparison also reframes the question of whose formation is at stake. Offshoring moved entry work to India, Brazil, the Philippines and elsewhere, and moved the formative rungs with it; the seniority that grew there is now part of the global supply. If agents absorb that work, the loss falls first on the places that had been receiving it, and on markets with less capacity to subsidise formation internally, which changes who pays for it first without changing the mechanism. The employment series used here are from the United States and Europe because that is where the measurement exists, which is a limitation of evidence rather than a choice of scope.

The episode also shows, and this weighs against this paper’s propositions, that if the pipeline is aggregate, whoever trains pays alone a cost whose benefit leaks to the whole market, and the firm that does not train hires ready-made from the one that did. Offshoring is the historical case of firms doing that arithmetic correctly, one by one, until the aggregate arithmetic changed, and since the propositions in Section 6 are all firm-level, they must answer why a firm would pay.

While seniority was abundant, buying was cheaper than forming and the free-rider’s arithmetic held; in a market where senior vacancies stay open for months and the wage premium rises (Table 3), forming internally begins to compete with buying externally, and the reason to pay is thus economic before it is moral. That reason answers only in part, because it depends on scarcity persisting and because the return on formation outlasts almost any business plan. The rest is a coordination problem, historically solved outside the firm by accreditation and sector rules, and without that design the propositions below depend on firms willing to pay for a good they do not capture whole.

5.3. Why literacy is not the answer

The current answer has two parts, and both sound sensible. The first is to keep hiring juniors, which Matt Garman put by calling the substitution of junior staff by AI “the dumbest thing I’ve ever heard”, because without them there will be nobody, in ten years, who has learned anything [50]. I agree with that conclusion, and my disagreement lies in the second part, which wants juniors to be AI-literate and makes literacy in the tool what renders them employable.

That second part confuses two different competences, for three reasons.

Literacy is competence in operation; formation is competence in judgement. Knowing how to orchestrate an agent, write a good “prompt” and chain tools is a real and teachable skill, and it is, by construction, a skill about the abstract layer. It does not produce the schema that allows one to assess whether the output is right, and so task performance rises while comprehension falls (Section 4), when it is comprehension that the organisation will need in five years.

Literacy is the shortest-lived skill in the set. The interfaces, models and orchestration patterns of 2026 will not be those of 2029. The foundation has a half-life of decades, and includes data structures, computational cost, concurrency semantics and domain modelling. Training the young primarily in what expires fastest is an allocation choice that is hard to defend, even in purely economic terms.

Literacy does not solve the problem it means to solve. The argument says literate juniors become employable again because they produce like mid-level engineers, and the evidence in Section 4.2 supports that they indeed do, but the firm would be buying present capacity obtained through the tool, which it could equally obtain by giving the same tool to someone who already has judgement. What would sustain hiring the junior is the price difference between the two options, a difference that narrows as the tool improves.

Since Section 4.3 concedes that the design of assistance changes the outcome, rejecting literacy entirely would be incoherent, and the objection is aimed at shallow literacy, understood as operational fluency: writing good “prompts”, chaining tools, knowing this year’s interfaces. From what I have examined, that is mostly what the short courses, corporate upskilling tracks and platform certifications teach, but this is observation rather than measurement, because a content analysis of AI-literacy curricula against the distinction proposed here does not exist, and only such an analysis would settle whether the target is as common as I claim. That uncertainty does not reach the short half-life of operational fluency, nor the fact that it does not produce judgement.

The literacy that would solve the problem is something else, and the literature already indicates its shape. Shen and Tamkin identify six interaction patterns with the assistant, and three of them preserve learning even with the tool available, because they keep the learner cognitively engaged in the path [28]. Teaching those three patterns is teaching metacognition about the use of the tool, not operation of the tool, which calls for a different curriculum with different assessment, such as the one Section 6 proposes. The objection therefore applies only to the literacy that mistakes fluency for formation.

5.4. The legitimate periphery is being dismantled

Lave and Wenger called legitimate peripheral participation the mechanism by which a craft reproduces itself: the novice enters through real, low-risk tasks and migrates toward the centre as the community trusts them with more [51]. It is a description of how communities of practice transmit what they know, including what they cannot articulate in a manual [52], [53]. Dreyfus arrives at the same place by another route, showing that experts stop operating by explicit rules precisely because they accumulated concrete cases [54].

The periphery is the mechanism, and it is exactly the first layer assistance absorbs, that of the simple “CRUD”, the missing test, the small “bug” and the obvious fix, so that absorbing it is the first effect of assistance.

It follows that continuing to hire juniors is necessary and insufficient, because hiring juniors and setting them to operate the agent preserves the post and destroys the trajectory. Reaching this conclusion requires only the dissociation between performance and comprehension, without the hypothesis refused in Section 4.2, since what interrupts the trajectory is the loss of the periphery, not selective harm to the youngest. What has to be preserved are real, consequential, low-abstraction tasks through which the novice enters. Allen has the simple formulation of the human mechanism behind this, written about another transformation, according to which humans learn from other humans by watching, asking questions and repeating [3], and the organisation has to decide what is left for the novice to watch, ask about and repeat.

5.5. The periphery is also being bought from outside

There is a second movement, discussed in the investment literature and not in the education one. Bek argues that the next generation of AI companies will sell outcomes to the end client rather than tools to the professional, capturing services budgets roughly six times larger than software budgets [1]. The entry wedge is already-outsourced work, high in volume and low in judgement, whose budget already exists and whose replacement requires no internal reorganisation.

Read from the perspective of formation, that list of targets coincides, in almost every profession, with the tasks through which the junior used to enter: the standard NDA for the lawyer starting out, procedure coding in hospital billing, top-of-funnel screening in recruitment. Bek’s thesis [1] does not address this and need not, but it describes, from the capital-supply side, the same phenomenon I described from the formation side, with the periphery being removed by two vectors, one internal through the adoption of assistants and one external through the replacement of the supplier who performed it.

Hence the same fact carries opposite valences, and what the investment case reads as structural scarcity of professionals the profession reads as depletion of the stock of those who can judge [1].

6. Propositions for formation

The propositions below follow from what has been argued, and are formulated to be testable, or at least falsifiable in the practice of whoever adopts them.

P1. Deliberate low abstraction at the start of the trajectory. The curriculum and the first year of a career should contain an explicitly unassisted core, defined by a formation objective: building the data structure, debugging without suggestions, reading someone else’s code without an automatic summary. The prohibition applies only to that core, which is reserved as a space where friction is itself the objective [7]. P1 carries the boundary condition set out in Section 8: if low abstraction merely selected the current seniors without forming them, reintroducing friction will also merely filter, and filter with known bias. Whoever adopts P1 should therefore measure attrition alongside the performance of those who remain.

P2. Assistance with safeguards, not free assistance. The finding in Section 4.3 is operationalisable: schools and firms that give tools to beginners should configure tutor mode as the default rather than as an option [25].

P3. Assess by reconstruction and transfer, not by delivery. If the delivery indicator improves while formation worsens, it has stopped being an indicator of formation. Assessment by reconstruction measures what delivery has come to hide, and is done by asking the learner to explain, to modify under a new constraint, or to debug a corrupted version of their own artefact.

P4. Preserve the legitimate periphery as engineering policy. It falls to the organisation to decide which real tasks stay human for formative reasons, and to sustain that decision as an explicit investment with a declared cost, since a decision named as investment survives the next bad quarter better than one named as generosity. Section 5.2 states the limit of this proposition: part of the return leaks to the market, and the firm that adopts it alone subsidises those that do not. It is justified by the cost of buying seniority in a scarce market, and is not justified if that market becomes slack again.

P5. Treat formation as the firm’s responsibility, not only the school’s. The thesis that the team you have is the team you need was formulated by Allen about the move to the cloud [55], [56]: capacity is built inside rather than bought ready-made. Now, however, training on the new tool is not enough, because the old foundation also has to be protected.

7. What AI compresses in the economics of software

7.1. Hard to do is no longer hard to copy

Competitive advantages in software used to split into two families, the hard to do and the hard to get (Figure 2). Compression attacks the first and leaves the second, because it compresses the time it takes to do things and not the time things take to happen: years of operation still take years, and a regulatory process does not accelerate because the applicant writes code faster. Engineering complexity has therefore stopped being a moat, and if a competitor reproduces in weeks what took two years, what remains defensible is not in the code.

Compression reduces the time to do, not the time to happen. What remains defensible is in the lower band, and Section 7.2 qualifies that claim.
Compression reduces the time to do, not the time to happen. What remains defensible is in the lower band, and Section 7.2 qualifies that claim.

Figure 2. Compression reduces the time to do, not the time to happen. What remains defensible is in the lower band, and Section 7.2 qualifies that claim.

7.2. The prize is not in software

The six-to-one ratio between services and software budgets, presented in Section 5.5, reorganises this discussion: whoever sells a tool competes for the smaller of the two budgets, and the next trillion-dollar company would be a software company masquerading as a services firm [1].

Bek predicts where this happens first with a distinction of interest to anyone designing organisations. In “intelligence” tasks, the answer follows from known rules, codes and patterns; in “judgement” tasks, it depends on deciding what should be done. In his formulation, writing code is mostly intelligence, and knowing what to build next is judgement [1]. The higher the intelligence ratio in a profession, the sooner the outcome-selling model beats the tool-selling one.

Two consequences follow, and the first weakens what I wrote in Section 7.1, because if proprietary data is the hard-to-get moat, and the way to accumulate it is to perform service work that is already outsourced, then the position of whoever generates the data can be bought even though the time the data takes to exist cannot be compressed. The second reaches the central premise and has its own section.

7.3. The bottleneck moved

The adoption is measured: 88% of organisations use AI in some business function [57] and 90% of developers use it, with more than 80% perceiving a productivity gain [58], [59]. More informative than adoption, however, is the internal contradiction, since confidence in the accuracy of outputs fell from 40% to 29% in one year [60] and 30% report little or no confidence in generated code [58]. When almost everyone uses a technology that few trust, the work it was meant to eliminate is displaced into verification.

On the business-outcome side, the MIT NANDA report states that 95% of generative AI pilots produce no measurable impact on results [61]. It is a preprint whose methodology has been criticised as widely as its figure has been cited, but even with a generous discount the distance between adoption and impact is too large to be attributed to model capability.

The displacement is measurable and converges across independent sources, beginning with the individual. The controlled trial by Becker and colleagues at METR contradicts every prediction that preceded it, including the participants’ own. Experienced “open-source” developers, on codebases they knew well, took 19% longer to complete tasks with AI, and estimated at the end that they had been about 20% faster [62]. The sample of 16 developers and 246 tasks is small and does not generalise, so the interest lies less in the size of the effect than in the sign of the perception error, since perceived gain and measured gain point in opposite directions and restructuring decisions are being made on the first.

In the artefact, review became more expensive. Across 470 real “pull requests”, AI-coauthored PRs show 10.83 issues against 6.45 for human ones, with up to 2.74 times more security issues and 40% more critical findings [63]. Across roughly 62,000 enterprise repositories, three to four times more code is produced and ten times more security issues appear, with a 322% increase in privilege-escalation paths [21]. The two studies have different methodologies and should not be fused into a single number, although that fusion already circulates (Online Resource 1).

In the working day, the composition changed. Vella and Blincoe followed professional engineers at two moments six months apart and found 82% reporting less time writing code, with a statistically significant shift from creation to verification, and with the share reporting a worse development experience nearly doubling, from 14% to 27% [64]. Although it is a preprint and measures self-reported perceptions, the result converges with the finding of Lee and colleagues that critical effort shifts from producing to verifying [29].

For the organisation, if generation accelerated and verification did not, hiring more generation capacity does not help, because the constraint now lies on the other side.

8. Today’s senior is a historical artefact

The current explanation for the senior engineer’s value in the age of AI is domain knowledge, since he understands the client, the regulation and the nuance of the product, and AI does not compress that. The explanation is right but incomplete, as shown empirically by the fact that professionals with deep domain knowledge and no engineering training have produced, with assistance, artefacts that compete with those of professional developers. A lawyer and a cardiologist placed among the winners of a hackathon with thirteen thousand applicants, to cite the best-documented case [65]. If domain sufficed to explain the senior’s value, that result would not be possible.

The professional literature has a better answer than domain, and still an incomplete one. It describes a convergence of profiles: specialists broaden because AI absorbs their narrow execution tasks, and generalists deepen because they gain specialist-grade tools [66], [67], [68], [69]. Sanfilippo observes that whoever controls the ideas of the software is in a different position from whoever controls only the code [70], and Ng gives the economic version: since building became fast, more time is spent deciding what to build [71].

The convergence of profiles, as the professional literature describes it, and what it leaves out.
The convergence of profiles, as the professional literature describes it, and what it leaves out.

Figure 3. The convergence of profiles, as the professional literature describes it, and what it leaves out.

These authors correctly describe what the valuable professional does (Figure 3), but none of them explains how that capability came to exist, and whoever has to produce the next one depends on that answer.

What distinguishes senior engineers in 2026 is historical, because they were formed before assistance, when they wrote the hard implementation, debugged without suggestions, read other people’s code without a summary, were beaten by concurrency and watched the system fall without understanding why until they understood. That trajectory produced schemas about the layer the agent now abstracts, and with them they judge a non-deterministic output. When they review the agent’s code, they do little comparing with what they would have written; what they do is recognise, in the output, failure patterns one recognises only having stood beneath the abstraction.

Three consequences follow.

The first is that the stock does not replenish itself, and the reason is one of position rather than intensity, per Section 4.2. While the organisation operates on it, everything works, and the productivity gains of senior teams with agents are real. The problem is temporal, and the claim has a quantitative form even where the numbers are missing. The stock is the number of engineers who can judge formulation, drained by retirement, by turnover out of engineering and by context obsolescence, and I dispute none of these. Replenishment comes from juniors who enter and form, and is the product of the rate of hiring into consequential work and the fraction of those hired who form judgement while doing it. The first rate is falling, as Section 5 reports, and nobody measures the second, although I hold that assistance lowers it while raising output. Any of the three terms can be wrong. A firm that hires as before, and whose juniors form as before, is not exposed to this argument at all; a firm that cuts entry hiring and hands the remaining juniors an assistant is exposed twice, and its delivery metrics will not show it.

The second is that the senior’s merit is partly an accident of era, mine included, because I was formed on fundamentals for lack of an alternative rather than because I chose well. For formation policy, the consequence is to move the question from individual virtue to system design. If what produced the competence was the absence of a shortcut, and the shortcut now exists and is free, the only way to reproduce the competence is to deliberately reintroduce the constraint that used to come for free.

The third consequence is an uncertainty I cannot resolve with what exists today and that changes the prescription depending on the answer. I said seniors were “formed” by low abstraction, but the data are equally compatible with the hypothesis that they were “selected” by it, in which case those who did not cross the barrier left the profession and what we observe would be the residue of a filter rather than the product of a formation.

The two stories produce exactly the same observations and recommend opposite interventions. If the mechanism is formation, reintroducing friction forms people; if it is selection, it merely filters, and filters with known bias, because tolerance for friction is unevenly distributed and correlates with background, available time and support network. In that second case, proposition P1 of Section 6 stops producing more judgement and instead produces fewer people, from where few already came.

This is an explicit boundary condition of the proposition, and the design that would distinguish the two hypotheses is the one in Section 13, extended by the requirement to measure attrition alongside the performance of those who stayed.

9. What if the frontier moves?

Judgement would be the durable human residue, what remains when execution is absorbed, and Bek denies that premise by asserting that “Today’s judgement will become tomorrow’s intelligence” [1]. If he is right, this prescription has an expiry date and protects the last position before it falls too.

His argument for the frontier moving is specific and does not rest on a promise about future model capability. As AI systems perform intelligence work inside a profession, they accumulate proprietary data about how that profession’s decisions are made, and it is that accumulation which converts judgement into intelligence, through a data-acquisition mechanism of the kind that has a track record of working. I accept the mechanism and disagree with the conclusion on three points, all checkable against his own argument.

The first is that the mechanism is parasitic on what it consumes, because the frontier moves while the system observes human judgement being exercised, and if the profession stops forming those who exercise judgement, the data source dries up before the conversion completes. Bek’s argument [1] implicitly assumes a continuous stock of observable human judgement, and the two theses collide precisely there, because I described that same assumption as at risk.

The second is that the mechanism works where there is a pattern to extract, which differs from there being a decision to make. The categories he picks as priorities, such as hospital billing coding and standardised contract issuance, are cases where the decision was already rule-governed and human judgement entered as a transaction cost without involving any choice, which is why their fall was predictable before generative AI.

The third is temporal and bears directly on whoever has to decide, because even if the conversion happens, it happens profession by profession, over years, and depends on data that must be generated before it can be learned. The window in which an organisation operates with a declining stock of human judgement and incomplete conversion coincides with the five-to-ten-year horizon under discussion here, and within it the thesis [1], if correct, postpones the problem and transfers it to whoever holds the data, without eliminating it.

He does convince me, however, on the distinction between intelligence and judgement, which is more operational than the one I had been using, and Section 10 incorporates it as a delegation criterion.

10. AI as tool, not as methodology

10.1. The distinction

The prescription depends on separating tool from methodology, a distinction usually treated as rhetorical that has operational consequences. Methodology is what decides how work is organised, how the problem is decomposed, and how decisions are made and revised, whereas a tool is what executes within an already decided method. The same technology can occupy both places, and the difference between them does not show up in the short run, only on the horizon in which the organisation has to form whoever decides (Table 4).

Table 4. The same technology in two distinct places in the system of work.

AI as toolAI as methodology
Problem decompositionHuman, explicit, reviewableEmergent from the session with the agent
Acceptance criterionDefined beforehand, by whoever answers for the systemDefined by the plausibility of the output
ArchitectureDocumented human decision, with the agent implementing within itBy-product of what the agent generated and what passed the tests
Role of reviewVerify against a prior criterionDiscover the criterion while reviewing
Effect on formationThe decision path stays human and observableThe path disappears and the result remains
Failure modeLocalised, attributable, correctable errorDiffuse erosion, in which nobody knows why the system is the way it is

The second column works while the stock described in Section 8 lasts, and depends on it to work. Contrary to what public discussion assumes, the risk does not follow from having stopped training juniors in AI, but from their not having been exposed to the low abstraction of the foundation, which leaves them out of position to occupy the first column when the current occupants leave.

Risky, in this case, does not mean unviable, because two exits remain open and both were raised here. The capacity may be institutionalised in continuous evaluation and executable policy, as Section 3.5 discusses, or the frontier between intelligence and judgement may move far enough for the first column to shrink, as Section 9 argues. All that is established is that the second column consumes a stock whose replacement rate is falling and that neither exit is demonstrated, so operating this way is a bet, and whoever makes it should know what they are betting on.

The distinction between intelligence and judgement [1] gives this rule its cut-off criterion, and the criterion applies to the task, regardless of one’s confidence in the model. A task whose answer follows from a known rule is a legitimate candidate for delegation, because delegating it does not transfer method, whereas a task whose answer depends on deciding what should be done is not, even when the model produces a plausible output, because there delegation moves methodology inside the model.

10.2. What the distinction implies in practice

Specification and architecture stay human and written, not from distrust of the model but because they are the artefact that preserves the decision path, and the path is what forms whoever comes next and what allows auditing later. The agent operates inside a declared envelope, with scope, permissions and acceptance criteria defined before execution and verified outside the inference loop (Figure 5, Section 11.2).

Two allocations follow from this. Review becomes a first-class activity with an owner and a budget, because if the bottleneck migrated to verification (Section 7.3), treating review as a residual expense optimises the wrong side of the constraint, and accepting assisted output in production comes to require sign-off from someone who has the fundamentals, in recognition that the competence to judge is now the scarce resource.

11. Structure and governance

11.1. Who builds, who operates, who sustains

The separation between an engineering group that builds and an operations group that runs was already an anti-pattern before AI: one team making critical technical decisions and, when something goes wrong, waking another team to fix it is not compatible with ownership [72].

Agents make the separation unsustainable, and not merely undesirable, because an agent’s behaviour is non-deterministic and debugging it requires understanding the system “prompt”, the context window, retrieval quality and the reasoning chain, while an operations group organised to restart services and file tickets sees only the symptom. The test that serves as a diagnostic is operational and consists in checking whether the operations team can read prompts, trace reasoning chains and assess retrieval quality, or whether it restarts services and files tickets.

For delivery, the pod format that works is known and consists of a small team of experienced engineers with agents, with end-to-end ownership of a workflow. Conway’s law operates in it [73], and it is what Skelton and Pais call a stream-aligned team with full ownership [74]. Shipper takes the trend to its limit by observing that the two-pizza team has already grown too large for building software with agents [75]. That format requires a platform layer, because end-to-end ownership with agents is impossible without infrastructure that treats them as first-class citizens [76].

The pod and the organisation do not share an optimal shape. The inverted pyramid describes a delivery pod well; the hourglass is the only one that also reproduces the input the pod depends on.
The pod and the organisation do not share an optimal shape. The inverted pyramid describes a delivery pod well; the hourglass is the only one that also reproduces the input the pod depends on.

Figure 4. The pod and the organisation do not share an optimal shape. The inverted pyramid describes a delivery pod well; the hourglass is the only one that also reproduces the input the pod depends on.

As a consequence, shown in Figure 4, the optimal shape of the pod is not the optimal shape of the organisation, because a senior pod with agents delivers well and forms nobody, and an organisation made only of such pods delivers well for a few years and then has nowhere to draw whoever forms the next pod. The hourglass sustains both because it keeps the base, not as a social concession but as the layer that produces the model’s scarce input, even though it costs more in the short run and is the first thing to be cut in a bad quarter.

On where this decision lives, Allen’s formulation is that speed is an executive choice [72], and so is the shape of the organisation, except that its effects, unlike those of speed, appear only after the term of whoever chose it.

When part of the intelligence work migrates to suppliers selling outcomes, the buying organisation needs fewer people to execute and more people able to specify and verify, so the pressure falls on the same scarce competence.

11.2. Governance as infrastructure

If the agent operates inside an envelope, the envelope has to be executable. The convergence between regulators and product stacks suggests the framing is structural: Singapore’s IMDA organises agentic governance into bounding risks upfront, making humans meaningfully accountable, implementing technical controls across the lifecycle and enabling end-user transparency [77], and the logging and traceability obligations of Regulation (EU) 2024/1689 ask for substantially the same [78]. The four questions and the mechanism each one demands are summarised in Online Resource 2.

The agent’s envelope. The decision about what may be done is evaluated outside the inference loop, at the authorisation point, and is therefore not reachable by content entering the context.
The agent’s envelope. The decision about what may be done is evaluated outside the inference loop, at the authorisation point, and is therefore not reachable by content entering the context.

Figure 5. The agent’s envelope. The decision about what may be done is evaluated outside the inference loop, at the authorisation point, and is therefore not reachable by content entering the context.

The layer most implementations lack is the second. Asking the system to behave differs from preventing it from acting in the position at which the control operates, not in its content, which is why authorisation languages with analysable semantics matter more here than in conventional applications [79], [80].

Technically, this materialises what Section 10 prescribes, with the method outside the model, written by humans, verifiable and auditable, and the model executing inside it.

12. Limitations

The central thesis is a grounded hypothesis, not a finding. No available study follows a cohort of developers long enough to measure the formation of judgement, which is the dependent variable of the argument.

The base is young and partly commercial, on both sides. Relevant findings come from vendor reports with only partly published methodology [21], [63] and from preprints [28], [30], [61], [64]. This holds for the evidence supporting the thesis as much as for the evidence against it, since the market bifurcation rests on vendor reports [40], [41] and the strongest finding on formation is a preprint [28], and a reader who discounts one should therefore discount the other in the same proportion.

There is evidence that qualifies the thesis and evidence that confronts the premise. Sections 4.3 and 9 deal with both, and my answer to the second, being argumentative rather than empirical, remains open.

The thesis has to say what would refute it. A claim predicting consequences outside the measurement horizon has the shape of a self-protecting proposition, and I therefore declare three observations that, should they occur, count against the argument:

  1. An interaction term between prior knowledge and exposure to assistance, estimated with adequate power on a transfer measure, with a non-negative sign.
  1. The rate of promotion to senior, at intensively adopting organisations, remaining within the historical range over five years, corrected for the economic cycle.
  1. No difference in judgement quality, under blind assessment of architectural decisions, between engineers formed before and after 2023, controlling for years of experience.

There is no Brazilian data. All the employment series cited are from the United States or Europe, and the structure of the Brazilian market, with a different proportion of outsourcing and service exports, may produce different dynamics, so treating the European curves as predictive would be extrapolation.

The scale argument has no aggregate evidence. No aggregate series supports what Section 3.6 claims, and for that reason no other part of the argument is as speculative.

13. Conclusion and agenda

The lighter was a good trade because you did not need to know how to make fire in order to judge fire, and because the compression was local. The trade we are making now satisfies neither condition, since the abstracted layer is the one that produced the competence to judge and the compression is simultaneous across nearly every sphere. Without those conditions, the reassuring analogy that it was always like this and always turned out fine does not bear the weight placed on it.

The same mechanism operates at different scales in the person and in the organisation, which occupy the two halves of the argument. Assistance delivers the result to the individual and removes the path, and the indicator improves while the capacity worsens; in the organisation, the senior pod with agents delivers more than it ever did while the pipeline that produces seniors is dismantled by quarterly decision, and in both cases the damage remains invisible on the horizon where it is measured, though severe on the horizon where it is decided.

The prescription is also the same at both scales, because at both what is protected is the capacity to judge, the first to be lost and the last to return. In formation, it consists in reintroducing the constraint that used to come for free, with assistance that guides and assessment by reconstruction, under the condition declared in Section 8, according to which the constraint forms if the mechanism is formation and filters if it is selection, and the safeguarded arm of Bastani et al. weighs in favour of the first hypothesis [25]. In engineering, the method decision stays on the human side, with the model executing inside an envelope verified outside the inference loop.

The objection of Section 9 still stands, because, if today’s judgement becomes tomorrow’s intelligence, this prescription protects a temporary position. Even so, protecting a temporary position differs from protecting nothing by the time it buys, and the decision about what to do with that time has to be made within it.

The central question of the agenda is whether assistance degrades the formation of judgement over a long horizon, and under what design of task and tool that degradation disappears. The minimum design to answer it follows an early-career cohort for twenty-four months or more, with arms differentiated by exposure regime and measures of conceptual transfer and debugging rather than task completion. Two further questions remain open: is there, under simultaneous compression, a compensating sphere where first-hand learning goes on happening? And, if the frontier between intelligence and judgement does move, what happens to the mechanism that moves it when the source of observable human judgement begins to dry up?

Until those studies exist, every organisation cutting its base is running the experiment without a control group and without recording the result, and I wrote this because mine is one of them.

Statements and Declarations

Funding. No funding was received for this work.

Competing interests. The author directs an engineering company and teaches at a school of computing. The thesis that human formation remains the scarce resource is consistent with both positions, which is declared in Section 1 and treated there as a reason for adversarial sourcing rather than as a disclosure that settles the matter.

Ethics approval. Not applicable. The paper is a work of synthesis and argument and involves no human participants.

Data availability. No new data were generated. Every empirical claim is traced to a published primary source listed in the references; Online Resource 1 records the claims in circulation that could not be traced.

Use of generative AI. A large language model was used for translation between Portuguese and English, copy-editing, reference formatting and structural revision, and it produced the first draft of Section 3.3, which the author then verified against the primary texts and rewrote. The argument, the selection of sources and the claims are the author’s. No text was accepted without verification against the cited source.

Online Resource 1 — Evidence hygiene

Tracing the sources of this paper, I found a pattern that repeats. Several of the most cited statistics in executive material on AI and engineering [2] do not survive checking. Not through bad faith, but through chain citation, in which each link cites the previous one and nobody returns to the source. The table below lists the claims verified directly. “Not located” does not mean false. It means the claim should not be used as a basis for decision unless whoever presents it supplies the reference.

Table 5. Claims in circulation and what the primary source supports.

Claim in circulationWhat the primary source supportsVerdict
2.74× more vulnerabilities across 62,000 repositoriesThe 62,000-repository scope belongs to a study reporting about ten times more security issues. The 2.74× factor is from another, over 470 PRs. Two studies fused into one citation [21], [63]Partial
73% of tasks with five or more steps hallucinate, attributed to Stanford HAII located no publication from the institute with that claim. Its material on hallucination concerns legal models, at a rate of one in six queriesNot located
17% less comprehension, n≈1,500, attributed to “Bedard et al.” via HBRThe finding belongs to another study, with n=52: 50% against 67% correct, that is, 17 “percentage points”. The speed difference was not significant [28]Partial
Nine months compressed into 76 days with five times fewer peopleThe project cited exists, and the located datum is roughly ten times in code throughput, with attribution and date different from those in circulation [81]Not located
13% of agents have strong visibilityIt is 13% of “companies” reporting strong visibility into how AI touches their “data”. Subject and object swapped [82]Partial
1,587% growth in agent skillsThe reference report lists “AI Engineer” as the fastest-growing role, and I did not locate that percentage in it [46]Not located
5.8% unemployment among recent graduates, “worst since 2013”The value corresponds to Q1 2025 and is described as the highest level since 2021 [83]Partial
1.3 million new AI posts, attributed to a WEF reportIt is an article signed by a LinkedIn executive and published by the WEF, based on LinkedIn data [44]Partial

Online Resource 2 — Four governance questions and their mechanisms

The table below expands the governance framing discussed in Section 11: the four questions an executable envelope has to answer, and the mechanism each one demands.

Table 6. Four governance questions and their mechanisms.

QuestionMechanism
Who is this agent?Workload identity bound to a human supervisor. The agent never exceeds the permissions of the human who answers for it
What is it allowed to do?Deterministic policy evaluated before tool access, outside the inference loop
Is it working correctly?Continuous evaluation in production and complete reasoning-chain traces
Can we prove it?A record of every “enforcement” decision, with cost, latency and error dashboards

References

[1] J. Bek, “Services: The new software.” Sequoia Capital, Mar. 2026. Available: https://sequoiacap.com/article/services-the-new-software

[2] Industry executive material, “Executive presentations and reports on artificial intelligence and software engineering.” 2026.

[3] J. Allen, “A 12 step program to get from zero to hundreds of AWS-certified engineers.” AWS Cloud Enterprise Strategy Blog, Jul. 2017. Available: https://aws.amazon.com/blogs/enterprise-strategy/a-12-step-program-to-get-from-zero-to/

[4] N. J. Slamecka and P. Graf, “The generation effect: Delineation of a phenomenon,” Journal of Experimental Psychology: Human Learning and Memory, vol. 4, no. 6, pp. 592–604, 1978, doi: 10.1037/0278-7393.4.6.592.

[5] J. Sweller, “Cognitive load during problem solving: Effects on learning,” Cognitive Science, vol. 12, no. 2, pp. 257–285, 1988, doi: 10.1207/s15516709cog1202_4.

[6] R. A. Bjork, “Memory and metamemory considerations in the training of human beings,” in Metacognition: Knowing about knowing, J. Metcalfe and A. P. Shimamura, Eds., Cambridge, MA: The MIT Press, 1994, pp. 185–206. doi: 10.7551/mitpress/4561.003.0011.

[7] R. A. Bjork and E. L. Bjork, “Desirable difficulties in theory and practice,” Journal of Applied Research in Memory and Cognition, vol. 9, no. 4, pp. 475–479, 2020, doi: 10.1016/j.jarmac.2020.09.003.

[8] J. Tai, R. Ajjawi, D. Boud, P. Dawson, and E. Panadero, “Developing evaluative judgement: Enabling students to make decisions about the quality of work,” Higher Education, vol. 76, no. 3, pp. 467–481, 2018, doi: 10.1007/s10734-017-0220-3.

[9] M. Bearman and R. Ajjawi, “Learning to work with the black box: Pedagogy for a world with artificial intelligence,” British Journal of Educational Technology, vol. 54, no. 5, pp. 1160–1173, 2023, doi: 10.1111/bjet.13337.

[10] L. Bainbridge, “Ironies of automation,” Automatica, vol. 19, no. 6, pp. 775–779, 1983, doi: 10.1016/0005-1098(83)90046-890046-8).

[11] R. Parasuraman and D. H. Manzey, “Complacency and bias in human use of automation: An attentional integration,” Human Factors, vol. 52, no. 3, pp. 381–410, 2010, doi: 10.1177/0018720810376055.

[12] N. Carr, The glass cage: Automation and us. New York: W. W. Norton, 2014.

[13] H. Braverman, Labor and monopoly capital: The degradation of work in the twentieth century. New York: Monthly Review Press, 1974.

[14] P. Attewell, “The deskilling controversy,” Work and Occupations, vol. 14, no. 3, pp. 323–346, 1987.

[15] S. Zuboff, In the age of the smart machine: The future of work and power. New York: Basic Books, 1988.

[16] H. Collins, Tacit and explicit knowledge. Chicago: University of Chicago Press, 2010.

[17] H. Collins and R. Evans, Rethinking expertise. Chicago: University of Chicago Press, 2007.

[18] E. W. Dijkstra, “How do we tell truths that might hurt?” E. W. Dijkstra Archive, University of Texas at Austin, 1975. Available: https://www.cs.utexas.edu/~EWD/transcriptions/EWD04xx/EWD498.html

[19] J. Backus, “The history of FORTRAN I, II, and III,” ACM SIGPLAN Notices, vol. 13, no. 8, pp. 165–180, 1978, doi: 10.1145/960118.808380.

[20] N. L. Ensmenger, The computer boys take over: Computers, programmers, and the politics of technical expertise. Cambridge, MA: The MIT Press, 2010. doi: 10.7551/mitpress/9780262050937.001.0001.

[21] I. Nussbaum, “4x velocity, 10x vulnerabilities: AI coding assistants are shipping more risks.” Apiiro Blog, Sep. 2025. Available: https://apiiro.com/blog/4x-velocity-10x-vulnerabilities-ai-coding-assistants-are-shipping-more-risks/

[22] D. M. Wegner, “Transactive memory: A contemporary analysis of the group mind,” in Theories of group behavior, New York: Springer, 1987, pp. 185–208. doi: 10.1007/978-1-4612-4634-3_9.

[23] E. Hutchins, Cognition in the wild. Cambridge, MA: The MIT Press, 1995. doi: 10.7551/mitpress/1881.001.0001.

[24] J. Hollan, E. Hutchins, and D. Kirsh, “Distributed cognition: Toward a new foundation for human-computer interaction research,” ACM Transactions on Computer-Human Interaction, vol. 7, no. 2, pp. 174–196, 2000, doi: 10.1145/353485.353487.

[25] H. Bastani, O. Bastani, A. Sungu, H. Ge, Ö. Kabakçı, and R. Mariman, “Generative AI without guardrails can harm learning: Evidence from high school mathematics,” Proceedings of the National Academy of Sciences, vol. 122, no. 26, p. e2422633122, 2025, doi: 10.1073/pnas.2422633122.

[26] Y. Fan et al., “Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance,” British Journal of Educational Technology, vol. 56, no. 2, pp. 489–530, 2025, doi: 10.1111/bjet.13544.

[27] J. Prather et al., “The widening gap: The benefits and harms of generative AI for novice programmers,” in Proc. 2024 ACM conf. On international computing education research (ICER), 2024, pp. 469–486. doi: 10.1145/3632620.3671116.

[28] J. H. Shen and A. Tamkin, “How AI assistance impacts the formation of coding skills,” Anthropic, Jan. 2026. Available: https://www.anthropic.com/research/AI-assistance-coding-skills

[29] H.-P. Lee et al., “The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers,” in Proc. 2025 CHI conf. On human factors in computing systems, 2025. doi: 10.1145/3706598.3713778.

[30] N. Kosmyna et al., “Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task.” arXiv:2506.08872, 2025. doi: 10.48550/arXiv.2506.08872.

[31] M. Stankovic, E. Hirche, S. Kollatzsch, and J. N. Doetsch, “Comment on: Your brain on ChatGPT.” arXiv:2601.00856, 2025. doi: 10.48550/arXiv.2601.00856.

[32] K. Z. Cui, M. Demirer, S. Jaffe, L. Musolff, S. Peng, and T. Salz, “The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers,” Management Science, 2026, doi: 10.1287/mnsc.2025.00535.

[33] E. Brynjolfsson, D. Li, and L. Raymond, “Generative AI at work,” The Quarterly Journal of Economics, vol. 140, no. 2, pp. 889–942, 2025, doi: 10.1093/qje/qjae044.

[34] F. Dell’Acqua et al., “Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality,” Harvard Business School, Working Paper 24-013, 2023. doi: 10.2139/ssrn.4573321.

[35] S. Noy and W. Zhang, “Experimental evidence on the productivity effects of generative artificial intelligence,” Science, vol. 381, no. 6654, pp. 187–192, 2023, doi: 10.1126/science.adh2586.

[36] S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer, “The impact of AI on developer productivity: Evidence from GitHub Copilot.” arXiv:2302.06590, 2023. doi: 10.48550/arXiv.2302.06590.

[37] S. Maier, M. Gunzenhäuser, J. Schweisthal, M. Schneider, and S. Feuerriegel, “A meta-analysis of the effect of generative AI on productivity and learning in programming.” arXiv:2605.04779, 2026.

[38] M. Alanazi, B. Soh, H. Samra, and A. Li, “The influence of artificial intelligence tools on learning outcomes in computer programming: A systematic review and meta-analysis,” Computers, vol. 14, no. 5, p. 185, 2025, doi: 10.3390/computers14050185.

[39] S. Kalyuga, P. Ayres, P. Chandler, and J. Sweller, “The expertise reversal effect,” Educational Psychologist, vol. 38, no. 1, pp. 23–31, 2003, doi: 10.1207/s15326985ep3801_4.

[40] Ravio, “Early career hiring is down 73%: Why it is time to redesign entry-level roles for the AI era.” Ravio Blog, Jul. 2025. Available: https://ravio.com/blog/early-career-hiring

[41] A. Bantock, “The SignalFire state of tech talent report 2025,” SignalFire, May 2025. Available: https://www.signalfire.com/blog/signalfire-state-of-talent-report-2025

[42] S. M. Hosseini Maasoum and G. Lichtinger, “Generative AI as seniority-biased technological change: Evidence from U.S. Résumé and job posting data,” Harvard University, SSRN Working Paper 5425555, 2025. Available: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5425555

[43] E. Brynjolfsson, B. Chandar, and R. Chen, “Canaries in the coal mine? Six facts about the recent employment effects of artificial intelligence,” Stanford Digital Economy Lab, Aug. 2025. Available: https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine

[44] D. Shapero, “AI has already added 1.3 million jobs, LinkedIn data says.” World Economic Forum, Jan. 2026. Available: https://www.weforum.org/stories/jobs-and-the-future-of-work/ai-has-already-added-1-3-million-new-jobs-according-to-linkedin-data/

[45] Implicator.ai, “Software engineering openings jump 30% even as 52,000 tech workers lose jobs.” Reportagem sobre dados do TrueUp (9.000 empresas), 1º trimestre de 2026, Apr. 2026. Available: https://www.implicator.ai/software-engineering-openings-jump-30-even-as-52-000-tech-workers-lose-jobs/

[46] LinkedIn News, “LinkedIn jobs on the rise 2026: The 25 fastest-growing roles in the U.S.” Jan. 2026. Available: https://www.linkedin.com/pulse/linkedin-jobs-rise-2026-25-fastest-growing-roles-us-linkedin-news-dlb1c

[47] P. Tambe and L. M. Hitt, “Now IT’s personal: Offshoring and the shifting skill composition of the U.S. Information technology workforce,” Management Science, vol. 58, no. 4, pp. 678–695, 2012, doi: 10.1287/mnsc.1110.1445.

[48] M. Y. Vardi, “Globalization and offshoring of software revisited,” Communications of the ACM, vol. 53, no. 5, p. 5, 2010, doi: 10.1145/1735223.1735224.

[49] J. Bessen, Learning by doing: The real connection between innovation, wages, and wealth. New Haven: Yale University Press, 2015. doi: 10.12987/9780300213645.

[50] S. Sharwood, “AWS CEO says using AI to replace junior staff is the ‘dumbest thing i’ve ever heard’.” The Register, Aug. 2025. Available: https://www.theregister.com/2025/08/21/aws_ceo_entry_level_jobs_opinion/

[51] J. Lave and E. Wenger, Situated learning: Legitimate peripheral participation. Cambridge: Cambridge University Press, 1991. doi: 10.1017/CBO9780511815355.

[52] J. S. Brown and P. Duguid, “Organizational learning and communities-of-practice: Toward a unified view of working, learning, and innovation,” Organization Science, vol. 2, no. 1, pp. 40–57, 1991, doi: 10.1287/orsc.2.1.40.

[53] M. Polanyi, The tacit dimension. Chicago: University of Chicago Press, 2009.

[54] S. E. Dreyfus, “The five-stage model of adult skill acquisition,” Bulletin of Science, Technology and Society, vol. 24, no. 3, pp. 177–181, 2004, doi: 10.1177/0270467604264992.

[55] J. Allen, “12 steps to get started with the cloud.” AWS Cloud Enterprise Strategy Blog, Jan. 2018. Available: https://aws.amazon.com/blogs/enterprise-strategy/12-steps-to-get-started-with-the-cloud/

[56] J. Allen, “The future waits for nobody: My Capital One journey to the AWS cloud.” AWS Cloud Enterprise Strategy Blog, Jun. 2017. Available: https://aws.amazon.com/blogs/enterprise-strategy/the-future-waits-for-nobody-my-capital-one-journey-to-the-aws-cloud/

[57] AI Index Steering Committee, “The 2026 AI index report, chapter 4: economy,” Stanford Institute for Human-Centered Artificial Intelligence, Apr. 2026. Available: https://hai.stanford.edu/assets/files/ai_index_report_2026_chapter_4_economy.pdf

[58] DORA and Google Cloud, “2025 DORA report: State of AI-assisted software development,” Google Cloud, Sep. 2025. Available: https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report

[59] Stack Overflow, “2025 stack overflow developer survey: AI,” Stack Overflow, 2025. Available: https://survey.stackoverflow.co/2025/ai

[60] E. May, “Mind the gap: Closing the AI trust gap for developers.” The Stack Overflow Blog, Feb. 2026. Available: https://stackoverflow.blog/2026/02/18/closing-the-developer-ai-trust-gap/

[61] A. Challapally, C. Pease, R. Raskar, and P. Chari, “The GenAI divide: State of AI in business 2025,” MIT NANDA, Jul. 2025.

[62] J. Becker, N. Rush, B. Barnes, and D. Rein, “Measuring the impact of early-2025 AI on experienced open-source developer productivity,” METR — Model Evaluation; Threat Research, Jul. 2025. Available: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

[63] CodeRabbit, “State of the AI vs. Human code generation report,” CodeRabbit, Dec. 2025. Available: https://www.coderabbit.ai/whitepapers/state-of-AI-vs-human-code-generation-report

[64] A. Vella and K. Blincoe, “The impact of AI coding assistants on software engineering: A longitudinal study.” arXiv:2605.23135, 2026.

[65] Anthropic, “Meet the winners of our built with Opus 4.6 Claude Code hackathon.” Anthropic (blog), Feb. 2026. Available: https://claude.com/blog/meet-the-winners-of-our-built-with-opus-4-6-claude-code-hackathon

[66] U. Joshi, G. Venkatraman, and M. Fowler, “Expert generalist.” martinfowler.com, Jul. 2025. Available: https://martinfowler.com/articles/expert-generalist.html

[67] J. Appelo, “Why ravens may rule the future of work.” jurgenappelo.com, Sep. 2025. Available: https://jurgenappelo.com/blogs/news/why-ravens-may-rule-the-future-of-work

[68] PwC, “2026 AI business predictions.” Dec. 2025. Available: https://www.pwc.com/us/en/tech-effect/ai-analytics/ai-predictions.html

[69] Amazon Staff, “Five tech predictions for 2026 and beyond, according to Amazon CTO dr. Werner Vogels.” About Amazon, Nov. 2025. Available: https://www.aboutamazon.com/news/aws/werner-vogels-amazon-cto-predictions-2026

[70] S. Sanfilippo, “Control the ideas, not the code.” antirez.com, Jul. 2026. Available: https://antirez.com/news/169

[71] A. Ng, “The product management bottleneck.” The Batch, DeepLearning.AI, issue 349, Apr. 2026. Available: https://www.deeplearning.ai/the-batch/issue-349

[72] J. Allen, “Untangling your organisational hairball: Well-controlled.” AWS Cloud Enterprise Strategy Blog, May 2024. Available: https://aws.amazon.com/blogs/enterprise-strategy/untangling-your-organisational-hairball-well-controlled/

[73] M. E. Conway, “How do committees invent?” Datamation, vol. 14, no. 5, pp. 28–31, 1968.

[74] M. Skelton and M. Pais, Team topologies: Organizing business and technology teams for fast flow. Portland, OR: IT Revolution Press, 2019.

[75] D. Shipper, “The two-slice team.” Every, Chain of Thought, Feb. 2026. Available: https://every.to/chain-of-thought/the-two-slice-team

[76] L. Galante, “Ten platform engineering predictions for 2026.” PlatformEngineering.org, Dec. 2025. Available: https://platformengineering.org/blog/10-platform-engineering-predictions-for-2026

[77] Infocomm Media Development Authority, “Model AI governance framework for agentic AI, version 1.0,” IMDA, Singapore, Jan. 2026. Available: https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/press-releases/2026/new-model-ai-governance-framework-for-agentic-ai

[78] European Parliament and Council of the European Union, “Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (artificial intelligence act).” Official Journal of the European Union, Jun. 2024. Available: https://eur-lex.europa.eu/eli/reg/2024/1689/oj

[79] J. W. Cutler et al., “Cedar: A new language for expressive, fast, safe, and analyzable authorization,” Proceedings of the ACM on Programming Languages, vol. 8, no. OOPSLA1, 2024, doi: 10.1145/3649835.

[80] Amazon Web Services, “Amazon Bedrock AgentCore now includes policy (preview), evaluations (preview) and more.” AWS What’s New, Dec. 2025. Available: https://aws.amazon.com/about-aws/whats-new/2025/12/amazon-bedrock-agentcore-policy-evaluations-preview

[81] J. Magerramov, “Amazon Bedrock Mantle and developing at the speed of AI.” AWS Podcast, episode 753, Jan. 2026.

[82] Cyera Research Labs, “2025 state of AI data security report,” Cyera, Sep. 2025. Available: https://www.cyera.com/research-labs/2025-state-of-ai-data-security-report

[83] Federal Reserve Bank of New York, “The labor market for recent college graduates.” Quarterly statistical series, 2025. Available: https://www.newyorkfed.org/research/college-labor-market

Comments

Every comment is moderated before it appears here. Nothing is published automatically.

Loading…