The Case Against "Smarter" AI

The Case Against "Smarter" AI

By Momar Lissa Ndiaye ("MLN"), Founder & CEO, weyoga Inc.

This essay is addressed to the research community, and it should open by disarming the title honestly: there is no case against smarter AI in the capability sense — the science is magnificent, the gains are real, and nothing here argues for slowing any of it. The case is against "smarter" as the field's definition of better — against a monoculture of optimization in which one axis of progress has absorbed the field's talent, capital, evaluation apparatus, and imagination so completely that a second axis, arguably worth more to more humans, does not merely lag. It does not exist as a discipline at all.

State the structural observation first, because everything else hangs on it. A field is defined by what it optimizes, and what it optimizes is defined by what it measures. Walk the measurement stack of contemporary AI — the benchmark suites, the leaderboards, the eval frameworks, the capability reports — and notice the invariant: essentially every instrument in the field's cabinet measures performance within a session, for an anonymous user. Reasoning depth, factual accuracy, task completion, harmlessness — all session-shaped, all subject-free. Now name what has no instrument anywhere in the stack: whether a system, deployed with one particular human for a year, made that human's decisions better. Longitudinal, per-subject benefit — the thing every previous essay in this series argued is the actual product of AI in a life — has no benchmark, no leaderboard, no eval, no workshop, no track at the conference. The field is not failing at the measurement. The field has not proposed it. And a discipline's silence about a measurement is a decision about what counts as progress, whether or not anyone made it on purpose. Let the target of this essay be exact, then, before anything else is argued: the case is not against researchers, whose incentives are set by instruments they inherited; it is against the missing instrument. The field has excellent capability benchmarks and almost no longitudinal human-impact benchmarks. Everything below is a footnote to that sentence.

The predictable reply is that this is application work — the models are general; longitudinal benefit is a product question for industry, not a research question for the field. This reply deserves to be taken apart carefully, because it is the load-bearing wall of the monoculture. Consider what is actually unsolved in building the recognition layer this series has specified: how to represent a person's behavioral sequence such that recurrence is detectable across surface variation (a representation-learning problem, and a hard one); how to calibrate confidence in a detected pattern against the ethical cost of a false positive delivered to its subject (a calibration problem with stakes the field has never priced); how to model reflexivity — Essay 7's core phenomenon, a forecast that alters its target upon delivery — which breaks the i.i.d. assumptions underneath most of the field's theory; how to evaluate any of it, when the treatment effect of recognition unfolds over years and the subject's seeing the measurement changes the measured. These are not product questions. They are open research problems of the first rank — representation, calibration, non-stationarity, evaluation under reflexivity — and they are unclaimed not because they are easy or applied, but because the field's prestige gradient, as Essay 11 observed of industry, points elsewhere. "That's just an application" is how a monoculture describes every axis it has declined to build instruments for.

Now the strongest counterargument, because a case addressed to researchers must pass their review. Generality is the bet, and the bet is winning: the scaling era's central lesson — the bitter one — is that general capability, scaled, eventually subsumes every specialized architecture; build the smart-enough model and the personal layer falls out as a downstream deployment detail. Three answers, in ascending order. First, the empirical record within the thesis's own terms: capability has scaled through multiple stunning generations while the longitudinal-benefit axis has moved approximately zero — if the personal layer were downstream of capability, some gradient should be visible by now, and Essay 17's universally felt "thin current" is the null result, reported by hundreds of millions of subjects. Second, the mechanism, for the last time in this series: sequence-level information is absent, not latent — no capability recovers information that was never observed, so the gap is not downstream of intelligence; it is orthogonal to it, which is precisely why scaling never touches it. Third, and aimed at the lesson itself: the bitter lesson counsels against hand-engineering what computation can learn — it says nothing about what data architecture permits observing. Continuity is not a clever feature to be subsumed by scale. It is the precondition for the observations scale would learn from. Misreading an observability constraint as an architecture choice is how a correct lesson becomes a category error.

There is also a human argument, and a research audience has earned the candid version. The monoculture is not only an allocation of measurement; it is an allocation of careers — a generation of the most capable technical minds alive, sorted by the prestige gradient into making the answer machine marginally better at answering, while the questions of representation, consent architecture, reflexive evaluation, and longitudinal benefit — questions whose solutions would compound in millions of lives — sit essentially unstaffed, waiting for the field to decide they are respectable. Fields have made this decision before. Medicine spent centuries prestige-ranked around heroic intervention before epidemiology — unglamorous, longitudinal, population-scale — became a science, and the sanitation engineers and cohort-followers saved more lives than the surgeons. The researchers who build the instruments for the second axis of AI will be that generation's epidemiologists: less cited at first, unfashionable at first, and correct.

So the case against "smarter" AI reduces to a proposal any researcher can act on, and it is deliberately modest: build one instrument. Propose the benchmark for pattern-recurrence detection in longitudinal behavioral sequences. Publish the evaluation framework for intervention timing under reflexivity. Define the consent-architecture standard the way differential privacy was defined — as math, not policy. The monoculture will not be argued out of existence, and it should not be: the capability axis deserves its brilliance. It will be ended the way monocultures always end — by the first credible measurement on a second axis, after which the field discovers it had been a monoculture all along.

The field's smartest systems still cannot tell a person the one sentence that would change their year. That is not a capability gap, and no amount of "smarter" closes it. It is a discipline that does not exist yet — waiting, as every discipline does, for its first instrument.

The next essay leaves the research hall for the operator's chair: what the people who build companies — and fund them, and burn out running them — systematically get wrong about the patterns this series describes.


Part of The Recognition Layer

Momar Lissa Ndiaye ("MLN") is the Founder & CEO of weyoga Inc., a Delaware company. — weyoga.ai · mln@weyoga.ai