There is no agreed scientific definition of artificial general intelligence (AGI), and therefore no scientifically fixed date for reaching it. At least six definitions are in use at the same time. They are not versions of one idea. Applied to the same AI systems of September 2026, they give opposite answers: “already achieved”, “about halfway”, and “not yet at the second of five levels”. The missing definition is not a gap that more research will close. It is a feature of the debate that every side can use, and several visibly do.

Summary of findings

  1. The definitions disagree by design. OpenAI’s charter defines AGI by economic value. Hendrycks and colleagues define it as the cognitive profile of a well-educated adult, and score GPT-4 at 27% and GPT-5 at 57% on that scale. Google DeepMind uses a six-level scale on which today’s leading models are still at level 1. Chollet defines it by how efficiently a system learns new skills. A February 2026 essay in Nature argues that current systems already qualify. The OpenAI–Microsoft contract reportedly defined it as a profit figure. Each sets a different condition for “achieved”.
  2. Concrete tests exist, but the measurements are fragile. On 3 September 2026, GPT-6 Astra scored 62.7% on the ARC-AGI-3 test under the test authors’ standard setup, and 99.9% under a setup supplied by OpenAI. The 37-point gap came from the software around the model, not from the model itself. The higher number spread. The test’s authors published both and said they are “not claiming that it is AGI”.
  3. The capability gains are real all the same. On the ARC-AGI-2 test, the best published result rose from about 54% at roughly $31 per task in late 2025 to 95% at about $1.12 per task by September 2026. The organisation that measured this publicly rejects the AGI conclusion.
  4. Almost no one says AGI is impossible. The best-known sceptics — LeCun, Sutton, Chollet, Marcus — dispute the method, the timing, or the usefulness of the word. Three of them have founded companies to build AGI another way.
  5. The commercial use of the word is documented. The OpenAI–Microsoft “AGI clause” was redefined, then placed under an expert panel, then made commercially irrelevant on 27 April 2026, when Microsoft’s revenue share was separated from OpenAI’s technical progress. Planned 2026 spending on AI infrastructure by the five largest cloud companies is roughly $660–800 billion, while the measured effect on the economy so far is small.

Basis, limits and what follows

This article rests on public sources only: test results read directly on the publishing organisations’ websites, original essays and papers, and dated statements. Eleven sources could not be accessed and are excluded rather than paraphrased.

Confidence is high for the definitional finding and the directly checked test scores, moderate for quotations relayed by news outlets, and low for spending figures. No independent peer review has taken place; the review that did occur is stated at the end of the article. The practical consequence: “Has AGI been achieved?” cannot be answered as asked. A claim becomes assessable only when it names the definition, the test conditions, the evaluator and the cost.

Key terms used in this article

Term Meaning here
AGI Artificial general intelligence: an AI system with broadly human-level ability across many kinds of task, rather than one specialised task. What exactly this means is the subject of this article.
Benchmark A standardised test: a fixed set of tasks and a scoring rule, used to compare AI systems.
Harness The software around a model during a test: how tasks are presented, what the model may keep in memory between steps, how many attempts it gets. The same model can score very differently under different harnesses.
Frontier model One of the most capable AI systems available at a given time.
Computing power (“compute”) The processing resources used to train or run a model. Allowing more computing power during a test usually raises the score and the cost.
Time horizon (METR) The length of a task, measured in the hours a skilled person would need, that an AI completes successfully half the time.
Percentile A rank within a group. “The 50th percentile of skilled adults” means “as good as or better than half of them”.
Doubling time How long a measured quantity takes to double.
Total factor productivity (TFP) An economist’s measure of how much an economy produces from a given amount of labour and capital.
Lean Software that checks a mathematical proof step by step, so that a proof can be verified by machine as well as by human experts.

1. The answer, and why it is a reframing

The question “what exactly is AGI” has no single correct answer today. The reason is not that the science is young. Each definition in use is coherent and can be measured. They simply define different things, and the organisations using them have practical reasons to prefer one over another.

This has a direct consequence. Two well-informed people can look at GPT-6 Astra in September 2026 and conclude, using published criteria, that AGI has arrived — and that it has not. Both can be right, because they are using different definitions. An argument about the word will therefore never end. Arguments about specific abilities, measured under stated conditions at a stated cost, can end, and several have.

Definition Author and date What has to be true Verdict on the systems of September 2026
Economic OpenAI charter, April 2018 Outperforms humans at most economically valuable work Not tested — no such trial has been run
Psychometric Hendrycks et al., October 2025 Matches a well-educated adult across ten mental abilities 57% of the way (GPT-5)
Levels Google DeepMind, 2023/2025 Level 2: at least as good as the median skilled adult across many non-physical tasks Not reached — still level 1, “Emerging”
Learning efficiency Chollet, 2019 Learns new skills efficiently, not by brute computing power Partly — and only with unlimited computing power
Behavioural Chen et al., Nature, February 2026 Broad, flexible competence; perfection not required Yes — already achieved
Contractual OpenAI–Microsoft, 2018–2026 Reportedly $100 billion in profits; later, an expert panel’s verdict No longer relevant — separated from the money in April 2026

The verdicts are this article’s application of each framework as its authors describe it. They are not statements by those authors about GPT-6 Astra.

The recommendation is therefore not to settle the term but to stop using it in analysis. Section 6 gives four questions to ask instead.

2. Key findings

Finding Evidence Confidence Main limit
No definition of AGI is accepted across the field; six are in use, each with a different condition for “achieved” Primary definitional texts, read directly: A Definition of AGI, Levels of AGI, On the Measure of Intelligence High The six are the prominent ones; others exist
Test results depend on the harness: 62.7% versus 99.9% for the same model on the same test ARC Prize leaderboard and analysis, read directly High Shown for ARC-AGI-3; assumed to apply elsewhere
Capability gains are large and confirmed by organisations that reject the AGI framing ARC-AGI-2: about 54% at about $31 per task (late 2025) to 95% at about $1.12 per task (September 2026) Moderate Different systems; possibly different versions of the test
Under strict computing limits, results fall sharply: 24.03% in the 2025 competition against an 85% prize threshold ARC Prize 2025 results High One competition year
The length of task an AI can complete alone doubles about every 3 to 4.5 months for recent models — on a task set its measurer calls nearly used up METR Time Horizon 1.1 High Software, machine-learning and cybersecurity tasks only
The ability to keep learning after deployment, and to store new long-term memories, is the most consistently named missing piece A Definition of AGI; AGI’s Last Bottlenecks Moderate The measured figure comes from one research group
Almost no prominent researcher says AGI is impossible in principle Positions of LeCun, Marcus, Chollet and Sutton, reviewed individually Moderate Covers the named people, not the whole field
The word has a documented commercial and contractual role alongside its scientific one Contract chronology; September 2026 launch reporting Moderate The interests are documented; intentions are not

3. Scope and method

The question has three parts — what the word means, what the evidence shows, and who benefits from which answer — so three kinds of evidence were used: the original definitional texts; test results read on the publishing organisations’ own websites; and dated statements by named people, read alongside the commercial arrangements those people are party to.

All material was retrieved on 9 September 2026. Where a website displayed its results through interactive scripts, the page was opened in a browser and read directly rather than summarised from search results. The ARC Prize figures used throughout were obtained this way. Search-engine summaries were used only to find material. One early summary reported ARC-AGI-2 scores that the leaderboard itself did not show; it was discarded, and the episode is recorded in the publisher’s case log.

A deliberate search was made for opposing evidence: sceptical positions, economic data on actual impact, and criticism of the term itself. Eleven sources could not be opened — the OpenAI charter page, three Nature items, and coverage by Axios and Bloomberg among them. They are listed as inaccessible and their content is not reproduced.

Outside the scope: existential risk, AI safety research, regulation, consciousness, robotics hardware, and any investment judgement.

4. Findings and reasoning

4.1 Six definitions that do not agree

The term was first used by Mark Gubrud in 1997 and reintroduced independently around 2002 by Ben Goertzel, Shane Legg and Peter Voss as the title of a book, to distinguish their ambition from the narrow, single-task AI of the time. It entered the field as the name of a research programme, not as a measurable threshold, and it has never acquired one agreed measurement. Six definitions are in active use.

(a) Economic. OpenAI’s charter (April 2018) defines AGI as “highly autonomous systems that outperform humans at most economically valuable work”. Forecasters use a similar wording: most desk work can be done by AI better, faster and cheaper than by people. This is the only definition whose fulfilment would show up in national economic statistics.

(b) Psychometric. Hendrycks and colleagues (October 2025) define AGI as matching “the cognitive versatility and proficiency of a well-educated adult”. They measure this with tests adapted from human psychology, across ten areas such as reasoning, memory and perception. The result is one percentage: GPT-4 at 27%, GPT-5 at 57%. The more important result is the shape behind the number. Performance is very uneven — strong in knowledge-heavy areas, close to zero in storing new long-term memories.

(c) Levels. Google DeepMind’s Levels of AGI (2023, revised 2025) does not set a single threshold. It grades systems on two things: how well they perform, and how broad their abilities are. Level 2, “Competent AGI”, requires performance at least as good as the median skilled adult across a wide range of non-physical tasks. In the latest revision, no general system has reached level 2. Today’s leading language models are placed at level 1, “Emerging AGI”. The paper’s six design principles are useful in themselves: judge abilities, not internal mechanisms; require both breadth and quality; count mental tasks, not physical ones; count what a system can do, not what it has been deployed to do; use realistic tasks; and describe a path with stages rather than a single finish line.

(d) Learning efficiency. Chollet (2019) defines intelligence as how efficiently a system acquires new skills. Measuring skill alone is not enough, he argues, because with unlimited training data or unlimited computing power any level of skill can be “bought” without the system becoming any more intelligent. The ARC family of tests is built on this idea. By this definition, a system that solves new problems only by using enormous computing power has not demonstrated the property in question.

(e) Behavioural. Four researchers at the University of California, San Diego — Chen, Belkin, Bergen and Danks — argued in Nature (vol. 650, pp. 36–40, February 2026) that AGI has always meant flexible, broad competence across mathematics, language, science, practical reasoning and creative work, without a requirement of perfection. By that standard, they say, today’s large language models already qualify. A reply in Nature by Quattrociocchi, Capraro and Marcus argued that this confuses very good statistical pattern-matching with intelligence, and that doing well on individual tasks is not evidence of general ability.

(f) Contractual. In the OpenAI–Microsoft agreement, “AGI” was a trigger for changing commercial rights. It began as the charter wording. According to reporting, a 2023 agreement redefined it as the point at which OpenAI’s systems could generate $100 billion in profits. From October 2025, any declaration of AGI had to be confirmed by an independent expert panel. On 27 April 2026 the clause lost its purpose: Microsoft’s revenue share was made independent of OpenAI’s technical progress through 2030.

Two further positions matter. Anthropic avoids the term. Its chief executive, Dario Amodei, writes about “powerful AI” instead — a list of properties, summed up as “a country of geniuses in a datacenter” — explicitly to avoid the science-fiction associations of “AGI”. Sam Altman of OpenAI has said both that AGI is “a very sloppy term” and that it is “something we’ve already reached without noticing”, and has acknowledged that the industry “made the mistake of not properly defining AGI”.

The essential point is not that these definitions emphasise different things. It is that each sets a different condition for when AGI counts as achieved. Under (e), AGI exists now. Under (c), it does not. Under (b), progress stands at 57%. Under (a), the test has never been run. Under (f), the question was settled by amending a contract, not by evidence.

4.2 Concrete criteria exist, but the measurements are unstable

Explicit thresholds exist today. Examples: 85% on the private ARC-AGI-2 test under strict limits on computing power (the ARC Prize Grand Prize); 100% on the Hendrycks index; the DeepMind levels at the 50th, 90th, 99th and 100th percentiles of skilled adults; METR’s time horizons at 50% and 80% reliability; and GDPval’s rate of matching or beating human professionals on real work products. Older informal tests — the Turing test, the “coffee test” (make a coffee in an unfamiliar kitchen), the “college student test”, the “employment test” — are still quoted, but no serious evaluator uses them as a decision rule.

Four problems recur.

Tests get used up. Once most systems score near the top, a test stops telling them apart. MMLU, a widely used knowledge test, is effectively exhausted above 88–90%. ARC-AGI-1 now stands at 98.5%. METR states that its own set of tasks is nearly used up, and that measurements above roughly 16 hours are unreliable — a caveat attached to its most-quoted result.

Results depend on the harness. The clearest example is days old. On ARC-AGI-3, GPT-6 Astra scored 62.7% under the standard harness written by the test’s authors, and 99.9% under a harness supplied by OpenAI. The OpenAI harness lets the model keep its intermediate reasoning between steps and compresses long sessions. Nothing about the model itself changed. The 99.9% figure is the one that circulated at launch.

Two horizontal bars on a scale from 0 to 100 percent. GPT-6 Astra on the ARC-AGI-3 test, September 2026: 62.7 percent under the ARC Prize standard harness and 99.9 percent under a harness supplied by OpenAI. Only the software around the model differed.

ARC Prize now publishes both, noting that only the standard-harness result allows a fair comparison between companies. It also states plainly what it does not conclude: “saturating the benchmark would not represent ‘proof of achieving AGI’. Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.” ARC Prize, analysis of GPT-6 Astra

Evaluators disagree. In the same week, two independent evaluation organisations reached different conclusions about the same model. One ranked GPT-6 Astra clearly first; the other placed it level with its predecessor. This is reported second-hand, and the exact scores are rated low-confidence. The disagreement itself is informative about the state of measurement.

It is unclear that the tests measure one thing. Two recent analyses question whether “general intelligence” scores measure a single ability. One recalculates the Hendrycks index so that a very uneven profile is penalised; GPT-5 falls from 58% to 24%. Another, by Krakauer (April 2026), examines 39 models on 14 tests. Among early models, a system that did well on one test did well on all of them. By 2024 that link had weakened markedly, as models specialised in reasoning appeared. This is evidence against treating a single AI “general intelligence” score as established. Krakauer, The Rise and Fall of G in AGI

One further gap receives little attention. Leading systems reach 90–95% on ARC-AGI-2 when allowed unlimited computing power. The best entry in the 2025 competition, which imposes strict limits, reached 24.03%. Learning efficiently — the property in Chollet’s definition — lags far behind learning with unlimited resources. ARC Prize 2025 results and analysis

4.3 Has science fixed a date? No

The forecasting literature says so about itself. A RAND review (March 2026) offers a framework “for interpreting forecasts under conditions of deep uncertainty” and documents gaps in method rather than endorsing any timeline. Four kinds of evidence exist, and none resolves the question. RAND, AGI forecasting and scenario analysis

Surveys of researchers. A 2023 survey of 2,778 published AI researchers gave a 50% chance of “high-level machine intelligence” by 2047 — thirteen years earlier than the same survey series had said in 2022 — and a 50% chance that all human jobs could be automated by 2116. That two similar questions produced answers 69 years apart is itself a warning about how much the wording shapes the result. A separate survey by the AAAI, the main professional body for AI research, found that 76% of 475 researchers consider it unlikely or very unlikely that scaling up current methods will produce AGI. AI Impacts, Thousands of AI Authors on the Future of AI · Tech Policy Press on the AAAI survey

Professional forecasters. Their estimates have swung rather than converged. Most moved their dates earlier between 2023 and 2025, later between 2025 and early 2026, and earlier again between January and April 2026. A field whose central estimate moves by years within eighteen months is not settling on a date. FutureSearch, AGI timeline tracker

Statements by industry leaders. These range from “already achieved” to “at best on a good path within five years”. They are set out in sections 4.5 and 4.6.

Measured trends. METR measures how long a task — in the working hours a skilled person would need — an AI can complete half the time. In January 2026, on 228 tasks, it found this length doubling about every 196 days over the whole record, about every 131 days for models released since 2023, and about every 89 days for models released since 2024. Extending that line into the future is tempting; the authors warn against it. The tasks are software, machine-learning and cybersecurity work; they resemble work given to a new contractor with little context; real jobs include interpersonal and judgement elements the tasks lack; and the task set is nearly used up. METR, Time Horizon 1.1

The honest summary: nothing converts a capability measurement into a date, because nothing fixes the target.

4.4 What would it look like, and what is the closest current example?

Each definition calls for a different demonstration.

Definition The demonstration it would need Where things stand
Economic An AI given a randomly chosen office job and performing it at or above the level of the typical employee for months, including the parts nobody wrote down, with no special engineering for that job Never attempted. GDPval tests single, well-specified work products. In its 2025 baseline, leading models matched or beat human professionals on roughly 40–50% of them
Psychometric Close to 100% on the Hendrycks index — which requires the system to remember and build on what it learns while in use Close to zero on storing new long-term memories
Learning efficiency Learning a genuinely new interactive environment as efficiently as a person, without a harness built for it The closest example to date. On ARC-AGI-3, GPT-6 Astra needed fewer moves than the human baseline on 96.0% of levels, and 51.7% fewer moves per level on average
Scientific Producing, on its own, a result that experts accept as a genuine advance Contested. See below

In September 2026 OpenAI reported that about 10,000 AI agents, working for about 88 hours, produced a result on the three-dimensional Navier–Stokes equations, which describe how fluids flow. Whether these equations can produce “blow-ups” — solutions that become infinite in a finite time — is one of the seven Millennium Prize Problems, each carrying a $1 million prize. The result was checked by Lean, proof-checking software, in a further 17 hours. This is the strongest candidate to date for “AI produces new science” — and it also shows why one case cannot settle the question.

Three things are in dispute at once. OpenAI says it will not claim the prize. It concedes it “cannot rule out” that anonymised data from other researchers’ use of its products “helped improve our models”. A mathematician working on a related problem disputes how credit was handled. No independent mathematical assessment was available for this article. Capability, origin and priority are all contested — and only the first is what “AGI” is supposed to be about. Simon Willison, On the Navier–Stokes Millennium Prize Problem

The ARC Prize authors’ own caution about their test applies here too: ARC-AGI-3 “has a tightly bounded scope and format … it does not represent the complexity and open-endedness of the real world”.

4.5 Who says it is close

Who Position and date What it rests on
Greg Brockman, OpenAI “70 to 80% there” (April 2026). At the GPT-6 Astra launch on 3 September 2026: “I think it’s not unreasonable to feel that we are now in the AGI era”, while calling AGI a “gray, fuzzy thing” Growth in model ability; test results reported by the company
Sam Altman, OpenAI Reportedly convinced OpenAI will have an internal system he would call AGI by the end of 2026. Separately: AGI is “a very sloppy term” and “something we’ve already reached without noticing” The charter’s economic definition, loosely applied
Mark Chen, OpenAI “80% of the way” Internal assessment
Dario Amodei, Anthropic Avoids the term. “Powerful AI” “could be as little as 1–2 years away, although it could also be considerably further out” (January 2026), adding “Nothing here is intended to communicate certainty or even likelihood” Extrapolation of scaling trends, heavily qualified
Demis Hassabis, Google DeepMind Around 2030, give or take a year; roughly a 50% chance by the end of the decade; one or two breakthroughs still needed. Has said he chose deliberately provocative language to create urgency Judgement about which components are still missing
François Chollet, Ndea and ARC Prize Previously around 2030. On 4 September 2026: “sooner, because progress is happening faster than I expected” ARC-AGI-3 results arriving about twice as fast as he expected
Elon Musk, xAI Probability that Grok 5 achieves AGI: 10% “and rising” (October 2025) Model size
Adam Khoja, co-author of the Hendrycks index 50% by end of 2028, 80% by end of 2030, on that index The index’s remaining gaps

Quotations from Brockman, Altman, Chen, Hassabis, Chollet and Musk are relayed by named news outlets and were not checked against transcripts. Amodei’s wording was read in his own essay. Amodei, The Adolescence of Technology

Three observations. The two boldest public statements come from the company with the largest direct commercial stake in the claim, and were made at a product launch. The most specific and most carefully qualified statement comes from a chief executive who refuses the word. And Chollet’s revision is the most interesting entry: he wrote a definition designed to be hard to meet, and built the test that produced the result, and he moved his estimate in the direction that does not favour his own commercial position.

4.6 Who says it is not achievable, and what they actually claim

The premise needs correcting. Almost no prominent researcher in this debate says AGI is impossible in principle. Gary Marcus, the most persistent critic, puts it directly: anyone who thinks AGI is impossible is wrong, and anyone who thinks it is imminent is equally wrong. Yann LeCun, François Chollet and Richard Sutton have each founded or joined organisations whose stated purpose is to build it. The disagreement is about route, timing and vocabulary. Four positions should be kept apart.

Method sceptics: achievable, but not this way. LeCun, speaking at Brown University on 1 April 2026, said current systems “fool us into thinking they are smart because they manipulate language. But in fact, they are completely helpless when it comes to the physical world”, and that “there’s literally hundreds of billions invested in an industry that basically is counting on the fact that LLMs [are] going to reach human-level intelligence. It’s complete BS.” He left Meta in December 2025 to build “world models” — systems that learn how the physical world behaves — at AMI Labs. Sutton argues that systems built on stored human knowledge will be overtaken by systems that learn continuously from their own experience. Chollet’s company Ndea combines deep learning with automated program-writing. The AAAI survey suggests this is the majority view among academic researchers: 76% doubt that scaling up current methods gets there. Brown University, report of LeCun’s lecture

Timing sceptics: this way, but not soon. Andrej Karpathy speaks of a “decade of agents”, placing genuinely useful autonomous AI years away. Hassabis, optimistic by public standards, still requires one or two breakthroughs.

Concept sceptics: the word does not name a real threshold. Melanie Mitchell has documented the disagreement about definitions as the central problem in itself. Krakauer’s analysis undermines the statistical basis for a single “general intelligence” score across AI systems. Amodei and Altman, by dropping the word for “powerful AI” and “superintelligence”, concede the point in practice.

Political-economy critics: the word does commercial work. Emily Bender and Alex Hanna argue that “artificial intelligence” is used precisely when its promoters profit from the belief that the technology is human-like. Eryk Salvaggio documents the resulting gap: most researchers doubt AGI is imminent, while policymakers legislate as though it were.

Marcus also supplies the empirical case against. Acemoglu estimates that AI will raise total factor productivity by no more than about 0.66% over ten years. Studies find that AI systems reliably perform only a small fraction of the tasks in typical occupations despite high test scores. Medical AI models stay accurate when key inputs are missing yet become unstable when conditions differ slightly from their training. And language models give confident classifications on thin evidence where people would withhold judgement. Marcus, Rumors of AGI’s arrival have been greatly exaggerated · Acemoglu, The Simple Macroeconomics of AI

4.7 Is it marketing? Partly, demonstrably, and not only

Evidence that the word does commercial work. The changes in definition run in one direction. The organisations that set a demanding bar in 2018 are the ones arguing in 2026 that the bar has been passed, or that it was never well defined. The contractual definition was reportedly a profit figure, not an ability. When the commercial terms were settled in April 2026, the AGI trigger was separated from the money rather than resolved by evidence. At the September 2026 launch, the number that travelled was the flattering one, and the organisation that built the test publicly declined the conclusion drawn from it. The capital at stake is very large: planned 2026 spending on AI infrastructure by the five largest cloud companies is roughly $660–800 billion, close to double 2025 — figures this article rates as unverified. A May 2026 study using several methods finds “localized bubble dynamics”, and notes that “investor narratives often capitalize future productivity gains before they have appeared in cash flows”. The measured effect on the economy so far remains disputed and, on the most-cited peer-reviewed estimate, small. Wang and Chen, Boom, Bubble, or Buildout?

Evidence against a pure-marketing explanation. The capability gains are measured independently and confirmed by organisations with no interest in confirming them. The rise on ARC-AGI-2 from about 54% at about $31 per task to 95% at about $1.12 per task was published by the ARC Prize Foundation, which rejects the AGI claim in the same breath. METR is an independent non-profit and publishes corrections to its own method, including corrections that lower earlier estimates. Costs on ARC-AGI-1 have fallen by orders of magnitude. And the vocabulary runs partly against the hypothesis: the two people with the strongest incentive to use the word are the ones dropping it. This point should be held loosely — “superintelligence” and “powerful AI” are not obviously less promotional — but it is not what a simple hype story predicts.

Assessment. “AGI” today does two jobs at once. It is a genuine research goal with several rigorous but competing ways of measuring it, and it is a marketing and contractual instrument whose flexibility suits the people who deploy it. Both are true. Treating the question as a choice between them is the main analytical mistake available here. Confidence: moderate. The interests are documented; intentions are not, and this article does not accuse any named person of acting in bad faith.

5. Uncertainty, alternatives, and what to watch

The strongest alternative to this article’s framing is the position of Chen and colleagues: that AGI always meant broad, flexible competence; that the critics, not the promoters, have moved the goalposts; and that demanding a sharp threshold is itself the mistake. This article does not adopt that view, because it does not explain why the split between definitions carries so much weight in contracts, product launches and disputes over test results. But it is a serious argument in a serious journal, and nothing gathered here refutes it. UC San Diego on the Chen et al. Comment

Seven indicators would change the assessment whichever definition a reader prefers.

  1. Learning after deployment. A credible demonstration of an AI that keeps learning while in use without losing what it already knows. This is the most consistently named missing piece across otherwise opposed camps, and the area where the Hendrycks index scores current systems near zero.
  2. Results that do not depend on the harness. Standard-harness and vendor-harness scores converging. If they keep diverging, headline numbers are measuring the wrapper.
  3. The efficiency gap. Whether the ARC Prize Grand Prize — 85% on the private ARC-AGI-2 test under strict computing limits, against 24.03% in 2025 — is won. This tests learning under constraint rather than learning bought with computing power.
  4. Economic transmission. Whether high scores on work-product tests such as GDPval start to appear in measured productivity and employment data, and whether estimates of the Acemoglu type are revised upward on evidence.
  5. Independent replication. Whether results like the Navier–Stokes claim are reproduced and assessed by people without access to the model or the vendor’s data.
  6. Robustness. Performance when conditions differ from training and on tests of physical-world understanding, where current systems perform close to chance on some tasks.
  7. Forecast stability. Whether expert estimates begin to converge. Convergence would suggest a shared criterion is emerging; continued swings suggest none exists.

6. How to read an AGI claim

Four questions turn an unfalsifiable claim into one that can be assessed.

  1. Which definition? Economic, psychometric, levels, learning efficiency, behavioural, or contractual. If the claimant will not say, the claim cannot be evaluated.
  2. Measured by whom, under what conditions? The vendor or an independent evaluator; which harness; which version of the test; how many attempts.
  3. At what cost, and does the cost count? A result obtained for $26,000 per evaluation run and a result obtained for $1 per task are different results. Under Chollet’s definition they differ in kind, not degree.
  4. What would prove it wrong? A claim with no stated failure condition is a positioning statement, not a finding.

7. Sources and disclosures

Research records. A register of 37 sources with access status and quality ratings, and a register of 26 material assertions with status and confidence, are held in the publisher’s research files. They are working records and are not published.

Material AI assistance. This article was researched and drafted by Claude Opus 5, an AI assistant, at the direction of the publisher. The plain-language revision and the adaptation for publication were performed by Claude Fable 5.1, a different Anthropic model, in the same working session. This is AI assistance, not independent validation, and no part of it replaces human review.

Interest disclosure. The AI assistants that prepared this article are made by Anthropic. Anthropic is named in this article as a party to the AGI debate, and its models appear in the test results cited. This is a material interest. It was handled by preferring original sources and independent test registers over company claims throughout, and by including Anthropic’s own leadership among the parties whose incentives are analysed in section 4.7. Readers should nevertheless weigh that section with this interest in mind.

Editorial review. The publisher read this text and approved its publication; that review is recorded in the article’s metadata. It is an editorial review, not independent peer review. No machine-learning researcher and no economist has reviewed the text. This article is not an audit, a systematic review, or a certified fact check.

Access limitations. The OpenAI charter page, three Nature items, Axios coverage of the September 2026 launch and of the credit dispute, the Bloomberg feature of 4 September 2026, and articles on Forbes, Fast Company, TechTimes and The Yale Review could not be opened. Where their content mattered, the claim was either supported from independent sources or marked as not established.

Principal sources. ARC Prize leaderboard · ARC Prize on GPT-6 Astra · ARC Prize 2025 results · METR Time Horizon 1.1 · A Definition of AGI · Levels of AGI · On the Measure of Intelligence · The Rise and Fall of G in AGI · RAND on AGI forecasting · Boom, Bubble, or Buildout? · The Adolescence of Technology · LeCun at Brown · History of the AGI clause · On Navier–Stokes · Chen et al. summarised · AGI’s Last Bottlenecks · AAAI survey reported · Rumors of AGI’s arrival have been greatly exaggerated · FutureSearch AGI timeline tracker

Edition note

This is the first public edition, dated 9 September 2026. It is based on internal working versions of the same date; a plain-language revision was made before publication at the publisher’s request, without changing any finding, figure or source. The most likely near-term triggers for an update are the publication of independent assessments of the September 2026 Navier–Stokes claim and of harness-controlled evaluations of GPT-6 Astra. Errors can be reported through the corrections page.