A Supplier Declared It. The Scoreboard Did Not
On Saturday the chief executive of a chip company announced the arrival of artificial general intelligence. The foundation whose benchmark he was leaning on had already declined to say the same thing.
By Katie Delaney · 2026-09-07 · 12 min read
Brand messaging starts with who said it#
additional GPUs NVIDIA said are coming online next, in the same post that declared AGI had arrived
On 6 September, NVIDIA's chief executive posted on X, replying to the head of an infrastructure company. The Free Press Journal and The Hans India both carry the wording in full: GPT-6 Astra, trained on around 100,000 NVIDIA Grace Blackwell NVLink72 systems, from ChatGPT to o1 to Astra in four years, AGI has arrived, congratulations to the OpenAI team, 400,000 GPUs coming online next.
Anyone responsible for brand messaging in a technology category had a decision to make within the hour, because the claim travelled faster than any check on it could.
Read that sentence twice, because the last clause is doing commercial work the first clause is getting credit for. A capability verdict and an order-book signal arrived in the same breath, from a company that sells the hardware.
GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.
The moonlit half of this is that none of it is hidden. The post is public, the hedges are on the record, and the refusal is published. It simply takes a patient prowl through four sources to assemble, which is more than a news cycle allows.
Nothing in that post is untrue as a statement of what its author believes. The problem for anyone writing brand messaging this week is that it was widely reported as a finding, and it is not one. It is an opinion, held by an interested party, about a term with no agreed definition.
The scent worth following is the hedge#
OpenAI did not say the same thing. Its president told Stratechery that maybe it was the previous model, maybe it is Astra, maybe it is the next model, but somewhere in there he thinks we cross most people's threshold. He went further in the same interview, and this is the most candid line anyone has offered: it is almost this blurry thing, and at the beginning they thought there would be a point in time everyone agreed on, and it has not played out like that at all.
TechRepublic quotes him more warmly, saying that for him personally he does think we are there. Both are real. A brand communications team that quotes only one of them is not reporting, it is casting.
The benchmark owner declined to follow the trail#
Here is the fact that should end the argument, and it has been almost entirely absent from the coverage.
The scores being cited as evidence include a near-perfect result on ARC-AGI-3. The organisation behind that benchmark, the ARC Prize Foundation, addressed the question directly. Reported by the Australian Computer Society's Information Age, it said that while it believes Astra represents meaningful progress towards generalisation, it is not claiming that it is AGI.
While we believe Astra represents meaningful progress towards generalisation, we are not claiming that it is AGI.
That single sentence should reset the brand messaging for every company planning to cite these results in a deck this month.
The people who built the measuring instrument, looking at the score on their own instrument, declined to draw the conclusion a hardware supplier drew from it. That is the whole story, and it is the sort of detail that separates brand credibility from noise.
CoinCentral lists the same three figures and notes plainly that it cannot say which organisation published them. That is the correct caveat and almost nobody printed it. These are the vendor's own numbers, and a vendor benchmark is vendor marketing until an independent evaluator repeats it.
A score is a configuration, not a capability#
There is a second reason to hold these numbers loosely, and the neatest illustration of it belongs to NVIDIA itself.
TechRepublic reports that NVIDIA achieved 100 per cent on ARC-AGI-3 using Claude Opus 5 with memory and tooling, against a baseline for that underlying model of roughly 30 per cent. Same benchmark, same model, seventy points of difference, and the only variable was how it was wired up.
This is the den most brand messaging gets built on and it is dug in sand. A number lifted from a launch post carries none of the conditions that produced it, and the conditions are where the capability actually lives.
Any marketing claims built on a single benchmark figure inherit that fragility. If a competitor can move the same model seventy points by changing the harness, the number was never describing the model on its own. Careful claims substantiation asks what the configuration was before it asks what the score was.
The sharpest omission is OpenAI's own. TechRepublic notes the company defines artificial general intelligence as outperforming humans at most economically valuable work, and that its own benchmark measuring exactly that, GDPval, was absent from the launch materials. A definition was offered. The test matching that definition was not reported.
The named critics, and the thicket they point at#
Scepticism here is not a vibe. It is specific, it is signed, and it is about definitions, which makes it usable in brand messaging rather than merely oppositional.
Gary Marcus, quoted by The Hans India, said it is unfortunate to see the claim made without evidence or definitions, and that by conventional definitions Astra does not yet reach that level. Toby Walsh of the UNSW AI Institute told Information Age he would be amazed if it really had matched all human cognitive capabilities, and that he would eat his hat if trivial things an eight year old can do were not found that Astra fails at.
| Speaker | What they actually said | Position |
|---|---|---|
| Jensen Huang, NVIDIA | "AGI has arrived" | Unhedged claim |
| Greg Brockman, OpenAI | "maybe it was the previous model, maybe it's Astra, maybe it's the next model" | Hedged |
| ARC Prize Foundation | "we are not claiming that it is AGI" | Declined |
| Gary Marcus, critic | "without evidence or definitions" | Rejected |
| Sam Altman, OpenAI | "a very poorly defined term" | Distanced |
The most useful objection is the quietest. Dr Rebecca Johnson of the University of Sydney asked what generally smarter than humans actually means: which humans, smarter at what, across which environments. Her verdict is the sentence any brand messaging lead should keep to hand. Artificial general intelligence, she said, has a philosophy of science problem masquerading as a benchmark problem.

Brand messaging that quotes a critic without quoting the claim is as lopsided as the reverse. Both belong on the page, which is why the table above carries all four positions rather than the two that suit an argument.
That is not an anti-technology position and folkfox is not taking one. Astra may well be remarkable. The point is narrower and it is about language: a term nobody has defined cannot be verified, and a claim that cannot be verified cannot be substantiated, whatever the scoreboard says.
Follow the incentive through the undergrowth#
None of this requires assuming bad faith. It requires noticing who is paid by which outcome, which is ordinary diligence rather than cynicism.
explainx.ai put it bluntly: the claim is rhetoric layered on top of a real hardware number rather than a new technical result, and its author has an unusually direct commercial interest in every major model release being read as validation that more spending on graphics processors is warranted.
And the sharpest dissent of all comes from inside the building. Digital Today reports OpenAI chief executive Sam Altman keeping a deliberate distance from the word, describing artificial general intelligence as a very poorly defined term and saying it is in some ways similar to an unimportant marketing term. The company's own chief executive is warier of the label than the company selling the chips.
The market context is not subtle either. CoinCentral notes NVIDIA shares up 23.7 per cent year to date with an average analyst price target of $325.23, and The Motley Fool frames the stakes as winner-take-all economics, citing RAND on an early advantage translating into a decisive economic edge.
Nobody here is hiding anything, and the fox does not need to outfox them. The hedges are quoted, the refusal is published and the incentive is declared on the same page as the claim. Reading them together is the entire discipline.
So the practical instruction for a brand communications team is short. Name the speaker, quote the hedge, carry the refusal, and date the claim. A capability sentence with no date is wrong within a quarter, and this field moves in weeks.
Good brand messaging in this category is mostly restraint: saying the smaller true thing rather than the larger unverifiable one, and letting the scoreboard catch up.
Do that and you can write about this week honestly without either cheering or sneering, which is the register most technology brand messaging cannot currently hold. The reward is not moral. It is that your marketing claims survive the correction when it comes, and your competitor's do not.
More each morning in the newsroom, positioning work on brand strategy, the model economics on AI consultancy, the writing on content marketing, published rates on pricing, and the companion piece on what an AI budget actually costs. The quarry here is not the model. It is the sentence, and brand messaging is where it gets caught or lost.
Frequently asked questions#
Did NVIDIA's chief executive actually say AGI has arrived?
Yes. Jensen Huang posted on X on 6 September 2026, replying to another executive, that GPT-6 Astra was trained on around 100,000 NVIDIA Grace Blackwell NVLink72 systems and that AGI has arrived, adding that 400,000 GPUs are coming online next. The wording is carried in full by several outlets.
Did OpenAI agree that AGI has arrived?
Not in those terms. Its president said maybe it was the previous model, maybe Astra, maybe the next one, and separately that it is a blurry thing rather than a single agreed point in time. He has also said that personally he thinks we are there. OpenAI as a company has not issued the flat claim.
What did the benchmark owner say?
The ARC Prize Foundation, whose ARC-AGI-3 benchmark supplied one of the headline scores, said that while it believes Astra represents meaningful progress towards generalisation, it is not claiming that it is AGI. That is the measuring organisation declining the conclusion drawn from its own measurement.
How should marketing claims handle an unverifiable capability claim?
Attribute it. Good brand messaging names the speaker rather than the field. Name the speaker, quote their exact words including any hedge, carry any refusal or denial in the same passage rather than later, and date the claim. If a term has no agreed definition, say so plainly instead of implying a measurement exists.
Why does claims substantiation matter for a benchmark score?
Because a score describes a configuration, not a capability. NVIDIA reached 100 per cent on ARC-AGI-3 using a rival model with memory and tooling whose baseline sat around 30 per cent. If scaffolding can move a result seventy points, the headline number was never describing the model alone.
What does OpenAI's own chief executive say about the term AGI?
Sam Altman has kept a deliberate distance from the label, describing artificial general intelligence as a very poorly defined term and saying it is in some ways similar to an unimportant marketing term. That is a notably cooler position than either the hardware supplier or his own company president has taken.
Is folkfox saying the model is not impressive?
No. The argument is about language rather than capability. Astra may well be remarkable, and the reported results are striking. The narrow point is that a term nobody has defined cannot be verified, and a claim that cannot be verified cannot be substantiated in your marketing.
Read more on this topic#
The Price Is Real. The Bill Is Somewhere Else
The companion piece: what an AI programme actually costs once the headline rate expires.
Read the pieceNobody Had Edited It in a Decade. Then Came Four Thousand Pages.
What autonomous agents did to a system nobody was watching, and why containment is governance.
Read the pieceThe Score Everyone Is Quoting Belongs to a Different Benchmark.
The same disease in a different vertical: a number travelling further than the test that produced it.
Read the pieceThey Studied the Audience Through the Wrong End of the Glass.
What audiences actually believe about AI claims, measured rather than assumed.
Read the pieceWriting about a claim nobody can check yet?
Our brand messaging work survives the correction: attributed, dated, and honest about what has actually been measured.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.