

Gemini 4 is getting closer. The loudest leak still has no name tag.
The official trail has moved from pre-training to post-training. Beside it runs a louder, stranger trail: a Flash-labelled Arena entry, dazzling demos and a scorecard with no named referee. The two trails should meet only where the evidence does.
By Katie Delaney / 2026-09-25 / 13 min read
Unverified leak
The viral scorecard, in full
Every number from the graphic that started this story, redrawn so it can actually be read. Claimed, not confirmed: neither the scores nor the prices have been independently verified.
- Gemini 4 Pro (leak)
- Claude Opus 5.5
- GPT-6 Astra
Benchmarks, higher is better
DeepSWE v1.1
Agentic software engineering score, 0 to 100 scale.
Terminal-Bench 2.1Rivals approximate
Terminal task score. Rival figures in the graphic are rounded with a tilde.
OSWorld 2.0Uneven test conditions
Computer use score. The graphic itself flags partial runs for both rivals.
Quality index, points
GDPval-AAVersions do not match
Gemini is scored on v2 while both rivals sit on v2.1, so treat this row as the shakiest comparison on the card.
Specs, separate scales
Context and price use different units, so they sit apart from the scored rows rather than sharing one axis.
Context windowWidest: Gemini
Maximum input size, in tokens.
Price per 1M tokensCheapest: Gemini
Input / output. Shorter bar means cheaper.
The folkfox read: as claimed, Gemini 4 Pro leads every row, doubles the context window and undercuts both rivals on price. None of it is verified, the GDPval versions do not match, and two OSWorld runs are partial. Figures transcribed from the viral graphic via the atoms.dev leak post; our full check of what is confirmed sits in the article below.

The Gemini 4 release date: what Google actually said#
The sharp scent of a launch is in the air, but the calendar remains blank. On 21 July, Google Gemini 3.6 Flash announcement said the company had begun what it called its most ambitious pre-training run for Gemini 4. That was an official training milestone, not a product listing. Google also said Gemini 3.5 Pro was then testing with partners. The claim did not supply a Gemini 4 release date, variant name, API route, price or performance table.
On 6 August, Google leadership note said Koray Kavukcuoglu would lead Google DeepMind as senior vice-president, while Demis Hassabis became chair of DeepMind and chief scientist of Alphabet. Hassabis explicitly referred to progress on Gemini 4. That matters because the project was named publicly before this week’s interview. It also matters because calling Hassabis a departed executive would distort the story: he changed roles and remained involved with the research leadership.
The new clue arrived at The Information’s AI Agenda Live summit on 23 September. The Information original report and The Verge account attribute to Kavukcuoglu the hope of releasing an early post-training output as soon as possible, much earlier than the end of 2026. 9to5Google report also says he did not name a launch date. Post-training is a real shift from the pre-training stage Google described in July, but it is a stage in development, not a launch notice. Safety work, evaluation, product integration and deployment can still change the calendar.
So when is Gemini 4 coming out? The careful answer is that Google wants to release it before the end of the year and has offered no day or month. An October forecast might prove right, but it is a forecast. The fox in this story does not pounce on a date that nobody has written down.
- 21 Jul
Google said Gemini 4 pre-training had begun
- 6 Aug
Google named its new DeepMind leadership
- 23 Sep
Kavukcuoglu discussed early post-training
- No date
Google has published no launch day
A separate Google AI podcast with Kavukcuoglu aired on 3 September and discusses the ambitions of the Gemini 4 pre-training run. It is useful context for what the team wanted to build, but it predates the 23 September summit. Treating that earlier recording as footage of the new announcement would turn a timeline into a trick mirror.
Is Gemini 4 Pro hiding under Gemini 3.8 Flash?#
One genuine model sits at the centre of the rumour: Gemini 3.8 Flash. Google 2 September launch post describes it as a workhorse model for reasoning, coding and longer agentic tasks. Its introductory list price is $0.75 per million input tokens and $3.75 per million output tokens. The public Gemini API model page presents 3.8 Flash as a product developers can call. These are checkable facts about a released model, not clues that its label secretly means another model.
On Arena’s text leaderboard dated 13 September, the Google entry named gemini-3.8-flash-high had a preliminary score of 1495 ± 9. The same snapshot put Claude Fable 5.1 Max at 1508 ± 8. A score gap of 13 points in that snapshot is not a universal capability gap, and the intervals and ranking method need to be read together. Arena ranking explanation says rank spread reflects uncertainty from confidence intervals. The platform measures preferences over model responses; a visually striking response is one observation, not an authenticated release label.
| Item | Value |
|---|---|
| Fable 5.1 Max | 1508 |
| 3.8 Flash High | 1495 |
Arena’s own 2 September announcement named 3.8 Flash High in several Arena modes. The public name predates the speculative X threads by roughly a fortnight. That leaves room for a later test configuration or routing change, but no public Arena or Google statement we found identifies a hidden Gemini 4 Pro behind that label. The gap between an observed screen and a backend model identifier is the whole mystery.
OfficeChai collection of original X posts shows how the inference spread. A clip from @hakmgpt declared that Gemini 4 Pro was under the Flash name; another tester shared a pelican on a bicycle; a third claimed an internal codename and a very large output limit. The clips show creative outputs and the posts show people making claims. They do not provide a signed model card, a reproducible routing log or confirmation from the platform.
Visual test: name tag
Repeat a public Arena claim without assuming the backend identity. · Save raw output and settings
Create an SVG diagram with two visible boxes, labelled Public model name and Confirmed backend identity. Put Gemini 3.8 Flash in the first box. Leave the second blank unless you can cite platform metadata. Show the missing connection clearly, without guessing.
A fresh Gemini 4 Pro checkpoint could be hiding in Arena. Gemini 3.8 Flash is behaving suspiciously well again.
The fresh post above uses “could”, which is the right hinge for this story. The X API found the 25 September post and its timestamp; it is a first-person claim about a demo, not a model-identification document. The crowded thicket of repeats around it is evidence of interest. Repetition does not turn a rumour into a release.

The purported leaked benchmarks, and what they cannot prove#
The viral scorecard is the most numerical and the least traceable part of the tale. Atoms reproduction of the circulating table labels the figures unverified and gives Gemini 4 Pro 88.7% on DeepSWE v1.1, 2064 Elo on GDPval-AA v2, 95.3% on Terminal-Bench 2.1 and 86.8% on OSWorld-2.0. It also lists $2.25 per million input tokens and $11.25 per million output tokens. These are purported leaked benchmark and price claims, not measured results that we can attribute to Google or an evaluator.

| Claimed measure | Viral Gemini 4 Pro figure | Evidence status |
|---|---|---|
| DeepSWE v1.1 | 88.7% | Unverified |
| GDPval-AA v2 | 2064 Elo | Unverified |
| Terminal-Bench 2.1 | 95.3% | Unverified |
| OSWorld-2.0 | 86.8% | Unverified |
| Input tokens, 1m | $2.25 | Unannounced |
| Output tokens, 1m | $11.25 | Unannounced |
- DeepSWE v1.188.7%Unverified
- GDPval-AA v22064 EloUnverified
- Terminal-Bench 2.195.3%Unverified
- OSWorld-2.086.8%Unverified
- Input tokens, 1m$2.25Unannounced
- Output tokens, 1m$11.25Unannounced
Visual test: scorecard audit
Turn a viral image into a list of checkable fields. · No scores added by inference
Draw a visual scorecard audit with rows for model identifier, evaluator, benchmark version, run date, harness, raw outputs and pricing page. Mark each field present or missing from the supplied graphic. Do not redraw the alleged scores as measured results.
The table looks precise because precision is cheap to typeset. It has no named evaluator, test harness, run date, model identifier, raw outputs or independent replication attached to it. The competitor figures in the same image are also claims, not a controlled comparison. A score without its test conditions is a number with no denominator; a price without a pricing page is a bill nobody can actually pay.
Here is a sounder comparison, confined to a released product. Artificial Analysis 2 September evaluation measured Gemini 3.8 Flash High at 59 on its Intelligence Index, up from 56 for Gemini 3.7 Flash High. Those are index points, not percentages, and they belong to Artificial Analysis’ test suite and effort setting. The same evaluator estimated $0.58 per task for 3.8 Flash High and $0.40 for 3.7 Flash High despite unchanged token list prices, because its 3.8 runs used more output and more agent turns. The lesson for buyers is as practical as it is prosaic: price per token is not cost per completed job.
| Item | Value |
|---|---|
| High reasoning | 56 to 59 |
Google’s own Gemini 3.8 launch materials claim 54.9% on HLE-Verified for 3.8 Flash. That is a vendor-reported number about a released model and a named evaluation, not independent proof of the secret checkpoint. Keep vendor tables, independent indices and unattributed screenshots in separate drawers. The den gets messy when a screenshot is promoted to the same status as a documented run.
The more extraordinary claims, including a ten-million-token input window, a 256,000-token output limit and persistent memory, are in the same rumour stream. None appears in a public Gemini 4 API model card. Building a product plan around them now would mean building against a ghost endpoint.
Three visual prompts that turn spectacle into a test#
A creative test can be useful without being a scientific benchmark. The pelican, pagoda and moving airship are memorable because they expose whether a model can keep structure, motion and detail together. Yet a polished clip hides the number of attempts, discarded outputs, human edits, runtime settings and selection decisions. One neat nest is not a census of the forest.
Google 3.8 Flash announcement itself includes interactive demos and says the released model can spend more tokens and make more tool calls on hard tasks. That alone offers an alternative explanation for some richer outputs: a different effort level, prompt, tool route or post-processing pass can change what a tester sees. It does not prove that every disputed clip came from ordinary 3.8 Flash, either. The evidence supports uncertainty in both directions.
Pelican under control
Test SVG structure and animation, then inspect the source. · Repeat three times per model
Create one self-contained SVG of a pelican riding a bicycle. Include anatomically distinct beak, wings and feet, rotating wheels and a labelled speed control. Return the complete SVG source, then explain which parts were simplified.
Pagoda in motion
Test spatial consistency and interaction. · Record time and failed attempts
Build a self-contained browser scene of a voxel pagoda in a garden with trees and blossom. Add camera orbit and one day-to-night control. Return runnable code, list any external assets and state what remains unfinished.
Airship with constraints
Test whether a beautiful demo also behaves. · Use the same harness and budget
Build a self-contained Three.js airship above a coastline. Include keyboard steering, visible speed and altitude, and a reset button. Report run time, token use, errors and every manual change needed before the demo works.
These visual prompts belong in the article because they let a reader repeat the task instead of admiring an edited highlight. Run the same prompt at the same effort setting and budget on each available model. Save the raw output, record retries and measure whether the controls actually work. Do not ask a model to disclose a hidden identity and treat its self-description as proof; models can confabulate their own routing.
The difference between preference and provenance matters here. Arena ranking method describes uncertainty around aggregate human votes, while the disputed posts are individual creative trials. A genuinely stronger public Arena score for a Flash-labelled entry would still identify the public label and the tested behaviour, not the confidential weights or future product name. The new model may indeed be near. This particular trail does not reveal the animal behind the hedge.

What a buyer should do before Google Gemini 4 ships#
The sensible buyer response is neither panic nor a pre-order for a phantom product. Keep the current workload on models whose identity, terms and cost you can inspect. Reserve a small, representative test set for the next Google release, and keep the exact prompts and scoring rubric frozen. The moment a model card and API endpoint arrive, compare completed work, latency, failure rate and total cost on that set.
That is the work behind AI consultancy: deciding where a model earns its place in a workflow. It connects with brand strategy when a new assistant changes who tells your story, content marketing when it changes the production process, and SEO and GEO when its answers mediate discovery. A frontier model release is an input to those decisions, not a strategy in its own right.
The evaluation card should have four columns. First, task success: did the answer or artefact work, and was it checked by a person who understands the job? Second, total cost: include input, output, retries, tools and review time. Third, reliability: count failures across repeated runs, not just the best demo. Fourth, governance: note data handling, retention, access controls and whether the released product is available in your region and contract. Those four columns will outlast any noisy leaderboard week.
If your team already uses multiple assistants, the recent folkfox account of changing assistant share is a useful reminder that audience and workflow mix move independently of any single model launch. The guide to prompt injection covers a different but related problem: agents that can act need controls beyond a benchmark score. The practical guide to Gemini citations shows why an apparently small model or product shift can still change the evidence your customers see.
Record the model card, API identifier, published limits and regional availability when Google posts them.
Use real tasks with a fixed rubric, equal effort settings, multiple runs and saved outputs.
Include tokens, tools, latency, retries and human review per completed task.
Promote a model only where it clears the quality, reliability and governance bar for that workflow.
Visual test: release dashboard
Build a decision sheet when the product is public. · Use verified release data only
Create a compact dashboard with four labelled columns: task success, full cost, reliability and governance. Compare the currently deployed model with the newly released Gemini version across five real tasks. Leave any unavailable metric blank and cite every populated figure.
The next verified signal to watch is specific: a Google announcement naming the released variant, a public model card, a callable API identifier and independently repeated tests. Until then, the Gemini 4 release date remains an open line on the ledger. Speculation is welcome in an editorial story; it becomes costly when it masquerades as procurement evidence.
Frequently asked questions#
When is Gemini 4 coming out?
Google has not announced a Gemini 4 release date. At a 23 September 2026 event, DeepMind chief Koray Kavukcuoglu said the company wanted to release an early post-training output as soon as possible, much earlier than the end of the year. That is a target, not a dated launch.
Is Gemini 4 Pro already in Arena?
Some X users think a model shown as Gemini 3.8 Flash is a hidden Gemini 4 Pro checkpoint. Arena publicly labels a 3.8 Flash entry, and no public Google or Arena statement confirms the proposed hidden identity.
Are the Gemini 4 Pro benchmark scores real?
A circulating graphic lists high DeepSWE, Terminal-Bench and OSWorld scores, but it has no named evaluator, run log or authenticated Google source. Treat its figures as unverified claims until a model card and reproducible evaluation appear.
What is the difference between Gemini 3.8 Flash and Gemini 4?
Gemini 3.8 Flash is a released and documented model. Gemini 4 is a publicly named development effort that, according to DeepMind’s chief, has reached early post-training. Google has not published a Gemini 4 product specification.
Does an impressive Arena demo prove a model is Gemini 4?
No. A demo shows output from a session, while model identity requires reliable routing information or official confirmation. Prompt changes, effort settings, retries and editing can also affect the result.
What should a business test on launch day?
Run the same representative tasks across existing and new models. Record success, total cost, latency, retries, reviewer time, data controls and regional availability. Choose by completed work for your use case.
Read more on this topic#
ChatGPT market share fell to 50 per cent. Here is where your content should live next.
A measured view of shifting assistant use and what that means for a content plan.
Read the pieceSEO & GEOGoogle fixed the citation drop in a day. Its own reports never would have seen it.
Why model and product behaviour need independent checks.
Read the pieceAI securityPrompt Injection Attacks Just Learned to Whisper in Cipher
A practical reminder that agent capability and agent safety need separate tests.
Read the pieceNeed a clear model decision when the moon rises?
folkfox helps teams test AI against real work, with evidence, costs and consequences on the same page.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.
