Skip to content
Skip to content
AI CONSULTANCY

Gemini 4 is getting closer. The loudest leak still has no name tag.

Summarise with

The official trail has moved from pre-training to post-training. Beside it runs a louder, stranger trail: a Flash-labelled Arena entry, dazzling demos and a scorecard with no named referee. The two trails should meet only where the evidence does.

Quick answerThe Gemini 4 release date is unannounced. DeepMind chief Koray Kavukcuoglu said on 23 September that Google wants an early post-training release as soon as possible, much earlier than year-end. Arena and benchmark rumours do not confirm a date or model identity.

Unverified leak

The viral scorecard, in full

Every number from the graphic that started this story, redrawn so it can actually be read. Claimed, not confirmed: neither the scores nor the prices have been independently verified.

  • Gemini 4 Pro (leak)
  • Claude Opus 5.5
  • GPT-6 Astra

Benchmarks, higher is better

DeepSWE v1.1

Agentic software engineering score, 0 to 100 scale.

Gemini 4 Pro (leak)88.7
Claude Opus 5.574.2
GPT-6 Astra74.1

Terminal-Bench 2.1Rivals approximate

Terminal task score. Rival figures in the graphic are rounded with a tilde.

Gemini 4 Pro (leak)95.3
Claude Opus 5.5~88
GPT-6 Astra~88

OSWorld 2.0Uneven test conditions

Computer use score. The graphic itself flags partial runs for both rivals.

Gemini 4 Pro (leak)86.8
Claude Opus 5.581.8 (partial)
GPT-6 Astra72.6 (offline partial)

Quality index, points

GDPval-AAVersions do not match

Gemini is scored on v2 while both rivals sit on v2.1, so treat this row as the shakiest comparison on the card.

Gemini 4 Pro (leak)2,064 (v2)
Claude Opus 5.51,846 (v2.1)
GPT-6 Astra1,542 (v2.1)

Specs, separate scales

Context and price use different units, so they sit apart from the scored rows rather than sharing one axis.

Context windowWidest: Gemini

Maximum input size, in tokens.

Gemini 4 Pro (leak)2M tokens
Claude Opus 5.51M tokens
GPT-6 Astra1.1M tokens

Price per 1M tokensCheapest: Gemini

Input / output. Shorter bar means cheaper.

Gemini 4 Pro (leak)$2.25 / $11.25
Claude Opus 5.5$4 / $20
GPT-6 Astra$10 / $50

The folkfox read: as claimed, Gemini 4 Pro leads every row, doubles the context window and undercuts both rivals on price. None of it is verified, the GDPval versions do not match, and two OSWorld runs are partial. Figures transcribed from the viral graphic via the atoms.dev leak post; our full check of what is confirmed sits in the article below.

Section 01

The Gemini 4 release date: what Google actually said#

The sharp scent of a launch is in the air, but the calendar remains blank. On 21 July, Google Gemini 3.6 Flash announcement said the company had begun what it called its most ambitious pre-training run for Gemini 4. That was an official training milestone, not a product listing. Google also said Gemini 3.5 Pro was then testing with partners. The claim did not supply a Gemini 4 release date, variant name, API route, price or performance table.

On 6 August, Google leadership note said Koray Kavukcuoglu would lead Google DeepMind as senior vice-president, while Demis Hassabis became chair of DeepMind and chief scientist of Alphabet. Hassabis explicitly referred to progress on Gemini 4. That matters because the project was named publicly before this week’s interview. It also matters because calling Hassabis a departed executive would distort the story: he changed roles and remained involved with the research leadership.

The new clue arrived at The Information’s AI Agenda Live summit on 23 September. The Information original report and The Verge account attribute to Kavukcuoglu the hope of releasing an early post-training output as soon as possible, much earlier than the end of 2026. 9to5Google report also says he did not name a launch date. Post-training is a real shift from the pre-training stage Google described in July, but it is a stage in development, not a launch notice. Safety work, evaluation, product integration and deployment can still change the calendar.

So when is Gemini 4 coming out? The careful answer is that Google wants to release it before the end of the year and has offered no day or month. An October forecast might prove right, but it is a forecast. The fox in this story does not pounce on a date that nobody has written down.

  • 21 Jul

    Google said Gemini 4 pre-training had begun

  • 6 Aug

    Google named its new DeepMind leadership

  • 23 Sep

    Kavukcuoglu discussed early post-training

  • No date

    Google has published no launch day

A separate Google AI podcast with Kavukcuoglu aired on 3 September and discusses the ambitions of the Gemini 4 pre-training run. It is useful context for what the team wanted to build, but it predates the 23 September summit. Treating that earlier recording as footage of the new announcement would turn a timeline into a trick mirror.

Section 02

Is Gemini 4 Pro hiding under Gemini 3.8 Flash?#

One genuine model sits at the centre of the rumour: Gemini 3.8 Flash. Google 2 September launch post describes it as a workhorse model for reasoning, coding and longer agentic tasks. Its introductory list price is $0.75 per million input tokens and $3.75 per million output tokens. The public Gemini API model page presents 3.8 Flash as a product developers can call. These are checkable facts about a released model, not clues that its label secretly means another model.

On Arena’s text leaderboard dated 13 September, the Google entry named gemini-3.8-flash-high had a preliminary score of 1495 ± 9. The same snapshot put Claude Fable 5.1 Max at 1508 ± 8. A score gap of 13 points in that snapshot is not a universal capability gap, and the intervals and ranking method need to be read together. Arena ranking explanation says rank spread reflects uncertainty from confidence intervals. The platform measures preferences over model responses; a visually striking response is one observation, not an authenticated release label.

What Arena actually labelled on 13 September
Watercolour ladder chart of Arena scores, Fable 5.1 Max at 1508 and Gemini 3.8 Flash High at 1495; the Google entry is preliminaryArena score05001K1.5KFable 5.1 Max: 1508Fable 5.1 Max15083.8 Flash High: 14953.8 Flash High1495
Watercolour ladder chart of Arena scores, Fable 5.1 Max at 1508 and Gemini 3.8 Flash High at 1495; the Google entry is preliminary
ItemValue
Fable 5.1 Max1508
3.8 Flash High1495
Arena’s public 13 September snapshot listed Gemini 3.8 Flash High at 1495 ± 9 and Fable 5.1 Max at 1508 ± 8. It did not label the Google entry Gemini 4. Source: Arena text leaderboard.

Arena’s own 2 September announcement named 3.8 Flash High in several Arena modes. The public name predates the speculative X threads by roughly a fortnight. That leaves room for a later test configuration or routing change, but no public Arena or Google statement we found identifies a hidden Gemini 4 Pro behind that label. The gap between an observed screen and a backend model identifier is the whole mystery.

OfficeChai collection of original X posts shows how the inference spread. A clip from @hakmgpt declared that Gemini 4 Pro was under the Flash name; another tester shared a pelican on a bicycle; a third claimed an internal codename and a very large output limit. The clips show creative outputs and the posts show people making claims. They do not provide a signed model card, a reproducible routing log or confirmation from the platform.

01

Visual test: name tag

Repeat a public Arena claim without assuming the backend identity. · Save raw output and settings

Create an SVG diagram with two visible boxes, labelled Public model name and Confirmed backend identity. Put Gemini 3.8 Flash in the first box. Leave the second blank unless you can cite platform metadata. Show the missing connection clearly, without guessing.

@MehdiCade
A fresh Gemini 4 Pro checkpoint could be hiding in Arena. Gemini 3.8 Flash is behaving suspiciously well again.
25 September 2026, 15:42 UTCView on X

The fresh post above uses “could”, which is the right hinge for this story. The X API found the 25 September post and its timestamp; it is a first-person claim about a demo, not a model-identification document. The crowded thicket of repeats around it is evidence of interest. Repetition does not turn a rumour into a release.

Section 03

The purported leaked benchmarks, and what they cannot prove#

The viral scorecard is the most numerical and the least traceable part of the tale. Atoms reproduction of the circulating table labels the figures unverified and gives Gemini 4 Pro 88.7% on DeepSWE v1.1, 2064 Elo on GDPval-AA v2, 95.3% on Terminal-Bench 2.1 and 86.8% on OSWorld-2.0. It also lists $2.25 per million input tokens and $11.25 per million output tokens. These are purported leaked benchmark and price claims, not measured results that we can attribute to Google or an evaluator.

The folkfox vixen weighs an unsigned model nameplate against a confirmed paper trail in the Gemini 4 investigation
A model nameplate needs a source. The viral scorecard has none.
Every value in this table comes from an unauthenticated viral graphic reproduced by Atoms. None is a verified Gemini 4 result or announced price.
Claimed measureViral Gemini 4 Pro figureEvidence status
DeepSWE v1.188.7%Unverified
GDPval-AA v22064 EloUnverified
Terminal-Bench 2.195.3%Unverified
OSWorld-2.086.8%Unverified
Input tokens, 1m$2.25Unannounced
Output tokens, 1m$11.25Unannounced
  • DeepSWE v1.188.7%Unverified
  • GDPval-AA v22064 EloUnverified
  • Terminal-Bench 2.195.3%Unverified
  • OSWorld-2.086.8%Unverified
  • Input tokens, 1m$2.25Unannounced
  • Output tokens, 1m$11.25Unannounced
01

Visual test: scorecard audit

Turn a viral image into a list of checkable fields. · No scores added by inference

Draw a visual scorecard audit with rows for model identifier, evaluator, benchmark version, run date, harness, raw outputs and pricing page. Mark each field present or missing from the supplied graphic. Do not redraw the alleged scores as measured results.

The table looks precise because precision is cheap to typeset. It has no named evaluator, test harness, run date, model identifier, raw outputs or independent replication attached to it. The competitor figures in the same image are also claims, not a controlled comparison. A score without its test conditions is a number with no denominator; a price without a pricing page is a bill nobody can actually pay.

Here is a sounder comparison, confined to a released product. Artificial Analysis 2 September evaluation measured Gemini 3.8 Flash High at 59 on its Intelligence Index, up from 56 for Gemini 3.7 Flash High. Those are index points, not percentages, and they belong to Artificial Analysis’ test suite and effort setting. The same evaluator estimated $0.58 per task for 3.8 Flash High and $0.40 for 3.7 Flash High despite unchanged token list prices, because its 3.8 runs used more output and more agent turns. The lesson for buyers is as practical as it is prosaic: price per token is not cost per completed job.

A measured improvement on a released model
A measured improvement on a released modelWatercolour trail chart showing Artificial Analysis Intelligence Index moving from 56 for Gemini 3.7 Flash High to 59 for Gemini 3.8 Flash High3.7 Flash3.8 Flash015304560High reasoning: 56 to 59High reasoning59+3
Watercolour trail chart showing Artificial Analysis Intelligence Index moving from 56 for Gemini 3.7 Flash High to 59 for Gemini 3.8 Flash High
ItemValue
High reasoning56 to 59
Artificial Analysis measured 3.8 Flash High at 59 index points, up from 56 for 3.7 Flash High on 2 September. These are independent 3.x scores, not Gemini 4 benchmarks.

Google’s own Gemini 3.8 launch materials claim 54.9% on HLE-Verified for 3.8 Flash. That is a vendor-reported number about a released model and a named evaluation, not independent proof of the secret checkpoint. Keep vendor tables, independent indices and unattributed screenshots in separate drawers. The den gets messy when a screenshot is promoted to the same status as a documented run.

The more extraordinary claims, including a ten-million-token input window, a 256,000-token output limit and persistent memory, are in the same rumour stream. None appears in a public Gemini 4 API model card. Building a product plan around them now would mean building against a ghost endpoint.

Section 04

Three visual prompts that turn spectacle into a test#

A creative test can be useful without being a scientific benchmark. The pelican, pagoda and moving airship are memorable because they expose whether a model can keep structure, motion and detail together. Yet a polished clip hides the number of attempts, discarded outputs, human edits, runtime settings and selection decisions. One neat nest is not a census of the forest.

Google 3.8 Flash announcement itself includes interactive demos and says the released model can spend more tokens and make more tool calls on hard tasks. That alone offers an alternative explanation for some richer outputs: a different effort level, prompt, tool route or post-processing pass can change what a tester sees. It does not prove that every disputed clip came from ordinary 3.8 Flash, either. The evidence supports uncertainty in both directions.

01

Pelican under control

Test SVG structure and animation, then inspect the source. · Repeat three times per model

Create one self-contained SVG of a pelican riding a bicycle. Include anatomically distinct beak, wings and feet, rotating wheels and a labelled speed control. Return the complete SVG source, then explain which parts were simplified.
02

Pagoda in motion

Test spatial consistency and interaction. · Record time and failed attempts

Build a self-contained browser scene of a voxel pagoda in a garden with trees and blossom. Add camera orbit and one day-to-night control. Return runnable code, list any external assets and state what remains unfinished.
03

Airship with constraints

Test whether a beautiful demo also behaves. · Use the same harness and budget

Build a self-contained Three.js airship above a coastline. Include keyboard steering, visible speed and altitude, and a reset button. Report run time, token use, errors and every manual change needed before the demo works.

These visual prompts belong in the article because they let a reader repeat the task instead of admiring an edited highlight. Run the same prompt at the same effort setting and budget on each available model. Save the raw output, record retries and measure whether the controls actually work. Do not ask a model to disclose a hidden identity and treat its self-description as proof; models can confabulate their own routing.

The difference between preference and provenance matters here. Arena ranking method describes uncertainty around aggregate human votes, while the disputed posts are individual creative trials. A genuinely stronger public Arena score for a Flash-labelled entry would still identify the public label and the tested behaviour, not the confidential weights or future product name. The new model may indeed be near. This particular trail does not reveal the animal behind the hedge.

Section 05

What a buyer should do before Google Gemini 4 ships#

The sensible buyer response is neither panic nor a pre-order for a phantom product. Keep the current workload on models whose identity, terms and cost you can inspect. Reserve a small, representative test set for the next Google release, and keep the exact prompts and scoring rubric frozen. The moment a model card and API endpoint arrive, compare completed work, latency, failure rate and total cost on that set.

That is the work behind AI consultancy: deciding where a model earns its place in a workflow. It connects with brand strategy when a new assistant changes who tells your story, content marketing when it changes the production process, and SEO and GEO when its answers mediate discovery. A frontier model release is an input to those decisions, not a strategy in its own right.

The evaluation card should have four columns. First, task success: did the answer or artefact work, and was it checked by a person who understands the job? Second, total cost: include input, output, retries, tools and review time. Third, reliability: count failures across repeated runs, not just the best demo. Fourth, governance: note data handling, retention, access controls and whether the released product is available in your region and contract. Those four columns will outlast any noisy leaderboard week.

If your team already uses multiple assistants, the recent folkfox account of changing assistant share is a useful reminder that audience and workflow mix move independently of any single model launch. The guide to prompt injection covers a different but related problem: agents that can act need controls beyond a benchmark score. The practical guide to Gemini citations shows why an apparently small model or product shift can still change the evidence your customers see.

A four-step release watch
Watch the official catalogue

Record the model card, API identifier, published limits and regional availability when Google posts them.

Freeze the test set

Use real tasks with a fixed rubric, equal effort settings, multiple runs and saved outputs.

Count full cost

Include tokens, tools, latency, retries and human review per completed task.

Decide by use case

Promote a model only where it clears the quality, reliability and governance bar for that workflow.

01

Visual test: release dashboard

Build a decision sheet when the product is public. · Use verified release data only

Create a compact dashboard with four labelled columns: task success, full cost, reliability and governance. Compare the currently deployed model with the newly released Gemini version across five real tasks. Leave any unavailable metric blank and cite every populated figure.

The next verified signal to watch is specific: a Google announcement naming the released variant, a public model card, a callable API identifier and independently repeated tests. Until then, the Gemini 4 release date remains an open line on the ledger. Speculation is welcome in an editorial story; it becomes costly when it masquerades as procurement evidence.

Questions

Frequently asked questions#

When is Gemini 4 coming out?

Google has not announced a Gemini 4 release date. At a 23 September 2026 event, DeepMind chief Koray Kavukcuoglu said the company wanted to release an early post-training output as soon as possible, much earlier than the end of the year. That is a target, not a dated launch.

Is Gemini 4 Pro already in Arena?

Some X users think a model shown as Gemini 3.8 Flash is a hidden Gemini 4 Pro checkpoint. Arena publicly labels a 3.8 Flash entry, and no public Google or Arena statement confirms the proposed hidden identity.

Are the Gemini 4 Pro benchmark scores real?

A circulating graphic lists high DeepSWE, Terminal-Bench and OSWorld scores, but it has no named evaluator, run log or authenticated Google source. Treat its figures as unverified claims until a model card and reproducible evaluation appear.

What is the difference between Gemini 3.8 Flash and Gemini 4?

Gemini 3.8 Flash is a released and documented model. Gemini 4 is a publicly named development effort that, according to DeepMind’s chief, has reached early post-training. Google has not published a Gemini 4 product specification.

Does an impressive Arena demo prove a model is Gemini 4?

No. A demo shows output from a session, while model identity requires reliable routing information or official confirmation. Prompt changes, effort settings, retries and editing can also affect the result.

What should a business test on launch day?

Run the same representative tasks across existing and new models. Record success, total cost, latency, retries, reviewer time, data controls and regional availability. Choose by completed work for your use case.

Keep reading

Read more on this topic#

Need a clear model decision when the moon rises?

folkfox helps teams test AI against real work, with evidence, costs and consequences on the same page.

Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.

The den

Where to?

Pricing

Choose a section. Enter opens it, Escape continues reading.

Cookie preferences

folkfox uses data the way we use strategy: only when it earns its place.