Google could not have picked a worse time to deliver a chronically underpowered Gemini 3.6 Flash, coming right on the heels of stunning open-source AI models from China, such as Moonshot's Kimi K3, which employs novel architectural feats to deliver notable cost savings, all the while offering frontier-level performance.
Google's Gemini 3.6 Flash performs worse than Meta Spark 1.1, GLM-5.2, GPT-5.6 Luna, Sonnet 5, Grok 4.5, and GPT-5.6 Terra
Google has just released the all-new Gemini 3.6 Flash, sporting a context window of 1 million tokens, which is equivalent to 1,500 A4 pages of text, 30,000 lines of code, or an hour of video, and priced at $7.5 per 1 million tokens of output. It also supports text, image, speech, and video inputs.
Critically, according to Google, the model consumes 17 percent fewer output tokens across multi-step workflows.
This brings us to the core of today's topic. Google's Gemini 3.6 Flash has achieved a score of 50 on the the Artificial Analysis Intelligence Index, which measures agentic, general, coding, and scientific reasoning performance of AI models.
In fact, this score is worse than what Meta Spark 1.1, GLM-5.2, GPT-5.6 Luna, Sonnet 5, Grok 4.5, and GPT-5.6 Terra have earned!
A deeper dive reveals that Google's Gemini 3.6 Flash is often superceded by cheaper and older models on most benchmarks. For instance, on SWE-Bench Pro (Public), it is beaten by Grok 4.5, which debuted on July 08, and by Claude Sonnet 5 on MLE Bench.
In fact, Gemini 3.6 Flash might be the most underwhelming model that Google has released to-date!
Of course, Google has also released Gemini 3.5 Flash-Lite, which is fairly cost effective, managing to deliver 350 output tokens per second, and Gemini 3.5 Flash Cyber, which is geared towards cyber security applications.
Follow Wccftech on Google to get more of our news coverage in your feeds.
