It doesn't matter how good the model is if you're (mostly) forced to use it in Antigravity - which turns any model into crap.
Wake me up when Antigravity doesn't suck.
It's time for them to start focusing on open-weight models and efficiency. Otherwise there's just a layer of marketing hype and "will it do this?" that has to be cut through for evaluation of each and every release cycle.
Models are getting easier and easier to create. The money, if there's any here, is in the harness the user interfaces with, and the data centers running them.
Google was late to coding agents and as-per-usual fucked it up with their crazy project management culture.
Usually Google gets away with it due to inertia, however this time they are paying a heavy price because they missed out on the training data that Anthropic and OpenAI have gathered with claude and codex.
Screw your government! US and Israeli governments should get the least access, but of course we all know they'll be the (only) ones to get full unfiltered access.
Gemini 3.6 Flash https://news.ycombinator.com/item?id=48993130
For me, Gemini models are the most usable. Claude Opus and Mistral always try to turn queries into one-shot enormous commits, which just burns tokens, time and annoys me for something which is still wrong more often than not.
Gemini seems far better at listening to instructions and giving me what I actually want, on top of using far fewer tokens and wasting my time. Fable is the only model that's come close to Gemini Pro for me.
And as this is about Flash, it's exciting, I find Flash can usually get the right answer pretty quickly and without too much nonsense.
Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive.
The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters
you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Charlie Kirk, it also does not know. I tried it multiple time to ensure that I didn't not get routed to older models!
> but google has search
irrelevant, without deeper knowledge about cutting edge technologies or latest libraries, all of it suggestions are crap. even you ask it to search it will still use outdated keyword thus only getting outdated information.
in other word, what a disaster!
[0] - https://www.ft.com/content/6049a031-9e9b-464c-97bb-414da04d5...
https://www.wsj.com/tech/ai/chinas-xi-touts-open-source-ai-a...
So who even knows.
Anyway, given that both Gemini and OpenAI have 3 sizes of models, one would think Google compares their medium size to OpenAIs.
we are stealing plutocracy from the jaws of emancipation.
i don't want to live in a world where abundance is guarded and shared among politicians and cronies, whilst the rest are left to rot.
1- no comparison with gemini 3.1 pro
2- no comparison with any other model
https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
My results [0] put Gemini 3.6 Flash at the top.
3.6 Flash high has same $1.5 input price as 3.5 Flash, but output is cheaper from $9.0 to $7.5.
Google said 3.6 Flash is more token efficient, but in my tests it's actually LESS token efficient[1] than 3.5 Flash, so despite the output price reduction, it still costs more.
[0]: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...
[1]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...
Front page is tedious these days.
Google watches over the last few months a flat out assault on the Pareto curve from American and Chinese companies. Release after release pushing the boundaries of frontier intelligence and price/performance.
And the response from arguably the biggest AI research labs in the world by headcount is Flash 3.6.
What do you do when you are given essentially unlimited resources and still find yourself falling behind?
Has anybody found any models better at image or audio analysis?
It is also cheaper than 3.5:
> This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
If it wasn't for Gemini/Antigravity I'd have to go with a Max Claude plan, as it stands now I can get by with just a Claude Pro plan to get Opus when I need it, whilst using Antigravity as my day-to-day workhorse.
Unfortunately Gemini Flash became too expensive to use as a general purpose model (i.e. for AI features in Apps), luckily there are plenty of cheaper Chinese models to fill that gap now.
gemini-2.5-flash-lite: $0.10 input / $0.40 output
gemini-3.1-flash-lite: $0.25 input / $1.50 output
gemini-3.5-flash-lite: $0.30 input / $2.50 output (a 6.25x increase over 2.5!)
Now watch them deprecate Gemini 2.5 Flash-Lite in the coming months...
Plus they are probably running these things on every Google search so saving tokens is a huge win for them.
Deepseek Pro: 0.435/m 0.87/m
That's wildly ambitious pricing by Google. You can maybe get away with spicy pricing at the SOTA edge but at the lower tiers everything is a lot more price sensitive.
I use 3.1 Flash Lite regularly to classify listings on eCommerce websites. It's great for this task - fast, cheap and accurate.
In fact, it was the single best model we tried in terms of the speed vs accuracy vs price tradeoffs - including the Chinese models.
Of course, 3.5 Flash was more accurate but the 5x cost increase couldn't be justified.
3.5 Flash Lite sounds like it could be a strict upgrade for our use case, without a significant increase in costs or drop in speed.
It's not GPT-6 but it's not trying to be. It's a completely different tool and great at what it does.
3.5 Flash Lite is only a hair cheaper than 3.0 Flash, but I think 3.0 Flash is a massively more capable model?
"3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)"
So which one is it? 65% or 49%?
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
Being able to ask questions to small open models seems.... obviously useful?
They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow-up. Antigravity IDE cannot even have poweruser subscriptions now from Google Workspace an Gemini Enterprise Agent Platform cannot be attached to Antigravity IDE.
Gemini Enterprise Agent Platform has an incredibly abysmal setup process, and if I want to limit spending per-user I have to create projects per user. The fact that you cannot activate Anthropic models on it if the billing still has free credits is almost a joke.
I was a big proponent of Google and Gemini, but they left us reeling with their abrupt product decisions. Forced us to buy $200 subscriptions directly from Anthropic/OpenAI.
The GCP team wants their slice, the other team wants some otjer slice, and so on. Everyone wants some crap for their promotion package.
It's no wonder Meta has shit the bed even worse.
It's also why Google still releases actually decent, useful models despite the product being such a hilarious mess. A lot of the time Gemini models have actually been better as production LLMs as part of LLM-based production applications than OpenAI and Anthropic models when it comes to the complete cost:quality:latency:adherence picture. And they still are. We have products in prod that use Gemini because they're better than any other model at the specific task. But we wouldn't dare use it for anything coding related, or even just as productivity tool to rely on, because as a consumer product it's a joke.
When they swapped Google Assistant for Gemini as the default voice provider in Android Auto it was so annoying. My wife's non-work space account can get Gemini to do the normal things like play music and what not, but my Workspace one can't do much of anything at all. I can talk about nearly any random topic with it, but getting it to change the playlist, nah, can't help you there.
It's no surprise to me to see them fumble actually supporting a lot of the consumer features of Gemini into Workspace.
They're not even benchmarking against other models now, just against themselves - which tells you everything you need to know.
There is absolutely no loyalty when it comes to coding. Nothing could be more common than people threating to jump ship whenever another frontier or open source model comes.
Google is clearly able to keep growing their free and consumer and small business use cases. Unlike corporate coding, we actually have evidence that solo and small businesses can actually see productivity gains.
Anthropic and OpenAI need to stay dancing like mad, because it's their source revenue which underpins their investments.
Why does Google need to shove something at the top at the same desperate cadence? Other than "recursive self improvement leads to AGI" it seems perfectly fine if they push out something dramatically better every year and half.
You should be happy for them.
This is still very early days. Who is "on top" has flipped back and forth many times already. The next frontier model release (from whomever) will change things again.
Claude Code has largely won individual developer mindshare and has been on top ever since it came out. The benchmarks change, but almost nobody opts to use anything other than Claude IME when I ask them. Enterprise is more competitive since they care about costs and other things, but developers leaning towards Claude puts a thumb on the scales there.
The product doesn't have much lock in, so it is possible to dislodge Claude, and Anthropic could (and some may argue is likely to) just shoot themselves in the foot again and again and again, but Google has never been particularly good at enterprise sales, and they have never actually been at the frontier of intelligence.
I think Google's incentives have mostly about building models for their products, which makes them focus more on the cheap end, and while they need that, it feels like the Innovator's Dilemma is biting them here.
I own a lot of Google stock from working there in the past and have been quite happy about their trajectory up until the last 6 months, but I am getting pretty antsy about their AI story these days.
At this rate, if Google has a flagship model, you're better off plugging it into a competitor's tooling than hope Google figures out how to use it.
They are not allowing me to hit their endpoints which agy hits - it's frustrating . i tried to hack it with gemini itself. what i love about gemini is it's so encouraging and ready to help you - even against the agy client : ) .
Even though im so frustrated with this - i still love Gemini for some reason ! Most encouraging model in the world!
And the real numbers could be better for Anthropic. It's feasible Opus models are actually cheaper to serve than GLM 5.2 because Anthropic have optimized the hell out of inference.
Likewise. This seems like a common feel. I have at least spent $4000 and likely a lot more on Gemini API because I really wanted them to win. I gave up.
Why do you care? Why would you spend your own money to a multi trillion dollar company so that they win their own "war" against another multi trillion dollar company?
Please don't get me wrong, I know the question can seem a bit negative, I am really just curious.
Google's biggest and most important customer for all this AI stuff is google. Do they actually want other customers, or is having other people use their AI just an annoyance at this point, where we use up compute that they'd rather use internally...
HN lives in a bubble.
I have German/Italian/Polish clients virtually all use Gemini and NotebookLM. Talking insurance, banking, consulting, legal.
The real world doesn't look at pointless benchmarks on writing react tailwind crap, they are already google suite users, get the tools, test them and adopt them, end of story.
It's going to be like with angular, never mentioned on the net, widely used in the real world.
Needless to say 3.5 was a disappointment. Curious to see 3.6.
That sounds awful.
For those of us who don't follow the AI hype cycle, what does that have to do with the topic of this thread: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber?
What is Google's recently released AI coding product?
Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public.
edit: looks like benchmarks are up on https://artificialanalysis.ai/models/gemini-3-6-flash. It's solidly middle-of-pack. However, if you want to be most fair to flash, look at the intelligence vs time per task and intelligence vs outputspeed benchmarks. This is a very fast model.
edit 2: I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it's good for. It's very good at frontend (much better than gpt 5.5) and it's fast, so it's a great tool for iteration. I expect 3.6 to be no different.
I have a very price sensitive workload that used to run on flash 2.5 lite - it's deprecated now.
The replacement 3.1 flash lite is a lot more expensive, but now also has a sunset date.
3.5 flash lite is even more expensive.
So the price is rising and you have no choice but to keep paying more and more.
It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
2.5 Flash: $0.3 / $2.5
3.0 Flash: $0.5 / $3
3.5 Flash: $1.5 / $9
3.6 Flash: $1.5 / $7.5
---
2.5 Flash-Lite: $0.1 / $0.4
3.1 Flash-Lite: $0.25 / $1.5
3.5 Flash-Lite: $0.3 / $2.5
As is, they are thoroughly outclassed for most usecases. I will say the one area where i do see Gemini punching above its weight class is in tasks that are effectively "Google this for me" / knowledge stuff. So it does have a role, and I do use it. So while I think Google is still in a strong position overall, they are really stuck as a tier 2 AI player right now with text models. They are tier 1 in bio, images, and video.
That being said with any open model we of course do know the total cost (or estimate)
EDIT: It less less verbose in final output though, but it reasons more.
I assume the optimization comes when you have long-running tasks with many tool calls, and by reasoning more, it reduces the number of tool calls needed.
Given the extremely competitive releases of GLM 5.2 and DeepSeek V4 (both pro and flash), I don't think there'll be appetite for it.
> Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready.
> We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
Hopefully 3.5 Pro is soon, and that Gemini 4 can be here end of year and finally have an updated knowledge cutoff.
Anyone have any good alternatives?
I tested Jules and while the idea is good in theory, I found the model's intelligence to be very lackluster.
It might be overkill features-wise, but there's a free tier and it likely won't be left for dead anytime soon.
https://artificialanalysis.ai/#intelligence-comparison-tabs
Differences in token "density" are accounted for by pricing per task
You can get decent open-weight models now. That's not difficult. The difficulty is 1) running them and 2) compliance.
My company runs Claude on GCP's Vertex AI solution. We're in the US healthcare IT space, so the models need to be from somewhere that American healthcare agencies and companies have traditionally been okay with sourcing code from - which means the US, Canada, and maybe Europe. The stuff that handles PHI/PII must be in the US. The expense of hosting is more of a PITA than most customers want to go through this early in the technology's lifecycle, and intelligence gains are simply a matter of degree for most business tasks.
In theory, we could find some open-weight model (likely from China) for our development agentic work and host it anywhere you can host AI models. We don't, though, and I think Google, OpenAI/Microsoft, and Anthropic see that as the core of their business.
I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next gen TPUs to train Gemini 4, which greatly limits how much intelligence they can increase and forces them to do cost efficiency increases.
That being said, it seems that Gemini is still the best image analysis model, so hopefully 3.6 flash builds on this even more.
1. Their AI efforts are very fundamental research oriented. They are really good at it.
2. Their productization sucks. The end products gets little attention compared to competition. It can be canceled at any time. You should never build anything around Google only APIs, AI or not.
Fable 5 still wins on detail with no visible errors, but it's close. And this isn't a memorized pelican;
https://playcode.io/blog/macbook-svg-benchmark#gemini-3-6-fl...
https://artificialanalysis.ai/models/gemini-3-6-flash?intell...
artificialanalysis.ai has it going from 165 tps -> 304 tps. openrouter.ai needs more data but it has it going from ~100 tps -> ~150 tps, though at peak 3.5 has reached 156tps.
Ever frontier lab lived it at least once : missing the frontier by a few months triggers extremly negative reactions, then you take back the lead for 2 weeks, and the hype cycle repeats.
Rough patch for google ai
Claude Code's success is not due to the agent but because the model is considered the best for programming and is very heavily subsidized, compared to pay as you go API prices. Consumers and Enterprise are not really locked in and will go where it makes the most sense.
I think they have almost no loyalty by actual developers.
Basically you may choose to drink brand A water bottle, brand B water bottle or tap water. Oh and you might choose the glass water bottle if you use API/Fable.
It's not true, it's just a play for margin.
Where does that 90% figure come from?
For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.
My view then was they are optimising the models for inference ability on their own hardware AND use cases, which is often speed and time to first token.
They've somehow seemed to end up with terrible compute shortages, which again is surprising given how good Google is at infra deployments AND have their own hardware. From rumors out there they are turning down enterprise deals for Gemini because they don't have the compute.
The problem is they're falling further and further behind on frontier class on coding especially, and since I wrote that article it's got even worse with open weights models undercutting them on price AND intelligence.
Speed as a differentiator has always been Google's thing. They (used to?) show the microseconds it took to query & rank web-scale search results. Chrome, notoriously, focused on speed at the expense of resource use. The very many efforts to efficiently speed up Android & its runtime since its inception, and so on...
> their big model underperforms chatgpt 5.6
Possible but TFA claims:
We have started our most ambitious pre-training run yet, for Gemini 4 ...It would be a shame if they cannot beat Kimi K3 or Qwen3.8 Max, both of which are claimed to be Fable-like. If that is true, it will be [or would be] the first time a major American lab falls behind a Chinese competitor.
Fast, light weight, ok intelligence. Perfect for serving 20B+ prompts per day mostly surrounding banal human things.
OAI and Anthropic's cloud spend can cover the revenue gap, as Google is already capturing a large chunk of those guy's revenue.
https://huggingface.co/microsoft/bitnet-embedding-0.6b
It’s a small multilingual embedding model designed for things like search, RAG, and semantic similarity. It supports a fairly large context window and is designed to run efficiently on a CPU in a GPU starved world.
The interesting part is that it builds on BitNet, using ternary weights of -1, 0, and 1 instead of the usual floating-point weights. That should make indexing and searching large amounts of text much cheaper without giving up too much accuracy.
From the outside they look like they're behind in terms of frontier models, but I think they might be the best positioned to not go out of business when the bubble pops.
Also look at the fact that they've been able to deploy AI-assisted search at google scale. It must be another order of magnitude larger (at least) than the model deployments for OpenAI and Anthropic.
Of course unless you're inside Google it's impossible to know for sure.
Yes, subs like codex are heavily subsidized. But API billing has massive margins and that's what enterprises pay.
Rumors say 4) it didn't perform well, especially in coding so has been delayed
Features stay in Beta for ages, whatever that actually means, and released ones get deprecated things fast.
Where some of the competitions treats deprecating entire services as "let’s not put it on your frontpage, put deprecation notices all over the doc, and politely ask new users not to start new project with them".
I presume you can't use deepseek?
google's inability or unwillingness to provide stable timelines for model deprecation makes it risky to build complex workflows using their models
For my use case, `gemini-3.1-flash-lite` is ~20% higher accuracy than the next best model of comparable cost (considering both proprietary and open-weight alternatives)
we are switching to Deepseek.
These days no company even has completion models where one controls the text input fully. Worthless.
You can also just write code like you did a year or two ago.
As far as I can tell it's slightly better than GLM 5.2.
GLM defaults to max effort btw
https://docs.together.ai/docs/glm-5.2-quickstart#reasoning-e...
Light. Lite is product marketing seepage.
They're easy enough to skip - click the little "-" icon and you'll collapse the entire sub-thread.
All of the pelicans so far have had really weird flaws / quirks so I am always a little interested to see how well these models perform at this task, since I've seen all the past pelicans and have some anchoring.
Seeing a truly flawless pelican would tell me that the model has true visual reasoning capabilities as well as good taste.
The 3.6 Flash pelican is just about the best I've seen.
Not only does it give you a super easy-to-grok understanding of the model quality just by looking at the image, but when you compare tokens and costs (both input and output), you really get a good, simple COST x QUALITY evaluation across models.
Simon explains it well: https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l...
Simon, you should put up a summary table page that you update after every release.
Sponsored blogs and paid newsletters are after all, notoriously poor at subsisting on silence :)
The fish and the cap where always added when I asked an llm to improve it's first attempt.
This continues the trend in LLM progress of better=more stuff
Although, I think saying 'wanted them to win' was not accurate. More like, I stuck with them hoping it will get better, and it did get better in many ways, coding was not one of them.
but the implementation will be up to your provider and harness, for deepseek, they expose some numbers: https://api-docs.deepseek.com/guides/kv_cache/ and Anthropic has a list of actions invalidating your cache: https://platform.claude.com/docs/en/build-with-claude/prompt...
Basically, you avoid anything dynamic: model change, tool change, etc it's also important that your system prompt or main prompt doesn't have non-static data like the date/time/place or someone's name (the person you interact with in a chatbot for example). That should be left to tool call or search.
If you’re constructing the prompt you don’t have to jam everything together you can arrange it appropriately.
All of the models, you need to have a consistent input to get the cache hit. So if you are chatting with a document, and change the system prompt, it will be a cache miss, even if the rest of the items are all the same. If you even pass in the document in not the same order as the prompts, it will be a cache miss. Or if you add tool calls or structured outputs, it will be a cache miss. (Since those generally go at the beginning of the prompt call, not at the end.)
Most of the time when reading documents from URLs directly it will never cache. (Need to typically pass in the bytes directly, or use the provider document store index.)
Gemini has a 4096 minimum token size with the 3 version models before even getting a cache hit. OpenAI it is lower (1024), and is automatic, but only happens in increments of 124. Anthropic can also get cache hits at 1024 tokens, but you need to explicit ask for it (and pay extra).
Caching by default typically lives for 5 minutes since the last cache hit across providers. But some of them you can ask for longer. AWS for Anthropic models can be tricky with multiple endpoint routing, so can get cache misses if it happens to route to a different endpoint.
They are not anywhere close according to pretty much every benchmark (even v4-flash is considerably ahead and its way cheaper than flash-lite). Maybe tuning prompts/tools/etc. might be useful?
I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.).
I have no idea why the Gemini models do so well.
I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemini Flash models leading in accuracy). I made a more complex coding/tool-usage test, that I expected it to fail, it did fail it once locally in my debug tests, but when I finalized the test and ran the entire testing suite for all models, somehow Gemini 3 Flash still got it right...
Gemini models are REALLY intelligent (and they are actually my favorite model to use via the chat app to ask questions), but they somehow fail in real-word coding tasks where they have to modify files, check results, debug, etc.
My tests harness provides a lot of mock data, and limits the number of actions a model can choose from. I am starting to think that maybe the models are not bad, just that the coding harness are not optimized for those type of models, and Google doesn't really provide their own "Codex".
But why is this an indication of literally anything else?
I agree to rednb that at this point it feels like rather obvious brand building, but also, I agree with you that some value is in it.
It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
> It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
I hope SimonW keeps them coming.
I get people burning out on the pelican SVG test alongside the rest of the AI burnout, but I guess for myself I'm just choosing to keep enjoying it while I still can.
If it weren't for the $10 GCP credit, I'd straight away cancel it. I don't see enough value in Gemini to justify the $20 subscription.
Build a whole new management tree - the current people all do a terrible job.
I have multiple anthropic and OpenAI max plans. For Gemini I just use my Cursor $200 a month plan (which also gives me the ability to try grok, conductor, etc)
It's a decent heuristic because the better models generate better pelicans. That's all. Nobody sane is going to make a bet on a model based on a pelican. But it's cool, it's tradition by now, and it's a semblance of a good first impression for new models.
thank you.