Pular para o conteúdo
← Back to Skalablog

Published article

How To Use Gemini 3.7 Flash Without Overpaying

Software EngineeringGeminiAnthropicClaude

Gemini 3.7 Flash is a strong cheap coding model with a time-limited price: 75 cents per million tokens in, 3.75 out, through 31 December 2026. After that, the input line doubles. Everything else in this article explains whether the model itself justifies the switch.

What Gemini 3.7 Flash costs, and what the price actually is

Gemini 3.7 Flash, released by Google on 13 August, is a flash-tier coding model priced at 75 cents per million input tokens and 3.75 dollars per million output tokens, according to the Gemini pricing documentation cited in the video. That is under half of what the video says Anthropic OpenAI charge for comparable models at their workhorse tiers.

The catch sits in the same pricing table. The video reports that Google's own page shows the 75-cent input rate running only through 31 December 2026, with the rate doubling to 1.50 dollars on 1 January 2027. The output line doubles on the same date. That makes the discount a promotion of roughly 140 days, with the expiry printed by Google itself.

For comparison, the video notes that Anthropic introductory pricing on Claude Sonnet 5 at 2 dollars in and 10 out, later raising it, and that Anthropic documentation now says a planned September increase will not occur. One lab removed its countdown; Google's is still running. Treat launch-day numbers as trial offers until the schedule says otherwise.

If you are experimenting rather than serving production traffic, the video points out that flash-class models are free in Google AI Studio, rate-limited but free. The free tier has a currency of its own: Google's pricing table states that prompts on the free tier are used to improve Google's products, while paid-tier prompts are not.

How strong are the coding benchmark claims?

Google's launch claims, as relayed in the video, are aggressive: the company states in its own comparison card that the model beats comparable Anthropic OpenAI models across nine separate benchmarks. The card itself, however, covers 20 benchmarks, and the video counts 10 of them going to competitors.

The headline movements the video cites are large. On what it calls a frontier coding benchmark, the score moves from 34.4 to 43.6. On a long-horizon software engineering test, it climbs from 49% to 65.3%. On the web development leaderboard run by LMArena, which the video calls Code Arena, the model lands at 1588 Elo, ahead of Claude 5 and other rivals.

Two cautions apply. First, these figures come from the video and Google's launch materials; treat them as vendor-reported until reproduced independently. Second, benchmark wins at the flash tier do not settle the frontier question, which the next section covers.

Why is Google's real flagship still missing?

The same day the flash model launched, the video reports, Forbes ran a story titled 'Gemini 3.5 Pro delay continues'. Google announced that Pro-class model on stage at its developer conference in May for rollout the following month, then missed June, missed July, and by the launch date had missed three deadlines.

Forbes attributed the delay to coding and reliability work, senior researcher departures, and the possibility of restarting the pre-training run. The video is explicit that Google has confirmed none of those causes. What Google did confirm, via Sundar Pichai on the July earnings call, is that the model was still in testing while the team built the next generation.

The last Pro-class model actually shipped, per the video, was 3.1 Pro on 19 February, leaving roughly 175 days without a new flagship while three flash models went out the door. Google's answer at the top tier is, for now, a model it has failed to ship three times.

The flash model is beating Google's own Pro model

The strangest scoreboard entry the video cites is internal. On the web development leaderboard, Google's own 3.1 Pro preview sits at rank 46 with an Elo of 1447, while the new flash sits at rank eight with 1588, a score the board still marks preliminary. That is 141 Elo points in favor of the cheap model.

Google's cheapest model is beating Google's most expensive shipped model at the exact job this article is about. The delay therefore becomes a decision, and Pichai stated it on the same earnings call: releasing models at close to a monthly cadence is part of the roadmap while Gemini 4 is built.

Tulsi Doshi, who leads product for Gemini, described the upgrade in agent terms, saying it better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity. That matters because flash-tier models are where agent tokens are actually spent, in loops that run all day.

Is this a price war or a token-subscription strategy?

The falling sticker prices hide rising bills, and the video offers the mechanism: token counts are multiplying faster than prices fall. It cites Anthropic own research putting a multi-agent system at roughly 15 times the tokens of a simple chat, and work from Microsoft and Stanford's Digital Economy Lab putting agentic tasks near a thousand times a chat message.

Two figures, two orders of magnitude apart, one direction. A price cut at the workhorse tier is an invitation to spend more tokens, and Google's own meter reflects that the invitation works. The video reports Google's model APIs growing from 16 billion tokens per minute to 22 billion in a single quarter, about 37% in three months.

Competitors are running variations on the same play. DeepSeek moves to peak and off-peak billing on 16 August, which the video says raises cached input costs about 12 times during peak hours. Meta lists a contributor tier roughly 16 times below its standard rate, in exchange for permission to train on your traffic. Cheap has a currency, and sometimes it is your data.

Where does Google actually stand against Anthropic, OpenAI and Chinese labs?

The honest answer is: at the top and at the workhorse tier, two different stories. On the web development leaderboard cited in the video, Claude maxes at 1691, followed by two Chinese models at 1674 and 1669, with four more entries before Google appears at rank eight with a flash model. If your work needs the best model available, the video's pick is Claude 5, and two of the three top names are Chinese.

On consumed traffic, the picture reads worse for Google. The video reports that the share of OpenRouter traffic going to American models fell from about 70% to about 30% in the 12 months to June, and that Google does not appear in that platform's ten most-used models.

At the workhorse tier, though, where most agent tokens are spent, the video's verdict is that Google is not behind anybody this week. What got it there is a 23-day release cycle rather than a benchmark table: three flash models at one tier in 86 days, each cheaper on the output line than the last.

Should builders switch to Gemini 3.7 Flash now?

For a builder shipping product against a deadline, the video's call is yes: 3.7 Flash is the model to reach for today, and the author would still reach for it at double the money, which is what it becomes on 1 January. The concession is the frontier. Cheap does not mean best-in-class, and the top of the market belongs to someone else this quarter.

The practical checklist before switching:

  • Confirm the current rate on Google's pricing documentation before budgeting, since the promotional window ends 31 December 2026.
  • Treat leaderboard scores as vendor-reported and pilot the model on your own agent workload.
  • Remember that the free tier in Google AI Studio trains on your prompts; paid prompts do not.
  • Plan for the January doubling now, and revisit if Google follows Anthropic making the promotional price permanent.

FAQ

  • What is Gemini 3.7 Flash priced at? Per the video's reading of Google's pricing documentation, 75 cents per million input tokens and 3.75 dollars per million output tokens through 31 December 2026, doubling on 1 January 2027.
  • Is the half price permanent? No. Google's own table prints the expiry date. Anthropic made its Sonnet 5 introductory price permanent; Google's countdown is still running, though the video bets Google will not let the January doubling stand.
  • Does Gemini 3.7 Flash beat Claude for coding? At the workhorse tier, per the benchmarks cited in the video, it is competitive and cheaper. At the frontier, the video names Claude 5 as the best available model, with two Chinese models close behind.
  • Why does my bill keep rising if model prices fall? Agentic workflows consume far more tokens than chat. Anthropic research puts multi-agent systems at roughly 15 times chat token use, and Microsoft-Stanford work near a thousand times, so volume outgrows the discount.
  • Is the free tier of Google AI Studio really free? It is rate-limited but free, with a trade-off stated in Google's own pricing table: free-tier prompts are used to improve Google's products; paid-tier prompts are not.
  • What happened to Gemini 3.5 Pro? Per the video, it has missed three deadlines since being announced in May, with Forbes reporting coding and reliability work and researcher departures as causes, none confirmed by Google.
  • How fast is Google shipping flash models? The video counts three flash releases at one tier in 86 days, in May, July and August, a cadence Pichai described on the July earnings call as part of the roadmap while Gemini 4 is built.
  • Is the 1588 Elo web development score reliable? It comes from the leaderboard the video calls Code Arena and is still marked preliminary there, so treat it as directional rather than settled.
  • Who should not switch to Gemini 3.7 Flash? Teams that need the best model on Earth regardless of price should stay at the frontier tier, where Google's answer, 3.5 Pro, has failed to ship three times per the video.

From launch videos to lasting write-ups

The central lesson of the Gemini 3.7 Flash story is that the fine print outlives the launch hype: a price with an expiry date, a benchmark with a preliminary tag, a flagship with three missed deadlines. The same is true of your own knowledge. If you have explained a pricing table, a model release or a migration in a YouTube video, that analysis deserves a written form readers can find and quote.

Skala Blog turns a video URL into a structured article: paste the link, transcribe the talk, and generate a draft you control. For more developer writing and tools, see Dev doido and the Crazystack typescript resources at crazystack.com.br.

Source video