Y'all, the AI market did something interesting this week. Instead of everybody racing in the same direction on price, the big players started splitting up. Some are getting cheaper. Some are getting more expensive. And one of them just figured out how to go stupid fast. Let's get into it.
Google Just Cut Coding AI Prices in Half
Google dropped Gemini 3.7 Flash this week and called it their smartest workhorse model yet for coding and agent work. Big jumps on the benchmarks that actually matter, one of them went from 34.4% to 43.6% on a coding test called FrontierCode. Another jumped from 49% to 65.3%. That's not a small bump, that's a real leap in three weeks since the last version shipped.
Here's the kicker though. Pricing is $0.75 per million input tokens and $3.75 per million output tokens through the end of the year. Half of what the previous Flash model cost.
Why it matters: this is the model most developers and small shops are actually going to touch, not the flagship show pony everybody talks about on Twitter. When the workhorse model gets better AND cheaper at the same time, that's when adoption actually moves. Cheap coding agents stop being a toy and start being something you build a business process around.
My take: I've said it before, the model wars stopped being about who's smartest a while back. It's about who's cheapest per unit of actual work done. Google playing the volume game on their bread and butter model is smart business. If you're building anything with AI coding agents right now, go check the pricing on whatever you're using. There's a real chance you're overpaying.
OpenAI and Anthropic Are Cutting Prices, DeepSeek Just Raised Theirs 1,100%
This one's wild. OpenAI cut pricing on GPT-5.6 Luna substantially. Anthropic priced Claude Opus 5 at about half of what their top end Fable 5 model costs. Meanwhile DeepSeek, the company that made its name by being dirt cheap, just rolled out V4 Pro and raised prices on some workloads by more than 1,000%.
DeepSeek's new flagship runs up to $1.32 per million input tokens and $3.96 per million output tokens at peak times. Still cheap next to the top shelf American models, but a long way from the "we'll undercut everybody forever" story that made them famous.
Why it matters: for two years the story was cheap Chinese models forcing American labs to compete on price. Now it's flipped a little. The big US labs are cutting prices to fight off that competition, and DeepSeek is realizing they can charge more for their best stuff because people will pay for reliability and performance. Nobody wins a price war forever. Eventually you gotta make money.
My take: don't build your whole business model around whatever's cheapest today. These prices are gonna keep moving around for a while as everybody figures out what the market will actually bear. Pick your model based on what it does for you, and keep half an eye on the invoice, but don't chase the bottom dollar every single month or you'll spend more time migrating than building.
OpenAI's New "Ultrafast" Mode Runs 14 Times Faster
OpenAI put out an early preview called Ultrafast, an API tier that runs their GPT-5.6 Sol model up to 14 times faster than the normal version. It's powered by Cerebras hardware and can spit out up to 750 tokens a second. Right now it's only open to a limited group of API customers.
Why it matters: everybody's been obsessing over how smart these models are and how cheap they are. Speed is the quiet third thing that decides whether AI actually works inside real products. A model that thinks great but takes eight seconds to answer is useless for a live customer chat, a trading tool, or an agent that needs to react in real time. Fast changes what's even possible to build.
My take: this is the one to watch longer term. Once you can run inference at 750 tokens a second without losing quality, whole categories of products open up that just weren't practical before. Real time voice agents, live coding copilots that don't make you wait, autonomous agents that can actually react instead of think-then-act in slow motion. Specialized inference chips like Cerebras are becoming just as big a deal as the model itself. Keep an eye on who else jumps on this train.
That's it for today. Prices splitting two directions and speed becoming the new battleground, that's where AI's headed this week. Catch y'all tomorrow.