Three things caught my eye today: a wild new AI chip, a leadership shakeup at Google, and AI prices dropping across the board. Let's get into it.
AMD Buys Taalas: Baking AI Models Straight Into Silicon
AMD dropped a big one after the market closed Thursday. They're buying a Toronto startup called Taalas that does something wild: instead of storing an AI model's weights in memory next to the chip, they etch those weights permanently into the silicon itself. The weights aren't stored near the compute cores. They are the compute cores.
Taalas ran a test chip serving Meta's Llama 3.1 8B model at almost 17,000 tokens per second, using about a tenth of the power an Nvidia H200 needs to do the same job. That's not a small improvement. That's a different category of chip.
Why it matters: the AI industry has been fighting the same bottleneck for years. GPUs spend more time shuffling data in and out of memory than actually computing. Skip that step by hard-wiring the model into the chip and you blow past the ceiling everybody else is stuck under. AMD folding this into its Helios systems next to Instinct GPUs tells you they're serious about owning inference, not just training.
Robert's take: this is the kind of move that doesn't pay off for a year or two but matters a lot when it does. Nvidia bought Groq for $20 billion seven months back for basically the same reason. Everybody sees inference cost as the next battlefield, because that's where the real money gets made once models stop changing every six weeks. I like this bet from AMD. Baking one specific model into a chip is inflexible as heck if you need to swap models, but for high volume, single purpose serving, it's hard to beat. Watch for more of these small chip startups getting scooped up before the year is out.
Google Shuffles Its AI Leadership Deck
Bloomberg reported Google is consolidating its AI leadership back in Mountain View. Koray Kavukcuoglu is taking the wheel on research and operations, and Demis Hassabis is stepping back from day to day work to become chairman of DeepMind and Alphabet's Chief Scientist.
Why it matters: this comes right after Google lost some serious talent this year. Noam Shazeer went to OpenAI, and John Jumper, a Nobel Prize winner, went to Anthropic. When a company reorganizes leadership like this while also losing top researchers to competitors, that tells you there's real pressure behind the scenes to move faster.
Robert's take: Hassabis moving to a chairman role isn't retirement, it's a promotion dressed up as a step back. But putting operations under one person in one location usually means somebody up top decided things were too spread out and too slow. Google's got the compute, the talent, and the data. What they've struggled with is shipping fast enough to keep pace with Anthropic and OpenAI. This reorg is a bet that fixing the org chart fixes the speed problem. Sometimes it does. Sometimes it just moves the bottleneck somewhere else.
The AI Price War Just Got Real
Google cut its entry level AI Plus subscription from $7.99 a month down to $4.99. Word is OpenAI is looking at slashing its own token pricing to protect its enterprise business from Anthropic eating its lunch.
Why it matters: for a while the big labs were competing mostly on capability, whoever had the smartest model won the headlines. Now they're competing on price too, which means the market is maturing. Cheaper AI is great news for anybody building on top of these models or just using a chatbot day to day.
Robert's take: I like this for the little guy. If you're a small business or solo developer building something on these APIs, watch the pricing pages closely over the next few months, because I think we're only getting started. But don't get too comfortable. Nobody's making real money in this price war yet, and eventually somebody has to. When that happens, prices go back up or free tiers get squeezed. Enjoy the cheap tokens while they last.