Article Brief

Key Takeaways

5 Points30s Read

  1. AvailabilityIronwood reached general availability on April 22, 2026, but Google publishes no on-demand price for it, while Trillium (v6e) and older TPUs still carry public rates.
  2. The estimateThe only Ironwood price in circulation is a SemiAnalysis estimate of about $1.60 per chip-hour for Anthropic’s negotiated capacity — a contract figure, not a list price.
  3. The hardwareEach Ironwood chip pairs 192 GB of HBM with 4,614 FP8 teraFLOPS; a full pod links 9,216 chips and 1.77 petabytes of shared memory.
  4. The squeezeWithout a rate card, smaller AI teams cannot model inference unit costs or compare Ironwood against Nvidia GPUs that do list prices.
  5. The signalWhether Ironwood ever appears on Google’s pricing page with an on-demand rate will show if Google means to sell its best inference silicon by list price or by private contract.

Pricing figures in this article include third-party estimates and list rates current as of July 2026; cloud and hardware pricing changes frequently. This is informational analysis, not purchasing or investment advice.

Google moved its most powerful AI accelerator into general availability on April 22. Three months later, you can rent an Ironwood tensor processing unit on Google Cloud — but you cannot look up what an hour of it costs. There is no rate on the public pricing page, no self-serve quote, no per-chip figure sitting beside the numbers Google still lists for every older TPU generation. The hardware reached production faster than its price did, and the gap is not an oversight.

Google built Ironwood for the serving side of AI — the phase where a trained model answers real user requests rather than the phase where it learns — and pitched it as the company’s first TPU designed for that job at scale. For a chip aimed squarely at inference, the missing number is the most revealing thing about it. A price is how a cloud market sets expectations, and cost-per-token is the metric that decides whether an AI product clears its own serving bill. Google has chosen to publish none of it.

A price list with one missing line

Open Google Cloud’s TPU documentation and the older generations behave normally. Trillium, the sixth-generation v6e, carries a published rate: about $2.70 per chip-hour on demand, dropping to $1.89 with a one-year commitment and $1.22 across three years. The fifth-generation v5p is listed too. Ironwood — which Google’s own documentation files under the name TPU7x — shows up with a full specification sheet and no rate at all.

The specification is not modest. Each Ironwood chip is a dual-chiplet design carrying two TensorCores and four SparseCores, 192 GB of high-bandwidth memory and roughly 7.37 TB/s of bandwidth. Google rates it at 4,614 FP8 teraFLOPS per chip, and a single pod links 9,216 of them through a three-dimensional torus interconnect with 1.77 petabytes of shared memory. That memory figure is the part inference buyers care about most: serving a large model is a memory-bandwidth problem before it is a raw-compute one, because every generated token has to stream the model’s weights and its growing key-value cache through the chip. A part built to move that much memory that fast is aimed at exactly the workload where a per-token price would be easiest to compare — and hardest for Google to take back once posted. Google rates the chip at roughly ten times the peak performance of TPU v5p and four times the per-chip performance of Trillium. The fastest tensor processing unit Google sells, in other words, is the one it declines to quote. Every slower part has a sticker; the flagship has a specification and a phone number.

That inversion is unusual even by cloud-pricing standards, where opacity is common but total silence is not. It is also a change of posture from Google’s earlier TPU push against Nvidia, which leaned on public performance-per-dollar claims to make the case for renting Google silicon instead of buying into the GPU stack.

What $1.60 an hour implies

There is exactly one Ironwood price in circulation, and it is an estimate. SemiAnalysis figures that Anthropic pays roughly $1.60 per chip-hour for Ironwood capacity on Google Cloud. Read against Trillium’s $2.70 on-demand sticker, that number tells you what it is: a negotiated hyperscale rate, not a retail one. A customer buying enough capacity, early enough, gets the newest and fastest chip below the list price of the older one.

The contrast with Nvidia sharpens the point. A B200 on a neocloud runs on the order of $9.36 per GPU-hour on demand, or about $5.35 on spot, which works out to roughly $2.60 per million tokens for a Llama-70B FP8 workload before you tune anything. Those are different accelerators with different memory and different software, so the comparison is not one-to-one. But the numbers exist. A buyer can pull them up, plug them into a model, and decide. Ironwood offers no equivalent starting point unless you are already in a sales conversation.

The people who do see Ironwood’s real economics are the ones large enough to negotiate for them. Anthropic is one; the hyperscale AI labs committing to multi-year capacity are the rest. That is the same dynamic visible in the economics of serving open-weight models, where the headline cost-per-call and the real cost of running the thing at scale diverge sharply, and where the gap is widest for whoever cannot commit up front. What Ironwood actually delivers in production is not a mystery — the 4.7x speedups Google has shown running Qwen on the chip are documented. The cost of getting them is the part left blank.

Why the silence works in Google’s favor

Withholding a price is a decision, and in a capacity-constrained ramp it is a rational one. When supply is short, a public on-demand rate does two unhelpful things: it invites more demand than Google can serve, and it anchors expectations Google may not want to defend later. Allocating scarce Ironwood to a handful of committed buyers on private contracts avoids both problems and adds a third advantage — it keeps competitive intelligence away from Nvidia and Amazon, who would happily reverse-engineer Google’s cost-per-token from a posted rate.

Structure reinforces the choice. Nvidia’s chips scatter across dozens of clouds and resellers, so street prices leak whether Nvidia likes it or not. Google keeps TPUs inside its own cloud, which means Google controls what the market gets to know. The same instinct shows up across the industry’s scramble over compute cost and supply, where the terms that matter most are the ones negotiated privately rather than posted. Google marketed Ironwood as its first TPU designed for inference at scale, the workload where cost-per-token decides whether an AI product has a business. A headline price would let every rival benchmark that number directly. Silence keeps it a private conversation.

The number a mid-size buyer actually needs

For the customer Google is not courting, the effect is concrete. A startup deciding where to serve its model needs a unit cost — dollars per million tokens, or per chip-hour — to build even a rough margin model. Ironwood’s absence from the rate card means that team cannot run the comparison. It has to either open a sales channel and wait, or default to an accelerator that publishes its price, which in practice means Nvidia GPUs on a neocloud or a competitor’s listed instances.

Walk the math a mid-size team would actually do. Suppose it serves 500 million tokens a day and wants to know whether Ironwood beats a GPU it can already price. On the Nvidia side it can estimate: at roughly $2.60 per million tokens on-demand for a comparable open model, that workload lands near $1,300 a day before optimization, a figure the team can pressure-test against its revenue. On the Ironwood side there is no on-demand rate to enter into the same spreadsheet — only a contract rate reported for a customer thousands of times its size. The comparison the team needs cannot be built, so the decision defaults to the platform that answered the question.

Opacity, then, functions as a filter. It routes the long tail of AI builders toward priced alternatives and reserves Ironwood for the customers who move first and commit large. That may be exactly what Google wants during a supply crunch, but it also cedes the smaller end of the inference market — the developers experimenting today who become the committed buyers of 2027 — to whoever will simply tell them what a run costs. The same buildout pressure driving the current data-center arms race is what makes that early-adopter cohort worth capturing, and a quote-only product is a strange way to capture it.

What would change this read

The estimate deserves its caveats. The $1.60 figure is SemiAnalysis’s, not Google’s, and Google has not confirmed it; negotiated rates move with commitment size and contract term, so one customer’s number is not a market price. Google could also list an Ironwood on-demand rate tomorrow — it eventually published prices for every prior TPU — at which point this stops being a story and becomes ordinary cloud pricing.

It is equally possible that general-availability capacity is effectively spoken for, which would make a public retail price close to meaningless in the near term. Google is not hiding everything, either: the Axion CPUs launched alongside Ironwood already carry partial public pricing, so the blackout is specific to the flagship accelerator rather than a blanket policy. The signal worth watching is narrow. If Ironwood eventually appears on the pricing page with an on-demand number, the market normalizes. If it stays quote-only through the back half of 2026, that confirms something more durable — that Google has decided its best inference silicon is a contract good, sold to buyers it chooses, at prices only they get to see.

A chip’s price is a small thing next to 4.6 petaFLOPS of compute. It is also the one figure that tells every other buyer where they stand. For three months, on its most important chip, Google’s answer has been to say nothing.