Over the past year, LLM API prices have changed almost every month. Many people assume the price cuts come from model compression, inference framework optimization, or vendors burning cash to grab market share. Those factors certainly exist, but after running the numbers internally, we found that what truly determines the long-term price floor may be something developers rarely discuss — electricity bills.
Simply put, LLM inference is a business of converting electricity into tokens. Whoever can do this conversion solidly with cheaper electricity and higher scheduling efficiency will be the one who lasts to the end in the price war. This is also the logic behind positioning platforms like Silicon Carbon Phase Change around green computing power.
How Much of Inference Cost Is Actually Electricity
Training a large model is a one-time investment; inference burns money every single day. A mainstream inference card draws between 300 and 700 watts at full load. For a mid-sized online serving cluster running several hundred cards simultaneously, the daily electricity bill runs into five figures. Add data center cooling, and actual energy consumption rises another 30 to 50 percent. The industry consensus is that electricity plus cooling accounts for roughly 30 percent of inference service operating costs, and the larger the scale, the more sensitive that ratio becomes.
Model capability matters, of course, but capability can be caught up with. Today you're half a version ahead; three months later, someone else catches up. Electricity is different — it's determined by physical location and energy structure, and it's very hard to change in the short term. That's why I place the next stage's deciding factor on electricity costs.
Where the East-West Computing Power Math Diverges
We did a rough internal calculation. Take the same batch of inference tasks: placed in a data center in a Tier-1 city in the east, with industrial electricity rates plus cooling overhead, the all-in cost per kilowatt-hour approaches one yuan. Placed in a computing center in the west where wind and solar resources are concentrated, with direct green power supply and favorable natural cooling conditions, the all-in cost per kilowatt-hour can be pushed down to half or even less. These aren't numbers from policy documents — we estimated them from actual quotes and PUE.
The difference comes from two things. First, energy structure: the west has more installed wind and solar capacity, so green power procurement prices are inherently lower. Second, climate: in places with low average annual temperatures, cooling electricity can save a considerable amount. Stack these two together, and the compute cost gap per token emerges. Silicon Carbon Phase Change has computing nodes in both the east and west. When scheduling, it prioritizes deferrable batch inference tasks to nodes with abundant green power — and the impact of this move on cost is bigger than many people imagine.
How Green Computing Power Translates into Lower API Prices
Lower electricity bills don't automatically mean cheaper APIs — there's a layer of procurement and scheduling in between. The logic goes like this: the computing center bulk-purchases green power at a unit price below market retail; the platform then aggregates scattered inference demand to fill that compute capacity and dilute unit costs; finally, it sells to developers in the form of LLM APIs. Bulk procurement plus green power cost reduction — two layers stacked — and only then is there room for prices to go lower.
This is also one of the reasons AI API aggregation platforms exist. Developers going it alone to connect directly to each model provider can't get bulk pricing, nor can they utilize idle compute capacity. After aggregation, the cost structure of token purchases is completely different. When we tested Silicon Carbon Phase Change's multi-model routing, we noticed that for the same request volume, the all-in cost after scheduling was noticeably lower than connecting directly to each provider — and the gap came mainly from compute-side optimization, not the models themselves.
What This Means for Developers
The most direct impact is that API prices will continue to fall, but the pace will diverge. Price cuts propped up by cash-burning subsidies aren't sustainable; price cuts supported by electricity costs and scheduling efficiency are stable. When choosing a provider, besides looking at model quality and latency, it's worth asking one more question: where does the compute behind this price come from?
Second is stability. Green power is volatile — wind and solar output fluctuates — and a good scheduling system uses energy storage and cross-regional scheduling to shave peaks and fill valleys. Not every provider has this capability, and it determines whether your requests get rate-limited during peak hours.
One pitfall to avoid: don't just look at the listed price. Some low-priced APIs downgrade to smaller models during peak hours or queue requests for processing, so the actual experience doesn't match the listed price. When choosing an AI API gateway, testing peak-hour latency and success rates together is more meaningful than simply comparing prices.
The LLM API price war has reached a point where the gap in model capability is shrinking, while the gap in compute power and electricity is widening. The winners of the next stage will most likely be the platforms that can push green power and scheduling efficiency to the extreme. If you want to keep discussing specific approaches to multi-model integration and cost optimization, feel free to check out the previous articles I've written.