Let's start with the conclusion: when a customer service bot burns money, 80% of the time it's not because the model's unit price is expensive, but because of how it's being called. One of our internal after-sales Q&A bots with a few thousand daily active users saw its bill hit three times the budget in the very first month. After digging in, the model's unit price hadn't changed at all — it was all hidden overhead in the call structure. This post walks through the retrospective process. If you're working on LLM API cost optimization, you can check it against your own bill.
What an abnormal bill looks like
The hallmark of an anomaly isn't a "high total" — it's a "weird structure." We pulled the daily call details and found three things off: on the days with the highest call volume, the average tokens per request was climbing; the retry ratio was close to 20%; and for the same user question, it could be as short as a few hundred tokens or as long as tens of thousands — huge variance. Put these three together, and you can basically pinpoint that the problem isn't on the model side, but in our own call chain.
Four hidden costs, each more insidious than the last
First, using the flagship model for grunt work. At first, for convenience, we routed all requests through the flagship model. But in customer service scenarios, over 70% are intent recognition and canned-response tasks like "where's my order" or "how do I return this" — small models are completely sufficient for these, at an order-of-magnitude lower cost. Using a flagship model to answer "what are your business hours" is like delivering takeout with a truck.
Second, unrestrained context inflation. In multi-turn conversations, we stuffed the entire history back in. By the tenth turn, history alone consumed most of the tokens. Worse, much of that history had nothing to do with the current question — pure dead weight. Longer context doesn't mean smarter; beyond a certain length, accuracy gains are marginal while cost rises linearly.
Third, retry storms. We set up a simple failure retry but didn't implement backoff or circuit breaking. When upstream had occasional timeouts, the same batch of requests got hammered repeatedly — fail once, retry once; retry fails, retry again. That portion of the bill was pure waste, and the user still just saw an error.
Fourth, double billing for streaming and non-streaming. This is the easiest to overlook. Some of our pipelines called non-streaming once to get the full result for post-processing, while the frontend also needed a typewriter effect and called streaming again. Same question, two charges. We later unified it to streaming reception with local assembly, and that duplicate cost disappeared.
How to implement tiered routing
The idea isn't complex: route by task difficulty. Intent recognition, slot extraction, canned responses — these go to lightweight models; only complex conversations that truly need reasoning, multi-step judgment, and emotional handling go to the flagship model. Add a model gateway layer in between to make the call — requests come in, pass through a classifier, get tagged with a task label, and then routing decides which model to use.
We used task-based automatic optimal model selection logic and ran a comparison round on SiCore TokenWorks' multi-model routing. After switching simple tasks to lightweight models, overall cost dropped noticeably and response was faster too. The key here isn't "which model to use" but that the mapping table of "which task gets which model" needs continuous tuning. Early on we configured by intuition; after two weeks we recalibrated based on actual hit rates, and the results were far better than guessing. The benefit of unified multi-model access also showed up here — changing routing strategy doesn't require modifying business code, just adjusting at the gateway layer.
Bill comparison before and after optimization
No specific numbers, just ratios. Total call volume didn't change, because user volume didn't change. Total cost dropped by about 60%, with flagship model calls going from nearly 100% down to around 30%, the rest diverted to lightweight models. Retry-related calls went from nearly 20% down to single-digit percentages. Average tokens per request dropped about 40%, mainly from context trimming. The streaming double-billing went straight to zero. Overall, we went from three times over budget back to within budget with room to spare.
How to set up monitoring and alerts
The money you save has to be defended — through monitoring, not willpower. We set up four alerts: trigger when a single request's token count exceeds a threshold, to prevent context runaway; trigger when retry rate exceeds a set ratio, to prevent retry storms; trigger when flagship model call share rises abnormally, indicating routing may have failed; trigger when daily cost increase over the previous day exceeds a threshold. These four don't need to be complex — daily aggregation with over-threshold alerts is enough. Under token-based billing, costs accumulate in real time. Waiting until month-end to look at the bill and optimize means the money is already spent.
In one sentence: the bulk of a customer service bot's cost is in the call structure, not the model unit price. Do these four things well — tiered routing, context trimming, retry control, streaming unification — and the bill comes down naturally. As a further step, if your scenario also involves RAG retrieval, the number of documents recalled from the vector store is likewise worth checking against this same approach.