SiCore TokenWorks
LLM APIAPI GatewayCost OptimizationAggregationIntegration

token8341 Engineer in Practice: Building a Smart Customer Service Prototype in One Week, Notes on Pitfalls in Multi-Model API Comparison

SiCore TokenWorks Team·2026-10-05

Last week I took on a job: building a smart customer service prototype for a team that makes a SaaS ticketing system. The requirement was to get it running within a week, and also to do a side-by-side comparison of the response quality of four models: GPT-4o, DeepSeek-V3, Qwen-Max, and Doubao. Sounds easy enough, but once I actually started, I realized that when it comes to unified multi-model access, all the pitfalls are hidden in the details. This article records the process, to save some time for peers who also need to do multi-model comparisons.

Key Management: 5 Platforms, 5 Sets of Backends, Get the Accounts Straight First

The first trouble wasn't writing code, it was managing keys. The four models came from four platforms, plus a backup one, so five backends and five consoles, each with different key formats, quota-checking methods, and rate-limiting rules. Some platforms display keys in plain text directly, while others require creating a sub-account first and then allocating permissions. During the prototype phase, for speed, I stuffed all the keys into a single .env file. As a result, the next day when test volume picked up, one platform's key got rate-limited, and the error message gave no clue which platform was the problem.

Later I switched to a configuration mapping layer, binding each platform's key to an alias and purpose tag, and only printing the alias in logs. An even easier approach is to use an AI API aggregation platform, where one key manages all models. During our comparison we tried token8341. Its model gateway unifies the authentication of several domestic large models, so switching models only requires changing the model name in the config, without touching the key. For the prototype phase, maintaining four fewer sets of authentication logic is what made one week enough.

SDK Compatibility: Every Vendor's Interface Looks Different

The SDK installation step alone wore out half my patience. The OpenAI SDK ecosystem is the most mature, and many vendors claim compatibility, but once you actually integrate it, you find the parameter names don't match. For example, some platforms call it temperature, some mix in top_p, and others changed max_tokens to max_output_tokens. Streaming switches are also inconsistent: some use stream=True, while others require passing a separate stream_options.

My approach was to abstract an adapter layer, exposing only a unified call function externally, with vendor-specific branches internally. This way the business code doesn't perceive the differences. If you don't want to write this layer yourself, an OpenAI SDK-compatible solution can save a lot of trouble—changing one line of base_url switches the model, moving the complexity of unified multi-model access directly from the code layer to the configuration layer. During the prototype validation phase, this trade-off is well worth it.

Streaming Output: SSE Protocol Implementations Differ Across Vendors

Smart customer service must do streaming, otherwise users wait three seconds before seeing text and the experience collapses immediately. The problem is that the implementation details of the SSE protocol differ across vendors. Some platforms include a complete event structure in each chunk, while others only push the data field; the end marker is sometimes [DONE], sometimes a finish_reason field being set; and some insert heartbeat packets midway, which the frontend can easily misread as content.

At first I wrote the parser according to OpenAI's format, and it garbled as soon as I connected the second vendor. The solution was to write a unified SSE parsing middleware that normalizes each vendor's chunks into the same event structure, so the frontend only recognizes that one. The pitfall I hit: don't trust the "fully compatible" claims in the docs. You must capture real responses and inspect them—docs and implementations often differ by a notch.

Exception Handling: When One Vendor Times Out, How to Automatically Switch

Once the comparison tests were running, the most annoying thing was a single vendor timing out. During one stress test, Qwen-Max's responses suddenly slowed down, the whole customer service pipeline got stuck, and the frontend just kept spinning. There was no fallback mechanism in the prototype phase, so if one vendor went down, everything went down.

Later I added a model routing layer, setting a timeout threshold for each request, automatically switching to a backup model on timeout, and logging the switch. Note here: switching can't be blind retrying. You have to distinguish network timeouts from content moderation blocks—the former can be switched, the latter is pointless to switch. The value of large model routing lies exactly here: turning availability from a single point into multiple points. In our project we used SiliconFlow's scheduling for similar validation, automatically selecting models by task type, and the timeout-degradation chain ran fairly stably.

Cost Monitoring: How to Aggregate Token Consumption

The most unexpected expense over the week was tokens. Running four models in parallel for testing, the daily call volume wasn't large, but because there was no aggregation, at month-end reconciliation I found one vendor's consumption was three times the estimate. The reason: under streaming output, the usage field returned by many platforms is empty, so you have to estimate by character count yourself, and the estimate isn't accurate.

My approach was unified accounting at the gateway layer: each call records the model name, input/output tokens, latency, and whether degradation occurred, all written into a single table. Under the pay-as-you-go model, you must calculate this account yourself; you can't rely entirely on the platform backend. When comparing API prices, also note that if a low-priced model has complex output-token billing rules, the actual cost may end up higher.

To sum it up in one sentence: the core of a multi-model comparison prototype isn't getting any single model working, it's turning access, streaming, degradation, and accounting into a unified layer. If you want to dig deeper into model gateway selection, you can look up more material on API aggregation.

Author: Zhou Mingzhe

Published: October 6, 2026