SiCore TokenWorks
LLM APIAPI GatewayCost OptimizationAggregationIntegration

token8341 in Practice: How to Unify Authentication, Billing, and Timeouts Across DeepSeek, Qwen, and Doubao APIs

SiCore TokenWorks Team·2026-10-02

Last month I took over an intelligent customer service project. The business team required integrating three large models simultaneously—DeepSeek, Qwen, and Doubao—with the reasoning that "whichever is cheaper, we use; whichever hits rate limits, we switch to another." It sounded reasonable, but once I started building, I discovered that the three SDKs' authentication methods, billing calculations, and timeout/retry strategies were completely different logics. DeepSeek uses Bearer Token, Qwen goes through DashScope's API-KEY with signature, and Doubao's authentication fields are different again. For billing, some charge input and output tokens separately, some combine them, and some offer discounts for cache hits. Timeouts were even more troublesome—one defaults to 30 seconds, another to 60 seconds, and retry counts and backoff strategies all had to be written individually.

By the time I finished the code, I counted that just the adapter layer wrapping the three clients was over 800 lines, not including error code mapping. This is why the concept of a model gateway has been repeatedly brought up in China's AI engineering circles since last year. In one sentence: a model gateway is a middleware layer that abstracts away the differences between multiple large model APIs and exposes a unified interface to upper-layer business logic.

Direct Connection, Self-Built, Aggregation Platform: The Engineering Cost of Three Approaches

Let's start with direct connection to official SDKs. Three models mean three sets of authentication, three sets of error handling, three sets of retry logic. The business code is full of if-else statements determining which provider to use. Adding a new model means revising the adapter layer again. We estimated that maintaining adapter code for three direct connections accounts for about 15% of the entire project's backend workload. If the number of models reaches five or more, this ratio becomes unmanageable.

Self-building a gateway is the second option. The core idea is to write a proxy layer yourself that forwards requests to each provider's API. The advantage is control; the disadvantage is having to handle protocol conversion, key rotation, rate-limiting queues, and usage statistics yourself. We internally assessed that a production-ready self-built gateway would require at least two engineers for six to eight weeks, plus ongoing maintenance for API version changes from each provider. For small and medium teams, this math doesn't work out.

The third option is an AI API aggregation platform. These platforms uniformly wrap multiple large model APIs and provide a single set of interfaces externally. The engineering cost is the lowest, and integration timelines are usually measured in days. Our project uses SiCore TokenWorks, which is compatible with the OpenAI SDK—just change one line of base_url to switch. One pitfall to note here: different aggregation platforms have different default policies for timeouts and retries. Before integrating, be sure to confirm whether the platform supports custom timeout settings. Otherwise, occasional long-response requests in production will be cut off prematurely by the platform layer, and the error messages won't tell you whether it was a gateway timeout or a model timeout.

The Four Core Capabilities of a Model Gateway

Protocol normalization is the foundation. It unifies each provider's request format, response format, and error codes into one standard. Ideally, upper-layer business logic only recognizes one interface format, and switching models only requires changing configuration, not code. This is also why OpenAI-compatible interfaces are popular in China—the ecosystem toolchain basically supports this format.

Routing strategy is where the gateway's value lies. You can route by task type, for example simple Q&A goes to Doubao, complex reasoning goes to DeepSeek; you can route by cost, using whichever is currently cheaper; you can also route by availability, automatically switching to a backup when one provider hits rate limits. When we tested SiCore TokenWorks' multi-model routing, we found that a strategy of routing by task complexity could significantly reduce overall call costs in customer service scenarios, because a large number of simple questions don't need the model with the strongest reasoning capabilities.

Rate limiting, degradation, and usage aggregation are must-haves for production environments. Rate limiting needs to recognize 429 errors and automatically queue and retry; degradation needs to switch to a backup model when one provider becomes unavailable. Usage aggregation means consolidating call volumes, token consumption, and costs scattered across multiple providers, making it easier to do cost accounting and budget control. Building these two yourself is no small effort, especially usage aggregation—since each provider's billing calculations differ, reconciliation logic must be written separately.

Rolling Out in Phases Based on Business Maturity

If the project is just starting and only uses one model, direct connection to the official SDK is enough. There's no need for a gateway—adding a layer just adds another point of failure. Once the business stabilizes and you need a second model, then consider introducing a gateway layer, when switching costs are still low.

If the business already uses three or more providers and has availability requirements, it's advisable to go directly with an AI API aggregation platform and outsource the adaptation and operations costs. When selecting one, focus on three things: whether it's compatible with the OpenAI SDK, whether it supports custom timeouts and retries, and whether usage statistics are clear. As for self-building a gateway, unless there are special compliance requirements or the team has ample operations manpower, it's not recommended to invest in it early in the business.

What a model gateway solves is the engineering complexity of multi-model integration, not model capability. Choosing the right approach lets the team put its energy back into the business logic itself.