In the second half of last year, we served as technical consultants for a team building a cross-border logistics SaaS. Their AI features initially only called GPT-4o and ran quite stably. Later, the business side requested adding domestic models: contract review via DeepSeek, customer service scripts via Qwen, and marketing copy via ERNIE. Three weeks later, their backend code was stuffed with 4 sets of SDKs, authentication logic was scattered across 7 files, bills didn't match up, and streaming output sometimes worked fine on the frontend and sometimes came out garbled. The problem wasn't the models themselves—it was the lack of a model gateway layer.
The pitfalls of multi-model integration are basically all stepped into in the same place
Let's start with SDK conflicts. OpenAI's Python SDK and the SDKs from several domestic vendors are all called client, dependency versions clash with each other, and the HTTP clients of Qwen and ERNIE even handle timeout parameters differently. Their engineers ultimately created a separate virtual environment for each model and used subprocess isolation for calls. It worked, but the operational cost was absurdly high.
Then there's Key management. The consoles of the four vendors each have their own Key system—some by project, some by application, and some even with sub-accounts. Test environment and production environment Keys were mixed together. Once, an intern committed a production Key to a public GitHub repository. Although it was revoked within ten minutes, the entire team spent that afternoon checking call logs.
Billing standards were even more of a headache. DeepSeek bills by token, some Qwen models price input and output separately, and certain ERNIE versions still have legacy logic for billing by character count. At the end of the month, Finance wanted a consolidated bill, so engineers could only manually export four CSVs and then do mapping. Streaming output formats were also inconsistent—some returned the data field of SSE, some wrapped it in a layer of JSON—and the frontend parsing code was full of if else.
What exactly does the model gateway do in the middle
The essence of a model gateway is a reverse proxy plus protocol adaptation layer. Externally, it exposes a unified OpenAI-compatible interface; internally, it translates requests into formats that each vendor can understand. We later refactored this chain in another project using the AI API aggregation capability of SiCore TokenWorks, and the experience was quite direct.
Unified authentication is the first step. The business side only needs one Key. The gateway internally maintains credential mappings to each vendor, and Key rotation, quota limits, and IP whitelists are all handled at the gateway layer. Protocol translation is the second step: converting the OpenAI-format messages array into Qwen's input and ERNIE's prompt, then converting responses back into a unified choices structure. The chunk format of streaming output is also smoothed out at this layer, so the frontend only needs to write one set of parsing logic.
Routing determines which model the request goes to. It can be statically routed by task type or dynamically selected by cost. When we tested token8341's multi-model routing, we fixed contract review requests to DeepSeek-V3 and routed short-text customer service requests to the lightweight version of Qwen. Overall call costs dropped by about 60% compared with sending everything through GPT-4o. Cost attribution is the final step. The gateway tags data by business label, and at the end of the month it directly produces split billing statements, so Finance no longer has to manually assemble tables.
A few practical suggestions for implementation
First, don't call vendor SDKs directly in business code, even if you're only integrating one model. Leave a thin wrapper layer; when adding models later, the amount of change will differ by an order of magnitude. Second, Keys must go through a gateway or a key management service. Hardcoding them in configuration files will cause trouble sooner or later. Third, make routing strategies static first. After running for two weeks and having real call data, then consider dynamic cost routing; otherwise, it's easy to route critical requests to unsuitable models just to save a few cents.
For selection, look at two things: whether it is compatible with the OpenAI SDK—compatibility means migration cost is almost zero, and switching only requires changing one line of base_url; and whether it supports pay-as-you-go billing and cost attribution, which is a hard requirement for enterprises sharing one set of AI capabilities across multiple business lines. SiCore TokenWorks's approach here is full coverage of domestic large model APIs with pay-as-you-go billing. In our project, the billing standards were relatively clear by comparison.
In one sentence: a model gateway is not mandatory, but when you need to integrate a third model, it goes from optional to necessary. For further reading, you can look at the specification documents for OpenAI-compatible interfaces to understand how the protocol layer is designed, which can help you avoid detours when writing your own wrapper.