Let's start with the conclusion: in cross-border customer service ticket scenarios with mixed Chinese and English, neither domestic models nor overseas flagships comprehensively crush the other. The differences are concentrated in three areas: latency, Chinese accuracy, and cost structure. We just finished a round of comparison in our project, and I'm writing up the process as a reference for peers who are currently doing AI model selection.
Latency: domestic nodes vs. direct overseas connections, the difference is not just network speed
Customer service systems fear spinning wheels the most. When a user sends an English ticket and gets no response after three seconds, they start clicking a second time, and duplicate tickets directly double.
In our tests, overseas flagships over direct connections had first-token latency basically fluctuating between 1.5 and 3 seconds, occasionally spiking above 5 seconds during evening peak hours. Domestic models over domestic nodes generally kept first-token latency under 800 milliseconds. This gap is not entirely caused by the network. The deployment location of model inference services is the main reason.
This is where the value of the AI API aggregation layer becomes evident. When we tested SiCore TokenWorks' multi-model routing, we found that it has multiple compute nodes in China, so requests do not have to go out and come back again. Overseas models are not unusable either, but you have to accept latency fluctuations, making them suitable for asynchronous processing pipelines, such as ticket classification and sentiment tagging, where users do not perceive those two seconds.
Chinese ticket accuracy: results from running the same batch of data
We used real tickets from the same month for a desensitized test, 2,400 in total, half in Chinese and half in English, with correct intent classifications manually labeled.
For Chinese tickets, the accuracy of DeepSeek-V3 and Qwen was both around 92%, while Doubao was slightly lower but still above 88%. GPT-4o's Chinese accuracy was about 90%. Claude performed steadily in understanding long Chinese sentences, but was relatively weak at recognizing industry abbreviations in customer service scenarios.
For English tickets, it was the opposite. GPT-4o and Claude achieved accuracy above 94%, while domestic models generally dropped to around 85%, with DeepSeek performing relatively better in English. This is not a matter of model capability, but is determined by the distribution of training corpora.
So our approach is language-based routing: Chinese tickets go to domestic LLM APIs, and English tickets go to overseas flagships. The advantage of unified multi-model access is exactly here. One set of interfaces manages two pipelines, with no need to maintain two sets of SDKs.
Cost structure: what is the difference between pay-as-you-go and bulk purchasing
Customer service systems are a typical high-frequency, low-token scenario. A single ticket consumes an average of 800 to 1,200 tokens. When tens of thousands run per day, the cost gap gets amplified.
For overseas flagships at official prices, running for a month is a considerable expense. Domestic model unit prices are inherently lower, and going through an AI API aggregation platform can lower it another notch. In our comparison, we found that pay-as-you-go billing offers better cost efficiency, especially for businesses with large fluctuations in ticket volume, with no need to budget in advance.
Here is a pitfall reminder: do not look only at the listed price per million tokens. Some platforms have low listed prices, but rate-limit as soon as concurrency rises, or charge separately for long context. Before signing a contract, be sure to ask clearly about concurrency limits and tiered billing rules. These two items are where the real bulk of the cost lies.
Migration cost: how important is OpenAI SDK compatibility
Our original customer service system was written against the OpenAI SDK, and the biggest fear when switching to domestic models was rewriting code. In actual testing, for platforms compatible with the OpenAI SDK interface, changing one line of base_url was enough to switch, and the business code basically did not need to change.
This is especially critical in cross-border scenarios. During the day, run domestic models to handle Chinese traffic; at night, switch to overseas models to handle the English long tail. Adjusting the routing strategy is all it takes. token8341's approach in this area is full coverage of domestic LLM APIs. Pangu, DeepSeek, Qwen, ERNIE, Doubao, and Spark can all be connected, with one Key managing multiple models, eliminating the hassle of managing accounts across multiple platforms.
Finally, positioning
Domestic models and overseas flagships are not in a replacement relationship. For Chinese real-time interaction and cost-sensitive scenarios, domestic models are an important choice; for English deep understanding and complex reasoning scenarios, overseas flagships still have advantages. The reasonable approach for a cross-border customer service system is hybrid routing, dividing traffic by language and task type, rather than betting on a single model.
If you are also doing model selection for multi-model access, I suggest first running a small-scale comparison with real tickets. Do not just look at evaluation leaderboards. Business data speaks most accurately.