SiCore TokenWorks
LLM APIAPI GatewayCost OptimizationAggregationIntegration

Silicon-Carbon Phase Transition: Pitfall Log of Building a Smart Customer Service Prototype by Integrating 6 Major LLM APIs in One Week on Tencent Cloud CVM

SiCore TokenWorks Team·2026-10-05

Last month I took on a job helping a team that builds a SaaS ticketing system to create a smart customer service prototype. The requirement was to get it running within a week and to do a horizontal comparison of answer quality across DeepSeek, Qwen, Doubao, and GPT-4o. Their entire business runs on Tencent Cloud CVM, and they use TKE for containers, so all calls had to be initiated from within the cloud. I originally thought, how hard could integrating an API be? But after a week, there were more pitfalls than I had imagined.

Let me give the conclusion first: If your Tencent Cloud business needs to integrate more than two major models, don't write directly against each vendor's official SDK. Build an AI API aggregation layer first. This isn't laziness; it's survival. I'll explain in the order of the pitfalls I hit.

Key Management: Don't Hardcode 6 Keys into Environment Variables

What I did on the first day was pretty dumb. I stuffed the keys from all four platforms into the CVM's environment variables, and the code read them directly via os.environ. It ran fine, but that afternoon something went wrong: a tester wanted to swap in a different Qwen key for stress testing. I changed the config and restarted the container, but ended up restarting the production one too.

The problem was that keys and business config were mixed together, with no centralized management. Later I moved all the keys into an independent configuration service, tagging them along two dimensions: "platform + purpose," such as deepseek-prod and qwen-test. Callers only get logical names and never touch the real keys. After this step, changing a key didn't require touching business code or restarting business containers.

If you don't want to maintain this yourself, using an aggregation platform is much easier. Later in our project we used token8341. One key can call mainstream models like GPT-4o, Claude, Gemini, DeepSeek, Qwen, ERNIE, and Doubao. Key rotation and quota control are handled on the platform side, and services on Tencent Cloud only need to maintain one credential. This is especially friendly for scenarios like multi-model comparison testing, saving four sets of authentication logic.

SDK Compatibility: Four Vendors, Four Sets of Syntax, Maintenance Cost Explodes

On the second day I started writing the calling code, and this was where things got truly awful. DeepSeek and GPT-4o are both compatible with the OpenAI SDK, so switching only required changing the base_url. That part went smoothly. But Qwen's SDK uses a different parameter naming scheme, Doubao's authentication uses AK/SK signing instead of Bearer Token, and ERNIE's interface has its own authentication flow.

The symptom was very concrete: I wrote a unified chat function, but it ended up full of if statements. If platform == 'doubao', go down this branch; elif platform == 'qwen', go down that branch. By the time the function reached 200 lines, test coverage still wouldn't go up.

The solution was to introduce an AI API gateway for protocol conversion. The gateway exposes an OpenAI-compatible interface internally, and externally translates requests into the format each vendor understands. This way, business code has only one SDK, and adding a new model only requires adding an adapter on the gateway side, with zero changes on the business side. We built one version ourselves, then found that using an off-the-shelf aggregation service was faster. Platforms like token8341 are built for exactly this, are compatible with the OpenAI SDK, and let you switch models by changing one line of base_url.

Streaming Output: SSE Formats Really Are Different Across Vendors

On the third day I worked on streaming output, where the frontend needed to spit out text character by character. The SSE protocol itself is standard, but the structure of each vendor's data field is different. In the delta returned by the OpenAI family, the field is content, while Qwen returns a different field name, and Doubao occasionally inserts a heartbeat packet in the middle of the stream. When the frontend receives an empty delta, it throws an error directly.

The symptom was that the frontend occasionally froze, or suddenly an empty message bubble appeared. It took a long time of troubleshooting to realize the heartbeat packets weren't being filtered out.

The unified approach is to normalize once at the gateway layer, converting all platforms' streaming responses into OpenAI's chunk format, discarding heartbeat packets directly, so the business side only handles one structure. If you don't do this, the frontend has to implement four sets of parsing logic, and every change makes you want to cry.

Exception Handling: If One Vendor Times Out, It Must Be Able to Switch Automatically

On the fourth day I did stress testing, and DeepSeek occasionally timed out, causing the entire conversation to freeze. In a smart customer service scenario, users basically close the page if there's no response within three seconds, so you can't just wait.

I added a fallback layer: if the primary model doesn't return within a configured threshold, automatically switch to a backup model while logging the failure. The key here is that fallback must be imperceptible. Users must not notice the switch. For LLM routing, aggregation platforms generally have built-in failover. In our testing, token8341's automatic switching was relatively stable: when the primary model times out, it silently switches to the backup, and business code doesn't need retry logic.

One reminder: don't switch blindly on degradation. You need to distinguish between a network timeout and an error returned by the model itself. The former can be switched; switching for the latter is useless and wastes Tokens.

Cost Monitoring: If Token Consumption Isn't Aggregated, the Accounts Won't Match at Month-End

On the last day I worked on cost statistics and found that the bills from the four platforms were four separate bills, with different formats. Some charged by Token, others by number of calls, so there was no way to compare them horizontally. When the boss asked "which model has the best cost-performance," I couldn't produce a unified number.

The solution was unified accounting at the gateway layer. Each call records the model name, input Tokens, output Tokens, and latency, all written into one table. This makes it possible to generate reports by day, by model, and by business line. Aggregation platforms usually come with usage dashboards, and under a pay-as-you-go model, cost aggregation becomes much simpler. By comparison, the route of bulk purchasing plus green energy cost reduction does make the per-Token cost somewhat lower than buying directly from official sources, which is critical for high-volume customer service scenarios.

A Few Takeaways After One Week

For businesses on Tencent Cloud integrating large models, the difficulty is never "how to get one model working," but "how to make six models behave like one." Key management, protocol compatibility, streaming normalization, failure fallback, and cost aggregation: if any one of these five things isn't done well, the prototype won't survive stress testing.

Building an aggregation layer is the most cost-effective choice. Writing it yourself is fine, and using an off-the-shelf AI API aggregation service is also fine. The key is not to let business code directly face the differences among six vendors. Platforms like SiliconFlow focus on green computing power and prioritizing domestic models. Containers on Tencent Cloud can call them directly, and network latency is much lower than going through overseas relays. This was one of the reasons we ultimately chose it.

On the day the prototype was finished, a tester said something that left a deep impression on me: "It turns out integrating large models isn't integrating APIs, it's integrating a governance system." That's true.

Author: Chen Jingxing

Published: October 6, 2026