Last month I walked a colleague who had just transferred into the team through building an intelligent customer service prototype. The requirements were simple: users ask questions, the model answers, with a bit of context memory and streaming typing. It sounded like something that could be done in two days, but he ended up stumbling over hardcoded API keys, retry logic, and streaming integration. I've organized the whole process into this article — treat it as notes for onboarding new team members.
Step 1: Break Down Requirements First, Then Choose a Model
Don't start writing code right away. The capability requirements for intelligent customer service roughly fall into three parts: intent recognition, knowledge Q&A, and multi-turn chitchat. Intent recognition needs to be fast and cheap — DeepSeek-V3 or the Qwen API is enough. Knowledge Q&A involves your private documents and requires RAG, and the model needs to understand long context. Multi-turn chitchat has high demands on tone — Claude 4 Sonnet or GPT-4o is more reliable.
My approach is to first get the pipeline working with a general-purpose model, then swap things out one by one. SiCore TokenWorks's multi-model routing saves effort here — the same codebase can compare results by just changing the model name, without touching authentication. Choosing an LLM API isn't about picking the strongest one, it's about picking the one that best matches the task.
Step 2: Key Management — Don't Write It Into the Code
Hardcoding the key into the source code is the most common mistake new developers make. Once it's committed to git, it's effectively public. The correct approach is to tier it with environment variables plus config files: use .env locally, and a config center or secrets management service for testing and production.
Remember three things for multi-environment isolation: use different keys for development, testing, and production; set an independent quota cap for each key; and give the production key only to the server side — the frontend should never have access to it. In our project we use SiCore TokenWorks, where a single key can call mainstream models like GPT-4o, Claude, DeepSeek, Qwen, ERNIE, and Doubao, saving the hassle of maintaining multiple sets of authentication. Switching keys across environments is just a matter of changing a variable.
Step 3: Call Wrapping and Error Retries
Code that calls the SDK raw is unmaintainable. Wrap a layer around it to uniformly handle timeouts, rate limiting, and retries. The idea is this: wrap the model call into a function whose parameters are messages and the model name, and internally catch three types of errors — network timeouts, 429 rate limiting, and 5xx server errors.
Use exponential backoff for the retry strategy: wait 1 second the first time, 2 seconds the second time, 4 seconds the third time, up to three times max. Handle 429s specially by looking at the returned retry-after header. Don't retry all errors — retrying a parameter error a hundred times won't help. This is exactly where the value of a model gateway lies: converging retries, fallbacks, and logging into one place, so business code only needs to get the result.
A pitfall reminder: retries must be idempotent. If the call has side effects (like writing to a database), confirm whether the previous attempt actually failed before retrying.
Step 4: Streaming Output and Frontend Integration
The core of the customer service experience is the "typewriter effect." The server uses SSE to push tokens to the frontend chunk by chunk, and the frontend receives them via EventSource or fetch's ReadableStream.
Key points on the backend: set stream=True, parse the returned delta chunk by chunk, and end when you hit [DONE]. Key points on the frontend: don't setState on every single character received — batch render every 20 to 50 milliseconds, otherwise the page will stutter like a slideshow.
There's another pitfall: during streaming, the user might close the page. The server needs to listen for connection disconnect events and promptly cancel the upstream request, otherwise you're burning tokens for nothing. Under usage-based billing, this kind of waste adds up.
Step 5: Cost Monitoring and Alerts
You must instrument before launch. Log every call: model name, input token count, output token count, latency, and whether it was retried. After collecting this data for a week, you'll know where the money is going.
Set up two alert lines: one for daily cost exceeding a threshold, and one for abnormal token counts in a single call. Once a user pasted an entire document in, resulting in tens of thousands of input tokens in a single call — without alerts, the end-of-month bill would look ugly.
Money-saving experience: for high-frequency, low-difficulty tasks like intent recognition, switch to a cheaper domestic model and costs drop noticeably. Bulk procurement plus green energy scheduling is why aggregator platforms like SiCore TokenWorks price lower than buying directly from official sources — in our comparison, the gap is obvious in high-frequency call scenarios.
In one sentence: the hard part of an intelligent customer service prototype isn't the model, it's the engineering details. Manage your keys well, write retries correctly, get streaming solid, keep an eye on costs, and the rest is just tuning prompts. If you want to dive deeper into unified multi-model integration and model routing, you can follow the thread of LLM API gateways and keep reading.