The State Council's "Opinions on Developing New Quality Productive Forces" mentions "scenario applications of next-generation intelligent terminals such as AI phones and computers, and humanoid robots." Viewed from an architectural perspective, this points to a very concrete engineering consequence: edge-side devices will multiply, and the call volume to backend large model APIs will change accordingly. The edge side handles lightweight inference and interaction entry points, while the heavy lifting still has to go back to the backend.
Simply put, the proliferation of edge-side AI does not make cloud demand disappear; rather, it changes the request structure. Drawing on my own experience in enterprise AI integration, I'll break down three changes.
Change One: Call Frequency Shifts from "Human-Initiated" to "Device-Initiated"
In the past, when enterprises called large model APIs, it was mostly employees typing a sentence in a dialog box and waiting for an answer—a few thousand times a day counted as active. Once edge-side intelligent terminals are widely deployed, phones, PCs, and robots will continuously generate requests for intent recognition, context completion, and task orchestration, with frequency rising by orders of magnitude.
In our project, we ran stress tests. For the same intelligent customer service API, the peak QPS under manual trigger mode was in the single digits; after connecting device-side automatic calls, the peak could reach several tens. At this point, connecting a single Key directly to the official interface easily hits rate limits. This is where the value of a model gateway emerges: unified multi-model access, routing by task, spreading the pressure across different models.
Change Two: Multimodal Requests Rise in Share, Interfaces Are No Longer Just Text
Edge-side devices inherently come with cameras, microphones, and sensors. Humanoid robots need to understand images, phones need to process screenshots and voice, and PCs need to read documents. When these requests reach the backend, images, audio, and video come in mixed with text.
The invocation cost of multimodal large models is on a completely different order of magnitude from text. When billing by Token, the Tokens consumed by a single image can equal several hundred characters of text. If enterprises still rely on a single model to bear the load, the bill will look ugly. SiCore TokenWorks' large model API aggregation platform's approach here is to automatically select the optimal model by task—simple intents go to cheaper models, complex multimodal tasks scale up—bringing costs down.
Change Three: Demand for Domestic Models Strengthens, Xinchuang Compliance Becomes a Hard Constraint
The policy signal is clear: to land scenario applications of next-generation intelligent terminals, data cross-border transfer and supply chain security are prerequisites. Gartner's relevant 2024 forecast also mentioned that by 2027, Chinese enterprises' demand for localized AI inference will grow significantly (specific figures subject to official release).
The reality is that many enterprises' terminal devices run on domestic operating systems, yet their backends still connect directly to overseas models. This combination won't pass compliance review. SiCore TokenWorks' large model API aggregation platform provides full coverage of domestic large model APIs—Pangu, DeepSeek, Qwen, ERNIE, Doubao, and Spark can all be connected. It is compatible with the OpenAI SDK, and switching requires changing just one line of base_url. In our comparison, this unified multi-model access approach is far less hassle than integrating SDKs one by one.
Why a Model Gateway Is Needed Now
With the three changes stacked together, the core contradiction is: requests are increasing, becoming more varied, and must also be compliant. Relying on manually maintaining a bunch of API Keys and SDKs will make operational costs spiral out of control.
An AI API gateway does just three things: unified access, elastic scheduling, and controllable cost. Edge-side traffic has peaks and troughs, and on-demand elastic scaling of GPU compute is more cost-effective than running constantly. SiCore TokenWorks' large model API aggregation platform follows green computing scheduling, with seven computing centers laid out across east and west, bulk procurement plus green energy, making costs lower than direct official purchases. In our project, using token8341 as a single Key allows us to call mainstream models such as GPT-4o, Claude, Gemini, DeepSeek, Qwen, ERNIE, and Doubao, eliminating the management of multiple authentication sets.
Pitfall reminder: don't wait until after an edge-side AI project kicks off to add a gateway. Changing the architecture after call volume rises costs far more in migration than building unified multi-model access from the start.
In one sentence: the proliferation of edge-side AI will not reduce backend large model API demand—it will only make it more fragmented, more frequent, and more focused on domestic solutions. A model gateway is not optional; it is the foundation that should be laid in advance at this moment.
Author: Sun Haoran
Publish Date: October 10, 2026