Pseudocode: core layering/routing logic
async def handle_message(user_msg, session):
Layer 1: lightweight classification model, outputs only a label
intent = await gateway.chat(
model="lite-classifier",
messages=[{"role": "user", "content": ROUTER_PROMPT + user_msg}],
max_tokens=10 # Just one label, costs almost nothing
)
if intent == "order_query":
order = query_order_api(extract_order_id(user_msg))
return render_template("order_status", order) # No LLM generation
if intent in ("policy", "product"):
docs = vector_search(user_msg, top_k=3)
return await gateway.chat(
model="small-chat", # Small model is sufficient
messages=build_rag_prompt(user_msg, docs)
)
if intent == "complaint":
emotion = detect_emotion(user_msg)
if emotion.risk > 0.7:
ticket = await gateway.chat(
model="large-chat", # Large model writes ticket summary
messages=summarize(session.history)
)
return escalate_to_human(ticket)
return await gateway.chat(
model="large-chat",
messages=session.history + [user_msg]
)
Recommended rollout order: start with intent classification and direct order API integration (biggest gain, smallest change), then introduce RAG to replace the full-context approach, and finally add emotion detection and the human escalation pipeline. My friend's two-person team completed the full workflow in about two weeks.
## Why a Unified AI API Gateway Reduces Maintenance Costs
This architecture requires calling at least three different models. Integrating each vendor's native SDK separately means three sets of authentication logic, three sets of error codes, three rate-limiting strategies, and three places to investigate every time a model changes pricing or its API changes. What independent developers lack most is time to maintain this kind of "glue code."
The value of a unified gateway lies in:
- **One integration, switch anywhere**. All calls go through the same protocol; switching models means changing a single model name string. When testing a newly released cheap small model, the change is one line of code, not one afternoon;
- **Unified monitoring and billing**. Call volume, latency, and cost across model layers are visible in one dashboard, making it obvious which layer is over budget—this is exactly the data foundation for continuously optimizing a layered architecture;
- **Unified fallback and retry**. If the large model layer times out, it automatically falls back to the small model for a backup response, without writing separate error-handling logic for each vendor in your business code;
- **Centralized key management**. Avoids API keys being scattered across multiple services and logs, so security audits only need to happen once.
## The Efficiency Math: Before and After
| Metric | Before (single model) | After (layered) |
|---|---|---|
| Average response time | ~11 seconds | ~2.4 seconds |
| Monthly API cost | ~8,000 yuan | ~3,900 yuan |
| Complaint escalation accuracy | Keyword matching, frequent misses | Emotion model, significantly fewer misses |
| Model switching cost | Rewrite prompts and call logic | Change one line of model name |
The more important hidden gains: since order queries no longer occupy large model throughput, the stability issues during peak hours disappeared along with them; and ticket summaries cut agent onboarding time from an average of three to four minutes of background reading down to a dozen or so seconds to skim the summary.
## Final Thoughts
Multi-model collaboration sounds complex, but for small teams, it's essentially about "reserving expensive capabilities for problems that deserve them." After layering, each layer's optimization goal becomes clear—make fast questions even faster, make hard questions more accurate, and the bill naturally comes down.
If you'd like to quickly build a multi-model architecture like this, consider starting with a unified AI API gateway—sign up and get started immediately, connect multiple models with a few lines of code, get the layered routing working first, then refine step by step: https://api.thistoken.ai/register
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.Хотите попробовать Token.AI?
Создайте API Key уровня проекта, включите каналы в консоли и настройте маршрутизацию, бюджеты и журналы аудита.
注册 ThisToken.AI 并获取 API Key