Business Background and Pain Points
My three-person team took on a project: a pet medical Q&A assistant where users upload descriptions of their pets' symptoms and the assistant provides preliminary analysis and veterinary advice. It sounds simple, but once it was running, we found three types of problems stacking on top of each other:
- Huge variance in question difficulty. "Should I fast my dog if it has diarrhea" and "Insulin dosage adjustment for a cat with diabetes" are two completely different difficulty levels. Using top-tier models for everything burned through 300+ yuan a day in token costs; using cheap models for everything meant serious medical questions often got irrelevant answers, or even dangerous advice.
- Hallucinations are costly in medical scenarios. A model confidently fabricating drug dosages is worse than admitting it doesn't know. We needed controllable fallback and human handoff mechanisms.
- We onboarded three model vendors (one primary, one backup, one small local model), each with different SDK versions, authentication methods, and error codes. Every time the primary vendor's API had issues, we had to dig through two projects' code to fix the adaptation layer, averaging two hours of hassle per incident.
After discussing as a team, we decided: rather than scattering if-else model selection throughout the business code, we should first do tiering, then consolidate all model calls.
Architecture Design: Three-Tier Classification + Gateway Consolidation
2.1 Tiered Model Strategy
Requests are classified into three tiers by risk and difficulty:
- L1 Lightweight tier: Greetings, common illness education, feeding basics. Handled by cheap, fast models, with a response requirement of under 2 seconds.
- L2 Analysis tier: Symptom descriptions + preliminary analysis. Handled by the primary large model, with system prompts constraining it to "not make definitive diagnoses, provide medical consultation advice."
- L3 Fallback tier: Emergency symptoms (swallowed foreign objects, seizures, difficult labor), or requests where model confidence is low or that fail quality checks twice in a row. Routed directly to the on-duty veterinarian queue, along with fixed safety notice copy.
Routing relies on a lightweight classifier: first, rule-based keywords ("swallowed", "seizure", "labor", etc.) hard-route to L3; then a small intent classification model distinguishes L1/L2. Classification itself takes under 200ms.
2.2 Why Use a Unified AI Gateway
Initially, each tier connected directly to vendor SDKs. After two iterations of development, we tallied up the maintenance costs:
- Before the change: Three vendors' SDKs were scattered across 4 services. Every key rotation, error code change, or rate-limiting policy adjustment took an average of 3~4 hours per week of synchronized changes; when a vendor's API hiccuped, troubleshooting required checking logs across two repositories.
- After the change: All model calls converged into a unified gateway, and business code only deals with one consistent request/response structure and set of error codes. Switching vendors became a one-line config change in the gateway admin panel; budget control, call logging, and per-key rate limiting were all handled uniformly at the gateway layer.
Rough calculation: for model integration maintenance alone, we went from about 3.5 hours per week to under 20 minutes, saving about 12 hours a month—for a three-person team, that's an extra day and a half of development time out of thin air. More importantly, the fallback logic became simpler: the gateway layer handles health checks and automatic switching. If the L2 primary model times out or errors, it automatically degrades to the backup model for one retry, completely transparent to the business code.
Key Implementation Steps
- Define classification rules (1 day): Pulled 500 real question samples from the testing period, manually labeled the tiers, and distilled the L3 keyword list and L1/L2 classification features.
- Integrate the unified gateway (0.5 day): Registered on the gateway and replaced the three vendors' model calls with a unified REST interface, with keys managed by the gateway.
- Implement routing and fallback (2 days): Classifier + gateway degradation strategy + human queue integration.
- Add quality-check guardrails (1 day): Ran a cheap rule-based review on L2 output (whether it contains specific dosage numbers, whether it includes a disclaimer); failures escalate to L3.
- Staged rollout observation (3 days): Monitored three metrics—tier hit rate, cost, and human handoff rate—before fully opening up.
Core routing pseudocode:
def route(question: str, ctx: dict) -> Route:
# 第一层:安全关键词硬路由,优先级最高
if hit_keywords(question, EMERGENCY_KEYWORDS):
return Route(level="L3", target="vet_queue")
# 第二层:意图分类决定 L1 / L2
level = classifier.predict(question) # "light" or "analysis"
if level == "light":
# 轻量模型,失败直接礼貌兜底,不升级
return Route(level="L1", target=gateway.model("cheap-fast"))
# L2:主力模型 + 网关自动降级备用模型 + 质检
return Route(
level="L2",
target=gateway.model("main", fallback="backup", timeout_ms=8000),
post_check=dosage_safety_check, # 不过则升级 L3
)The corresponding fallback workflow checklist:
- Request arrives → safety keyword hit? → Yes: hand off to human + fixed safety notice
- No hit → intent classification → L1: light model; on failure, return default educational copy
- L2: primary model → gateway detects error/timeout → automatically switch to backup model for one retry
- L2 output → dosage/disclaimer quality check → fail → escalate to L3 for human handoff
- L3 queue unclaimed for 5 minutes → SMS notification to the on-duty vet
Results Comparison
Data from one month after stable launch (our own testing and staged rollout statistics, not customer data):
| Metric | Before | After |
|---|---|---|
| Average cost per Q&A | ¥0.11 | ¥0.038 (L1 accounts for 62% of traffic) |
| Model integration maintenance time | 3~4 hours/week | 20 minutes/week |
| Vendor failure recovery | Manual switch, ~2 hours | Gateway auto-degradation, users barely notice |
| High-risk question misanswer rate | ~3% during early rollout | Near 0 after quality checks + hard routing (all escalated to humans) |
A 65% cost reduction and 90% maintenance time reduction—for indie developers, these two things are the difference between staying sustainable or not.
Lessons Learned
- Define risks before defining models. The L3 hard keyword list is the cheapest and most effective part of the whole system—one afternoon of work, yet it blocks almost all dangerous outputs.
- Fallback should be split into "model fallback" and "process fallback". The gateway solves model availability; the human queue solves the boundaries of model capability. Neither can be missing.
- The real value of a unified gateway is "change isolation". Any vendor-side change only affects gateway configuration, not business code. The smaller the team, the more you should front-load solutions for these cross-cutting concerns, rather than hoping to refactor later when you have time.
If you're also building a similar tiered model application and want to consolidate the messy work of multi-vendor integration, degradation, and usage control, you can try a unified AI gateway service—register here: https://api.thistoken.ai/register
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Vous voulez essayer Token.AI ?
Créez une API Key au niveau du projet, activez les canaux dans la console et configurez le routage, les budgets et les journaux d'audit.
注册 ThisToken.AI 并获取 API Key