## First, a Few Common Failure Scenarios
First, a Few Common Failure Scenarios
When many indie developers and small teams get a request to "build a customer service bot," their first instinct is: grab an open-source framework, load in the FAQ docs, ship it, done. This usually ends in one of these ways:
Outcome 1: Brute-forcing with keyword matching. Response accuracy looks fine in the first week after launch. By week two, users start asking "what's the refund process," "how do I return this," "I don't want it anymore, what do I do"—the same intent phrased three different ways, and the keyword bank misses them all. The ops person manually adds keywords until they question their life choices, and three months later the knowledge base has become a steaming pile of legacy junk nobody dares to touch.
Outcome 2: Calling a large language model raw. A different approach: stuff all the FAQ docs into the prompt and let the LLM improvise. The short-term results are impressive, but problems surface quickly: the model confidently fabricates return policies that don't exist, and its answers don't match what human agents say; plus, every message carries tens of thousands of tokens of context, and the monthly bill doubles.
Outcome 3: Multi-turn dialogue handled entirely with if-else. To handle chained conversations like "check order → follow up on shipping → then ask about refund," you write a full screen of branching logic. Every new scenario requires refactoring the state machine, until nobody can understand the code and feature requests get scheduled for next quarter.
The common problem with all these approaches: treating "being able to answer a question once" as the goal, when what's actually hard in customer service is making accurate intent recognition, controllable answers, coherent context, and predictable costs all hold true simultaneously.
The Right Path: A Three-Layer Architecture of Intent Recognition + RAG + Controlled Generation
After a retrospective, the architecture I set for my team is layered, with each layer doing one thing only:
用户消息
│
▼
┌─────────────────────────────────────┐
│ 第一层:意图分类(小模型,便宜快速) │
│ 输出:意图标签 + 置信度 │
│ - 高置信 + FAQ命中 → 直接返回标准答案 │
│ - 低置信 / 业务操作类 → 进入第二层 │
└─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────┐
│ 第二层:RAG 检索增强 │
│ 知识库切片 → 向量检索 top-k │
│ → 组装带引用的上下文 │
└─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────┐
│ 第三层:受控生成(强模型) │
│ 系统提示词限定:只基于给定资料回答, │
│ 资料里没有就说不知道并转人工 │
│ 携带多轮对话历史(滑动窗口管理) │
└─────────────────────────────────────┘The key trade-offs in this structure:
- Use a cheap small model for intent classification. Most requests are actually high-frequency FAQ hits—when matched, return the cached standard answer directly, spending zero on generation costs.
- Trigger RAG only when needed. Only low-confidence or business-operation questions go through vector retrieval and generation, keeping token consumption under control.
- Handle multi-turn dialogue with conversation history management, not a state machine. Maintain a sliding conversation window, stuff the summary of the most recent N turns into the context, and the model figures out what "it" refers to on its own—without writing a single line of if-else.
- Fall back to human agents. The controlled generation prompt explicitly requires "guide the user to a human agent when the material doesn't cover their question," avoiding fabrication. After this went live, complaints about inconsistent answers from customer service essentially disappeared.
Key Implementation Steps (Process Checklist)
- [ ] Organize the FAQ, split it as "one intent per entry," and clearly write out the standard answer and synonymous phrasings (this is the foundation of quality—more important than model selection)
- [ ] Intent classification: put intent labels and a few examples into the system prompt, and have a small model output structured JSON
- [ ] Knowledge base processing: chunk long documents by semantic paragraphs (300-500 characters), vectorize and index them, with metadata on each entry (update time, applicable scope)
- [ ] Retrieval and generation: take top-k of 3-5 entries, assemble into the prompt, explicitly stating "answer only based on the provided material"
- [ ] Multi-turn management: maintain a message array within the session; when it exceeds the window, compress via summarization
- [ ] Instrumentation and evaluation: log intent hits, retrieval hits, and whether each request was escalated to a human; run a weekly regression on a batch of labeled samples
- [ ] Gradual rollout: open to 10% of traffic first, comparing resolution rates between the bot and human agents
Why I Route Through a Unified AI API Gateway
With the architecture settled, there was still a practical issue: intent classification wants a cheap model, generation wants a strong model, and later you might add embedding and content moderation—each new capability means another SDK, another set of keys, another set of rate-limiting and billing logic. An indie developer's time shouldn't be wasted on this.
After integrating a unified AI API gateway, I maintain just one endpoint and one key, switching models via parameters. The reduction in maintenance overhead is real:
- Switching models requires no code changes. If the small model used for intent classification stops being cost-effective someday, changing the model name switches it out—no need to redo authentication or SDK integration.
- Unified billing and monitoring. All calls go through the same gateway; cost reports, failure rates, and latency are visible at a glance, so cost reviews don't require stitching together bills from various vendors.
- One set of fallback and retry logic. Model timeouts automatically fall back to a backup model; this logic lives in the gateway config, keeping business code clean.
I really don't want to go back to the days of three models, three sets of keys, and three SDKs.
Conclusion
A customer service bot is neither a toy that "just runs once you load an FAQ" nor magic that "just calls an LLM raw." A layered architecture plus a unified gateway lets a small team maintain a multi-model conversational system with the effort of maintaining a single project.
If you're about to build one yourself, you can start by registering with a gateway that supports multi-model switching: https://api.thistoken.ai/register. Get the foundation right first, and the road ahead will be much smoother.
---
Tired of juggling provider integrations? Register at https://api.thistoken.ai/register and call every model through one base_url.
Bạn muốn thử Token.AI?
Tạo API Key cấp dự án, bật kênh trong bảng điều khiển và định cấu hình định tuyến, ngân sách và nhật ký kiểm tra.
注册 ThisToken.AI 并获取 API Key