Business Pain Point: The Search Box Became the Product's Weak Spot
I maintain a knowledge base tool for small teams. The user base isn't large, but retention is good. Over the past six months, the most common feedback I've received is: "The search is too dumb." When a user searches for "customer complaint handling process," the system can only do exact keyword matching, so searching "how to handle customer complaints" returns nothing. As a result, users would rather browse the directory than search, greatly diminishing the value of the knowledge base.
As an indie developer, I knew I needed semantic search, but the workload estimate was a headache:
- Option A: Build my own vector retrieval—I'd need to pick an embedding model, set up a vector database, write retrieval logic, implement rerank, and maintain it all myself. Estimated 3-4 weeks, and every component could break later.
- Option B: Integrate a large model as an understanding layer—Controllable results, but every model swap or parameter tweak means code changes and retesting.
Before this, I had done an audit: AI-related code in the project was scattered across 6 files, with API addresses, model names, and timeout configurations duplicated everywhere. Last time I switched from the GPT series to another model for testing, I changed 3 places in the code and ran 2 rounds of regression testing, spending an entire afternoon. For a one-person team, this kind of repetitive work is the most expensive cost.
This time, I decided to take a different approach.
Architecture Design: Unified Gateway + Lightweight Retrieval Pipeline
The overall architecture has three layers, with the core idea being: completely decouple "model selection" from business code.
User query
│
▼
[Application layer] Query preprocessing (noise removal, intent detection)
│
▼
[AI Gateway layer] Unified API entry (ThisToken.AI)
├─ embedding model → text vectorization
├─ chat model → query rewriting / intent clarification
└─ rerank model → result reranking
│
▼
[Data layer] Vector store (pgvector) + original document index
│
▼
Search results returnedI chose pgvector because the product already uses Postgres, so no new components are introduced. The truly critical piece is the gateway layer in the middle: all model calls go through the same base_url and the same authentication, and model names are just strings in a config file.
Why does a unified gateway significantly reduce maintenance costs? My before/after comparison is very telling:
| Item | Before Gateway | After Gateway |
|---|---|---|
| Multi-model integration effort | Separate registration and adapter code per model, ~2-3 days each | Change one config item, ~10 minutes |
| Model swap testing | Code changes + regression, ~half a day | Config change + smoke test, ~30 minutes |
| API Key management | Multiple platforms and accounts, 3 keys scattered around | Single key, centrally managed |
| Billing | One bill per platform across 3 platforms, ~1 hour/month reconciliation | One bill, reviewed in minutes |
For an indie developer, this saved time is not trivial—the difference between shipping in three days versus four weeks is three extra weeks to polish product details.
Key Implementation Steps
Step 1: Data Vectorization (Day 1 Morning)
I chunked and vectorized the 8,000+ documents in the knowledge base and wrote them into pgvector. For the chunking strategy, I used the simplest approach: split by heading hierarchy, 300-500 characters per chunk, keeping the heading path as metadata.
Step 2: Unified Gateway Configuration (Day 1 Afternoon)
from openai import OpenAI
client = OpenAI(
api_key=os.environ["THISTOKEN_API_KEY"],
base_url="https://api.thistoken.ai/v1"
)
def embed(texts: list[str]) -> list[list[float]]:
resp = client.embeddings.create(
model="text-embedding-3-small", # switch models by only changing this
input=texts
)
return [d.embedding for d in resp.data]
def rewrite_query(query: str) -> str:
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "把用户的口语化搜索改写为适合检索的关键词组合,只输出改写结果。"},
{"role": "user", "content": query}
]
)
return resp.choices[0].message.contentNote that this uses the standard OpenAI SDK—no new dependencies needed, and switching to the gateway only required changing the base_url line.
Step 3: Wiring Up the Retrieval Pipeline (Day 2)
Complete flow checklist:
- [ ] Query preprocessing: strip meaningless words, truncate length
- [ ] Query rewriting: call the chat model to convert colloquial input into retrieval-friendly phrasing, with a 3-second timeout; fall back to the original query on failure
- [ ] Vector retrieval: after embedding, fetch Top 20 candidates from pgvector
- [ ] Keyword fallback: if vector results are fewer than 5, merge in traditional keyword matching results
- [ ] Result reranking: rerank model rescores candidates, take Top 5
- [ ] Return formatted results with links to original documents
Step 4: Degradation and Testing (Day 3)
One easily overlooked point: AI calls must be degradable. I set up a switch and timeout for every AI step—if query rewriting fails, use the original query; if rerank times out, use the raw vector ordering. This way, even if the model service fluctuates, the search feature doesn't go down entirely. Day 3 was mostly spent running evaluations: I prepared 120 real query samples and compared hit rates between the old and new search.
Results: A Ledger You Can Actually Balance
Comparison one month after launch:
- Development cycle: Original plan for a full self-built pipeline was ~3 weeks; core functionality was actually completed in 3 days, saving ~80% of the schedule
- Search hit rate (self-built evaluation set): keyword matching ~41%, semantic search ~78%
- Per-search cost: embedding + rewriting + reranking totals ~$0.002, monthly incremental cost under $20
- Maintenance time: previously averaged 2-4 hours per model-related config adjustment, now ~20 minutes
The more important gain is user behavior: search usage rose from ~15% before launch to 47%, and "can't find content" feedback has essentially disappeared.
Three Tips for Indie Developers
- Unify the entry point first, then talk about models. Don't scatter model names and API addresses throughout your business code. A unified gateway turns "switching models" from an engineering problem into a configuration problem.
- Every AI step needs a degradation path. For core features like search, no single model timeout should ever show the user an error.
- Build the evaluation set before tuning. Without one, you can't know whether swapping models or adjusting parameters actually makes things better or worse.
If you want to make the first step even easier, you can start by registering an account at ThisToken.AI and getting embedding, chat, and rerank models all running with a single key and unified API entry: https://api.thistoken.ai/register —for indie developers, every bit of infrastructure you don't have to maintain is more time to polish the product itself.
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Хотите попробовать Token.AI?
Создайте API Key уровня проекта, включите каналы в консоли и настройте маршрутизацию, бюджеты и журналы аудита.
注册 ThisToken.AI 并获取 API Key