Why Managers Need to Care About Model Switching
Last month our team needed to replace the model used for customer support replies. I thought it was just a one-line code change, but at the review meeting I got stumped: Why this model? Who tested it? What was the test data? Can the conclusions be reproduced?
I couldn't answer any of it. The developer, Xiao Zhang, said "I tried a few prompts and felt the new model's replies were smoother." That kind of "feeling" won't hold up when things go wrong. The moment a customer complaint comes in, you can't even produce the comparison evidence from back then.
So I set a rule: Any model switching decision must first go through a traceable local A/B comparison test. The test script and data go into the repository, and the conclusions are attached to the review record. This article explains how we implemented this process—it's ready for solo developers and small teams to copy directly.
Process Design: Three Roles, a Four-Step Closed Loop
Small teams don't need complex divisions of labor, but responsibilities must be clear. Our approach:
- Owner (usually me): Define evaluation dimensions (accuracy, response speed, cost share, content safety), prepare the test set, and sign off on conclusions.
- Executor (developer): Write the script, run the tests, and produce the comparison report.
- Reviewer: Look at the quantitative data in the report and decide whether to switch.
The four-step closed loop is: prepare test set → run comparison through a unified gateway → produce quantitative report → archive for traceability. Each step produces a deliverable, and none can be skipped. The test set is especially important—we extract anonymized samples from real business logs, and only consider it a valid sample once it exceeds 100 entries, to avoid drawing conclusions from just a handful of prompts.
Why You Must Go Through a Unified Gateway Instead of Direct Connections
This is the risk control point I want to emphasize. If your script connects to each vendor's native SDK separately, three things will happen:
- When one vendor adjusts their API, your test script breaks along with it, and your historical comparison data loses comparability;
- API Keys end up scattered across multiple scripts, and the leak risk spikes when people leave;
- Every time you swap in a candidate model, you have to rewrite the integration code.
A unified gateway solves all three at once: all requests go through the same OpenAI-compatible protocol, there's only one Key, and switching models is just a matter of changing one model name parameter. We use ThisToken.AI's unified interface—after registering, you can get your Key from the console, and the official model list on the website tells you which models you can call.
Step 1: Register and Get Your Key
Open the ThisToken.AI website, register an account, go to the console, and create an API Key. Best practice: create a separate Key for each purpose (e.g., one dedicated to A/B testing), which makes billing and permissions easier to manage. Once the Key is created, don't hardcode it—put it in an environment variable or a local .env file, and make sure .env is in .gitignore—this is a mandatory check item in our code reviews. Don't agonize over costs in advance; refer to the pricing page on the website. Consumption during testing is usually minimal.
Step 2: Get Your First Piece of Code Running
The Python script below is our standard archived version: for the same batch of test questions, it calls two models in sequence, records the replies and timing, and outputs a comparison summary.
import os
import time
import json
from datetime import datetime
from openai import OpenAI
client = OpenAI(
api_key=os.getenv("THISTOKEN_API_KEY"),
base_url="https://api.thistoken.ai/v1"
)
TEST_SET = [
{"id": 1, "question": "你们的退货流程是怎样的?"},
{"id": 2, "question": "订单显示已发货但查不到物流,怎么处理?"},
{"id": 3, "question": "发票多久能开出来?"},
]
MODELS = ["model-a-name", "model-b-name"] # 替换为控制台里可用的模型名
def ask(model: str, question: str):
start = time.time()
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": question}],
)
elapsed = round(time.time() - start, 2)
return resp.choices[0].message.content, elapsed
def main():
report = {"run_at": datetime.now().isoformat(), "results": []}
for case in TEST_SET:
for model in MODELS:
answer, elapsed = ask(model, case["question"])
report["results"].append({
"case_id": case["id"],
"model": model,
"latency_s": elapsed,
"answer": answer,
})
print(f"[{case['id']}] {model} | {elapsed}s | {answer[:50]}...")
# 归档留痕,评审时直接引用
with open("ab_report.json", "w", encoding="utf-8") as f:
json.dump(report, f, ensure_ascii=False, indent=2)
if __name__ == "__main__":
main()How to run it:
pip install openai
export THISTOKEN_API_KEY="你的Key"
python ab_test.pyOnce it finishes, ab_report.json will be generated in the current directory, containing the complete comparison record with timestamps. This file is our archived decision-making evidence—who ran it, when it was run, what each model answered for each question, and how long it took—all fully traceable.
Step 3: Turn Conclusions into Decision Evidence
After getting the raw report, we require the executor to add three things before the review meeting:
- Average latency comparison: mean and percentile latencies for each model;
- Manual scoring: the owner rates reply quality (1-5), with scoring criteria attached;
- Risk notes: any hallucinations, off-topic answers, or content safety concerns.
At the review meeting, we only look at these three sets of numbers plus sampled raw text, and give a "switch / keep / run another round" verdict within ten minutes. The conclusion is written into the review record and archived together with ab_report.json. Three months later, if someone asks "why did we switch models back then," we just pull the file—no need to rely on anyone's memory.
A Few Disciplines We Established After Hitting Pitfalls
- Version your test sets: Put test set files under version control, and note which version was used for each test run; otherwise comparisons across batches are meaningless.
- Key rotation: When anyone involved in testing leaves or changes roles, revoke their Key in the console immediately.
- Control variables: Within the same test batch, both models must use identical prompts and parameters (temperature, etc.), or the conclusions won't hold up.
- No testing, no switching: No matter how good the vendor's marketing is, anything without local data doesn't make it to review.
Final Thoughts
Once this process was running smoothly, our model decision-making went from "arguing all afternoon" to "reading a ten-minute report." The foundation of the whole process is that unified gateway plus a reproducible test script. If you also want your team's model selection to be evidence-based, start by registering an account, getting your Key, and getting the script above running: https://api.thistoken.ai/register —once your first comparison report is generated, you'll understand why I wrote it into our team's standards.
---
Every example in this post runs with a single API key — get yours at https://api.thistoken.ai/register and start in minutes.
Ready to try Token.AI?
Create a project-level API Key, enable channels in the console, and configure routing, budgets, and audit logs.
注册 ThisToken.AI 并获取 API Key