## I
I. The Business Pain Point: One Photo, Three Roles Waiting
I lead a five-person team building a SaaS tool for a second-hand trading platform, with small and mid-sized sellers on the platform as our customers. When I took over this project last year, the first thing I did wasn't reviewing the code—it was sitting next to our operations colleague for an entire afternoon, watching how she helped sellers list products.
The process went roughly like this: a seller uploads a product photo, the operations colleague opens it, guesses the category, then handwrites a title plus three to five lines of description, noting condition, size, and defects. A skilled person takes two to three minutes per image; unfamiliar categories (like fishing gear or collectible figurines) require looking up references, doubling the time. During peak periods, hundreds of images would pile up—sellers pushing, operations complaining, and me, the manager, caught in the middle.
The engineers' first reaction was "just use AI—a multimodal model reads the image and generates the description, a single API call." The prototype was working within two days, and the results were genuinely impressive. But as a manager, I quickly discovered three new categories of problems:
- Unstable quality: The model occasionally described women's clothing as men's, or described items with obvious scratches as "in good condition"—in second-hand trading, inaccurate descriptions directly translate to customer complaints and refunds.
- Collaboration gaps: AI-generated text went straight into the database, while the versions edited by operations never flowed back. We had no idea what the AI's error rate was, and we couldn't accumulate improvements to the prompts.
- Risk of runaway costs: Each image description generation call costs money. If sellers uploaded images and retried without limits, the bill would quietly balloon—and at the time we had no monitoring or quota mechanisms at all.
The biggest lesson this project taught me: the hard part of AI adoption isn't getting the first API call to succeed—it's bringing quality, collaboration, and cost under a controllable process.
II. Architecture Design: Three Checkpoints Plus One Unified Gateway
I reworked the architecture around a core principle: all AI output must pass through human checkpoints, and all API calls must go through a unified gateway.
The overall pipeline:
Seller uploads image
↓
【Checkpoint 1: Pre-screening】 Image compliance check (size/clarity/prohibited content) → rejected if failed
↓
【Gateway】 Unified AI Gateway
├── Routes to vision model (image → structured JSON: category/color/condition candidates/defect areas)
├── Unified auth, rate limiting, billing, logging
└── Primary/backup model automatic fallback (switch to backup if primary times out)
↓
【Checkpoint 2: Template assembly】 Structured JSON + category-specific templates → generate description draft
↓
【Checkpoint 3: Human spot-check】 Operations sees a three-pane interface in the admin console: "original image + AI draft + confidence"
├── High confidence → one-click accept
├── Low confidence → forced manual edit before publishing
└── Accept/edit records flow back for monthly prompt optimization
↓
Product listedFrom a manager's perspective, the value of this design is that every stage has a clear owner and acceptance criteria. Pre-screening is engineering's responsibility, model calls are the architecture's responsibility, and final publication is operations' responsibility—when a complaint about an inaccurate description comes in, we can trace the logs to determine whether the model made a mistake or an editor missed a fix.
III. Key Implementation Steps
Step 1: Define the Output Contract Before Writing the Prompt
Don't have the model "write a description"—require it to output a strict JSON structure:
VISION_PROMPT = """
分析这张二手商品图片,以JSON输出:
{
"category": "一级品类/二级品类",
"brand_guess": "品牌(不确定则为null)",
"color": "主色调",
"condition": "全新/几乎全新/明显使用痕迹",
"defects": [{"type": "划痕/污渍/缺损", "location": "位置描述", "severity": "轻/中/重"}],
"confidence": 0-100
}
要求:看不清的信息宁可null,不要编造。
defects为空数组时,condition不得标注"几乎全新"以上。
"""This rule of "prefer null over fabrication" was drilled into me by customer complaints. Trust in second-hand trading is built on accurate descriptions—AI hallucination is fatal here.
Step 2: Encapsulate Everything at the Gateway Layer, Invisible to Callers
We did three things at the gateway layer: persist request logs to the database (recording the model response and latency for each image), rate limiting per seller (a maximum of N generations per person per day), and primary/backup model switching (automatic fallback if the primary model times out for 3 seconds). Business code calls a single internal function with no idea which model is behind it—this saved enormous coordination costs when we later switched models.
Step 3: Make Human Feedback a Process, Not a Slogan
The key design in the operations console is "confidence-linked behavior": descriptions with confidence below 80 force the interface into edit mode, and the system tracks weekly "AI acceptance rate," "manual edit rate," and "edit type distribution." These three numbers are the reports I check in every monthly meeting—they tell us whether the prompts need optimization and which categories need their own dedicated templates.
IV. Why a Unified AI Gateway Reduces Maintenance Costs
Three months after launch, we went through one model vendor price increase and one interface response format change. Both were handled by a single colleague at the gateway layer within half a day, with zero changes to business code. I did the math: if five business modules each connected directly to the model API, every change would mean notifying the owners of five modules and scheduling five rounds of integration testing—at our team's pace, at least two weeks. With a unified gateway, such changes collapse into "one adaptation inside the gateway," and the manager doesn't need to re-coordinate staffing every time a model fluctuates.
Beyond consolidating changes, the gateway brings two hidden benefits: first, cost visibility—all token consumption across calls is aggregated in one place, so I can view the bill by feature line instead of receiving one confusing total at month's end; second, risk control—API keys are stored centrally on the gateway side rather than scattered across code repositories, which is the lowest-cost safety net for a small team.
V. Results and Lessons Learned
Once the process was up and running, the average time for operations to process an image dropped from two to three minutes to under 40 seconds (mostly confirmation and fine-tuning), and peak-period backlogs essentially disappeared. But I believe the more valuable outcome was establishing two team rules: "AI-generated output must pass through a human" and "all calls must go through the gateway"—these guarantee long-term project stability far better than any single-point optimization.
Three pieces of advice for managers of small teams working on AI adoption: map out the complete process before writing code; AI output must have a human checkpoint as a safety net; and unify your API gateway from day one—don't wait until the bill and the breaking changes show up on your doorstep together.
If you're looking for an AI gateway service that provides unified access to multiple models, it's worth consolidating your integration entry point into one place before you start. Here's a solution you can register and try directly: https://api.thistoken.ai/register
---
Every example in this post runs with a single API key — get yours at https://api.thistoken.ai/register and start in minutes.
Хотите попробовать Token.AI?
Создайте API Key уровня проекта, включите каналы в консоли и настройте маршрутизацию, бюджеты и журналы аудита.
注册 ThisToken.AI 并获取 API Key