## First, a Failure Story
First, a Failure Story
Last year, an indie developer friend of mine asked me to review his project: an internal AI code assistant with a simple requirement — a user selects a piece of code in an IDE plugin, and the assistant provides modification suggestions and generates a patch. He spent two weeks building it, but after a two-week internal trial, no one used it anymore.
The problem wasn't model capability — it was that he built the whole system as "one straight line":
Mistake #1: The plugin calls the model API directly. He hardcoded an OpenAI key into an Electron plugin (even after packaging, it could be extracted by unpacking), manually distributed one key per user, and when someone left the company, the key leaked, forcing a full key rotation and re-release.
Mistake #2: Prompts scattered everywhere. The system prompt "You are a code assistant" was written once in the plugin and again in the backend log analysis script. After two rounds of requirement changes, the two copies were inconsistent, and the same question got completely different response styles in the plugin versus the logs.
Mistake #3: No middle layer — model output displayed as-is. When a user asked "help me refactor this function," the model returned a huge chunk of Markdown explanation plus code blocks, which the plugin rendered verbatim. What users actually wanted was a "directly usable diff" — and diff formatting, context window truncation, and timeout retries were all left unhandled. When the model occasionally returned incomplete JSON, the UI simply went blank.
Mistake #4: Switching models equals a rewrite. Later, he wanted to try Claude for better code handling, but discovered that the two vendors' SDKs, parameter names, streaming protocols, and error codes were all different. Integrating a second model would cost roughly as much as rewriting all his business logic, so he gave up.
This is a very typical pattern: mistaking "getting the model API to work" for "having built an AI application." The model call is just the easiest piece of the system — the real complexity lies beyond the call.
The Right Architecture: Split the "Straight Line" into Four Layers
During the rewrite, we restructured the architecture into four layers, each with a single responsibility:
IDE plugin (thin client)
│ Only handles interaction: collecting selected code, displaying diffs, one-click patch application
▼
Unified AI API Gateway (self-built or managed service)
│ Unified auth, routing, rate limiting, billing, logging
│ Exposes a model-agnostic interface to the layer above
▼
Orchestration service (business core)
│ Prompt template management, context assembly, output parsing, patch generation
▼
Model layer (replaceable)
│ GPT / Claude / open-source models, routed by task through the gatewayThe corresponding key workflow checklist:
- User selects code in the plugin + describes intent → the plugin sends only these two fields to the orchestration service
- The orchestration service loads versioned prompt templates and assembles context (relevant file snippets, project language conventions)
- The orchestration service calls the unified gateway, which routes to the appropriate model (strong models for code generation, cheap models for simple explanations)
- Model output is structurally validated (JSON Schema); parse failures trigger automatic retry and degradation
- Validated modification suggestions are converted to unified diff format and returned to the plugin
- The plugin shows a diff preview; the user confirms and applies the patch locally — the AI only ever produces suggestions; humans decide whether to apply them
Key Implementation Steps
Step 1: Centralized prompt management. All prompts live in one directory in the repository, versioned, with changes going through code review. The plugin and backend reference the same copy — no more "two places maintaining their own prompts."
Step 2: Enforce structured output. Require the model to return fixed JSON:
{
"summary": "修改说明",
"changes": [
{"file": "src/utils.ts", "diff": "@@ -10,3 +10,7 @@\n..."}
],
"confidence": 0.85
}The orchestration service validates against a schema; invalid output triggers one retry with the error message attached, and a second failure degrades to plain-text suggestions. After this step, the blank-screen problem never happened again.
Step 3: Route all model calls through a unified gateway. This was the highest ROI step of the entire refactor. The reasons are straightforward:
- Protocol normalization: The orchestration service integrates with a single OpenAI-compatible API. Switching models, adding models, or routing by task are all just configuration changes on the gateway — zero changes to business code. My friend's original "switching models equals a rewrite" problem was completely eliminated here.
- Centralized key governance: Keys live only on the gateway side; the plugin authenticates with short-lived tokens, drastically reducing leak risk and rotation costs.
- Visibility into observability and costs: Token counts, latency, and costs of every call are recorded uniformly. Which features burn money, which model offers the best value for a given task — you see it in the data instead of guessing.
- Unified rate limiting and degradation: When one model gets rate-limited, the gateway automatically switches to a backup model with no impact on the business layer.
For small teams, a self-built gateway can start with open-source solutions like LiteLLM; if you don't even want to handle ops, managed unified AI API gateway services can be used right after signing up, offloading maintenance costs to specialists — for projects with only one or two people, this is almost the only sensible choice.
Step 4: Close the loop on patch application. The plugin parses unified diffs; when application fails (e.g., the user's local code has changed), it prompts about the conflict and requests regeneration, rather than silently applying a misaligned patch.
The Result
In the refactored version, the plugin's codebase was cut in half, and all model-call-related code was consolidated into the orchestration service alone. Later, when they switched generation models from one vendor to another, the entire change was a single line of routing configuration in the gateway backend — not a single line of business code changed. The internal trial went from "nobody uses it" to "people file feature requests daily" — because responses were stable and outputs were usable, the tool entered a virtuous cycle.
A Few Takeaways
For applications like AI code assistants, model capability is rented — engineering capability is what you own. These four things — thin client, centralized prompts, structured output, and a unified gateway — each work to isolate "model uncertainty" outside the system boundary. If your team is building something similar, I recommend stabilizing the model integration layer first — you can start with a multi-model unified AI API gateway; register an account and get your first pipeline running in half an hour: https://api.thistoken.ai/register
---
Tired of juggling provider integrations? Register at https://api.thistoken.ai/register and call every model through one base_url.
Vous voulez essayer Token.AI ?
Créez une API Key au niveau du projet, activez les canaux dans la console et configurez le routage, les budgets et les journaux d'audit.
注册 ThisToken.AI 并获取 API Key