## I
I. First, the Failures—Then How We Recovered
Last year, two friends and I took on a project: adding a "smart form-filling" feature to a website that sells legal document template downloads. The user picks a "Residential Lease Agreement" template, fills in a few key pieces of information (Party A, Party B, rent, term), and the system automatically fills in dozens of related fields throughout the entire document.
Sounds easy, right? Our first version was crude and straightforward:
Failure #1: Throwing the entire template at the LLM and asking it to "fill in the blanks for us."
The results were a disaster. The word "Lessor" appeared 23 times in the contract; the model filled in 22 instances and missed 1. Converting amounts to Chinese uppercase numerals often went wrong ("壹万贰仟" written as "一万二千"). Even worse, the model would occasionally "helpfully" rewrite the wording of contract clauses—an absolute red line in legal documents, where not a single character may be changed.
Failure #2: Switched between three models for repeated testing, rewriting the call code each time.
Model A's JSON output was unstable, so we wrote a compatibility layer. Switching to Model B meant re-tuning the prompt format all over again. Model C was cheap but poor at understanding Chinese monetary amounts. Three weeks went by, half of them wasted on "switching models," while not a single line of business logic moved forward.
Failure #3: The frontend stuffed the user's raw input directly into the prompt.
A user typed "Zhang San, and by the way, make it sound harsher" in the "Party B name" field—and the model actually complied. We had treated the seriousness and safety of legal documents as an afterthought.
II. The Right Path: The Model Only Understands, Rules Do the Generating
After our post-mortem, we threw everything out and started over. The core principle was just one thing: the LLM never touches the template body—it only does what it's good at.
A Four-Layer Architecture
- Template Layer: Templates are pre-processed into structured fields. Each template is parsed in the backend into a "field list + body text with placeholders." The body is read-only and cannot be modified at any stage.
- Understanding Layer: The user fills in natural language (e.g., "two-year lease, 8,000 per month, one month deposit, three months rent upfront"). The LLM has exactly one task—extract this into structured fields:
{lease_months: 24, monthly_rent: 8000, payment_terms: deposit_one_pay_three}. Output is strictly validated against a JSON Schema; failures trigger a retry.
- Rules Layer: Uppercase numeral conversion, date formatting, and title linkage ("Lessor" ↔ "Party A") are all hard-coded with deterministic logic. These are utility functions you can write in a few hundred lines—ten thousand times more reliable than praying the model gets it right.
- Generation Layer: Placeholder replacement, followed by a full-document validation before output—every placeholder must be gone, and clause text must match the template character by character. Any discrepancy triggers an error and blocks the output.
Key Implementation Steps (Process Checklist)
1. 模板入库:人工标注字段清单,生成占位符版本正文
2. 用户输入 → 发送到AI网关 → 抽取提示词
3. 模型返回JSON → Schema校验(失败重试最多2次)
4. 字段补全与规范化(本地规则:大写金额、日期、联动字段)
5. 占位符替换 → 残留占位符检查(有残留则告警)
6. 正文diff校验(条款区域与原文比对,零容忍)
7. 渲染输出 PDF / WordThe diff validation in step 6 was the final gate we added after learning the hard way. It's less than fifty lines of code, but since launch it has intercepted several instances of the model "improvising."
III. Why the API Gateway Saved Our Maintenance Costs
The root problem behind Failure #2 was this: the call code was too tightly coupled to specific models. Later, we consolidated all model calls behind a unified AI API gateway, and three benefits became immediately apparent:
- Switching models requires no code changes. The gateway layer handles protocol adaptation. We configured different model endpoints for A/B testing, while the business code only knows its own abstract interface. We later discovered that a domestic Chinese model actually performed better at extracting Chinese monetary information at half the cost—switching over took only half an hour of config changes.
- Failure retries and fallbacks are handled uniformly. Timeouts, rate limits, JSON format errors—all the retry logic is written once at the gateway side and shared by all template types, instead of being reimplemented for every feature.
- Centralized call logging. Legal scenarios require an audit trail of "how this document was generated at the time." The gateway's unified logging satisfies this directly, with no need to add instrumentation in the business code.
For a small team, this means any change in the model layer—price hikes, deprecations, quality degradation—is just a configuration change, not a code refactor. The three of us maintain a dozen-plus template types, and model-related code accounts for less than 5% of the entire project.
IV. Three Lessons for Independent Developers
- First ask "what should the model do," then ask "which model to use." Structured extraction is one of the most reliable capabilities of current LLMs. Let it work within boundaries—don't let it improvise.
- Deterministic logic always comes first. Anything that can be hard-coded should not be left to probability.
- Consolidate the integration layer from day one. Even with just one call site, route it through a gateway first. When you later add templates, switch models, or add features, your costs grow linearly, not exponentially.
The project ended up running very smoothly. The templates grew from the initial 3 to over 20, and the extraction layer barely changed. If you're building a similar AI application and want to manage model calls in a unified way, give this gateway service a try: https://api.thistoken.ai/register
---
Tired of juggling provider integrations? Register at https://api.thistoken.ai/register and call every model through one base_url.
Bạn muốn thử Token.AI?
Tạo API Key cấp dự án, bật kênh trong bảng điều khiển và định cấu hình định tuyến, ngân sách và nhật ký kiểm tra.
注册 ThisToken.AI 并获取 API Key