## The Pitfalls We Hit First
The Pitfalls We Hit First
Last year in Q3, our team's monthly AI API costs more than tripled. My boss asked me to review it, and the first thing I did was wrong: I opened the billing page, saw the total, and started "guessing."
My guesses went something like: maybe the new intern's usage is high? Maybe some campaign script ran wild? Maybe the model prices went up? So I sent out a notice: "Everyone, please be frugal with AI API usage." The result? Costs kept climbing the next month.
This is the first common failed approach: governing based on total spend. The total only tells you "it hurts," not "where the bleeding is."
The second failure mode is more insidious: indiscriminate downgrading. We decided to switch all calls from large models to cheaper small models. As a result, the quality of code explanation tasks plummeted, colleagues secretly hooked up their keys to personal accounts, we barely saved any money, and data ended up flowing outside the company's boundary.
The third: cutting off caching across the board. Someone suggested that cached results might be stale and risky, so we should just turn it all off. That's equivalent to throwing away a discount already in hand—many providers charge extremely little or nothing for cache hits. Caching isn't a nice-to-have optimization; it's the biggest "discount zone" in your bill structure.
By the end of the review, I finally understood: without usage attribution and cache hit-rate visibility, any cost optimization is just guesswork.
The Right Path: Data First, Strategy Second
We later redid the entire governance process, centered on three actions: attribution, routing, and whitelisting. Let me break them down.
Method 1: Usage Attribution—Find the Owner of Every Penny
The real turning point came after we routed all calls through a unified gateway. Previously, twenty-plus keys were scattered across various scripts, services, and individuals—the billing was a mess. After connecting through ThisToken.AI's gateway, every call carried tags: which project, which team, which feature module, which model used, and whether the cache was hit.
With this data, I was stunned the first time I looked at the reports:
- An internal weekly-report generation script contributed nearly 30% of costs, because it re-summarized all the raw data from scratch every time, never reusing anything;
- The overall cache hit rate was only 11%, while peers doing this well can reach over 40% with similar workloads;
- A "test environment" was still running traffic on weekends—it was a forgotten debug loop.
Each of these three findings corresponded to a specific fix, rather than "sending a notice telling everyone to save money." The essence of attribution is turning costs from "everyone's responsibility" into "a list of specific problems."
Method 2: Cache-Friendly Routing—Don't Pay Twice for Repeated Work
The keys to improving cache hit rates are understanding which calls are "worth caching" and "how to increase hit probability":
- Stable prompt prefixes. Many providers' caching matches by prefix. Placing system prompts, few-shot examples, and long document context at the very beginning and keeping them literally stable significantly increases hit probability. We once put a dynamic timestamp at the start of a prompt—one character's difference, and the entire cache was invalidated.
- Decide whether to cache based on the scenario. High-frequency, stable-result scenarios like codebase interpretation, documentation Q&A, and internal knowledge base retrieval have extremely high cache value; scenarios with strong real-time requirements (like market analysis) shouldn't force it.
- Configure uniformly at the gateway layer. ThisToken.AI's managed channel capabilities let us configure cache policies for different projects at the gateway level, without each script having to implement it themselves. A policy change takes effect once, and the whole team benefits.
A month after these adjustments, our overall cache hit rate climbed from 11% to 38%, and the reduction in this portion of spending showed up directly on the bill—not by downgrading models, but by not paying twice for the same work.
Method 3: Model Whitelists and Conditional Routing—Save Expensive Models for Worthy Problems
Another truth revealed by the attribution data: a large number of low-value calls (format conversion, simple classification, log summarization) were using flagship models, while tasks that truly needed deep reasoning were being restricted out of fear of cost.
Our approach was to establish role-based whitelists + conditional routing:
| Team/Scenario | Allowed Model Tiers | Cache Policy | Monthly Budget Cap |
|---|---|---|---|
| Core algorithm team | Flagship + standard tier | On demand | High |
| Business dev team | Standard + lightweight tier | Mandatory | Medium |
| Internal tools/scripts | Lightweight tier | Mandatory + long TTL | Low |
| Experiments/sandbox | Lightweight only | Mandatory | Hard cap with circuit breaker |
Setting this up on ThisToken.AI is straightforward: define channels and model whitelists on the gateway side, bind rules to project keys, and automatically trigger circuit breaking or downgrading when budgets are exceeded—rather than assigning blame after the monthly bill comes out.
A Copy-Paste-Ready Budget Governance Checklist
| Check Item | Done? | Notes |
|---|---|---|
| All AI calls go through a unified gateway with project/team tags | ☐ | The prerequisite for attribution; without this step, nothing else matters |
| Review multi-dimensional usage and cost reports weekly | ☐ | Look at dimensions, not just totals |
| Cache hit rate has a clear baseline and is included in weekly reports | ☐ | A drop in hit rate often means prompts were accidentally changed |
| Prompt prefix stability is standardized (dynamic content like timestamps moved to the end) | ☐ | One prefix change wipes the cache |
| Each team has a model whitelist, not everyone using all models | ☐ | Expensive models are a scarce resource |
| Low-value scenarios routed to lightweight models/managed channels | ☐ | Don't use a flagship model to count logs |
| Each project has a budget cap and over-limit policy (alert/downgrade/circuit breaker) | ☐ | Ex-ante control beats ex-post blame |
| Regularly (monthly) clean up zombie keys and forgotten scheduled tasks | ☐ | Our weekend-running loop was a lesson learned |
Final Thoughts
After that review, I came to a conclusion: the ceiling of cost optimization depends on whether you can answer "who spent every penny, in what scenario, and at what hit rate." If you can't answer that, all you can do is downgrade models, send notices, and cut budgets—all blunt instruments that hurt the innocent. If you can answer it, your optimization actions become precise and verifiable.
If your team is still using scattered keys and making decisions based on billing totals, consider building out the gateway and attribution layer first. ThisToken.AI's gateway, model whitelists, managed channels, and routing governance capabilities cover exactly the first half of the checklist above—sign up and get started: https://api.thistoken.ai/register
---
Tired of juggling provider integrations? Register at https://api.thistoken.ai/register and call every model through one base_url.
Token.AI を試してみませんか?
プロジェクトレベルの API Key を作成し、コンソールでチャネルを有効にして、ルーティング、予算、監査ログを設定しましょう。
注册 ThisToken.AI 并获取 API Key