Cutting API Costs by 40%: How I Used Caching and Tiered Routing to Eliminate Duplicate Prompt Spending
Last month, while doing usage attribution, I noticed a glaring number: roughly 30% of the API requests we send out every day are "nearly identical" prompts—repeated compliance review scripts, fixed translation templates, and industry briefing summaries that run on a schedule every morning. Every one of these requests gets billed at full price. Converted to an annual figure, that's enough to hire half an intern.
This batch of "duplicate spending" requests was the first thing I targeted in this round of budget governance. This article shares exactly how I did it, along with before-and-after efficiency comparisons.
First, Do the Math: How Much Budget Do Duplicate Prompts Actually Eat?
Before taking action, I did the attribution first. I pulled a month's worth of our team's call logs, clustered them by prompt prefix, and roughly broke them into three categories:
| Request Type | Share | Duplication Level | Priority |
|---|---|---|---|
| Fixed templates (briefings, translation, compliance checks) | ~30% | 100% identical prefixes | Highest |
| Semi-fixed (templates with a few variables) | ~40% | >80% identical prefixes | Medium |
| Truly one-off requests | ~30% | Basically no duplication | Low |
In other words, 70% of our traffic has a duplicate "skeleton." Without caching this portion, we were paying repeatedly for the same inputs every month. Based on our call volume at the time, even if cached hits only saved half the cost, we could cut nearly 40% of total annual API spending—this wasn't achieved by cutting features; it was pure efficiency gain.
Method 1: Application-Layer Semantic Caching—The Most Direct Way to Stop the Bleeding
The first approach was the simplest: add a caching middleware layer into the call chain. For exactly identical prompts (or prompts normalized via hashing), first check local Redis; on a hit, return the previous response directly. Only on a miss do we call the upstream API, and on writing back to the cache, set a reasonable TTL.
For that 30% of fixed-template requests, the effect of this change was immediate: in the first week after launch, billed call volume for briefing-type requests dropped by more than 60%. And response latency dropped from one or two seconds to tens of milliseconds—saving not just money, but user wait time too.
Two things to watch out for: first, the cache key must include the model name and key parameters; otherwise, returning stale results after switching models will pollute your output. Second, be careful with TTL settings for prompts involving real-time information (like "today's news summary"). I stepped in that trap once—a client came asking why they received yesterday's briefing.
Method 2: Gateway-Level Caching and Routing Governance—Managing Policy Centrally
Application-layer caching has a weakness: with three teams and five or six services each writing their own caching logic, hit rates were inconsistent, and the attribution numbers didn't reconcile. So in the second step, I moved caching policy up to the gateway layer—one of the core reasons I chose a unified gateway like ThisToken.AI.
After configuring it at the gateway, several changes were immediately apparent:
- Unified caching policy: All downstream services automatically inherit the same set of caching rules. No more reinventing the wheel per team—integration time went from two or three days of individual development down to a single config change;
- Clear usage attribution: Gateway logs directly distinguish "cache hit returns" from "actual forwarded billable calls." At month-end review, it's obvious at a glance which money went to duplicate requests—no more writing my own clustering scripts;
- Model whitelist as a safety net: For requests that miss the cache and must make real calls, they can only go through whitelisted models and hosting channels, preventing someone from temporarily switching to an expensive model and blowing through the budget.
Method 3: Tiered Routing—Cheap Models Handle the Repetitive Work
The third method goes one step further: not every request deserves the flagship model. I configured tiered routing on the gateway—
- Cache hits: returned at zero cost;
- Fixed templates with expired cache: routed to a low-cost model tier;
- Genuinely complex one-off requests: these are the only ones that go to the flagship model.
Combined with the whitelist and channel quotas, every model now has a monthly spending cap. One month after this rule set went live, our team's total API spending dropped about 40% month-over-month, with pure caching contributing roughly half and tiered routing the other half. Average response latency also dropped by more than 30%, since a large volume of requests never left the internal network.
A Budget Governance Checklist
I've distilled this experience into an actionable checklist you can apply directly:
| Check Item | Action | Expected Benefit |
|---|---|---|
| Prompt clustering analysis | Pull one month of logs and measure the share of duplicate prefixes | Understand your cacheable potential |
| Application-layer hash caching | Add Redis caching + TTL for highly duplicated templates | Reduce billed volume for that category by 50%+ |
| Unified gateway caching | Configure global caching policy on the ThisToken.AI gateway | Unified policy, clear attribution |
| Model whitelist | Restrict available models and hosting channels | Prevent budget overruns |
| Tiered routing | Route repetitive tasks to low-cost models | Another notch down in unit price |
| Hit-rate monitoring | Weekly review of cache hit rates and per-channel spend | Continuously iterate on strategy |
Conclusion
When budget governance is all said and done, you'll realize the most expensive thing is often not complex tasks, but the duplicate calls nobody pays attention to. Caching isn't a minor optimization—it's a fundamental discipline that ensures every dollar goes toward new information. If your team is still writing separate caches for each service and pulling their hair out over month-end reconciliation, try consolidating caching, routing, and attribution onto a single gateway—you can start validating with just one account registration: https://api.thistoken.ai/register
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Token.AI を試してみませんか?
プロジェクトレベルの API Key を作成し、コンソールでチャネルを有効にして、ルーティング、予算、監査ログを設定しましょう。
注册 ThisToken.AI 并获取 API Key