From Leaderboards to Bills: An Efficiency-First Look at Model Selection for AI Applications
As AI application developers, when we discuss model selection we tend to stare at leaderboards, but what actually determines our monthly bill is often the details hidden in API parameters and response headers. This post takes an efficiency perspective and runs through some concrete numbers: how tokens are spent, how caching is used, how requests are saved. The figures come from before-and-after measurements of a mid-sized conversational application on our team (roughly 8 million calls per month), offered here for reference.
Trend 1: Structured Output No Longer Relies on "Begging the Model"—It's a Native API Capability
In the past, to get a model to reliably output JSON, the mainstream approach was to repeatedly emphasize the format in the prompt, then write hundreds of lines of fault-tolerant parsing code. A single failed retry would double the token cost.
Today, mainstream LLM APIs generally offer native support for structured output / JSON mode, with format constraints enforced server-side. Here's our before-and-after comparison of the migration:
- Before migration: format failure rate around 7%, averaging 1.07 retries per successful request, roughly 400 lines of fault-tolerant parsing code
- After migration: format failure rate near zero, retries eliminated, parsing code shrunk to a few dozen lines
At the prices we were paying at the time, eliminating retries alone cut monthly token spend by about 6%. Even more important was debugging time: format-related bugs used to take developers half a day on average to track down; now they've essentially vanished.
Trend 2: Prompt Caching Goes from "Easter Egg" to "Standard Playbook"
Context caching, now supported by API providers one after another, may be the optimization with the highest ROI today. The principle is simple: repeated long prefixes (system prompts, tool definitions, few-shot examples) get cached, and hit portions are billed at steeply discounted rates—typically 10%–50% of the original price depending on the provider, with noticeably lower response latency as well.
Our application has a fixed system prefix of about 3,000 tokens. After enabling caching:
- Input token costs dropped by about 55% (cache-hit portions billed at the discounted rate)
- Time-to-first-token (TTFT) fell from an average of 1800ms to about 900ms—users felt it was "twice as fast"
- Integration effort: half a day, mostly verifying that the prefix structure was stable and dynamic content was placed at the end
The direct impact on developers is architectural: prompt organization shifts from "casual concatenation" to the discipline of "static prefix + dynamic suffix." Whoever establishes this discipline first gets the discount first.
Trend 3: Batch APIs—For Jobs That Aren't Rushed, 50% Off at Minimum
Any task whose results aren't returned to users synchronously (content moderation, log summarization, bulk labeling, offline evaluation) can go through batch APIs. Most providers offer a 50% discount, with the tradeoff being a completion window of several hours.
After we migrated our nightly batch log analysis:
- Cost for that workload: ~¥4,200/month → ~¥2,000/month
- Migration effort: one day, mostly converting synchronous calls into task submission + polling for results
Impact on model selection: batch pricing lets you "upgrade tiers"—tasks that only dared use lightweight models in synchronous scenarios may cost less with flagship models in batch scenarios than lightweight models in synchronous ones. Offline evaluation can therefore run more thoroughly, which in turn improves the quality of online model selection.
Trend 4: Finer-Grained Model Tiering Allows More Aggressive Routing
The capabilities of lightweight-tier models keep rising, while the price gap remains at 1/10 of flagship models or even less. We built a simple routing layer: lightweight models handle requests first, escalating to flagship models only when confidence is insufficient. About 80% of requests are completed by the lightweight tier.
- Overall per-request cost dropped by about 65%
- Quality spot-check scores declined by only 0.3 points (on a 10-point scale)
This changes the nature of "model selection": from a one-time decision to a runtime, per-request decision. What developers need to focus on is no longer "which model to choose," but "how to set routing thresholds and how to build fallback chains."
Recommendations for Developers
- Audit first, then optimize. Pull a month of call logs and track three things: format-failure retry rate, the proportion of fixed prefixes in input tokens, and the synchronous/asynchronous request ratio. These three numbers essentially bound how much you can save.
- Implement in this order: structured output → prompt caching → batch task migration → tiered routing. The first two show results within half a day to a day—prioritize them.
- Make token cost an observable metric. Log the token composition of every request at the gateway layer (cache hit/miss, input/output), otherwise optimization results can't be quantified.
- Set routing thresholds using offline evaluation. Batch APIs make large-scale evaluation cheap—don't just guess.
- Abstract away vendor differences. Cache-hit rules and batch discounts vary across providers; use a unified gateway layer to encapsulate them, so switching and price comparison stay cheap.
The Bottom Line
Taken together, these four measures cut our application's monthly API spend from about ¥31,000 to about ¥11,500—a 63% reduction—while user-facing latency actually improved. None of these optimizations depended on "switching to a better model"; they all came from fully utilizing existing API capabilities.
Efficiency optimization presupposes clear usage data and a basis for comparing multiple providers. If you're looking for a unified entry point to manage token costs, caching policies, and multi-model routing, you can start with https://api.thistoken.ai/register —once you register, you can get these numbers straight.
---
Ready to try it yourself? Sign up at https://api.thistoken.ai/register to get your API key and start building.
Bạn muốn thử Token.AI?
Tạo API Key cấp dự án, bật kênh trong bảng điều khiển và định cấu hình định tuyến, ngân sách và nhật ký kiểm tra.
注册 ThisToken.AI 并获取 API Key