The changing cost of AI: a year of API price prints
IFX Research reconstructed every observable LLM API price change in the PricePerToken history — 2,108 of them. The data does not support the one-line story that AI prices simply fall. It shows three pricing regimes.
Why price history, not price
API price is now a model property in its own right — it sits next to context length, latency, and quality when a team picks a model. A current price table answers one question: what does this model cost today. A price history answers a different one: how stable is that number when the budget has to survive the next quarter.
This note condenses an IFX Research study of the PricePerToken daily history from 2025-07-28 to 2026-07-07. All prices are USD per 1M tokens; a changed field (input, output, cache read, cache write) counts as one event.
Two markets in one table
Open-source and open-weight models produced 1,635 metric-level events against 256 for proprietary models. Among numeric changes, decreases are 50.3% of open-source events — but 66.2% of proprietary ones. Proprietary prices move rarely; when they move, the dollar effect on an output-heavy workload can be larger.
The higher open-source count is not one model repricing every day. It is breadth: many variants, many providers, active price discovery. The busiest month of the window is January 2026, with 241 events.
Most prices end where they started
The first-to-last view is calmer than the event count. 81% of proprietary model-metric pairs end the window at the price they started at; open-source pairs split into cheaper, flat, and costlier thirds. The median net change is +0.0% in every group.
Plotting each pair by its first and last observed price makes the shape of the market visible. The corpus mixes sub-dollar open-weight APIs with expensive frontier and reasoning models, so the scale is logarithmic.
A 50% cut on a $0.10 API is small money. A 10% move on a $20 output price is not.
Percent change and dollar exposure have to be read together. Cheap open-weight APIs show large percentage moves on small bases; a smaller move in an expensive proprietary output price can matter more for a workload that generates many tokens.
Cache pricing is a separate layer
A single "price per token" is no longer enough. Input and output carry most events, but cache read and cache write prices move independently — and they were introduced or removed 268 times in the window. A model with a stable headline price can still get cheaper or dearer for workloads that reuse long prompts.
Where the activity lives
Qwen alone accounts for 549 events — more than every closed frontier line combined. The closed families behave like managed product lines: prices sit still within a variant, then move in steps when a new variant or cache term ships. Claude logged 26 events in the whole window; Gemini 41.
Net moves are just as uneven. Three families stand out against the flat median:
The single most active models are all open-weight — including OpenAI's own gpt-oss-120b, the busiest row in the corpus:
The practical rule
The study is deliberately narrow: one public source, daily granularity, no private contracts, no usage-share estimates. Within those limits, the window shows three regimes — stepwise proprietary lines, actively repricing open-weight families, and a growing cache layer that headline prices miss entirely.
For buyers, the rule is simple: record price history, not only current price. A current price is enough for a one-off comparison. A history is needed for a budget, a routing policy, or any model choice that has to survive more than one billing cycle.
Trade the cost of intelligence.
PricePerToken — provider pricing history · api.pricepertoken.com/api/provider-pricing-history
Data retrieved 2026-07-07 22:20 UTC · IFX Research
The complete PricePerToken pricing-history study, including the full methodology notes and supporting tables.