GPT-5.5$30.00/ 1M OUT CLAUDE FABLE 5$50.00/ 1M OUT GEMINI 3.1 PRO$12.00/ 1M OUT QWEN3.7 MAX$3.75/ 1M OUT DEEPSEEK V4 PRO$0.87/ 1M OUT SAKANA FUGU ULTRA$30.00/ 1M OUT OPENAI OTPI$1.02/ 1M OUT ANTHROPIC OTPI$1.81/ 1M OUT GEMINI 2.5 FLASH$2.50/ 1M OUT GPT-5.4 MINI$4.50/ 1M OUT GEMINI 2.5 FLASH-LITE$0.40/ 1M OUT GPT-5.4 NANO$1.25/ 1M OUT GPT-4.1$8.00/ 1M OUT GPT-5.5$30.00/ 1M OUT CLAUDE FABLE 5$50.00/ 1M OUT GEMINI 3.1 PRO$12.00/ 1M OUT OPENAI OTPI$1.02/ 1M OUT ANTHROPIC OTPI$1.81/ 1M OUT GEMINI 2.5 FLASH$2.50/ 1M OUT GPT-5.4 MINI$4.50/ 1M OUT GEMINI 2.5 FLASH-LITE$0.40/ 1M OUT GPT-5.4 NANO$1.25/ 1M OUT GPT-4.1$8.00/ 1M OUT GPT-5.5$30.00/ 1M OUT CLAUDE FABLE 5$50.00/ 1M OUT GEMINI 3.1 PRO$12.00/ 1M OUT OPENAI OTPI$1.02/ 1M OUT ANTHROPIC OTPI$1.81/ 1M OUT GEMINI 2.5 FLASH$2.50/ 1M OUT GPT-5.4 MINI$4.50/ 1M OUT GEMINI 2.5 FLASH-LITE$0.40/ 1M OUT GPT-5.4 NANO$1.25/ 1M OUT GPT-4.1$8.00/ 1M OUT
All notes
Desk Notes JUL 08 2026 · 8 MIN READ · IFX RESEARCH

The changing cost of AI: a year of API price prints

IFX Research reconstructed every observable LLM API price change in the PricePerToken history — 2,108 of them. The data does not support the one-line story that AI prices simply fall. It shows three pricing regimes.

The Changing Cost of AI — IFX Research cover
111,326
Daily observations
591
Model histories
2,108
Price changes
344
Days covered
00

Why price history, not price

API price is now a model property in its own right — it sits next to context length, latency, and quality when a team picks a model. A current price table answers one question: what does this model cost today. A price history answers a different one: how stable is that number when the budget has to survive the next quarter.

This note condenses an IFX Research study of the PricePerToken daily history from 2025-07-28 to 2026-07-07. All prices are USD per 1M tokens; a changed field (input, output, cache read, cache write) counts as one event.

Method Events reconstructed from daily rows per provider + model. Open / proprietary labels from the catalog's is_open field; unmatched histories stay Unknown. Same-day input and output moves count as two events — the buyer sees two unit prices.
01

Two markets in one table

Open-source and open-weight models produced 1,635 metric-level events against 256 for proprietary models. Among numeric changes, decreases are 50.3% of open-source events — but 66.2% of proprietary ones. Proprietary prices move rarely; when they move, the dollar effect on an output-heavy workload can be larger.

Fig 01 · Price-change events by model class AUG 25 – JUL 26
Open-source
1,635 197 models · 50.3% decreases · median abs move 25%
Proprietary
256 161 models · 66.2% decreases · median abs move 35%
Unknown
217 233 histories without a catalog match — kept visible, excluded from binary claims
One event = one changed price field in the daily PricePerToken history. Class taken from the catalog's is_open flag.

The higher open-source count is not one model repricing every day. It is breadth: many variants, many providers, active price discovery. The busiest month of the window is January 2026, with 241 events.

Fig 02 · Monthly events by class
OPEN PROP UNKN
241
AUG SEP OCT NOV DEC JAN FEB MAR APR MAY JUN JUL
2025 2026
Monthly totals; class split is proportional to the window shares. Peak activity: January 2026, 241 events.
02

Most prices end where they started

The first-to-last view is calmer than the event count. 81% of proprietary model-metric pairs end the window at the price they started at; open-source pairs split into cheaper, flat, and costlier thirds. The median net change is +0.0% in every group.

Fig 03 · First-to-last direction by model-metric pair
▼ CHEAPER FLAT ▲ COSTLIER
Prop · Input
13.1 · 81.2 · 5.6
Prop · Output
14.4 · 80.6 · 5.0
Open · Input
24.0 · 43.4 · 32.7
Open · Output
19.4 · 45.9 · 34.7
Share of pairs, %. Flat = last observed price equals the first observed price in the window.

Plotting each pair by its first and last observed price makes the shape of the market visible. The corpus mixes sub-dollar open-weight APIs with expensive frontier and reasoning models, so the scale is logarithmic.

Fig 04 · First vs last observed price $/MTOK · LOG–LOG
FLAT 0.01 0.1 1 10 100 0.01 0.1 1 10 100 FIRST OBSERVED PRICE LAST PRICE
Stylised view of input / output model-metric pairs. Points below the dotted line ended cheaper; above it, costlier. Grey points ended flat.

A 50% cut on a $0.10 API is small money. A 10% move on a $20 output price is not.

Percent change and dollar exposure have to be read together. Cheap open-weight APIs show large percentage moves on small bases; a smaller move in an expensive proprietary output price can matter more for a workload that generates many tokens.

03

Cache pricing is a separate layer

A single "price per token" is no longer enough. Input and output carry most events, but cache read and cache write prices move independently — and they were introduced or removed 268 times in the window. A model with a stable headline price can still get cheaper or dearer for workloads that reuse long prompts.

Fig 05 · Events by price field 2,108 TOTAL
Input
809
Output
743
Cache read
432
114 introduced · 51 removed
Cache write
124
69 introduced · 34 removed
Cache fields (orange) are a separate economic layer: retrieval-heavy and agentic workloads pay a different effective price than short chat on the same model.
04

Where the activity lives

Qwen alone accounts for 549 events — more than every closed frontier line combined. The closed families behave like managed product lines: prices sit still within a variant, then move in steps when a new variant or cache term ships. Claude logged 26 events in the whole window; Gemini 41.

Fig 06 · Price-change activity by family EVENTS
Qwen
549
DeepSeek
263
Z-AI GLM
222
Moonshot Kimi
161
Meta Llama
120
Mistral
105
Google Gemma
97
MiniMax
91
OpenAI GPT/o
85
Google Gemini
41
Anthropic Claude
26
One event = one changed price field. High counts reflect broad families — base, instruct, reasoning, coder, VL, and size tiers — not one model repricing daily.

Net moves are just as uneven. Three families stand out against the flat median:

Qwen
▼ −30.4% IN
▼ −20.0% OUT
Meta Llama
▲ +100.0% IN
▲ +42.9% OUT
Google Gemma
▲ +33.3% IN
▲ +104.2% OUT

The single most active models are all open-weight — including OpenAI's own gpt-oss-120b, the busiest row in the corpus:

Most active models EVENTS
gpt-oss-120bOpenAI GPT/o43
kimi-k2.5Moonshot Kimi40
deepseek-v3.2DeepSeek35
kimi-k2.6Moonshot Kimi34
glm-5.1Z-AI GLM32
05

The practical rule

The study is deliberately narrow: one public source, daily granularity, no private contracts, no usage-share estimates. Within those limits, the window shows three regimes — stepwise proprietary lines, actively repricing open-weight families, and a growing cache layer that headline prices miss entirely.

For buyers, the rule is simple: record price history, not only current price. A current price is enough for a one-off comparison. A history is needed for a budget, a routing policy, or any model choice that has to survive more than one billing cycle.

Trade the cost of intelligence.

Sources
PricePerToken — pricing catalog · api.pricepertoken.com/api/pricing
PricePerToken — provider pricing history · api.pricepertoken.com/api/provider-pricing-history
Data retrieved 2026-07-07 22:20 UTC · IFX Research
Full research PDF

The complete PricePerToken pricing-history study, including the full methodology notes and supporting tables.

Read full study