Silicon Data’s token-price index just printed $0.96 per million tokens — the first reading under a dollar and the cheapest rate on record [1]. Five days earlier, OpenAI put a $500-a-month price tag on its fastest ChatGPT seat [3]. My read: these are the same event, and neither shows up in the per-token budget a chief financial officer (CFO) signed off.

The big picture:

The index gives the cleanest single number for what the market pays for a Large Language Model (LLM) token, and by that measure the rate is near zero. It crossed below a dollar this month, more than half below its summer high [1][2].

The easy read is that Artificial Intelligence (AI) costs are falling. The bills say otherwise.

In Futurum’s second-half survey of tech buyers, 46.9% of enterprises blew past their AI budgets, while only 5.6% came in under plan [4]. CloudZero’s panel shows the same strain: the top quarter of firms now spend more than a tenth of their cloud bill on AI, a line crossed in a single month [5].

What strikes me is how the input price and the final bill move in opposite directions. The reason sits on the vendor’s pricing page: providers no longer meter the bill with the one number buyers watch.

OpenAI makes the split plain. Its seats cost $100, $200 and $500 a month, but the fastest mode — called Ultrafast — is locked to the $500 plan [3].

That same week, the $200 plan cut the usage it includes for new customers and held the price flat [3]. Vendors stopped competing on the cost of a token and started charging for speed and access. A plain token meter sees neither.

By the numbers

  • $0.96 — Token price low: Silicon Data’s index fell under a dollar for the first time on Oct. 4, sitting more than half below its summer peak [1][2].
  • 46.9% — Over-budget enterprises: Only 5.6% of firms came in under plan, and 10% had no formal AI budget to measure against [4].
  • 47.6% — Extra-funding requests: The most common answer to an overrun was to ask for more money, not to slow the work [4].
  • 11.12% — AI share of the cloud bill: The top quarter of firms crossed a tenth of cloud spend, up from 9.63% a month earlier [5].

What I’d watch:

The part I keep circling is what firms do when the overrun lands. Futurum found that 47.6% went back for more money and 43.3% pushed the gap into the next planning cycle [4]. Only about one in six cut or paused the work [4].

Nearly four in ten moved funds from elsewhere in the tech budget, and almost a quarter parked the spend in a non-tech business unit [4]. That funds the growth and hides the gap, but it does not control the cost. It also shifts who owns the call, since the unit that pays part of the bill tends to shape the next contract.

Three signals I am watching next:

  • The meter shift: Vendors now bill on speed tiers, usage allowances and volume. I want to see which firms rework those lines instead of chasing a cheaper token.
  • The hard ceiling: A fixed cap per team, agent or finished task is still the one lever that moves the total. I am watching how many firms set one, and how they pick the limit.
  • The work-unit: Cost per finished job — a closed ticket, a processed invoice — gives a finance team a number it can act on. Raw token totals do not.

The catch

I could be wrong that the token price and the final bill are pulling apart everywhere. A team running short, simple prompts really does save money when the rate falls, and some of those 46.9% overruns are just healthy new use.

The catch is the direction of the pull. When a unit gets cheaper, people buy more of it. An agent re-sends its whole context on every step, so one job can burn many times the tokens of a single question.

A lower rate makes those heavy jobs look cheap up front, and each new job is a line finance never planned. Cheaper tokens can lower the unit cost and still raise the total. Only a ceiling tied to finished work tells you which one is playing out.

At a glance

  • The Big Shift: Token prices hit a record low this month while enterprise AI budgets ran over — 46.9% of firms blew past their plans in the latest survey. The two are the same story, because vendors now charge for speed and usage, not raw tokens.
  • Why It Matters: If the rate card falls but the bill climbs, a budget built on cheaper tokens is a forecast built on the wrong number. The gap lands first on the tech budget, then on the business units.
  • What I’d Watch: I am watching how firms change the meter, not the price.
  • The ceiling: A hard cap per team or per agent — the one lever that changes the total.
  • The work-unit: Cost per finished job, such as a ticket closed, which a finance team can act on.
  • The allowance line: What vendors charge for speed and volume once the token gets cheap.
  • The Catch: The split is not the same everywhere. Short, simple work still gets cheaper as the rate falls, and some overruns are healthy use. The risk is that a falling rate quietly pays for tests nobody budgeted for.

Related reading

Sources

[1] Silicon Data — “LLM Token Expenditure Index” (reading as of October 4, 2026) — https://www.silicondata.com/products/silicon-index/llm-token-expenditure-index [2] CNBC — “Artificial intelligence token prices are hitting new record lows” (September 1, 2026) — https://www.cnbc.com/2026/09/01/ai-token-prices-lows.html [3] OpenAI — “About ChatGPT Pro tiers” (DevDay, September 29, 2026) — https://help.openai.com/en/articles/9793128-about-chatgpt-pro-tiers [4] Futurum Research — “Enterprise AI Overruns Hit 46.9%: Is the Reckoning in FY2027?” (September 14, 2026) — https://futurumgroup.com/insights/enterprise-ai-overruns-hit-46-9-is-the-reckoning-in-fy2027 [5] CloudZero — “Your AI Economics Pulse for September 2026” (September 8, 2026) — https://www.cloudzero.com/blog/ai-economics-pulse-september-2026