Cheaper AI tokens keep raising advisory tech bills
A single 50,000-token planning task explains why advisory firms should watch the meter more than the unit price.
Kevin Hughes, president of financial planning at Advyzon Investment Management, ran a planning agent through a test recently and got a number he was not looking for: one action consumed 50,000 tokens, a developer told him, as Financial Planning reported. The per-token price keeps falling, and the invoice keeps growing.
Token costs fell from $20 per million to as little as seven cents through 2024, according to a McKinsey report citing data from Stanford's Institute for Human-Centered AI, and enterprise spending on large language models tripled in the twelve months after. A survey cited in the Financial Planning article found 93% of respondents had blown through their token budgets, and one in five had reined in AI use because costs were climbing.
Tokens are the units vendors use to price AI, and the meter does not behave like any software license an advisor has bought. A model breaks text into small fragments before processing, counts everything sent in—the prompt, an attached client document, the conversation history—as input tokens, and prices its answer as output tokens, which is often the expensive direction.
Tokenization does not map cleanly onto words, as an explainer from Sentisight cited in the article shows using Google's Gemini: 1,000 tokens run roughly 750 words, four characters make one token, and "fantastic" might become "fan," "tas," "tic." An advisor writing a prompt in plain English is sending a pre-sliced bill to the meter.
Subscription software has a ceiling: one CRM seat costs the same whether it is opened twice a day or twice a month. Token-based AI has no ceiling, because each request is a separate transaction drawing on input and output buckets with separate prices.
Hughes describes the model the way anyone who grew up in an arcade will understand. "You put your quarters in it, and you keep playing until you're officially done," he told Financial Planning, "or you stop and come back at a later time to finish the project when the clock resets."
Efficiency cuts the other way. As models become more intuitive and efficient, the number of tokens needed for a task falls and computing costs fall with it, as Sentisight notes. Cheaper units invite more use, which is how falling prices and tripled spend are compatible, and why advisory firms need to watch the meter rather than the rate.
The rate is still the first thing vendors quote. Ray Wu, founding managing partner of Alumni Ventures, notes that the model price spectrum is wide, with OpenAI currently offering models from roughly 20 cents per million tokens. A number like that tells an advisor almost nothing; the number that matters is tokens per completed job.
Hughes's team surfaced exactly that number: a planning agent consumed 50,000 tokens doing one thing. At headline rates the dollar cost of a single run looks small; across a practice doing that work daily, with output often priced higher than input, it is a different story.
Vetting an AI vendor should therefore start with three questions: how input and output tokens are counted, what a representative task consumes, and whether the account has a usage cap before an overage bill arrives. A meter with no cap is a budget with no floor.
Last month this publication argued that AI is the capacity answer to the advisor shortage. That thesis survives Financial Planning's numbers, but the billing reality adds an accounting requirement: the capacity comes metered. A firm using AI to stretch a lean team has to count tokens the way it counts headcount, because the efficiency is only captured if the invoice stays below the cost of the labor saved.
The adoption gap between heavy and light AI users is turning into a cost-control gap. Heavy users who ignore the meter hand the efficiency gains back to the vendor; light users who start with a meter in place can scale without learning the lesson on a blown budget.
Hughes's test run is the right starting budget for any practice adopting an agent: ask the vendor to run one representative task and show the token count before signing. If the count is anywhere near that figure, the price per million tokens barely matters; the number to negotiate is the cost of completing the job at the volume the firm will actually run.