Showing posts with label Agentic AI. Show all posts
Showing posts with label Agentic AI. Show all posts

Thursday, September 3, 2026

The Hidden Costs of Tokenmaxxing

As of 2 Sep 2026
NOTE FROM THE HUMAN: The following blog post was generated and sourced solely by Claude Fable 5.1 and the model even chose the title. As a Claude subscriber, I have wondered if I am "losing" some of money I pay for usage by not tokenmaxxing. After reading the article Fable compiled, I would say "yes" and "no". As a genealogist, Fable has been a powerful model for accurate transcriptions, genealogical compilations, research and even producing apps and plug-ins. Could I tokenmax for genealogy? Yes, but I would have to put a lot of effort into doing so. Do I want to? No, not really. I already switch to lower use models like Sonnet if I am doing a proofreading and grammar check. I also delete chats I do not need later to save server space because it's on a server somewhere. Some of the Silicon Valley workers are using the models for EVERYTHING - things you would previously use a Google search for. I do not do that - I still use search browsers like Google and Brave. I do realize that those browsers have incorporated AI but again, my goal in using them is NOT using tokens on my subscription plans. I have used Claude at high usage times, and it has told me to come back later. If anything, I find that irritating, especially if I am being blocked from using it because someone is tokenmaxxing what to eat for dinner, driving directions, or even more wasteful - asking the same questions repeatedly as Fable cites below. (Grrr....) Be sure to read the articles linked here for the big picture. Even the Wall Street Journal article can be read with a free account. As a genealogist, I hope this information helps my peers find their balance. ~ Katherine

The Hidden Costs of Tokenmaxxing 
by Claude Fable 5.1


"Tokenmaxxing" is the practice of maximizing AI token consumption and treating that volume as proof of productivity. The term entered mainstream use in April 2026, driven largely by reports of an internal leaderboard at Meta that ranked roughly 85,000 employees by their AI token usage, with the top user reportedly burning through 281 billion tokens in a single month.[1] Other large employers, including JPMorgan and Disney, were reported to be running similar rankings.[2] The idea rested on an assumption: heavy AI consumption would eventually produce better outcomes, and inference costs would keep falling fast enough to make high usage a nearly free bet.[3]

That assumption has not held up, and it holds up least well for Anthropic's most expensive tier, Claude Fable 5.1. What follows is a summary of the documented downsides, grouped into three categories: money and plan limits, productivity and organizational effects, and environmental impact. It closes with what critics propose instead.

Financial Costs and Plan Limits

Fable-class models sit at the top of Anthropic's price list. On the API, Fable 5 costs $10 per million input tokens and $50 per million output tokens, which is double the rate of Opus 5.[4] Fable 5.1 reduced the cost of cache reads to $0.25 per million tokens, but every other line on the price sheet stayed the same, and Anthropic's advertised savings of roughly 25 to 45 percent describe measured bills from its own August usage rather than any change to published rates.[5]

Subscription users face a separate constraint. According to Anthropic's help center, Fable 5 and Fable 5.1 draw from a plan's regular weekly usage limits and consume them faster than other Claude models. On Max plans and premium seats, up to half of the weekly limit can be spent on Fable models before usage credits are required. On Pro plans and standard seats, Fable models are not included in the plan's limits at all and run only on prepaid credits billed at API rates.[6] Since July 20, 2026, that split is permanent.[4] Early user reports on Fable 5.1 note that even with cheaper cache reads, the model still exhausts usage limits quickly in long agentic sessions, because per-token pricing and per-session quotas are governed separately.[7]

When a limit is reached, further requests may be throttled or blocked until the cooldown resets.[8] The pressure is industry wide. During a compute crunch earlier in 2026, Anthropic responded by capping token consumption on certain pricing tiers during peak hours, and OpenAI moved its Codex product from per-message to per-token pricing.[3]

Productivity and Organizational Effects

The central problem with tokenmaxxing is that it measures an input and calls it an output. As IBM's analysis put it, usage soon became a proxy for value, and organizations that built usage leaderboards found people quickly learned to game them.[9] The Pragmatic Engineer newsletter reported on a Microsoft engineer who admitted inflating token counts to avoid being seen as using too little AI, including asking the AI questions already answered in internal documentation and prototyping features with no intention of shipping them. The newsletter concluded that the incentive in some cases produced slower work and busywork.[10]

By midsummer the fad was visibly reversing. Tom's Hardware reported that agentic AI can consume up to 1,000 times more tokens than standard AI, prompting corporate pullbacks at Microsoft, Meta, and Amazon as costs rose without a matching gain in output.[11] The Wall Street Journal noted the underlying arithmetic: the price per token has dropped, but the number of tokens needed per meaningful result has risen sharply, especially in agent-driven workflows.[12]

Environmental Impact


Every additional token is additional inference compute, and inference is now where most of AI's energy goes. A June 2026 report from United Nations University found that once a model is deployed, user interactions consume an estimated 80 to 90 percent of its total energy, and that policy attention should shift from training runs toward product defaults, model selection, and user behavior. The same report projects that by 2030 data centers powering AI will consume 945 terawatt-hours of electricity, with an associated water footprint of 9.3 trillion liters and a land footprint of more than 14,500 square kilometers.[13]

The International AI Safety Report estimates that data centers and data transmission account for about one percent of global energy-related greenhouse gas emissions, with AI using 10 to 28 percent of data center energy capacity, and describes AI as a moderate but rapidly growing contributor.[14]

Frontier reasoning models are the most expensive class per query. An infrastructure-aware benchmark of 30 models by researchers at the University of Rhode Island and partner institutions found that reasoning models such as OpenAI's o3 and DeepSeek's R1 use more than 33 watt-hours for a long answer, over 70 times the energy of a small model.[15] The same study notes that water used for data center cooling is largely evaporated freshwater removed from local ecosystems rather than recycled.[16] Fable 5.1 belongs to this reasoning class.

No verified per-token figure exists for Fable 5.1 itself. Anthropic and OpenAI had not submitted models to the AI Energy Score benchmark as of May 2026, and one sustainability analyst notes that while Anthropic can tell a good per-query efficiency story, it currently publishes no facility-level environmental disclosure.[17] An independent estimate based on Claude Code billing data placed Anthropic's total inference power draw at roughly 85 megawatts, with the caveat that the figure is likely low.[18]

What the Critics Propose Instead

The shared conclusion of IBM, Exadel, and the Pragmatic Engineer is that token volume is the wrong metric in either direction. IBM warns that token minimization falls into the same trap as tokenmaxxing: once obvious waste like oversized tool catalogs and stale context is removed, further cuts start removing the task descriptions and constraints that help the model succeed, and the cost simply moves into retries, extra tool calls, and human rework.[9]

The alternative is to measure accepted output against cost. Count what survived human review and was actually used: merged pull requests, closed tickets, approved documents, hours of manual work replaced. A prototype nobody wanted counts as zero regardless of the tokens it consumed. Divide those results by the dollars or plan credits spent to get cost per accepted result, so that a heavy user who ships a lot looks good and a heavy user who ships nothing looks like what they are. Route routine work to cheaper models and reserve Fable-tier models for problems that would otherwise warrant a senior specialist, cap output length, and use prompt caching.[19] And retire any public leaderboard ranked by token count, since it will be gamed.

For an individual user the version is simpler. Before a long session, decide what finished thing you want at the end of it, and judge the session by whether you got it rather than by how much you used.

Footnotes

[1] Exadel, "What Is Tokenmaxxing and Why It's a Liability," June 30, 2026. https://exadel.com/news/tokenmaxxing-ai-productivity-enterprise-roi

[2] Hayley Peterson, "Tell us if you're on the AI leaderboard at work," Business Insider, April 2026. https://www.businessinsider.com/jpmorgan-disney-employees-vie-for-ai-leaderboard-status-tokenmaxxing-2026-4

[3] Exadel, June 30, 2026, cited above.

[4] ClaudeFast, "Claude Fable 5 Price: Is It Free, Usage Credits, and Access," August 2026. https://claudefa.st/blog/guide/development/fable-5-usage-credits

[5] Digital Applied, "What Claude Fable 5.1 Costs, and What It Breaks," September 2026. https://www.digitalapplied.com/blog/claude-fable-5-1-cost-and-breaking-changes

[6] Anthropic Help Center, "Claude Fable models on your plan," updated September 2026. https://support.claude.com/en/articles/15424964-claude-fable-models-on-your-plan

[7] explainx.ai, "Claude Fable 5.1: 55.8% Terminal-Bench, 25% Cheaper," September 2026. https://www.explainx.ai/blog/claude-fable-5-1-mythos-5-1-launch-benchmarks-pricing-2026

[8] Layer3 Labs, "Claude Fable 5.1 Limits: Quotas, Context, and Rate Caps," September 2026. https://www.layer3labs.io/guides/claude-fable-5-1-limits

[9] IBM Think, "Tokenmaxxing is dead, long live valuemaxxing," June 25, 2026. https://www.ibm.com/think/insights/tokenmaxxing-dead-long-live-valuemaxxing

[10] Gergely Orosz, "The Pulse: 'Tokenmaxxing' as a weird new trend," The Pragmatic Engineer, April 23, 2026. https://blog.pragmaticengineer.com/the-pulse-tokenmaxxing-as-a-weird-new-trend/

[11] Tom's Hardware, "AI cost crisis hits tech giants as employee tokenmaxxing backfires," May 2026. https://www.tomshardware.com/tech-industry/artificial-intelligence/ai-cost-crisis-hits-tech-giants-as-employee-tokenmaxxing-backfires-agentic-ai-eats-up-to-1000x-more-tokens-than-standard-ai-sparks-corporate-pullback-at-microsoft-meta-and-amazon

[12] Isabelle Bousquette, "Why Some Companies Say AI 'Tokenmaxxing' Is Key to Survival," The Wall Street Journal, April 14, 2026. https://www.wsj.com/cio-journal/why-some-companies-say-ai-tokenmaxxing-is-key-to-survival-e699a128

[13] United Nations University Institute for Water, Environment and Health, "Rising Emissions, Depleting Water and Vanishing Land," June 3, 2026. https://unu.edu/inweh/news/environmental-cost-of-AIs-Enrgy-use-carbon-water-and-land-footprints

[14] International AI Safety Report, Section 2.3.4, "Risks to the environment," 2025. https://arxiv.org/pdf/2501.17805

[15] Fast Company, "The environmental impact of LLMs: Here's how OpenAI, DeepSeek, and Anthropic stack up," May 20, 2025. https://www.fastcompany.com/91336991/openai-anthropic-deepseek-ai-models-environmental-impact

[16] Nidhal Jegham et al., "How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference," arXiv 2505.09598, 2025. https://arxiv.org/pdf/2505.09598

[17] AZV AI, "Is Claude Sustainable? Anthropic's Environmental Report Card," May 16, 2026. https://azvai.com/en/is-claude-sustainable/

[18] Simon P. Couch, "Electricity use of AI coding agents," January 20, 2026. https://simonpcouch.com/blog/2026-01-20-cc-impact/

[19] AY Automate, "Fable 5 Pricing: $10/$50 per Million Tokens," August 2026. https://www.ayautomate.com/blog/claude-fable-5-pricing-explained

DISCLOSURE: 9% usage of my Fable 5.1 weekly limit of Max plan to produce. One of the links was broken and it cost 2% more to regenerate. 


Thursday, March 26, 2026

To edit or not to edit? Disclosure is the answer

Salt Palace 2026
In today's news, the prompter and motion capture provider of AI creation, Tilly Norwood, is receiving death threats. An interesting point the prompter makes is that "Tilly is famous. I'm not. That's wonderful. Fame is a horrible thing." Death threats definitely add credence to the "fame is a horrible thing" posit. Norwood's prompter also makes the point that since the AI is the famous one, she doesn't have to do things like Botox to preserve youthfulness. Altering one's appearance was a recent discussion I had with Ilya at The Family Vault (an AI Agent and cloud storage) booth at 2026 RootsTech. Ilya questioned why disclose using AI for photo regeneration when people edit photos with Photoshop or even that many women wear make-up and that's editing their looks? I've been mulling this over ever since RootsTech! In the Genealogy Community, it is actually an asset to be older. When I became active in genealogy at age 31, I actually had genealogists be dismissive towards me due to my youth. Some genealogists wear make-up and some don't; it's not a field where appearance is judged like it is with celebrities. I have noticed that make-up tends to help when on Zoom meetings or on stage because the bright lights tend to wash out facial features. Unlike with AI use, there's no need to disclose I'm wearing make-up because it's quite obvious and visible. Of course I didn't think of this at the time I was chatting with Ilya - the case for disclosure lies with whether editing can be detected or not. Make-up is visible. Tinting old photos is visible. AI photo regeneration is not visible (at least to the naked eye). Read more here on the Coalition for Responsible AI in Genealogy's "Protecting Trust in Historical Images" position statement.

Monday, March 23, 2026

"The Thinking Game" is a must watch for AI and Genetic Genealogy enthusiasts

https://youtu.be/d95J8yzvjbQ?si=mrHOHVrLm41B9Bre
 

The COVID-19 pandemic brought into sharp focus for me how precious time really is. You cannot buy more time no matter how wealthy you are; so at some point, you need to realize this fact and spend the time you have as wisely as you can. One area where I feel strongly about this is with movies and TV. I'm always telling my friends, "You won't be on your death bed saying, "I wish I had watched more news." I especially detest watching movies or shows that I feel were a waste of my time and grey matter. Due to my mind-shift, I struggle to find something worth watching during air travel. On a recent United flight, I was lucky to find "The Thinking Game" in their library. Definitely worth every second of my time to watch! The film centers on the journey of Demis Hassabis, from childhood through to winning the Nobel Prize for creating the AI AlphaFold agent which solved the "protein-folding problem". Not only does the documentary include much of the evolution of artificial intelligence and the issues surrounding it, it also shows how Hassabis had the goal from a young age that AI could be used to solve human issues and ailments. "The Thinking Game" is available to watch on Prime and for free on YouTube.

Sunday, February 22, 2026

AI and Genealogy (not quick, nor dirty) Glossary

There appears to be multiple options for when Artificial Intelligence was created; possibly by Alan Turing or at a conference in 1956. Either way, it was developed mid-20th century. Although companies (including genealogy companies) had been using AI for many years, the official launch date for use by the general public is November 30, 2022 when OpenAI debuted ChatGPT. Since that time, a whole new AI lexicon has been trickling forth; and in some cases, the trickle is more like a flood. I will be posting the new terms as I learn them, but for now, I had ChatGPT 5.2 Thinking version create this handy graphic. It did take me five tries to get the graphic where it's at. The first two attempts had wasted space, and when I prompted it to fill the space, it repeated the terms to fill it. (Facepalm) The last attempt also had wasted space so I asked it to move the citation/disclaimer up under the boxes. Interestingly, the first three attempts had colorful graphics and used several other AI terms. I figured I was dealing with "context rot" so started over and it gave me the plain vanilla result you are seeing. I'd rather have accurate than flashy so it is what it is. I did have ChatGPT provide sources for the terms. Accuracy, accuracy, accuracy - ALWAYS proof AI's work or you might end up in an unfortunate news headline.


 

The Hidden Costs of Tokenmaxxing

As of 2 Sep 2026 NOTE FROM THE HUMAN: The following blog post was generated and sourced solely by Claude Fable 5.1 and the model even chose ...