Full Breakdown
Amazon’s “Tokenmaxxing” Sparks Debate Over AI Usage Metrics
5/13/2026, 11:39:45 AM
Tokenmaxxing at Amazon: Inflating AI Usage Scores
In early 2024 Amazon’s internal AI platform, MeshClaw, became the focus of a new internal competition. Engineers created AI agents to perform unnecessary or trivial tasks solely to increase the number of tokens their prompts consumed. The practice, dubbed *tokenmaxxing*, was driven by company-wide targets that required over 80 % of developers to use AI each week and by public leaderboards that displayed individual token consumption. Employees reported that managers monitored these leaderboards, turning token volume into a proxy for performance.
Background & Context
MeshClaw, inspired by the open-source OpenClaw agent creator, lets users automate code deployments, triage emails, and interact with apps such as Slack. Amazon rolled out the tool in the spring, promoting it as a way to “automate repetitive tasks each day.” Simultaneously, the firm introduced internal dashboards that aggregated token usage across teams, a move intended to showcase AI adoption but which inadvertently created perverse incentives.
Key Figures & Groups
- Amazon – developer of MeshClaw and sponsor of the AI usage targets.
- MeshClaw – internal AI agent platform used for task automation.
- Amazon engineers – participants in tokenmaxxing.
- Managers – alleged reviewers of token-usage leaderboards.
- Financial Times – reporter of the phenomenon.
- PYMNTS – analyst of token-based metrics and proponent of alternative measures.
- Salesforce – competitor offering the Agentic Work Unit (AWU) metric as a contrast.
Data & Statistics
- Target: >80 % of Amazon developers to log AI usage weekly.
- Leaderboards tracked token consumption per employee.
- Salesforce’s AI platform generated 2.4 billion AWUs to date, with 771 million in Q4 2025 alone; service-agent usage grew 106 % QoQ, and AI search in Slack rose 116 %.
- Salesforce’s early $2-per-conversation pricing yielded 5,000 deals, of which 3,000 paid, prompting multiple pricing revisions.
Official Statements & Responses
Amazon told the Financial Times that MeshClaw “enabled thousands of Amazonians to automate repetitive tasks each day” and affirmed its commitment to the safe, secure and responsible development and deployment of generative AI for our customers. In response to the token-centric metric, Salesforce introduced the Agentic Work Unit, a unit that counts completed AI tasks—prompts processed, reasoning chains finished, or tools invoked—rather than raw token volume.
Criticism & Opposition
Employees expressed concerns that tokenmaxxing creates security risks and dilutes genuine productivity. One engineer warned that “when they track usage it creates perverse incentives and some people are very competitive about it.” Observers noted that focusing on token counts may cause staff to neglect important work and pressure teammates, potentially lowering overall collaboration efficiency.
Conflicting Reports & Gaps
Amazon maintains that leaderboard data will not be used in performance evaluations, yet multiple staff members reported that managers do review these statistics. The company has since restricted access to usage metrics so that only individual employees and their managers can view them, but the extent of tokenmaxxing across the organization remains unquantified.
Why It Matters
Finance leaders now confront unpredictable AI costs because token consumption does not correlate with business outcomes. The PYMNTS analysis emphasizes that “the token was never designed to measure business value,” prompting a shift toward metrics like AWUs that aim for a high inference-to-work ratio—output tokens that produce actual results. Gartner forecasts that agentic AI will account for 30 % of enterprise-software revenue by 2035, surpassing $450 billion, underscoring the financial stakes of accurate measurement.
Verbatim Quotes
- “Managers are looking at it,” an employee told the FT. “When they track usage it creates perverse incentives and some people are very competitive about it.” — Amazon employee, cited by Financial Times
- “In a statement to the paper, Amazon said that the tool enabled “thousands of Amazonians to automate repetitive tasks each day,” adding that it is “committed to the safe, secure and responsible development and deployment of generative AI for our customers”.” — Amazon spokesperson, Financial Times interview
- “Metric That Only Measures Noise The token was never designed to measure business value.” — PYMNTS analysis
- “The goal is a high inference-to-work ratio: output tokens that produce actual results.” — PYMNTS conclusion
- “Agentic AI will account for 30% of enterprise application software revenue by 2035, surpassing $450 billion, up from roughly 2% in 2025, Gartner forecast.” — PYMNTS conclusion
What’s Next
Salesforce is expanding the AWU framework as enterprises demand predictable AI spend and demonstrable outcomes. Amazon has limited internal visibility of token metrics and may consider alternative usage indicators to curb gaming. Industry observers will watch whether tokenmaxxing fades as firms adopt work-completion-based metrics, aligning AI adoption with tangible business value.
