Full Breakdown
Companies Deploy “Caveman” Plugin to Cut AI Token Costs
6/30/2026, 8:47:21 PM
Rising Token Costs Prompt a Terse Plugin
Large language models such as Claude Code, Codex, and Gemini are routine in corporate workflows, but their verbose output drives high token consumption. Accenture flagged “soaring token spend” as a major cost, especially for PDF-to-presentation conversion. In early April, developer Julius Brussee released the open-source “caveman” plugin to strip pleasantries and hedging while preserving code, URLs, and numbers. OpenAI’s director of engineering Shayne Sweeney added Codex support, and engineers at OpenAI, Nvidia, GitHub, and Legrand have begun testing it.
Token Savings Demonstrated
Benchmarks show caveman cuts output tokens by roughly 65 %–75 % versus default verbose replies. In a private test the plugin saved about 5,800 tokens, a 65 % reduction, while still outperforming a generic “be concise” prompt. The GitHub repository advertises ~2× fewer tokens than Codex on identical coding tasks.
Enterprise Reactions
Legrand circulated an internal memo urging staff to “be mindful of our usage of AI so we don’t use up our entire budget allowance too quickly,” and listed using the caveman skill as a high-impact measure. GitHub announced a shift to per-token billing. Uber’s CTO said the firm exhausted its AI budget in four months, prompting a cap, and Walmart imposed similar limits. Accenture’s leaked audio framed token economics as a new consulting opportunity, despite earlier promotion of rapid AI adoption. OpenAI CEO Sam Altman warned that unnecessary pleasantries cost the company tens of millions in electricity.
Tensions and Concerns
Accenture’s pivot from championing rapid AI deployment to selling token-economics consulting creates a perceived conflict of interest, highlighting tension between AI adoption incentives and cost-containment pressures.
Developer Feedback
Brussee says he has heard from many developers and engineers at OpenAI, NVIDIA, GitHub, and DEPT testing caveman, and notes Shayne Sweeney’s contribution of Codex support.
Looking Ahead
With token-based pricing now common, more firms are likely to adopt cost-saving plugins like caveman to preserve AI budgets, making token efficiency a strategic priority.
Verbatim Quotes
- “I made Caveman back in early April because I was using Claude Code heavily and noticed a lot of my token spend was going to unnecessary prose: pleasantries, hedging, transitions, and chatty language that does not really matter inside an agent loop,” — Julius Brussee, creator of Caveman
- “It makes the model speak less like a polite chatbot and more like a terse tool,” — Julius Brussee
- “caveman-code shrink everything — full terminal coding agent, caveman top to bottom. ~2× fewer tokens than Codex on identical tasks. 20+ providers · plan mode · autopilot goal loop · MIT,” — Caveman GitHub repository
- “I’ve heard from many individual developers and engineers inside companies using or testing it, including people at OpenAI, NVIDIA, GitHub, and DEPT,” — Julius Brussee
