Anthropic AI bills: Glean finds an 81% cost gap
Glean says enterprise teams can spend materially less on AI tasks by routing work to different models and supplying better context. Its comparison with Claude Cowork found 81% lower cost per task, while The Information reports the company says Anthropic customers' bills are about 80% above a reasonable level.
Enterprise AI bills are becoming a routing and context problem, not only a model-price problem. The Information reports that Glean says Anthropic customers' bills are about 80% higher than a reasonable level. In Glean's own published comparison, its assistant cost 81% less per task than Claude Cowork across more than 180 enterprise tasks, which gives operators a concrete reason to inspect token usage before choosing a single frontier model.
Definition: AI cost routing means selecting a model and reasoning level for each task instead of sending every request to the same expensive model.
Example: Glean says its assistant averaged $0.58 per task, compared with $2.98 for Claude Cowork in its 180-plus-task comparison.
Key takeaway: Glean's numbers suggest that enterprise context, tool orchestration, and model selection can change the economics of an AI task as much as the headline model price.
Business impact: Operators budgeting for AI agents should measure cost per completed task and correction work, not only the provider's token tariff.
What is the reported Anthropic bill problem?
The Information's report puts the headline figure at roughly 80%: Glean says Anthropic customers are paying materially more than they need to. The claim describes a gap between observed enterprise spending and what Glean considers a reasonable level, but it does not establish that every Anthropic invoice contains an error or that every customer would see the same reduction. The practical takeaway is to audit actual task costs before treating a model choice as a fixed commitment.
Glean's public benchmark supplies a more specific comparison, but it is not the same statistic as the report's 80% figure. Glean says its assistant averaged $0.58 per task against $2.98 for Claude Cowork, an 81% lower cost in that test. The difference comes from the benchmark's selected tasks, model configuration, token volumes, and prices, so the figures should be read as attributed performance claims rather than as a market-wide Anthropic price survey.
How large was Glean's Claude Cowork comparison?
Glean says it compared Glean Assistant with automatic routing enabled against Claude Cowork running Claude Sonnet 5 at high reasoning across more than 180 enterprise tasks. Glean reports 1.3 million tokens for its system versus 4.4 million for Cowork, a 70% reduction, and says its responses were preferred 78% of the time. The comparison used synthetic queries on Glean's own production data, so teams should treat the result as a vendor benchmark to reproduce with their own workload.
| Measure | Glean Assistant | Claude Cowork |
|---|---|---|
| Average cost per task | $0.58 | $2.98 |
| Tokens used in the test | 1.3 million | 4.4 million |
| Preferred response rate | 78% preferred | Comparison baseline |
The comparison's cost result is therefore not explained by a cheaper token price alone. Glean says its assistant used fewer tokens and a blended mix of models, while Cowork was held to Claude Sonnet 5 at high reasoning. Operators comparing AI systems should reproduce both sides under equivalent task definitions and count review time, failed attempts, and retries alongside provider charges.
Why does context affect AI cost?
Glean attributes its lower token volume to an enterprise context layer that starts from a pre-indexed, ranked view of company information, rather than searching each connected system independently for every task. Glean says that approach reduces over-fetching, repeated normalization, and unnecessary reasoning loops. For a business evaluating AI agents, the implication is specific: measure how much context an agent retrieves and re-reads before measuring the model's nominal price.
Glean also says its harness keeps tool outputs and intermediate state in sandbox files, progressively loads only the tools and schemas a task needs, and isolates sub-agents from unrelated context. Those choices fit the wider AI automation stack distinction between the model, orchestration, integrations, and data layer. A lower bill can come from reducing unnecessary work around the model, not only from switching to a cheaper model.
What should enterprise operators measure next?
Enterprise operators should use Glean's benchmark as a prompt to build a task-level cost ledger, not as a reason to assume an 81% saving. For each representative workflow, record model calls, input and output tokens, retries, latency, human review, correction time, and successful completion. A structured AI ROI estimate is more decision-useful when it includes those operational costs rather than multiplying a token price by request volume.
The central signal in this story is that enterprise AI economics are moving up the stack. Glean's reported 70% token reduction and 81% lower per-task cost are vendor-published results from a defined comparison, while The Information's 80% figure is a broader claim attributed to Glean. Companies should verify both the cost and the quality on their own data before changing a production system, because the cheapest answer is not useful if it creates more correction work.
Frequently asked questions
What does Glean say about Anthropic customers' AI bills?
Glean says Anthropic customers' bills are about 80% higher than a reasonable level, according to The Information's report. Glean's own published comparison uses a related but different measure: across more than 180 enterprise tasks, Glean says its assistant averaged $0.58 per task versus $2.98 for Claude Cowork, an 81% lower cost. Both figures are vendor claims, not an independent audit of every Anthropic customer's invoice.
How did Glean's comparison with Claude Cowork measure cost?
Glean says it compared its assistant, with automatic model routing enabled, against Claude Cowork using Claude Sonnet 5 at high reasoning across more than 180 enterprise tasks. Glean reported 1.3 million tokens for its system versus 4.4 million for Cowork, then combined token volume with blended token prices to calculate $0.58 versus $2.98 per task. The evaluation used synthetic queries against Glean's own production data, so the result is a controlled vendor benchmark rather than a universal estimate for every company.
Why does Glean say its assistant uses fewer tokens?
Glean attributes the difference to two design choices: a pre-indexed enterprise context layer and a harness that stores intermediate state in sandbox files instead of repeatedly reloading it into the model context. Glean also progressively loads only the tools and schemas a task needs and routes tasks to different models and reasoning levels. These are Glean's explanations for its benchmark result, not proof that the same savings will appear in every deployment.
Should companies replace Claude Cowork based on this benchmark?
No single benchmark is enough to decide an enterprise AI purchase. Glean's result is useful as a cost-control hypothesis because it reports both token volume and output preference, but Glean designed and published the comparison. Companies should test representative private tasks, measure total cost including human correction, and check quality, permissions, data access, and reliability before changing a production workflow.
Alex
Founder & Lead AI Writer
Alex is the founder of Yowox and lead AI writer since 2024, breaking down complex information into clear, actionable insights for thousands of readers every day. Alex has built AI automation systems for businesses since 2024, focusing on AI agents, workflow automation, and business process optimization.
Save hours. Save thousands.
Practical guides, real workflows, and the latest AI and automation news that matters — straight to your inbox.