Article written and provided by, Glean. 

Executive summary

Enterprise AI is becoming more capable and more computationally demanding. The simple prompt-response model no longer describes how many enterprise tasks are completed. Modern AI systems may retrieve information, call tools, execute code, reason through intermediate steps, create artifacts and coordinate multiple agents before returning a result.

That changes the economics of AI. Token consumption is shaped by architecture: the quality of context retrieval, the number of tool calls, the amount of state carried across a workflow, the model selected for each step and the system's ability to learn from prior execution.

The central idea of the Token Economy is token yield: the amount of useful, work-ready output produced for every token consumed. Lower token usage alone is not enough. An answer that uses fewer tokens but requires more correction, rework or follow-up is not truly more efficient.

For enterprise leaders, this creates an actionable path forward:

  • Build a strong context layer that retrieves precise, permission-aware information.
  • Match the right model and reasoning level to the task instead of defaulting every step to a frontier model.
  • Capture execution patterns so the system becomes more targeted over time.
  • Use an agent harness that orchestrates context, tools, models and state without allowing complexity to compound unnecessarily.

The economics of enterprise AI are changing

Early generative AI interactions were relatively easy to understand: a user entered a prompt and a model returned a response. Enterprise AI systems now do much more. A request such as "Analyze churn risk for these accounts and create follow-up tasks" can involve account retrieval, business analysis, data interpretation, sub-agent work and structured artifact creation.

The user's original prompt may be only a small portion of the total execution. The rest can include system instructions, tool schemas, retrieved documents, tool calls, code execution, intermediate outputs and reasoning traces. An eleven-word request can expand into thousands of tokens during execution.

As AI moves into heavier agentic workflows, organizations need a way to evaluate both cost and utility. The relevant unit is not tokens in isolation. It is useful work per token.

Token efficiency is an architecture problem

A model's price matters, but it is only one part of the cost surface. The architecture determines how much context the model receives, how many times it must call a tool, how much reasoning it must perform and whether it repeats work that the system has already learned how to do.

Glean identifies four architectural capabilities that help improve token yield.

1. Context: retrieve the precise data

Better context replaces broad, token-expensive exploration with narrow, informed execution.

In a federated approach, an AI system reaches out to multiple sources at query time. It may issue separate calls to business applications, retrieve raw results, normalize and deduplicate them, reconcile conflicts and then reason across the combined output. Every step adds latency and consumes tokens, while the model carries the burden of stitching together the enterprise context.

An indexed approach works differently. Centralized, precomputed indexes and cross-application signals allow the system to begin with a stronger understanding of the organization. Enterprise search can account for relevance, permissions, recency, ownership and relationships across system, not just similarity to a query.

That distinction has measurable implications. In a benchmark of roughly 175 work-oriented queries, Glean was preferred about 2.5 times as often as off-the-shelf MCP configurations, while the off-the-shelf approach used about 30% more tokens. Across the benchmark, the comparison included content generation, meeting preparation, inbox and communication workflows, large-scale analysis and information seeking. 

For enterprise teams, the lesson is straightforward: better retrieval can improve answer quality and reduce the amount of context the model must process. Indexing is not simply a search convenience; it is an economic control point for AI execution.

2. Intelligent routing: match intelligence to the job

Frontier models are valuable when a task requires advanced reasoning, but not every step of an enterprise workflow needs frontier intelligence. Using the most expensive model for every retrieval, classification, transformation or routine action can create unnecessary cost and latency.

Intelligent routing matches the model—and the reasoning level—to the task. It reserves more expensive reasoning for the steps that benefit from it and uses more efficient models where they can deliver the required result.

This also requires model choice to remain flexible. A model-agnostic platform can evaluate multiple open-source and proprietary models, optimize for price-performance and adapt as new models become available. Evaluation should consider more than benchmark intelligence; it should measure correctness, completeness, utility and efficiency for the types of work the organization actually needs to perform.

The goal is not to use the cheapest model. The goal is to use the right model for the job while preserving the quality of the outcome.

3. Continual learning: get smarter with every execution

Organizations already generate a large amount of latent process knowledge. It is embedded in tickets, approvals, handoffs, edits, escalations and recurring workflows. These signals show how work is actually completed, but many AI systems do not use them.

A platform designed for continual learning can turn those patterns into usable context. When the system encounters a familiar task, it can begin with a stronger prior instead of exploring every possible path from scratch.

The execution trace itself is also valuable. Tool calls, queries, retrieved documents, branches explored and the path ultimately taken can reveal which routes are direct, which steps are redundant and where a workflow can be improved. Over time, each completed task can help make the next similar task faster, more targeted and less expensive.

This is where token efficiency becomes a compounding advantage. A system that learns from its work can improve both capability and economics over time.

4. The agent harness: orchestrate context at scale

As workflows become more agentic, the active context can grow quickly. Naive architectures may leave too many tools and instructions in the system prompt or carry accumulated state across every step. The result is a cost surface that expands with task length and complexity.

A well-designed harness keeps context scoped and compact. It progressively discloses the skills and tools relevant to the current task, moves intermediate outputs into external memory when appropriate and reintroduces information only when it is needed. It also coordinates an orchestrator agent, sub-agents, a skills index, a model hub and the execution environment.

The harness should scale with the work itself: adding complexity when the task demands it and removing unnecessary complexity when it does not. It should also remain model-agnostic so that every step does not inherit the economics and limitations of one model.

The result is a more disciplined execution pattern—one that can support simple questions and complex, multi-step workflows without allowing context, reasoning and tool usage to compound unchecked.

From token cost to token yield

The right metric for enterprise AI is not simply the number of tokens consumed. It is the useful, work-ready output produced for those tokens.

That means leaders should evaluate AI systems across several dimensions:

  • Does the system retrieve the right information the first time?
  • Does it respect permissions and current enterprise context?
  • Does it select the right model and reasoning level for each step?
  • Does the output require substantial human correction or rework?
  • Does the system improve as it observes more work?
  • Can the architecture scale across departments, workflows and data sources without creating a parallel support burden?

The benchmark results in Glean's Token Economy materials illustrate why quality and efficiency must be considered together. Glean's indexed context and intelligent routing are positioned to produce stronger results with fewer tokens—not by reducing the usefulness of the response, but by reducing unnecessary exploration and matching execution to the task. Glean reports roughly 23% fewer tokens with enterprise context and approximately 25% fewer tokens with intelligent routing while maintaining high-quality outcomes.

Building a more sustainable AI future

The most important shift is conceptual. Token efficiency is not about making AI cheaper at the expense of quality. It is about designing systems that do more useful work with the same unit of AI spend.

The more an AI system understands the organization—its data, people, processes, goals, permissions and patterns of work—the less it needs to explore blindly. Better context reduces unnecessary retrieval. Intelligent routing avoids using expensive reasoning where it is not required. Continual learning makes future execution more targeted. A disciplined harness keeps context and orchestration manageable as workflows expand.

Together, these capabilities create a foundation for enterprise AI that is more measurable, more governed and more economically sustainable. Organizations that make these architectural decisions early can improve not only cost control, but also response quality, adoption and the ability to scale AI into the workflows that matter most.

The future of enterprise AI requires strategy, architecture and implementation to work together. WWT helps organizations evaluate AI approaches, validate architectures, connect enterprise systems and plan adoption. Glean provides an enterprise context layer and AI capabilities that help organizations ground responses in company knowledge, orchestrate work and improve the yield of AI execution.

Learn more about Workforce AI and Glean Contact a WWT Expert 

Technologies