08/31/2026 | Press release | Distributed by Public on 08/31/2026 18:09
Enterprise AI has entered a new phase. The first was about adoption: companies bought chabots, launched copilots, distributed licenses and encouraged employees to experiment. The next is about economics.
In a recent McKinsey survey, 93% of organizations reported exceeding AI budgets, while spending increased nearly fourfold as companies scaled deployment. Walmart has introduced fixed token allowances for an internal AI agent. Atlassian now gives some employees monthly AI wallets. Companies are asking a different question: how much intelligence can they afford?
The instinctive response is to impose budgets, alerts and model restrictions. But companies cannot throttle their way to AI transformation. If the economics of enterprise AI depend on people using it less, the architecture is wrong.
Coinbase CEO Brian Armstrong recently described a different approach: use better defaults, model routing and caching to keep spending relatively flat even as token usage grows exponentially. Enterprises do not need less intelligence. They need to architect it more intelligently.
Most companies focus on the price of a model or token. But the cheapest model is not always the most economical. A low-cost model that produces errors or requires repeated prompts and human review can cost more than a premium model that completes the task correctly. Conversely, using the most powerful model for routine classification is like using a supercomputer as a calculator.
The more useful metric is cost per successful outcome: resolving a customer case, reviewing a contract or identifying a patient who needs follow-up care.
This shifts the question from "How much AI are we using?" to "How much valuable work is AI completing?" It is also the difference between deploying a chatbot and building an AI-native enterprise. Creating lasting value requires a new architecture built across four connected layers, each with a distinct role.
Large companies hold enormous amounts of information across systems such as SAP, Salesforce, Workday and Microsoft 365. AI needs a shared foundation (e.g., a data lake or lakehouse) that connects this information and reveals relationships across customers, transactions, operations and decisions.
It can also draw on authorized external or real-time sources and create reusable insights by identifying new patterns and correlations. But broad access does not mean using everything for every task. Irrelevant information increases cost and can reduce reliability, so AI should retrieve only the most relevant and trustworthy context.
Morgan Stanley's internal assistant, for example, can search approximately 100,000 documents, but the firm continually refines how it surfaces the right knowledge for each advisor question.
Data captures what happened. It rarely explains how work gets done: how a case moves between teams, where approvals occur or how exceptions are handled.
This operational context is the missing layer in many AI strategies.
Skan AI, an example from the Cathay Inno portfolio, is building what it calls the Context Graph of Work: a continuously evolving, machine-readable record of how work happens across people, systems, decisions and exceptions.
Rather than relying on documents or static process maps, Skan observes workflows across applications. It captures the variations and judgment calls that determine how work is completed in practice.
Enterprises can use this understanding to identify AI opportunities, translate workflows into agent-ready procedures and deploy agents with human oversight and auditability. Instead of starting from zero, AI acts on a continuously updated model of the business.
There is no reason for an enterprise to standardize every workload on one model. Different tasks require different combinations of intelligence, speed, privacy, reliability and cost.
Complex coding, strategy or high-stakes clinical reasoning may justify a frontier model. Classification, summarization and extraction can often run on smaller, specialized or open-weight models. Kimi K3 from CI-backed Moonshot AI shows how quickly the landscape is shifting: open-weight models can now compete with leading proprietary systems, giving enterprises greater flexibility to balance capability and cost.
Harvey recently demonstrated this flexibility with Tenet, its Kimi K3-based model developed with Fireworks. Tenet completed nearly twice as many held-out legal tasks as the base model at effectively flat cost per task.
The right architecture begins with an economical default and escalates only when complexity or risk requires it. Caching requests, reusing context and evaluating new models can further improve performance per dollar.
Reasoning alone creates limited value. Economic value appears when AI becomes part of the workflow and helps complete real work.
Looking again to our Inno III companies, Qualified Health offers a prime example in healthcare. Its platform uses Claude to analyze clinical data, identify gaps in recommended care and surface relevant patients within clinicians' workflows. At the University of Texas Medical Branch, it is helping identify patients with heart failure who might otherwise go undetected. Clinicians determine the intervention; AI performs analysis humans could never conduct continuously at the same scale.
The model applies across industries. Compliance teams can investigate prioritized exceptions rather than every transaction. Lawyers can focus on clauses flagged for risk. Customer support agents can begin with a complete case history. Investment teams can evaluate opportunities synthesized from thousands of documents rather than review every source.
The value does not come from generating an answer. It comes from closing the gap between knowing and doing.
An economically scalable architecture does not remove people indiscriminately or require review of every routine step. AI can handle predictable work while escalating ambiguity and high-stakes decisions.
The balance differs by workflow. A clinician remains accountable for patient care. A compliance officer determines whether an alert requires action. In lower-risk processes, people may intervene only when the system identifies uncertainty.
AI concentrates human judgment where it creates the most value.
The first generation of enterprise AI placed a chatbot in front of every employee. The next will place intelligence inside every important workflow.
Enterprise Data
↓
Operational Context
↓
Right Model for the Task
↓
AI Execution
↓
Human Judgment
Implementing this architecture is not like installing conventional software. Software is typically connected to the technology stack, configured and run. AI reaches into the business stack: how tasks are performed, decisions are made and people work with machines.
Nothing will be optimal on day one. Companies must observe workflows, automate predictable steps, preserve human judgment and continually refine their processes. Over time, AI may also propose-and eventually implement-new ways of working, although the pace and degree of that autonomy remain uncertain.
The durable advantage will belong to companies that enable people and machines to improve outcomes together. That is the difference between deploying AI tools and building an AI-native enterprise.