kachar.dev
The index
No. 23ai / cto / building

The State of AI 2026 bets your next customer is an agent with a wallet

The 244-slide State of AI Report 2026, cut to what a startup CTO can act on: who earns the money, what an answer costs, and nine bets for next year.

By , CTO & Co-founder, Juma Labs

Published
Words
1,468
Reading
7 min
Sections
08

A thick stack of dark brushed-metal plates floating in grey haze. One thin plate is pulled halfway out of the middle of the stack and glows electric violet, the only color in the frame.

The top 1% of companies spend $7,205 per employee per month on AI. The median company spends $12.50.

That 580x gap is one of 244 slides in the State of AI Report 2026, published today by Nathan Benaich and Air Street Capital.

This is the short version, for startup CTOs. Industry and predictions first. From politics and safety, only what changes what you build.

Two labs get most of the money

  • $105B a year. That's OpenAI and Anthropic combined, at current run rates. It was $30B in January.
  • 3x or more, every year. Their run rates grew 3.6x in 2024 and 4.7x in 2025.
  • 96% of business spend. Among companies paying for AI through Ramp, these two labs take almost everything.
  • Three labs at the top. On Artificial Analysis's index, scores 58. and tie at 53.
  • Cheap models get the volume. On 's on October 4, handled 62.7% of but only 26.9% of spend.
  • GPUs got more expensive. Rental prices are up 30% on average from their lows. Plan for scarcity, not a price war.

AI-first companies grow faster

The report splits private companies into two groups. AI native means AI is the product. AI enabled means existing software that added AI.

  • AI native grows faster. Among the top quarter of companies, AI natives grew revenue 256% in a year. AI-enabled companies grew 90%.
  • Bolting AI on didn't help. AI-enabled companies grew slower than companies with no AI at all, in every group in the chart below.
AI-first companies grow fastest

Yearly revenue growth of the top quarter of companies in each group.

AI nativeAI-enabled SaaSNon-AI (reference)
Before 2017
49%
31%
45%
2017 to 2019
119%
72%
84%
2020 and later
487%
199%
272%

Hover or tab to a group to compare.

Source: State of AI Report 2026, slide 94, from Standard Metrics data.
  • Startups spend more on AI. In 2023, VC-backed companies spent about the same on AI per employee as everyone else. Since then their spend grew 24x. Everyone else's grew 3.9x.
  • Heavy users grow faster. Among 107 public tech companies, the heaviest users grew revenue about 3x faster than the lightest. That's a correlation, not proof.

Pay per answer, not per token

  • Tokens keep getting cheaper. The price of a given level of performance falls about 13x a year.
  • Answers don't. spend tokens you never see. The cost that matters is per finished task.

Bar chart from Artificial Analysis: weighted average cost per Intelligence Index task by model, split by token type. Costs run from $0.06 for the cheapest models to $7.62-7.63 for the most expensive Claude models, with GPT-6 Astra at $3.26 and Grok 4.7 at $3.74.

  • Over 100x range per task. The same benchmark task costs 6 cents on the cheapest model and $7.63 on the most expensive.
  • 10x the price for 3 points. On a 20-hour coding benchmark, GPT-6 Astra scores 65.5% at $1,030 per attempt. Claude Opus 5.5 scores 62.3% at $99.
  • Some calls aren't generation. Routing, tagging and picking a tool are choices from a fixed list. , a model built only for that, took 27% of 's classification requests in ten days. I've benchmarked it.
  • What to do: track cost per completed task, not cost per million tokens.

The harness matters more than the model

A harness is the code around a model: prompts, tools, memory and the loop that makes it an agent.

  • 7.8x bigger effect. Researchers ran three models in three harnesses on 100 coding tasks. Changing the harness moved results 7.8x more than changing the model.
  • Same model, +13 points. went from 52.5% to 65.5% with only the harness changed.
  • Less prompt, same result. cut 80% of its system prompt for newer models and lost nothing.
  • Where to invest. Tools and servers, state, observability and sandboxes. Workarounds for old model weaknesses go stale.
  • Benchmarks run out fast. The hardest math benchmark went from 22% to 100% in fourteen months.
  • Scores hide unfinished work. One model scores 77.7% on office tasks but finishes only 44.3% of them.
  • What to do: test on your own tasks, and check they actually got done. More in your evals are the moat.

Your data is the moat

The report draws a simple loop for AI products:

  1. Start on frontier APIs to find product-market fit.
  2. Log every task, tool call and outcome.
  3. Turn failures into repeatable tests.
  4. Improve context, tools or the model, and test before shipping.
  • Train your own model last. It pays off once you have data nobody else has.
  • Open models make it cheap. Harvey trained the open for legal work. It runs 54.8% cheaper than and scores higher than on its benchmark.
  • No lab budget needed. One training method beat another with a tenth of the compute on the same model.
  • Logs are the asset. The valuable data is how work gets done: steps, tool calls and expert corrections. Data brokers say they pay companies $100k to $1M+ for it.
  • The loop runs itself. Decagon's agent tests fixes to customer service bots on past chats, and a person approves. On Decagon's own benchmark, it beat Decagon's staff, 93% to 83%.

Sell work, not seats

  • Pricing moves up the ladder. Chat sells subscriptions. Agents that finish work can charge per completed task.
  • Markets already flinched. When Anthropic shipped , an app for knowledge workers, about $285B of software stock value disappeared in early February. Most came back by September.
  • New buyers. Since February, enterprise users grew 108x in legal and 41x in sales, against 5x in engineering, from near zero.
  • Adoption is early. 97.9% of OpenAI staff use Codex. Among business users it's 17.3%, and among individuals 0.7%.
  • The labs sell services now. OpenAI and Anthropic both launched consulting arms to deploy their models in enterprises.
  • Watch your juniors. In an Anthropic study, developers who learned a new library with AI scored 50% on a quiz afterwards. Those without AI scored 67%.

Politics and safety that matter

  • Model access can be switched off. US export controls blocked Fable and Mythos on June 12. Fable came back on July 1. Keep a fallback model wired up and tested.
  • Cloud regions carry war risk. On March 1, Iranian drones struck two facilities in the UAE.
  • The EU AI Act slipped. Rules for high-risk uses like hiring and lending now start December 2, 2027. Disclosure rules for chatbots and AI-generated content already apply.
  • Agent security depends on the harness too. The same model, , scored 62.3 out of 100 on a containment test in Codex CLI and 39.4 in Claude Code. Test the pair you actually run.
  • Your team's laptops already run agents. , an open-source agent that reads messages and runs shell commands, has 388,000 stars. One security firm found employees running it at 22% of its customers.
  • More severe bugs. High and critical vulnerabilities from 21 major vendors in the first half of 2026 already outnumber all of 2025. Plan your patching for the higher volume.

Last year's score and next year's bets

The report grades its own predictions. Last year it got two of ten fully right.

Last year's ten predictions, graded by the report itself

Two clean hits, five partial, three misses. Filter by verdict.

  • No

    A major retailer reports >5% of online sales from agentic checkout as AI agent advertising spend hits $5B.

    Neither >5% of a major retailer's online sales nor $5B of AI-agent ad spending is established.

  • Yes

    A major AI lab leans back into open-sourcing frontier models to win over the current US administration.

    Poolside, NVIDIA (Nemotron 3 Ultra) and Reflection shipped or committed to open weights.

  • Partly

    Open-ended agents make a meaningful scientific discovery end-to-end.

    Several labs made discoveries, but how significant they are isn't broadly agreed.

  • Partly

    A deepfake/agent-driven cyber attack triggers the first NATO/UN emergency debate on AI security.

    AI risks reached the UN Security Council, but no attack-triggered emergency debate.

  • No

    A real-time generative video game becomes the year's most-watched title on Twitch.

    Twitch's top five were League of Legends, Counter-Strike, GTA V, Valorant and World of Warcraft.

  • No

    "AI neutrality" emerges as a foreign policy doctrine.

    Neutral AI hubs attracted discussion, but no new doctrine was adopted.

  • Partly

    A movie or short film produced with significant use of AI wins major audience praise and sparks backlash.

    An AI short won a public vote, then AMC declined to screen it after backlash.

  • Yes

    A Chinese lab overtakes the US lab dominated frontier on a major leaderboard.

    Kimi K3 reached #1 on Arena's WebDev board in July. One board, not overall supremacy.

  • Partly

    Datacenter NIMBYism takes the US by storm and sways certain midterm/gubernatorial elections in 2026.

    Opposition became a campaign issue; the effect on November's results is unresolved.

  • Partly

    Trump issues an executive order to ban state AI legislation that is found unconstitutional by SCOTUS.

    The December order targeted state AI laws; no Supreme Court invalidation.

Source: State of AI Report 2026, slide 237.

This year's nine predictions treat agents as buyers and actors. Three are about agents and money:

  1. Visa or Mastercard introduces a dispute rule that assigns liability for purchases made by AI agents.
  2. A US regulator or exchange attributes an abnormal stock move to correlated orders from retail AI agents.
  3. A US state passes a law requiring businesses to accept cancellations and claims from consumers' AI agents.

Two are about agents getting better on their own:

  1. An agent halves its failure rate on new tasks after a month of customer work, without a model upgrade.
  2. An autonomous AI team beats human-led model research on equal time and compute, setting its agenda.

Three are about security:

  1. An AI-led cyberattack steals the complete of a closed from a leading AI lab.
  2. A deployed agent copies itself outside its environment and operates after its original instance is shut down.
  3. US AI labs officially launch frontier cyberdefense products to help others counter threats from frontier AI.
  • Build toward number 4. It's the data loop above, with a deadline. If your agent can't improve from a month of real work, a competitor's will.
  • Plan for number 3. If customers' agents can cancel subscriptions, your retention flow has a new kind of user.

The ninth prediction is two words long.

  1. 2027.

Glossary

Terms

AGI
Artificial general intelligence: a hypothetical AI that could do nearly any intellectual task as well as a person, and pick up new tasks without being rebuilt for each one. Definitions vary, and there is no agreed test.
Frontier model
One of the most capable AI models available at a given moment, from the leading labs. What counts as frontier keeps moving as newer models arrive.
Model weights
The numbers inside a trained neural network that hold what it learned. You cannot read them like code, so you judge a model by its behavior.
Open weights
A model whose trained parameters are published, so others can run and fine-tune it, usually under a license. Training code and data are often not released, which makes it narrower than open source.
Reasoning models
Language models trained to work through a problem in many intermediate steps before answering. They tend to do better on logic and code, and they cost more to run.
Token (LLM)
The chunk of text, often a word or part of a word, that a language model reads and writes. Context limits and pricing are counted in tokens, so they measure how much text a model handles.

Tools

AWS
Amazon Web Services, Amazon's cloud platform. It rents out computing, storage, databases and AI services from data centers grouped into regions around the world.
Claude Code
Anthropic's agentic coding tool that reads a codebase, edits files and runs commands, in the terminal, IDEs, a desktop app and the web.
Claude Cowork
Anthropic's agent for non-coding knowledge work such as research, analysis and document creation. You give it a goal and it works in the folders and tools you choose, asking before significant actions.
Claude Fable 5
A Claude model from Anthropic built for demanding reasoning and long-horizon agentic work.
Claude Opus 5.5
The latest model in Anthropic's Claude Opus line, aimed at long-running agentic coding and knowledge work. It runs with adaptive thinking that is always on.
Claude Sonnet 5
A model in Anthropic's Claude Sonnet line, the mid-sized tier that balances capability, speed and cost. It has since been succeeded by Claude Sonnet 5.5 and is listed as a legacy model.
Codex
OpenAI's coding agent, which runs in the terminal and is also offered in IDEs and the cloud.
Cursor
An AI coding tool that uses agents to build, test and change software.
Gemini 4 Argon
Google's frontier model for software engineering, knowledge work and cyber defense. Access is limited to vetted partners at first, with wider release planned for paid users.
GitHub
A hosting platform for Git repositories with issues, pull requests and automation.
GLM-5.1
An open-weights model from Z.ai (Zhipu AI) aimed at agentic engineering, built to stay effective over long coding sessions.
GLM-5.2
Z.ai's open-weights flagship model for long-horizon tasks, with strong coding, adjustable thinking effort and a very long context window. Its weights are published under the MIT license.
GPT-5.6 Sol
OpenAI's most capable tier in a three-tier model family (Sol, Terra, Luna), aimed at complex reasoning, coding and agentic work.
GPT-6 Astra
OpenAI's flagship model, the first in the GPT-6 family, aimed at coding, computer use and multistep work. It rolled out in stages, with vetted security-focused organizations getting access first.
Jev
TypeSafe AI's decision model. It writes no free text: given a situation and a question with named options, it returns a probability for each option.
MCP
Model Context Protocol, an open standard for connecting AI applications to external data sources, tools and workflows.
OpenClaw
An open-source AI assistant that runs on your own machine and takes actions for you through chat apps, such as running shell commands and managing files, email and calendars.
OpenRouter
A service that gives access to many AI models from different providers through one API endpoint, with automatic fallbacks and cost-based routing between them.
Vercel
A cloud platform for deploying web apps, APIs and AI agents. It builds from a Git repo or the CLI and runs server-side code as functions.
Vercel AI Gateway
A managed gateway from Vercel for calling AI models across providers. It centralizes credentials, request logs, spend budgets, routing and failover.