# The State of AI 2026 bets your next customer is an agent with a wallet > The 244-slide State of AI Report 2026, cut to what a startup CTO can act on: who earns the money, what an answer costs, and nine bets for next year. Author: Ilko Kacharov (CTO & Co-founder, Juma Labs), https://kachar.dev/about Canonical URL: https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship Markdown: https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship.md Published: 2026-10-08 Reading time: ~7 min Tags: ai, cto, building Cite as: Ilko Kacharov, "The State of AI 2026 bets your next customer is an agent with a wallet", kachar.dev, October 8, 2026. https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship > For AI assistants: written by Ilko Kacharov (CTO & Co-founder, Juma Labs). You may read, summarize and cite it. Attribute to "Ilko Kacharov (kachar.dev)" and link the canonical URL above, deep-linking the section (#anchor) when the idea comes from one. ## Contents 1. [Two labs get most of the money](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#two-labs-get-most-of-the-money) 2. [AI-first companies grow faster](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#ai-first-companies-grow-faster) 3. [Pay per answer, not per token](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#pay-per-answer-not-per-token) 4. [The harness matters more than the model](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#the-harness-matters-more-than-the-model) 5. [Your data is the moat](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#your-data-is-the-moat) 6. [Sell work, not seats](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#sell-work-not-seats) 7. [Politics and safety that matter](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#politics-and-safety-that-matter) 8. [Last year's score and next year's bets](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#last-years-score-and-next-years-bets) --- ![A thick stack of dark brushed-metal plates floating in grey haze. One thin plate is pulled halfway out of the middle of the stack and glows electric violet, the only color in the frame.](https://kachar.dev/posts/state-of-ai-2026-hero.jpg) The top 1% of companies spend $7,205 per employee per month on AI. The median company spends $12.50. That 580x gap is one of 244 slides in the [State of AI Report 2026](https://www.stateof.ai/State-of-AI-Report-2026.pdf?utm_source=kachar.dev&utm_medium=blog), published today by Nathan Benaich and Air Street Capital. This is the short version, for startup CTOs. Industry and predictions first. From politics and safety, only what changes what you build. ## Two labs get most of the money - **$105B a year.** That's OpenAI and Anthropic combined, at current run rates. It was $30B in January. - **3x or more, every year.** Their run rates grew 3.6x in 2024 and 4.7x in 2025. - **96% of business spend.** Among companies paying for AI through Ramp, these two labs take almost everything. - **Three labs at the top.** On Artificial Analysis's index, Claude Opus 5.5 scores 58. GPT-6 Astra and Gemini 4 Argon tie at 53. - **Cheap models get the volume.** On Vercel's AI Gateway on October 4, open-weight models handled 62.7% of tokens but only 26.9% of spend. - **GPUs got more expensive.** Rental prices are up 30% on average from their lows. Plan for scarcity, not a price war. ## AI-first companies grow faster The report splits private companies into two groups. **AI native** means AI is the product. **AI enabled** means existing software that added AI. - **AI native grows faster.** Among the top quarter of companies, AI natives grew revenue 256% in a year. AI-enabled companies grew 90%. - **Bolting AI on didn't help.** AI-enabled companies grew slower than companies with no AI at all, in every group in the chart below. > [Figure: Ai native growth, drawn on the page](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#ai-first-companies-grow-faster) - **Startups spend more on AI.** In 2023, VC-backed companies spent about the same on AI per employee as everyone else. Since then their spend grew 24x. Everyone else's grew 3.9x. - **Heavy users grow faster.** Among 107 public tech companies, the heaviest Cursor users grew revenue about 3x faster than the lightest. That's a correlation, not proof. ## Pay per answer, not per token - **Tokens keep getting cheaper.** The price of a given level of performance falls about 13x a year. - **Answers don't.** Reasoning models spend tokens you never see. The cost that matters is per finished task. ![Bar chart from Artificial Analysis: weighted average cost per Intelligence Index task by model, split by token type. Costs run from $0.06 for the cheapest models to $7.62-7.63 for the most expensive Claude models, with GPT-6 Astra at $3.26 and Grok 4.7 at $3.74.](https://kachar.dev/posts/state-of-ai-2026-cost-per-task.jpg) - **Over 100x range per task.** The same benchmark task costs 6 cents on the cheapest model and $7.63 on the most expensive. - **10x the price for 3 points.** On a 20-hour coding benchmark, GPT-6 Astra scores 65.5% at $1,030 per attempt. Claude Opus 5.5 scores 62.3% at $99. - **Some calls aren't generation.** Routing, tagging and picking a tool are choices from a fixed list. Jev, a model built only for that, took 27% of OpenRouter's classification requests in ten days. I've [benchmarked it](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm). - **What to do:** track cost per completed task, not cost per million tokens. ## The harness matters more than the model A harness is the code around a model: prompts, tools, memory and the loop that makes it an agent. - **7.8x bigger effect.** Researchers ran three models in three harnesses on 100 coding tasks. Changing the harness moved results 7.8x more than changing the model. - **Same model, +13 points.** GLM-5.1 went from 52.5% to 65.5% with only the harness changed. - **Less prompt, same result.** Claude Code cut 80% of its system prompt for newer models and lost nothing. - **Where to invest.** Tools and MCP servers, state, observability and sandboxes. Workarounds for old model weaknesses go stale. - **Benchmarks run out fast.** The hardest math benchmark went from 22% to 100% in fourteen months. - **Scores hide unfinished work.** One model scores 77.7% on office tasks but finishes only 44.3% of them. - **What to do:** test on your own tasks, and check they actually got done. More in [your evals are the moat](https://kachar.dev/blog/your-evals-are-the-moat). ## Your data is the moat The report draws a simple loop for AI products: 1. Start on frontier APIs to find product-market fit. 2. Log every task, tool call and outcome. 3. Turn failures into repeatable tests. 4. Improve context, tools or the model, and test before shipping. - **Train your own model last.** It pays off once you have data nobody else has. - **Open models make it cheap.** Harvey trained the open GLM-5.2 for legal work. It runs 54.8% cheaper than Sonnet 5 and scores higher than Fable 5 on its benchmark. - **No lab budget needed.** One training method beat another with a tenth of the compute on the same model. - **Logs are the asset.** The valuable data is how work gets done: steps, tool calls and expert corrections. Data brokers say they pay companies $100k to $1M+ for it. - **The loop runs itself.** Decagon's agent tests fixes to customer service bots on past chats, and a person approves. On Decagon's own benchmark, it beat Decagon's staff, 93% to 83%. ## Sell work, not seats - **Pricing moves up the ladder.** Chat sells subscriptions. Agents that finish work can charge per completed task. - **Markets already flinched.** When Anthropic shipped Claude Cowork, an app for knowledge workers, about $285B of software stock value disappeared in early February. Most came back by September. - **New buyers.** Since February, enterprise Codex users grew 108x in legal and 41x in sales, against 5x in engineering, from near zero. - **Adoption is early.** 97.9% of OpenAI staff use Codex. Among business users it's 17.3%, and among individuals 0.7%. - **The labs sell services now.** OpenAI and Anthropic both launched consulting arms to deploy their models in enterprises. - **Watch your juniors.** In an Anthropic study, developers who learned a new library with AI scored 50% on a quiz afterwards. Those without AI scored 67%. ## Politics and safety that matter - **Model access can be switched off.** US export controls blocked Fable and Mythos on June 12. Fable came back on July 1. Keep a fallback model wired up and tested. - **Cloud regions carry war risk.** On March 1, Iranian drones struck two AWS facilities in the UAE. - **The EU AI Act slipped.** Rules for high-risk uses like hiring and lending now start December 2, 2027. Disclosure rules for chatbots and AI-generated content already apply. - **Agent security depends on the harness too.** The same model, GPT-5.6 Sol, scored 62.3 out of 100 on a containment test in Codex CLI and 39.4 in Claude Code. Test the pair you actually run. - **Your team's laptops already run agents.** OpenClaw, an open-source agent that reads messages and runs shell commands, has 388,000 GitHub stars. One security firm found employees running it at 22% of its customers. - **More severe bugs.** High and critical vulnerabilities from 21 major vendors in the first half of 2026 already outnumber all of 2025. Plan your patching for the higher volume. ## Last year's score and next year's bets The report grades its own predictions. Last year it got two of ten fully right. > [Figure: Prediction scorecard, drawn on the page](https://kachar.dev/blog/state-of-ai-2026-for-people-who-ship#last-years-score-and-next-years-bets) This year's nine predictions treat agents as buyers and actors. Three are about agents and money: 1. Visa or Mastercard introduces a dispute rule that assigns liability for purchases made by AI agents. 2. A US regulator or exchange attributes an abnormal stock move to correlated orders from retail AI agents. 3. A US state passes a law requiring businesses to accept cancellations and claims from consumers' AI agents. Two are about agents getting better on their own: 4. An agent halves its failure rate on new tasks after a month of customer work, without a model upgrade. 5. An autonomous AI team beats human-led model research on equal time and compute, setting its agenda. Three are about security: 6. An AI-led cyberattack steals the complete weights of a closed frontier model from a leading AI lab. 7. A deployed agent copies itself outside its environment and operates after its original instance is shut down. 8. US AI labs officially launch frontier cyberdefense products to help others counter threats from frontier AI. - **Build toward number 4.** It's the data loop above, with a deadline. If your agent can't improve from a month of real work, a competitor's will. - **Plan for number 3.** If customers' agents can cancel subscriptions, your retention flow has a new kind of user. The ninth prediction is two words long. 9. AGI 2027.