In January a bank approves an AI transformation pilot built on the best model it can buy. By June, when the pilot is ready for review, that model has been replaced twice.
That is the pace in 2026. Almost every large company now uses AI somewhere, far fewer can find it in their profit, and the difference lies in how the work around the model is organized.
What is AI transformation?
AI transformation is the work of changing how a company operates so that AI does part of the job. A claim is checked by software agents before an adjuster sees it. A doctor's note drafts itself during the visit. A payment is scored against the history of billions of others in the time it takes to authorize it.
It is a step beyond AI adoption, which is people using an AI tool in their daily work. Transformation means a process redesigned around what the model does well, with a number that shows whether it worked.
Is a company that bought chatbot licenses for every employee transformed? Not yet.
It has the tool. The transformation starts when a workflow, the data behind it and the people who run it change because of what the tool can do.
Where frontier AI models stand in October 2026
The current frontier was set by three releases in a single September.
| Model | Released | What the lab emphasizes | Who can use it |
|---|---|---|---|
| Claude Opus 5.5 (Anthropic) | 22 Sep 2026 | Agentic coding, computer use and long tasks such as codebase-wide migrations; 40% cheaper to run than Opus 5 | Anthropic's apps and API, AWS, Google Cloud and Microsoft Azure; biology and security work through vetted programs |
| GPT-6 Astra (OpenAI) | 3 Sep 2026 | A "new capability level", in Sam Altman's words; OpenAI's first model to reach its "Critical" cybersecurity threshold | ChatGPT Plus, Pro, Business and Enterprise, the API and AWS, with partners in OpenAI's cybersecurity program first |
| Gemini 4 Argon (Google) | 30 Sep 2026 | Coding, research and writing, with a focus on defensive cybersecurity | So far only Google's partners in its Fairwind security program |
Each lab says its model leads, on the tests it chose to publish. Anthropic's table puts Opus 5.5 at 66.4% on Terminal-Bench 4.0, a test of multi-step work in a command line, against 57.9% for GPT-6 Astra as reported by OpenAI. Google points to the Vals index for Argon. Anthropic also adds a line worth reading twice: "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences."
For a business, that means the brand of the model matters less than it did two years ago. The leaders are close, and the lead changes hands within weeks.
What the benchmarks say about the last year
Stanford's 2026 AI Index, an independent yearly review, gives the wider picture. On SWE-bench Verified, a coding test built from real software issues, performance rose "from 60% to near 100% in a single year". On OSWorld, which asks AI agents to complete real tasks on a computer, success went from about 12% to 66.3%, close to the 72.4% people manage. On Humanity's Last Exam, a test written to be hard for AI, frontier models gained 30 percentage points in a year.
The gaps that used to shape buying decisions have narrowed too. In March 2026 the best closed model led the best open-weight model by 3.3%, and the top US model led the top Chinese one by 2.7%.
The same report carries a warning. Gemini Deep Think earned a gold medal at the International Mathematical Olympiad, yet the top model reads an analog clock correctly only 50.1% of the time. Researchers call this the jagged frontier: a model can be brilliant at one task and unreliable at the next one over.
That is why every use case needs its own test, run on your own examples.
How long an AI agent can work on its own
For a business, a more telling measure is the length of task an AI agent can finish. METR, a nonprofit that evaluates AI models, measures the length of software tasks, counted in the time a skilled person needs, that a model completes half the time.
In March 2023 GPT-4 managed tasks of about four minutes. In April 2026 an early version of Claude Mythos Preview managed tasks METR puts at 16 hours or more, which is also where its test suite stops giving reliable readings. Since 2023 the length has doubled about every 129 days, a little over four months.
Half the time is a low bar for production work, and METR's tasks are self-contained software problems rather than the messy processes of a real company. The direction still matters for planning: tasks you rule out for AI today may be within reach before your project ends.
The most capable models now come with access controls
September brought something new. OpenAI gave its cybersecurity program partners first access to Astra. Google released Argon only to partners in its Fairwind security program. Anthropic ships Opus 5.5 with safeguards, and offers its biology and security uses through vetted programs.
If your plan depends on a model's newest abilities, check that you can actually use them, and on what terms.
How fast new frontier models are released
In 2023 a company could choose a model and expect it to stay current for most of a year. That window has closed.
Counting launches of new Claude, GPT and Gemini models in Epoch AI's database of AI models, the three labs announced new models on 12 separate days in 2023, 16 in 2024 and 31 in 2025. In the first nine months of 2026 they had already reached 27.
Anthropic's Opus line shows the rhythm: Opus 4.5 in November 2025, then 4.6 in February, 4.7 in April, 4.8 in May, Opus 5 in July and Opus 5.5 in September. That is six versions of one flagship in ten months, and OpenAI and Google kept a similar pace.
What does that mean for a company planning an AI project? Two things.
First, build so the model can be swapped. Keep your prompts, your test cases and the results of each evaluation in your own repository, so trying a new model means rerunning a test suite.
Second, expect prices to fall. Opus 5.5 costs 20% less per token than Opus 5, and Anthropic puts the saving on typical workloads at 40%. A business case that works at this year's prices will look better next year. A business case that only works with next year's capabilities is a bet.
How businesses are adopting AI in 2026, and how often it pays off
Adoption figures depend on who is asked.
Among the organizations in the surveys collected by the AI Index, which lean toward larger firms, 88% used AI in 2025 and 70% used generative AI in at least one business function. AI agent deployment was in the single digits across nearly all business functions.
The US Census Bureau asks every kind of business, including the smallest. Between December 2025 and May 2026, between 17% and 20% of US businesses said they had used AI in the previous two weeks. Among firms with at least 250 employees the figure was 37%; in the information sector 39.7%, and in finance and insurance 33.9%.
In short, most large companies have started, and most small ones have not.
Where AI delivers measurable gains
The gains are clearest where the work is structured and the output is easy to count. Studies collected by Stanford report productivity gains of 14% to 15% in customer support, 26% in software development and 50% in marketing output, with smaller gains on tasks that need deeper reasoning.
Software teams are the subject of our guides to the best AI coding tools and to AI-assisted application modernization, and of our look at the future of software engineering.
Where AI transformation falls short
At the level of the whole company, the results are thinner.
MIT NANDA's report The GenAI Divide, based on 52 interviews, 153 survey responses and a review of more than 300 public initiatives, found that only 5% of organizations got task-specific AI tools into production with a lasting effect on productivity or profit. Sixty percent had evaluated such tools and 20% had run a pilot. General-purpose chatbots fared better, mostly as aids to individual productivity.
The report's explanation is practical. Most tools "do not retain feedback, adapt to context, or improve over time", so they break on the exceptions every real workflow has. It also found that companies working with an outside partner reached deployment about twice as often as those building alone, roughly 67% against 33%. We are such a partner, so read that figure with the caution it deserves.
Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, because of rising costs, unclear value or weak risk controls. It also estimates that only about 130 of the thousands of vendors selling agentic AI offer the real thing; the rest is what Gartner calls "agent washing", chatbots and automation tools under a new name.
Two public cases show what going wrong looks like.
In February 2024 Klarna announced that its AI assistant handled two-thirds of its customer service chats in its first month, the work of 700 full-time agents, and estimated a $40 million profit improvement. In May 2025 its chief executive, Sebastian Siemiatkowski, told Bloomberg that the push toward AI had led to lower quality, and Klarna began recruiting human agents again.
In Australia, Deloitte agreed in October 2025 to partially refund the A$440,000 it was paid for a government report that had to be corrected. The revised version disclosed that a tool chain based on GPT-4o had been used to prepare it; Deloitte did not attribute the errors to the tool.
Neither problem was the model's quality. Klarna measured cost and not the customer's experience, and a report went out without the checking a junior analyst's draft would get.
Three AI transformation examples that paid off
The successes share a shape: a narrow task, a number measured before and after, and a person who stays accountable for the outcome. Here are three, from three industries.
Allianz: storm claims settled in hours by seven AI agents
After a storm cuts the power, thousands of households claim for spoiled food. Each claim is small and simple, and together they bury a claims team for days.
Allianz built Project Nemo for exactly these claims in Australia, in under 100 days, and launched it in July 2025. Seven AI agents split the work: a planner, a cyber agent for data security, and agents that check coverage, confirm a matching weather event, screen for fraud, calculate the payout and write an audit summary.
For claims under AUD 500, a claim reaches a human reviewer in under five minutes, and processing and settlement time fell by 80%, from four days or more to a day or less. "Payout decisions are never automated": a claims professional approves every payment.
Allianz plans to extend the same setup to travel delays, simple car claims and property damage. Insurers following the same path will find more on the industry in our overview of digitalization in insurance.
Stripe: a foundation model trained on payments
Card testing is a quiet kind of fraud. Criminals check which stolen card numbers still work by hiding small charges among the legitimate traffic of a large online shop.
Stripe trained its Payments Foundation Model on tens of billions of transactions, turning each payment into a compact numerical summary of its context. Reading those summaries in sequence, Stripe's detection rate for card-testing attacks on large users rose from 59% to 97% overnight, in the company's own words.
Stripe adds that card-testing attacks are down 80% on its network while they rise across the industry. These are company-reported figures, from a company that sells the protection. The lesson travels anyway: the advantage came from data only Stripe holds, not from a model anyone can rent. Our fintech trends post follows where payments go next.
Kaiser Permanente: AI scribes for 7,260 physicians
Doctors spend a large share of their day on notes, often in the evening after the last patient has gone home.
The Permanente Medical Group in Northern California gave its physicians AI scribes that record the visit, with the patient's consent, and draft the note for the doctor to check. Over the evaluation period, 7,260 physicians used them in more than 2.5 million patient encounters, and the analysis in NEJM Catalyst estimates the time saved at 1,794 working days in one year.
Physicians said the scribes had a positive effect on their interactions with patients (84%) and on their work satisfaction (82%). Patients noticed too: 47% said their doctor looked at the computer less.
Use was uneven. The top third of users accounted for 89% of activations, so the next gains depend on adoption rather than on a better model. The tool does not recommend treatment; the doctor signs the note.
Look at the three together and the pattern is plain. None of them asked AI to run a department. Each picked one frequent, well-understood task, measured it, and kept a professional in the loop where the stakes were real.
How to prepare your business for AI transformation: a checklist
The list below follows the order in which the decisions depend on each other. Work through it before you sign a contract or hire an AI team, and come back to it before every new use case.
Choose the work
- Pick one to three processes with a cost or a time you already measure: small claims, first-line support questions, document intake, invoice matching.
- Write down the baseline now (cost per case, handling time, error rate), so the pilot has a number to beat.
- Look at back-office work first. MIT found that budgets lean toward sales and marketing, while back-office automation often pays back more.
- Agree in advance which result ends the pilot and which result funds the rollout.
Data and systems
- List where the data for each process lives, who owns it and whether it may leave your environment.
- Check that the systems involved can be reached through an API. Older applications often need work first; our guide to AI-assisted modernization covers what that involves.
- Make sure an assistant can see only what the person asking is allowed to see.
People and accountability
- Name the person accountable for each AI-assisted decision, and keep a human approval wherever money, health or legal rights are at stake.
- Give staff approved tools. MIT found workers at over 90% of companies using personal AI tools, while only 40% of the companies had bought an official subscription.
- Show the people who will use the system working results every week or two, and change course on their feedback.
Models, costs and contracts
- Build an evaluation set from real cases, including the awkward ones, and rerun it whenever you change the model.
- Keep prompts, test cases and evaluation results in your own repository, so you can switch provider when a better or cheaper model ships.
- Track the usage cost per completed task, not only the monthly bill.
- Read the provider's data terms: retention, use of your data for training, and where it is processed.
Regulation
- If you serve customers in the EU, map each use case against the AI Act. Since the Digital Omnibus took effect in July 2026, obligations for the high-risk systems listed in Annex III, such as credit scoring and hiring, apply from 2 December 2027.
- Log what the system did and why, for each decision. Auditors and regulators will ask, and so will your own team the first time a result looks wrong.
Decide
- Compare the pilot against the baseline, then roll out, change or stop. A stopped pilot with a clear reason is a result.
- Decide per use case whether to buy a product, build in-house or work with a partner, and say why in writing.
Many of these processes run on Java systems that were never designed to call a model. Connecting the two is what our AI integration team does: a four-week pilot on one use case, with an evaluation set, measured usage costs and the code in your repository from the first commit. We built and ran Recostream, a recommendation engine that answered in 20 to 30 ms, before GetResponse acquired it in December 2022. A scoping call gets you a written ballpark with its assumptions before any work starts.