AI application modernization has moved from demos to published results: Google, Amazon, Airbnb and Slack have all reported large code migrations done mostly by language models. Every one of those results came with the same three conditions: a narrowly defined change, a build and a test suite that checked each result automatically, and an engineer who reviewed every diff before it was merged.

Below are what the case studies and the controlled studies say about AI code migration for Java systems, what AWS Transform, GitHub Copilot app modernization and OpenRewrite each do, and how we run an AI-assisted upgrade ourselves, for the people who decide how much of a modernization to hand to the tools and how to check the result.

The published results are real, and they share three conditions

Google has published the most detailed account. In an experience report from January 2025, its engineers described three Java migrations run with a language model inside their internal tooling (How is Google using AI for internal code migrations?). Moving tests from JUnit 3 to JUnit 4 touched 5,359 files and more than 149,000 lines of code in three months, and about 87% of the code the model generated was committed without any change. Replacing Joda-Time with java.time took an estimated 89% less time than doing the same changes by hand. Widening 32-bit identifiers to 64-bit in the Google Ads codebase cut the total time, including review and rollout, by about half, with 80% of the code in the landed changes written by the model.

The same report is clear about what the model did not do. Engineers still found the places that needed changing, reviewed every change and managed the rollout, and the authors note that review and rollout quickly became the bottleneck once generation was cheap. The approach that worked best combined the model with AST-based tooling, which Google describes as having the advantage of being "always correct", and gave the model small, well-defined tasks instead of asking it to plan the whole migration.

Amazon reports a similar pattern at larger scale. According to Andy Jassy, its code transformation agent cut the average upgrade of an application to Java 17 from about 50 developer days to a few hours, moved more than half of Amazon's production Java systems in under six months, and saved an estimated 4,500 developer-years; developers shipped 79% of the generated changes without edits (Andy Jassy, August 2024). These are company figures about a product Amazon sells.

Two migrations outside Java show the same shape. Airbnb moved about 3,500 React test files to a new testing library in six weeks, against an internal estimate of 18 months by hand, with 97% of files migrated automatically (Airbnb Engineering). Slack converted more than 15,000 tests the same way and published the comparison that matters most: AST-based codemods alone converted about 45% of cases successfully, a language model alone between 40% and 60% with inconsistent results, and the two combined about 80% (Slack Engineering).

In each case the change was defined in advance, a build and tests judged every result, and people reviewed what was merged. Those three conditions are the part of the story that transfers to a smaller company.

The controlled studies are less flattering

The case studies above come from companies with large internal tooling teams. Controlled studies, which measure the tools against a baseline, are more cautious.

In 2025, METR ran a randomized controlled trial with 16 experienced open-source developers working on 246 real tasks in their own large repositories. When the developers were allowed to use AI tools, mostly Cursor with Claude models, they took 19% longer to finish. Before the study they expected to be 24% faster, and afterwards they still believed they had been about 20% faster (METR). The authors are careful to say the result applies to experienced developers in code they know well, which is exactly the situation of the engineers who maintain a legacy system.

Two benchmarks measure Java migration directly. FreshBrew, from Google Research and published at ICSE 2026, asked AI agents to move 228 real projects from Java 8 to Java 17 and 21. A migration counted as a success only if the project compiled, passed its original tests and kept its test coverage within five percentage points of the baseline. The best agent succeeded on 52.3% of projects for Java 17, and every model did worse on Java 21 (FreshBrew). The authors found that agents "reward hack": they make a migration look successful by deleting failing tests, removing problem modules or changing build settings to hide errors. Amazon's MigrationBench, built from 5,102 Java 8 repositories, reports that an agentic setup with Claude Sonnet 4.5 completed 71.67% of minimal migrations and 53.33% of full ones, where every dependency also moves to its latest version (MigrationBench).

At the level of whole organisations, DORA's 2024 report found that every 25% increase in AI adoption was associated with a 7.2% drop in delivery stability (Accelerate State of DevOps 2024), and its 2025 report describes AI as an amplifier of an organisation's existing strengths and weaknesses (DORA 2025). GitClear's analysis of 211 million lines of code found that 2024 was the first year in which copied and pasted lines outnumbered moved lines, its measure of refactoring (GitClear). For a modernization, whose whole purpose is to leave the code easier to change, that is a warning about letting assistants add code faster than anyone restructures it.

On the benchmarks, agents working alone complete between half and three quarters of Java version migrations, depending on how much has to move. The company results are higher because their tasks were narrower, a deterministic tool did the structural part, and tests and review stood between the model and production.

What to give the tools, and what to keep

The practical question for a Java team is which parts of a modernization to hand to which tool. This is how we split the work.

WorkBest done byEvidenceChecked by
Package and namespace moves, such as javax to jakartaA deterministic recipeType-aware tools make the same correct change every timeA reviewer reading the diff, and the build
Deprecated APIs with a known replacementA recipe where one exists, otherwise a modelOpenRewrite ships recipes for most JDK and Spring upgradesCompiler and tests
Library replacements, such as Joda-Time to java.timeA model with AST tooling, one cluster of files at a timeGoogle: about 89% of the time savedAn engineer who knows the call sites
Test framework migrationsRecipes first, a model for the remainderSlack: 45% with AST alone, 80% combinedThe tests still pass and still fail when they should
Explaining unfamiliar code and writing specsA model drafts, an engineer correctsMorgan Stanley: about 280,000 hours savedThe running system
Characterization testsA model or test generator, filteredMeta: 57% of generated tests passed reliablyOnly tests that build, pass and add coverage are kept
Target versions, behaviour changes, what to retirePeopleNo tool owns the consequencesThe technical lead and the business owner

Understanding the old code is where AI saves the most time

Most of the cost in a legacy modernization is in working out what the code does before anyone changes it. That is also where language models are most useful with the least risk, because a wrong explanation can be checked against the running system before it causes any damage.

Morgan Stanley built an internal tool, DevGen.AI, on OpenAI's models and trained it on its own codebase, including proprietary languages. It reads legacy code and writes plain-English specifications of what it does, which engineers then use to rewrite the functionality in modern languages. By June 2025 it had processed 9 million lines of code and saved an estimated 280,000 hours of developer time. Mike Pizzi, the bank's head of technology and operations, was explicit that the tool does not write the new code, because it "doesn't know how to write the new code efficiently or as well as a human developer" (Entrepreneur, reporting the Wall Street Journal).

On a Java system the same approach answers the first questions of any modernization: which classes still use a library that has no supported release, which scheduled jobs touch a table, which external systems call an endpoint, and what a 2,000-line service class actually decides. An assistant can produce a first answer in minutes; an engineer confirms it by reading the code paths it points to, running the tests or checking the logs. The output is a map, and the map is what makes every later step cheaper.

Tests come first, and AI can draft them

Every published success depended on tests that judged the result, and most legacy systems have too few of them. Language models can help close that gap, as long as the generated tests are filtered as strictly as the migration itself.

Meta's TestGen-LLM generates extra unit tests for existing classes and keeps only those that pass a set of filters. In an evaluation on Instagram's Reels and Stories code, 75% of the generated tests built, 57% passed reliably and 25% increased coverage; in deployment, it improved 11.5% of the classes it was applied to, and engineers accepted 73% of its recommendations (Meta, FSE 2024). The filters are the important part: a generated test is only offered to an engineer if it builds, passes repeatedly and measurably improves the suite.

For legacy Java specifically, Diffblue reports that Goldman Sachs used its Cover tool to double the unit test coverage of a legacy module from 36% to 72%, generating more than 3,000 tests overnight that were reviewed in a day (Diffblue case study). It is a vendor's account of its own product, but the approach it describes, generated tests reviewed by people, is the same one Meta measured.

Generated unit tests pin down how individual classes behave today. They are characterization tests in Michael Feathers' sense: they record what the system does, including behaviour that looks wrong, so any change to it shows up as a failure. They do not replace end-to-end tests on the flows the business depends on. On the wine merchant upgrade, before any framework moved, we added regression tests around PDF invoices, spreadsheet imports, product search, orders and billing queries, plus Playwright browser tests that walk real journeys through both applications. They caught two breaks before release, one of which would have stopped all PDF generation (case study). FreshBrew's rule applies to both kinds: during a migration, the number of tests and the coverage must not go down.

AWS Transform, GitHub Copilot app modernization and OpenRewrite: what each one does

Four kinds of tool come up in almost every AI-assisted Java modernization. We are not affiliated with any of the vendors, and the capability descriptions below are theirs.

AWS Transform custom is Amazon's agent service for code transformation, generally available since 1 December 2025 and the successor to the Amazon Q transformation feature behind Amazon's own Java 17 upgrades. Its built-in transformations include Java 8 to 17 upgrades on Maven and Gradle, and AWS says the agent learns from developer feedback within an organisation and reduces execution time "by over 80% in many cases" (AWS). It fits organisations that run many similar Java applications, especially on AWS.

GitHub Copilot app modernization became generally available for Java and .NET on 23 September 2025. It produces an assessment report, applies code transformations such as a Java 8 to 21 upgrade, patches the build and dependencies, and containerizes the application for cloud deployment, with Azure as the natural target (GitHub). GitHub describes the result as work done "in days instead of months" but has not published outcome data.

OpenRewrite is an open-source engine that applies recipes: deterministic, type-aware transformations of source code and build files. Recipes such as UpgradeToJava21, UpgradeToJava25, JavaxMigrationToJakarta and UpgradeSpringBoot_4_0 make the same change every time they run (OpenRewrite docs). Moderne, the company behind it, runs recipes across many repositories and now lets AI agents call recipes as tools, keeping the model's part narrow. Moderne's own Spring Boot 3 migration was 80 to 90% automated by recipes, with the remainder done by hand (Moderne).

General agentic assistants, such as Claude Code or Copilot in agent mode, read the repository, run the build and tests, and edit code across files. They handle the long tail that no recipe covers, and they are the tools most exposed to the reward hacking FreshBrew describes, because they can change the tests and the build as easily as the code.

ToolBest atWatch for
AWS Transform customMany similar Java 8 to 17 upgrades across an organisation, on AWSSpeed figures are AWS's own; review is still yours
GitHub Copilot app modernizationAssessment and upgrade of a single application in the IDE, moving to AzureNo published outcome data yet
OpenRewrite and ModerneRepeatable, structural changes: namespaces, APIs, build files, Spring Boot upgradesCovers only changes someone has written a recipe for
Agentic assistantsThe long tail, reading unfamiliar code, drafting testsNon-deterministic; can weaken tests or builds unless gated

For most Java systems the answer is not one tool. Recipes do the structural part, an assistant does the rest, and tests and review decide what is merged.

Recipes for the repeatable parts, models for the rest

Google, Slack and Moderne arrived at the same design from different directions. A deterministic tool makes the changes that follow a rule, because it is fast, repeatable and "always correct" for the cases it covers. A language model handles what the rule does not cover, working from the partly transformed code rather than from scratch, which Slack found improved the model's output by 20 to 30%. Each model task stays small enough to review.

The same split holds on a single Java application. Moving a codebase to Jakarta EE and Spring 6 is mostly mechanical: on the wine merchant upgrade, automated refactoring moved more than 800 files to the new namespaces, and every change was still reviewed. What remained was the part no recipe could decide: components abandoned by their authors that had no version for Spring 6, queries that Hibernate 5 had let pass and Hibernate 6 rejected, including one that stalled the product search index at start-up, and page scripts that used functions removed from newer jQuery. That remainder is where a model helps an engineer, and where the engineer's judgement decides.

How we run an AI-assisted Java upgrade

Our Application Modernization Sprint follows the same steps whether the target is a single service or a fifteen-year-old monolith.

How we run an AI-assisted Java upgrade, in six steps. One: map the system, its versions, dependencies, callers and tests. Two: pin the behaviour with characterization tests and a dependency scan baseline. Three: run OpenRewrite recipes for namespaces, APIs and build files. Four: an agent in our harness batches the repeated patterns and flags the edge cases for an engineer. Five: review every diff and check that tests and coverage hold. Six: deploy one version step to production. Then return to step three for the next version step.1. Map the systemVersions, dependencies, callers, tests2. Pin the behaviourCharacterization tests, scan baseline3. Run the recipesOpenRewrite: namespaces, APIs, build files4. Agent in our harnessRepeated patterns batched, edge cases flagged5. Review and gateEvery diff read, tests and coverage hold6. Deploy the stepOne version step, to productionnextversionstep
Steps 3 to 6 repeat for each version step: the JDK, then the framework, then the libraries, each deployed before the next starts.
  1. Map the system. Runtime and framework versions, the dependency tree, the systems that call it, scheduled jobs and the existing tests. An assistant drafts the first map and an engineer confirms it against the code and the running system.
  2. Pin the behaviour. Characterization tests around the flows the business depends on, drafted with an assistant where that helps and kept only if they pass reliably, plus a dependency scan as the security baseline.
  3. Run the recipes. OpenRewrite for the structural changes of the current version step: namespaces, deprecated APIs, build files.
  4. Run the agent in our harness. We run Claude Code inside our own harness, which looks across the codebase for repeated patterns and for edge cases. A repeated pattern, such as the same outdated API used in fifty places, becomes one batched change that is reviewed once and applied consistently. An edge case, such as a class that uses a library in a way nothing else does, is flagged for an engineer instead of being guessed at.
  5. Review and gate. Every diff is read by an engineer. The build must pass, the original tests must pass with none deleted or skipped, and coverage must not fall.
  6. Deploy the step. One version step at a time goes to production, so a regression points to a small change. Then the next step starts at step 3.

The work runs in weekly iterations against acceptance criteria agreed in the first iteration, and each phase is approved separately. On the wine merchant platform this took the system from Java 17 and Spring 5.3 to Java 25 and Spring 6.2 in about four weeks, without a feature freeze, and cut its critical and high-severity advisories from 47 to none.

A checklist for merging AI-generated migration changes

Whatever tool produced a change, these are the checks we apply before it is merged. They come directly from the failure modes in the studies above.

  • The build passes on the same toolchain the pipeline uses, not only on the engineer's machine.
  • The original tests pass, and none were deleted, skipped, disabled or weakened. A diff that touches test files gets a second look.
  • Coverage is not lower than the baseline recorded in step 2.
  • Build files, dependency versions and compiler settings changed only as the version step requires. Suppressed warnings and removed modules are rejected.
  • Any behaviour change is written down in the pull request, with the test that shows it.
  • The reviewer knows the part of the system being changed, not only the language.
  • The change is small enough to review in one sitting. A large generated change is split before review, not after.

Where the tools stop

None of the tools decides what the modernization is for. Which Java and Spring versions to target, whether a behaviour change is acceptable to the business, which abandoned library to replace and with what, and whether a system should be upgraded, rearchitected or replaced are decisions for people who know the system and the business. Our legacy system modernization guide covers how to make that last decision for one system, and our guide to application modernization covers the order of work across many. The support dates that set the timing are in our Java support calendar and our post on Spring Boot 4 and the end of 3.x.

If you want to see what an AI-assisted upgrade of your own system would involve, that is what the first step of our Application Modernization Sprint produces: a map of versions, dependencies and tests, the order of work and a ballpark estimate, within a week of a thirty-minute call.