Java virtual threads, delivered by Project Loom and final since Java 21, let a server run one thread per request without running out of threads. They do not make code faster. They make waiting cheap, so a service that spends most of its time waiting on databases and other services can handle far more requests at once with the same simple, blocking code.
The failures reported from production come from assumptions that platform threads used to hide: thread pools that doubled as rate limits, locks that held an OS thread, and caches kept in thread locals. The sections below cover how virtual threads work, five code patterns that cover most uses, the production stories that shaped current advice, where they fit and where they do not, and a checklist for rolling them out.
A virtual thread borrows an OS thread only while it runs
A platform thread is a thin wrapper around an operating system thread. Each one reserves a stack of around a megabyte and costs a context switch in the kernel, so a typical server keeps a pool of a few hundred and makes requests wait for a free one. A virtual thread is an ordinary Java object managed by the JVM. When it runs, the JVM mounts it on a carrier thread, a platform thread from a small pool sized to the number of cores. When it blocks on I/O, a lock or a sleep, the JVM unmounts it, stores its stack on the heap, and the carrier picks up another virtual thread. When the I/O completes, the virtual thread is mounted again, on whichever carrier is free.
Little's law explains why this matters, and JEP 444 uses it to state the goal of the feature: for a given latency, the number of requests a server handles at the same time has to grow in proportion to its throughput (JEP 444). Take a service on Tomcat's default pool of 200 threads, where each request takes 100 ms, most of it waiting for a database. It can never complete more than 200 / 0.1 = 2,000 requests a second, however idle its CPUs are. With virtual threads the thread count stops being the ceiling, and the limit moves to whatever the requests are actually waiting for.
From Project Loom to Java 25: what changed in each release
Virtual threads arrived in stages, and the release you run decides which pitfalls apply.
| Release | What changed | What it means in practice |
|---|---|---|
| JDK 19 and 20 | Virtual threads in preview | For experiments only |
| Java 21 (LTS) | Virtual threads final (JEP 444) | Production-ready, but blocking inside synchronized pins the carrier |
| JDK 24 | Synchronized without pinning (JEP 491) | The main cause of pinning is gone |
| Java 25 (LTS) | Scoped values final (JEP 506); structured concurrency in fifth preview (JEP 505) | The first long-term support release with the pinning fix, and a supported replacement for thread locals |
For a team choosing a runtime today, that makes Java 25 the natural target: it is the first LTS release that includes JEP 491. The support dates for 21 and 25 are in our Java support calendar.
Five code patterns cover most uses
Most code does not need to create virtual threads directly, because frameworks do it. Where you do, five patterns cover nearly everything. All of them compile and run on Java 25.
1. One virtual thread per task. The executor creates a new virtual thread for every submitted task and closes itself at the end of the try block, after all tasks have finished. This is the standard way to fan out independent calls.
List<String> fetchAll(List<URI> uris) throws Exception {
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
List<Future<String>> futures = uris.stream()
.map(uri -> executor.submit(() -> get(uri)))
.toList();
List<String> bodies = new ArrayList<>();
for (Future<String> f : futures) {
bodies.add(f.get());
}
return bodies;
}
}
2. Named virtual threads. Names make thread dumps readable. The builder numbers each thread it starts.
Thread.Builder builder = Thread.ofVirtual().name("importer-", 0);
for (int i = 0; i < 3; i++) {
builder.start(work); // importer-0, importer-1, importer-2
}
3. A semaphore, not a pool, to protect a downstream system. A fixed thread pool used to limit how many calls reached a database or an API at once, as a side effect. Virtual threads should never be pooled, so the limit has to be stated explicitly. JEP 444 recommends semaphores for exactly this.
static final Semaphore PRICING_API = new Semaphore(20);
String quote(URI uri) throws Exception {
PRICING_API.acquire();
try {
return get(uri);
} finally {
PRICING_API.release();
}
}
4. ReentrantLock where a lock guards blocking work on Java 21 to 23. Before JDK 24, a virtual thread that blocked inside synchronized stayed pinned to its carrier. java.util.concurrent locks never pinned. On JDK 24 and later this rewrite is no longer needed.
private final ReentrantLock lock = new ReentrantLock();
void recordAndFlush() {
lock.lock();
try {
counter++;
flushToDisk(); // blocking call inside the lock
} finally {
lock.unlock();
}
}
5. Scoped values for request context. A scoped value carries data such as a request ID or the current user down the call stack for a bounded scope, is immutable, and is inherited by subtasks. It replaces most uses of ThreadLocal for context, and it is final in Java 25.
static final ScopedValue<String> REQUEST_ID = ScopedValue.newInstance();
void handle(String requestId, Runnable handler) {
ScopedValue.where(REQUEST_ID, requestId).run(handler);
}
void log(String message) {
String id = REQUEST_ID.isBound() ? REQUEST_ID.get() : "-";
System.out.println("[" + id + "] " + message);
}
Structured concurrency, in preview. StructuredTaskScope treats a group of subtasks as one unit: if one fails, the others are cancelled, and none outlives the scope. In Java 25 it is the fifth preview and needs --enable-preview to compile and run, so it belongs in code you can change when the API is finalized.
// Java 25 preview: javac --release 25 --enable-preview
Response handle() throws InterruptedException {
try (var scope = StructuredTaskScope.open()) {
Subtask<String> user = scope.fork(() -> findUser());
Subtask<Integer> orders = scope.fork(() -> countOrders());
scope.join();
return new Response(user.get(), orders.get());
}
}
Spring Boot virtual threads take one property
In Spring Boot 3.2 and later, including Spring Boot 4, setting spring.threads.virtual.enabled=true on Java 21 or newer moves most of the application onto virtual threads. According to the release notes, Tomcat and Jetty process requests on virtual threads; RabbitMQ, Kafka and Pulsar listeners get virtual-thread executors; and the application task executor behind @Async and asynchronous MVC becomes a SimpleAsyncTaskExecutor on virtual threads (Spring Boot 3.2 release notes). The scheduler becomes a SimpleAsyncTaskScheduler that ignores the pool size properties (Spring Boot reference).
Two details follow from that. Any concurrency limit that used to come from a pool size, such as server.tomcat.threads.max or the task executor's pool, no longer limits anything, so it has to be replaced with an explicit one. And an application whose only non-daemon work runs on virtual threads needs spring.main.keep-alive=true to keep the JVM running. If you are also planning the move to Spring Boot 4, our post on Spring Boot 4 and the end of 3.x covers the rest of that upgrade.
Netflix's deadlock shows what pinning looks like in production
The most detailed public account of virtual threads failing in production comes from Netflix. In July 2024 its JVM ecosystem team described services on Java 21, Spring Boot 3 and embedded Tomcat that, after virtual threads were enabled for request handling, suffered intermittent timeouts and instances that stopped serving traffic while the JVM stayed up. The clue was a growing number of sockets in CLOSE_WAIT: clients had given up and closed their connections, but the application never got round to closing its side (Netflix Technology Blog).
A thread dump taken with jcmd Thread.dump_to_file showed thousands of virtual threads that were doing nothing, and a heap dump showed why. Virtual threads had blocked on a lock inside synchronized code, which on Java 21 pinned each of them to its carrier. On the affected instances, every carrier thread was pinned by a virtual thread waiting for the same lock, while the virtual thread that could release the lock had no carrier left to run on (InfoQ). It is a deadlock between one lock and a scheduler with a fixed number of carriers, and nothing in the application code looked wrong.
JDK 24's JEP 491 changed the implementation so that a virtual thread blocked in synchronized releases its carrier, which removes this class of failure (JEP 491). Netflix's later talks record the same arc: virtual threads rolled back for safety, then re-evaluated once Java 25 shipped the fix (JavaOne 2026 talk notes). Pinning has not disappeared entirely. It still happens while a virtual thread runs native code through JNI or the foreign function API, during class initialization, and during local file I/O on Linux (InfoQ, July 2026). On any version, the JFR event jdk.VirtualThreadPinned shows where it happens, and on Java 21 the system property -Djdk.tracePinnedThreads=full prints the stack of each pinned thread.
The bottleneck moves to your connection pool
Once the thread pool stops being the limit, the next limit takes its place, and it is usually a database connection pool, an HTTP client's connection limit or the rate limit of another team's API. The InfoQ analysis of virtual threads after JDK 24 found this to be the most consistent pattern after an upgrade: pinning went away, but load that the old thread pool had held back now arrived at the database all at once, and a test that failed with platform threads failed at the same rate with virtual threads (InfoQ).
The old thread pool did two jobs at once: it ran the work and, without anyone deciding it, it limited how much work reached each downstream system. Virtual threads separate the two, so the limits have to be written down. In practice that means a connection pool sized for what the database can serve, not for the number of request threads; a semaphore in front of each external API whose limit you know; and timeouts on every call, because ten thousand virtual threads waiting on a stalled dependency are cheap for the JVM and still ten thousand stuck requests for your users.
ThreadLocal caches stop working
Libraries have long used ThreadLocal to cache expensive objects per thread, relying on a pool of a few hundred long-lived threads to reuse them. With a new virtual thread for every task, each cache is created, used once and thrown away.
Jackson, the JSON library behind most Spring applications, is the best-documented case. It recycled its internal buffers through a ThreadLocal, and under virtual threads that produced a steady stream of buffers that were never reused. Jackson 2.16 introduced a pluggable RecyclerPool so the strategy could be changed, and Jackson 2.17 made a shared lock-free pool the default (jackson-core issue 919). The InfoQ article measured the same effect in application code: a SimpleDateFormat kept in a thread local was initialized 200 times under platform threads and 443,267 times under virtual threads for the same requests, in the author's test (InfoQ).
The fix depends on what the thread local holds. Immutable, thread-safe objects such as DateTimeFormatter can be shared as a single constant. Expensive, non-thread-safe resources belong in an explicit pool. Request context, such as a trace ID or the authenticated user, belongs in a scoped value, which also solves a second problem the same article describes: values in an InheritableThreadLocal did not reach subtasks forked in a structured task scope, and tracing and security context were lost.
Millions of threads, but only with shallow stacks
Demonstrations of virtual threads often start a million of them doing nothing. Webtide, the company behind the Jetty web server, measured what happens when the threads do something. With stacks 1,000 frames deep, each virtual thread used about 152 KB of heap, more than the 114 KB a kernel thread used in the same test, and garbage collection pauses grew to 35 seconds at 55,000 threads (Webtide). Virtual threads are cheap because most of them are parked with small stacks most of the time. A workload where thousands are busy at once, deep inside a framework, still costs real memory.
Webtide also reported no performance benefit from adopting virtual threads in the core of Jetty, which already uses a carefully tuned asynchronous design, and kept virtual threads as an option for application code instead (Webtide, Introducing Jetty 12). Oracle took the opposite route with Helidon 4, whose web server was written from scratch around virtual threads and blocking code; the Helidon team reports performance comparable to a minimal Netty server with a much simpler, blocking programming model (Helidon). Both are consistent with the same rule: virtual threads help where code waits, and they add nothing where a system is already built not to wait.
Where virtual threads fit, and where they do not
| Application | Fit | Why |
|---|---|---|
| REST and web services on JDBC and HTTP clients | Strong | Most of each request is waiting; thread-per-request code scales without a rewrite |
| Aggregators and backends-for-frontends that fan out to several services | Strong | One virtual thread per call, with simple blocking code and structured cancellation |
| Message consumers that call databases or APIs per message | Strong | Listener concurrency stops being tied to a thread pool size |
| Batch and integration jobs calling slow external systems | Strong, with semaphores | High concurrency is cheap; the external system's limit must be explicit |
| Reactive code written only to avoid blocking threads | Good candidate to simplify | Blocking code with virtual threads is easier to read, debug and profile |
| CPU-bound work: pricing engines, compression, analytics | None | There is nothing to wait for; use platform threads, a ForkJoinPool or parallel streams |
| Heavy native calls or local file I/O on Linux | Weak | These still pin the carrier on current JDKs |
| Ultra-low-latency paths | Weak | A scheduler sits between the code and the CPU; dedicated threads and careful design still win |
| Systems already reactive and performing well | Little gain | The benefit is simpler code, not more throughput |
That last row answers the question most teams on Spring WebFlux ask. Virtual threads do not make a well-built reactive service faster. They make it possible to get similar concurrency from ordinary blocking code, which is easier to write, test, debug and profile. For a team that adopted reactive programming only to avoid running out of threads, they are a reason to write new services in the simpler style; for a system that already performs well, rewriting it is rarely worth the cost. For latency-critical systems, our guide to low-latency Java applications covers the techniques that apply there.
A rollout checklist
These are the checks we would run before switching a production service to virtual threads. Each one comes from a failure described above.
- Run Java 25, or at least JDK 24. On Java 21, audit every
synchronizedblock that can block, in your code and in your libraries, or replace it withReentrantLock. - Update the libraries that cache per thread or lock inside. Jackson 2.17 or later, current JDBC drivers, connection pools, HTTP clients and tracing libraries.
- Turn on JFR and watch for
jdk.VirtualThreadPinnedevents in a load test before any production traffic. - Write down every downstream limit. Size connection pools for the database, put a semaphore in front of each rate-limited API, and set a timeout on every call.
- Audit
ThreadLocaluse. Share thread-safe objects, pool expensive ones explicitly, and move request context to scoped values. - Load test with production-sized pools and dependencies, so the new bottleneck appears in the test and not in the first busy hour.
- Learn the new thread dump.
jcmd <pid> Thread.dump_to_file -format=json <file>includes virtual threads, which the classic dump does not show. - Roll out to a canary first, with the property ready to switch off, and compare latency, error rate and connection pool wait time against the rest of the fleet.
How we evaluate virtual threads on a client system
We treat virtual threads the way we treat any performance change: measure first, change one thing, and verify it under load. The first question is what the service is actually waiting for, from a profile and from the connection pool and dependency metrics, because virtual threads only help if requests are queuing for threads while the CPU is idle. If they are, we check the JDK and library versions against the checklist above, add the explicit limits the thread pool used to provide, and run the same load test with the property off and on. The result is a number for throughput, tail latency and pool wait time under each setting, and a decision based on those numbers rather than on a benchmark from somewhere else.
If a Java service of yours is running out of threads, or you want to know whether virtual threads would help it, that is what the first step of our Performance Rescue produces: a measured bottleneck and a verified fix.