A checkout service misses its p99 target, and the team spends a sprint replacing streams with loops. A wall-clock profile, taken afterwards, shows each request waiting on SQL queries issued one at a time.
To improve Java performance, measure first, then pull the levers in order of payoff: the JDK version, data access, caching, memory, concurrency, and last the code a profiler points at.
Define the performance goal before you change any code
"Faster" is not a goal. Pick the metric that costs you users or money and put a number next to it:
- Throughput: requests or records per second on a fixed hardware budget.
- Latency percentiles: p50, p99 and p99.9 at a stated load. An average hides the slow requests that users notice and that time out upstream.
- Startup and warmup: time to the first request and time to full speed, which decide how quickly an autoscaling group absorbs a spike.
- Memory footprint: heap plus native memory per instance, which decides how many instances fit on a node.
- Cost per request: the cloud bill divided by traffic, the number finance sees.
A usable goal reads like "p99 under 150 ms at 2,000 requests a second on four vCPUs". It also tells you when to stop.
The metrics pull against each other. A concurrent collector cuts pauses and spends more CPU and memory; a bigger heap means fewer collections and fewer instances per node. Choosing one metric first decides which trade you accept.
The measurement loop: JFR, async-profiler and JMH
The method is four steps, repeated: measure the system under realistic load, find the one place where the time or the memory goes, change that one thing, and verify under the same load. Each tool below answers a different question in that loop.
Profile the running system with JFR and async-profiler
JDK Flight Recorder ships with the JDK and is designed to stay on in production. One recording holds CPU samples, allocation samples, GC pauses, lock contention, socket and file I/O and thread states, and the jfr tool summarizes it without a GUI.
# record from startup, or attach to a running JVM
java -XX:StartFlightRecording=duration=120s,settings=profile,filename=app.jfr -jar app.jar
jcmd <pid> JFR.start duration=120s settings=profile filename=app.jfr
# summaries on the command line
jfr view hot-methods app.jfr
jfr view allocation-by-site app.jfr
jfr view contention-by-site app.jfr
jfr view gc-pauses app.jfr
Java 25 added two event types worth knowing. jdk.CPUTimeSample samples threads by CPU time rather than elapsed time; it is experimental and Linux-only (JEP 509). jdk.MethodTiming and jdk.MethodTrace record exact call counts, timings and stack traces for the methods a filter selects, by bytecode instrumentation, with no code change (JEP 520).
async-profiler samples Java, native and kernel frames together, GC and JIT compiler threads included, and does not suffer from safepoint bias. It writes flame graphs, and its modes map onto the questions you will ask:
asprof -e cpu -d 60 -f cpu.html <pid> # where CPU time goes
asprof -e wall -t -d 60 -f wall.html <pid> # where threads wait, per thread
asprof -e alloc -d 60 -f alloc.html <pid> # which code allocates
asprof -e lock -d 60 -f lock.html <pid> # which locks are contended
The wall-clock mode is the one teams skip, and the one that finds database and remote-call waits. On an I/O-bound service a CPU profile shows only a fraction of each request.
Which profiler to buy or install is a separate question, answered in our guide to Java profilers. The failure patterns profiles usually reveal (leaks, GC pressure, contention, slow queries) are catalogued in common Java performance problems.
Benchmark a single method with JMH
When the profile points at one method, JMH compares alternatives in isolation. It handles warmup, forking and the dead code elimination that makes hand-written timing loops report nonsense. This skeleton benchmarks the allocation fix shown later in this post:
@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.MICROSECONDS)
@Warmup(iterations = 5, time = 1)
@Measurement(iterations = 5, time = 1)
@Fork(3)
@State(Scope.Benchmark)
public class LedgerTotalBenchmark {
List<String> lines;
@Setup
public void setup() {
lines = Fixtures.ledgerLines(10_000); // production-shaped input
}
@Benchmark
public long before() {
return LedgerParser.totalBefore(lines); // return the result, or the JIT may drop the work
}
@Benchmark
public long after() {
return LedgerParser.totalAfter(lines);
}
}
java -jar target/benchmarks.jar LedgerTotalBenchmark -prof gc
Three rules keep the numbers honest. Return the result or pass it to a Blackhole, so the JIT cannot delete the work. Keep inputs in @State fields, so it cannot fold them into constants. Run several forks, because JIT decisions differ from one JVM to the next. The -prof gc profiler adds gc.alloc.rate.norm, bytes allocated per operation, which turns "allocates less" into a number.
Treat a JMH result as a hypothesis about the system until the load test confirms it. A method made three times faster saves nothing if it accounted for 2% of the request.
Verify under production-like load
Run the same load test before and after each change, with production-shaped data (payload sizes, key cardinality, cache hit rates), and compare percentiles. A load generator that waits for each response before sending the next request hides queueing and under-reports the tail; our post on low-latency Java applications covers measuring the tail correctly.
Correctness comes before speed in that comparison. On a schedule calculation system for a global flight information company, our team first proved with functional tests that a proof of concept produced the same daily output as production, and only then ran both on identical input. The new design, an open-source high-performance Java framework in place of DB2 in-database processing and proprietary Java, cut calculation time by 65% (case study).
Keep a continuous JFR recording running in production, for example -XX:StartFlightRecording=disk=true,maxage=6h,settings=default, so the next regression arrives with its evidence attached.
Upgrade the JDK: what Java 21 to 27 gives you without code changes
Moving to a current JDK is usually the cheapest lever, because the runtime work in each release applies to unchanged application code. Some of it is on by default; some needs one flag.
The timeline of performance JEPs from Java 8 to 27 is in our post on how Java enables high performance; the three changes below matter most when you upgrade.
Compact object headers: smaller objects, less GC
Compact object headers shrink every object header from 96 to 64 bits on 64-bit platforms. They arrived as an experiment in JDK 24, became a product feature in Java 25 behind -XX:+UseCompactObjectHeaders (JEP 519) and are the default in JDK 27 (JEP 534).
The JEP reports SPECjbb2015 using 22% less heap and 8% less CPU time in one setting and running 15% fewer collections in another, with G1 and Parallel; a highly parallel JSON parser benchmark ran in 10% less time. Before it became a product feature, the layout was tested at Amazon by hundreds of production services, most of them on backports to JDK 17 and 21. Applications holding many small objects gain the most.
Garbage collectors that improved underneath you
Generational ZGC (JEP 439, Java 21) collects young objects separately, which lets ZGC keep up with higher allocation rates on smaller heaps. It became ZGC's default mode in JDK 23 (JEP 474) and its only mode in JDK 24, where the non-generational mode was removed (JEP 490). Shenandoah's generational mode became a product feature in Java 25 (JEP 521).
G1, the collector most services run, got cheaper in JDK 26. JEP 522 reworked how application threads and GC threads share the card table: the JEP reports throughput gains of 5 to 15% in applications that heavily modify object-reference fields, and up to 5% elsewhere, because the x64 write barrier shrank from around 50 instructions to 12.
JDK 27 makes G1 the default in all environments (JEP 523). Until then, the JVM silently chose the Serial collector on machines with a single CPU or less than 1792 MB of memory, which describes many small containers.
Faster startup and warmup with the AOT cache
Project Leyden's AOT cache stores the work a JVM repeats at every start. JDK 24 caches loaded and linked classes (JEP 483); Java 25 adds method profiles, so the JIT compiles hot code immediately (JEP 515), and creates the cache in one step (JEP 514); JDK 26 lets the cache work with any collector, ZGC included (JEP 516).
# Java 25+: run a representative workload, then stop the app; the cache is written on exit (JEP 514)
java -XX:AOTCacheOutput=app.aot -jar app.jar
# every production start
java -XX:AOTCache=app.aot -jar app.jar
The cache is only as good as the training run: JEP 515 measured a short program warming up in 73 ms instead of 90 ms once profiles were cached, and the profiles reflect whatever the training run exercised. Build the cache in the same pipeline step that builds the image, with the same JDK and class path.
Treat the upgrade as a tested change
Java 25 is the current long-term support release and the natural target; JDK 26 and 27 are short-lived feature releases. Upgrade with the same load test you use for any other change, and check for removed flags (-XX:+ZGenerational is obsolete since JDK 24). Support dates for every version are in our Java end-of-life calendar, and the JVM mechanisms behind these gains are explained in how Java enables high performance.
Fix the database and I/O layer: N+1 queries, pools and batching
In the business systems we are asked to speed up, the database round trip is the most common bottleneck, and it rarely shows in a CPU profile. Count queries per request before tuning anything else.
Before and after: removing an N+1 query
A report lists the day's orders with the customer name and the number of lines. With lazy associations, the obvious code issues one query for the orders and then up to two more per order:
// Before: 1 + 2N round trips for N orders
@Transactional(readOnly = true)
public List<OrderSummary> summaries(LocalDate day) {
return orders.findByDay(day).stream()
.map(o -> new OrderSummary(
o.getId(),
o.getCustomer().getName(), // lazy to-one: a query per customer
(long) o.getLines().size())) // lazy collection: a query per order
.toList();
}
For 500 orders that is up to 1,001 round trips; at 1 ms each, a second of waiting before any business logic runs. A projection asks the database for exactly the three values, in one query:
// After: one round trip
public record OrderSummary(Long id, String customer, Long lineCount) {}
public interface OrderRepository extends JpaRepository<Order, Long> {
@Query("""
select new com.example.orders.OrderSummary(o.id, c.name, count(l))
from Order o
join o.customer c
left join o.lines l
where o.day = :day
group by o.id, c.name
""")
List<OrderSummary> summariesFor(@Param("day") LocalDate day);
}
When the code needs the entities themselves, join fetch or @EntityGraph loads a to-one association in the same query, and hibernate.default_batch_fetch_size (or @BatchSize) loads a lazy collection for many parents with one IN query. Fetch-joining two collections at once multiplies rows into a Cartesian product: join one, batch the other.
To catch the next one, set hibernate.generate_statistics=true in integration tests and assert on the statement count of the endpoints that matter. In a wall-clock flame graph an N+1 shows up as a wide tower of JDBC frames.
Size the connection pool for the database, not for the threads
The HikariCP wiki starts from the PostgreSQL project's formula, connections = (core_count * 2) + effective_spindle_count, counted on the database server, and describes an Oracle Real-World Performance demonstration in which cutting the pool from 2,048 to 96 connections took response times from about 100 ms to about 2 ms (HikariCP, About Pool Sizing). A pool larger than the database can serve moves the queue into the database, where it costs context switches and lock contention.
This matters more after a move to virtual threads, which remove the thread pool that used to limit concurrent queries; our guide to Java virtual threads covers why the bottleneck moves to the connection pool.
Batch writes, page by key, and read the plan
- Set
hibernate.jdbc.batch_sizeto a value between 10 and 50, as the Hibernate user guide recommends, and flush and clear the session per batch. Hibernate disables insert batching for entities withIDENTITYids, so bulk-inserted tables need a sequence with a pooled optimizer. - Replace
OFFSETpagination on large tables with keyset pagination (where id > :last order by id limit 50). An offset makes the database read and discard every skipped row; a key condition reads only the page. - Take the top queries by total time from
pg_stat_statementsor your database's equivalent, runEXPLAIN ANALYZEon each, and add the index that matches the filter and the sort order. - Read paths that only display data should load projections, not managed entities: no dirty checking, no lazy proxies, fewer columns.
- Give every remote call a connect and a read timeout, and reuse HTTP connections through a pooled client.
Caching: where to cache, and how to keep it correct
A cache turns repeated work into a memory lookup and creates a second copy of the data that can be wrong. Decide three things for each cache: where it lives, how stale it may be, and what invalidates it.
An in-process cache such as Caffeine answers in nanoseconds with no network hop, but every instance holds its own copy and warms up separately. A distributed cache or data grid (Redis, Hazelcast) is shared across instances at the cost of a network round trip; a near cache puts a small local copy in front of it. Good candidates are reference data, results of expensive computations and responses from slow or metered APIs; per-user data with frequent writes is a poor one.
LoadingCache<String, FxRate> rates = Caffeine.newBuilder()
.maximumSize(50_000)
.expireAfterWrite(Duration.ofMinutes(10)) // hard limit on staleness
.refreshAfterWrite(Duration.ofMinutes(1)) // reload in the background, serve the old value
.recordStats() // hit rate as a metric
.build(currency -> rateClient.fetch(currency));
A loading cache computes each missing key once, however many threads ask for it, which protects the source from a stampede after expiry; refreshAfterWrite reloads entries asynchronously while readers keep getting the old value (Caffeine wiki). Invalidation follows the data: a TTL where bounded staleness is acceptable, an event from the write path (a message or change data capture) where readers must see writes quickly.
At the scale of a whole system, caching becomes architecture. A UK hotel bookings aggregator ran availability search on dozens of SQL Server machines, about 200 million queries a day, 99% of them reads, and each added machine raised cost by about 2% and throughput by about 0.1%. We moved availability search into an in-memory data grid holding the complete data needed to answer a query, filled by a custom fast loader, with a message queue pipeline carrying intraday updates; SQL Server stayed the system of record.
The engine now serves 300 million queries a day, with response time down 60% and infrastructure cost down 80% (case study). The update pipeline is the invalidation design; without it the grid would be a fast way to serve wrong prices. Our explainer on what Hazelcast is used for shows how a grid partitions and replicates data.
Memory and garbage collection: heap sizing, collector choice, allocation
Container-aware heap settings
On Linux the JVM reads its container's memory and CPU limits, but the default heap is only 25% of the container's memory (MaxRAMPercentage, java command reference). A 4 GB container gets a 1 GB heap and leaves most of what you pay for unused.
-XX:MaxRAMPercentage=75 # heap as a share of the container limit (default 25)
-XX:InitialRAMPercentage=75 # commit it at startup on long-running services
-XX:+AlwaysPreTouch # touch heap pages at startup, not during the first requests
-XX:ActiveProcessorCount=4 # override the CPU count derived from cgroup quotas
-Xlog:gc*:file=gc.log:time,uptime:filecount=5,filesize=20m
-XX:+HeapDumpOnOutOfMemoryError
-XX:NativeMemoryTracking=summary # then: jcmd <pid> VM.native_memory summary
75% is our usual starting point, not a rule. Metaspace, thread stacks, the code cache, direct buffers and the GC's own structures live outside the heap; Native Memory Tracking shows how much they take, and the container limit has to cover both. On JDK 26 and earlier, check which collector ergonomics picked (java -XX:+PrintCommandLineFlags -version inside the container): a pod with one CPU gets Serial.
Choosing a garbage collector
G1 balances throughput and pause time and suits most services; size the heap before touching pause goals. Parallel GC gives the highest throughput for batch work that tolerates pauses. Generational ZGC keeps pauses short regardless of heap size, paying with CPU for concurrent work and extra memory. Netflix switched its default from G1 to generational ZGC on JDK 21 and reported that, for a given CPU target, ZGC improved both average and p99 latency with equal or better CPU utilization in its services (Netflix Technology Blog, March 2024). Shenandoah targets the same pause profile and ships in most OpenJDK builds except Oracle's.
Find the allocation, then cut it
Allocation rate drives GC frequency with any collector. asprof -e alloc or jfr view allocation-by-site names the code that allocates; the JMH -prof gc output confirms a fix. A typical find is a parser that compiles a regular expression and boxes a number for every line it reads:
// Before: String.matches compiles the regex on every call; every amount is boxed into a list
static long totalBefore(List<String> lines) {
List<Long> amounts = new ArrayList<>();
for (String line : lines) {
if (line.matches("\\d{4}-\\d{2}-\\d{2};.*")) {
amounts.add(Long.parseLong(line.substring(line.lastIndexOf(';') + 1)));
}
}
return amounts.stream().mapToLong(Long::longValue).sum();
}
// After: one compiled Pattern, a primitive accumulator, no substring
private static final Pattern DATED = Pattern.compile("\\d{4}-\\d{2}-\\d{2};");
static long totalAfter(List<String> lines) {
long total = 0;
for (String line : lines) {
if (DATED.matcher(line).lookingAt()) {
total += Long.parseLong(line, line.lastIndexOf(';') + 1, line.length(), 10);
}
}
return total;
}
The second version still creates a Matcher per line, and escape analysis may or may not remove it; gc.alloc.rate.norm tells you which.
Concurrency: virtual threads, lock contention and parallel streams
Virtual threads for I/O-bound services
Virtual threads (JEP 444) let a service run one thread per request while it waits on databases and other services, and since JDK 24 synchronized no longer pins them to a carrier (JEP 491). They raise throughput only for work that waits; CPU-bound code gains nothing. Rollout patterns and the production failures to avoid are in our Java virtual threads guide.
Finding and reducing lock contention
Contention looks like idle CPUs and slow requests at the same time. jfr view contention-by-site and asprof -e lock show which monitor or lock threads wait on, and from where. The fixes, roughly in order of effort:
- Move I/O and slow calls out of the critical section, so the lock guards memory only.
- Replace a hot
AtomicLongcounter withLongAdder, which spreads updates across cells and trades memory for throughput under contention. - Swap a synchronized map for
ConcurrentHashMap, and publish read-mostly configuration as an immutable snapshot behind a volatile reference. - Partition shared state by key so that threads working on different keys never meet.
When parallel streams make things slower
Every parallelStream() in the JVM shares one common ForkJoinPool, sized from the number of available processors. A blocking call inside a parallel stream ties up those workers for every other parallel stream in the process. In a web server, request threads already keep the cores busy, so splitting each request's work across the same cores adds coordination without adding capacity. Small inputs pay the split and merge overhead and gain nothing.
Parallel streams pay off for CPU-bound work on large in-memory arrays, typically in batch jobs, and only after a JMH run on realistic sizes confirms it.
Data structures and serialization on the hot path
The data structure decides memory layout, and memory layout decides cache misses. A List<Long> is an array of references to separate objects scattered over the heap; a long[] is one contiguous block. For large numeric collections, primitive arrays or primitive collection libraries (fastutil, Eclipse Collections) cut both memory and GC work.
ArrayListandArrayDequebeatLinkedListfor almost every access pattern, because a node per element defeats the CPU cache.- Presize maps and lists whose size you know:
HashMap.newHashMap(n)(Java 19 and later) allocates fornentries and skips the resizes. JFR's method tracing in Java 25 can show where resizes come from. EnumMapandEnumSetreplace hash lookups on enum keys with array indexing.
In the services we profile, serialization is often among the largest CPU costs left once the I/O is fixed. Create Jackson's ObjectMapper once and share it: it is thread-safe once configured and expensive to build, and a mapper per request also throws away its internal caches. Stream large payloads with the streaming API instead of building a tree. For high-volume traffic between your own services, a binary format such as Protocol Buffers is smaller and cheaper to parse than JSON. Built-in Java serialization belongs nowhere near external input, for security reasons covered in our secure Java guide.
JIT-friendly code, where the profile says it matters
Most code-level tips (short methods, ordering if-else branches, final on locals) change nothing measurable: the JIT inlines, unrolls and eliminates branches from runtime profiles. A few patterns do matter in code that a profile shows is hot.
Warmup
HotSpot interprets a method first, compiles it with C1, then recompiles it with C2 using the profile it collected, so a fresh instance is slow for its first seconds or minutes. Send synthetic traffic before the readiness probe passes, or cache the profiles with the AOT cache described above.
Megamorphic call sites
C2 inlines an interface or virtual call when the profile shows one or two receiver types at that call site. A site that sees more types becomes megamorphic: it stays a virtual call, and the optimizations that inlining enables are lost. Aleksey Shipilev's analysis of Java method dispatch measures the effect. In a hot loop over mixed types, splitting the work by type or keeping the hot path to one implementation restores inlining.
Huge methods
HotSpot does not JIT-compile methods larger than 8,000 bytes of bytecode (DontCompileHugeMethods, on by default; HotSpot source). Hand-written code rarely gets there; generated parsers, mappers and template code do, and then they run interpreted. That is the grain of truth in "avoid long methods". -XX:+PrintCompilation shows what was compiled, and -XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining shows what was inlined and why not.
Logging costs: what to switch off on hot paths
Logging is I/O on the request path, and it costs more than the disk write:
- Use parameterized messages (
log.debug("Order {} priced at {}", id, price)) so disabled levels skip string building. Arguments are still evaluated, so pass an expensive one as a supplier lambda, which Log4j 2 and the SLF4J 2 fluent API both accept. - Drop caller location (
%L,%M,%C,%F) from production patterns. Log4j takes a snapshot of the stack and walks it for every event to find the caller, which its own documentation calls expensive (Log4j performance guide). - Write through an asynchronous logger or appender with a bounded queue, and decide what happens when the queue is full: block the request or drop debug events.
- Count volume too. Log ingestion is billed per gigabyte on most hosted platforms, so every line logged per request is also a line on the cloud bill.
Cutting cloud cost: right-sizing and fast startup for autoscaling
Cost per request falls when each instance does more work at the latency goal, and when instances start fast enough that you do not need to run spares.
Right-size from the load test, not from dashboards: find the throughput one instance sustains at the p99 target, divide peak traffic by it, and add headroom for a zone failure. CPU limits deserve a second look. A JVM throttled by a cgroup quota pauses mid-request, and the pause lands in the tail; give latency-sensitive services a CPU allowance that covers GC and JIT threads, or cap what the JVM sees with -XX:ActiveProcessorCount. Memory headroom from compact object headers and a correctly sized heap lets more instances share a node.
For autoscaling, startup time is capacity you cannot use yet. The AOT cache shortens startup and warmup on any standard JDK from 24 on. CRaC (Coordinated Restore at Checkpoint) goes further: it checkpoints a warmed-up JVM and restores it in place of a start. It needs a CRaC-enabled JDK on Linux, offered by vendors such as Azul and BellSoft, plus the org.crac library; Spring Framework can take the checkpoint automatically with -Dspring.context.checkpoint=onRefresh (Spring Framework reference). The checkpoint files are an image of process memory and may contain secrets, so they need the same protection as credentials, and open connections must be closed before the checkpoint and reopened after restore.
Licenses are part of the bill too. Part of the flight schedule work above was moving calculation logic out of licensed in-database processing, where adding capacity also added license fees.
If a Java service misses its latency target, or costs more per request than the business can justify, our Performance Rescue engagement measures it under load, finds the bottleneck and verifies the fix against the same test. When the profile shows a configuration change your own team can make in a day, we say so.