Performance engineering services for Java systems under load
Find the causes of slow responses, limited throughput and rising infrastructure costs. Start with a measurement phase, then agree the changes and the results to test.
Measured results from platform engineering projects
Our work includes production platform development, architecture improvements and a framework proof of concept tested on production data.
- 60%
- lower response time on a hotel metasearch engine serving 300 million queries a day, with infrastructure costs reduced by 80%.Hotel metasearch case study
- 65%
- less flight schedule calculation time in a framework proof of concept, benchmarked against production using identical input data.Flight schedule case study
- 1B / hour
- financial transactions processed by FastPost, the accounting platform we developed for Legerity.FastPost case study
The features implemented have received overwhelmingly positive feedback from end-users. Stratoflow has an incredible technical expertise and a high degree of flexibility when it comes to changing project requirements.
Adam HillChief Technology Officer, Legerity- First iteration
- Measurement, usually 1 to 2 weeks, within the engagement.
- Implementation
- Usually 4 to 8 weeks after measurement, subject to scope.
- Commercial model
- Time and materials with a ballpark estimate, against agreed SLO targets.
- Acceptance
- Results tested against the workload and criteria in the proposal.
Follow the evidence to the bottleneck
Data access and caching
Measure query cost, database round trips and cache behaviour. Consider an in-memory data grid where the workload and consistency requirements justify it.
JVM and resource use
Profile allocation, garbage collection, CPU and memory. Check thread pools, connection pools and runtime settings against observed load.
Concurrency and scaling
Find contention, queues and serial stages that limit throughput. Evaluate partitioning and failure behaviour before adding more nodes.
Streaming and integrations
Trace delays across Kafka pipelines, external calls and integration flows. Establish where backpressure, repeated work or slow dependencies affect the user-facing result.
The technology behind the throughput
Hazelcast and Apache Ignite
Distributed caching and compute next to the data, when a database round trip is the bottleneck; Ignite when the workload needs SQL over the grid. Most of the gain, or the trouble, is in cluster design: partitioning, backup counts and failure behaviour.
Oracle Coherence
For enterprises already on Oracle: cache topologies and near-cache strategies tuned to the access pattern, and upgrades tested for serialization changes before they reach production.
Apache Kafka
Where the bottleneck sits between services: partitioning, consumer parallelism and back-pressure, with ordering and replay kept intact.
Profilers and load harnesses
Java profilers to find where the time goes, JMeter or a custom harness to reproduce the load, and the same harness run before and after every change, then handed to you.
Make each improvement testable
Establish the baseline
Agree the workload, test environment and metrics that matter: response-time percentiles, throughput, resource use or cost. Profile and trace the application using representative data.
Agree the changes
Rank the findings and propose a feasible scope. Define performance objectives, acceptance conditions, dependencies and the implementation ballpark estimate before starting changes.
Implement and compare
Make changes in weekly iterations and rerun the agreed tests. Check functional correctness alongside performance and review tradeoffs with your team every week.
Hand over the evidence
Deliver the code changes, test harness, results and operating guidance. Record remaining bottlenecks and plan deployment and monitoring around the application.
Separate diagnosis from the implementation decision
The measurement phase has its own fixed fee. You receive the findings and recommended next steps, including cases where the proposed target requires a larger architectural change. Implementation is a separate decision.
The implementation proposal defines the agreed scope, ballpark estimate and schedule, and the phase is billed time and materials. We agree how changes to the workload, environment or scope affect the acceptance criteria before additional work starts.
You receive the measurements and implementation findings even when a target is not met. A demonstrated result in a test environment is accompanied by a rollout plan and production monitoring requirements.
Before we start
Do you need access to production?
Not always. Existing telemetry and a representative test environment may be enough to start. We agree access, data handling and any production measurements with your team; load tests need an approved environment and operating limits.
Can you guarantee a particular improvement?
A target needs a baseline and a feasible implementation plan. Measurement establishes what is likely to be achievable. The proposal records the target, test conditions and commercial consequences if it is not met.
Will a Java upgrade or data grid solve the problem?
It depends on the measured constraint. A query fix, allocation change or simpler cache may be sufficient. We evaluate runtime upgrades and distributed technologies against the workload, support requirements and operational cost.
How long does the work take?
Plan for a 1 to 2 week measurement phase followed by a scoped implementation, usually 4 to 8 weeks. Access, test data, external dependencies and release windows affect the schedule.
What happens after delivery?
Your team receives the changes, test harness and operational notes. If ongoing maintenance or incident response is needed, we can agree a separate managed support engagement.
Is this Java performance tuning or something bigger?
It often starts as Java performance tuning: garbage collection, thread and connection pools, query plans. Tuning rarely finishes the job, because high performance Java is mostly a matter of design: where the data lives, how work is partitioned and what crosses the network. We measure first, so the change you pay for is the one the evidence points to.
Continue improving the application
Start with the performance problem you can observe
Bring the symptoms, current measurements and the business outcome you need. We will discuss a suitable measurement scope.