performance.service · rig armed
Performance Testing Services
Load, stress and scalability testing that finds your breaking point before launch day does — JMeter, k6 and Gatling rigs, p95 targets per journey, and bottleneck guidance your engineers can act on the same week.
what it is
Your launch is a load test. The only question is who runs it first.
Performance testing measures how your system behaves under real traffic — and where it breaks: load testing at expected volume, stress testing beyond it to find the ceiling, and endurance runs that surface the leaks only hours of sustained pressure reveal. The output isn't a graph for a slide — it's p95 targets per journey, a measured breaking point, and a ranked list of bottlenecks with fixes.
is this you?
Check what's true. The page will be honest back.
Tick anything that sounds familiar — we'll tell you honestly whether a load test is your next move.
how we work
Four phases. An artifact at the end of each.
Model
We profile your real traffic — journeys, ratios, peaks, geography — and define the targets that matter: expected load, surge load, and the p95 each journey must hold.
→ artifact: workload model + target table your team signs offRig
JMeter/k6 scripts built per journey with realistic data at scale, parameterized and versioned in your repo — plus a staging environment check for production parity.
→ artifact: re-runnable test rig, owned by youRamp
Baseline at expected load, stress past it to the breaking point, soak for endurance. Live dashboards while it runs — your engineers can watch the saturation happen.
→ artifact: baseline report — p95/p99 per journey + measured breaking pointHarden
Bottlenecks ranked by impact with concrete guidance — pool sizes, query plans, cache strategy, autoscaling thresholds. Fix, then we re-run and verify the new ceiling.
→ artifact: ranked fix list + verified re-test of the new headroomsee it run
Find the breaking point. Live.
A 15-second simulation of a stress ramp to 10,000 users — including the moment that justifies the entire engagement: the ceiling.
the rig
Open-source rigs. Production-grade discipline.
Scripts live in your repo, dashboards in your Grafana — so the baseline isn't a one-time PDF, it's a capability your team re-runs before every launch, forever.
deliverables
Numbers you can plan launches against.
Your real traffic, modeled: journey mix, peak ratios, growth scenarios — signed off before anything runs.
JMeter/k6 scripts in your repo, parameterized, documented — a permanent capability, not a one-off.
p95/p99 per journey at expected and surge load — the numbers your SLAs should be written from.
The measured ceiling, what saturates first, and how it fails — graceful or cliff-edge.
Pool sizes, query plans, cache strategy, scaling thresholds — concrete, owned, prioritized by impact.
After your fixes, we re-ramp and certify the new headroom — in writing.
proof
Ceilings found in staging, not in headlines.
A pre-Black-Friday stress ramp found connection-pool saturation at 3,200 concurrent users — far below the projected peak. Two ranked fixes later, the re-test certified 11,500 users of headroom and the peak weekend passed without a single slowdown.
"The QACraft team continued to improve in both functional QA and processes. I'd highly recommend them."
pricing logic
Fixed price for the baseline. A program for the fast-moving.
Load-Test Baseline
Scoped against journeys and target load — one quote, 2–3 weeks, rig + report + verified re-test.
the numbers your launch plan is missingPerformance Program
Baselines twice yearly inside a quality pod, regression-checked in CI — headroom as a habit, not an event.
included in Growth Pod engagementsstraight answers
Asked on every performance call. Answered here.
How much does performance testing cost?
A load-test baseline is fixed-price against your journey count and target load — typically 2–3 weeks with one quote, verified re-test included. Continuous programs fold twice-yearly baselines into a quality pod. A 30-minute scoping call gets you one exact number.
Load vs stress vs endurance — what do we actually need?
Load testing answers "are we fine at expected traffic"; stress finds the ceiling; endurance finds what breaks over hours. Pre-launch, you need all three once — the workload model in Phase 1 tells us the proportions, and the rig runs them all.
What metrics do you report?
p95/p99 latency per journey, throughput (RPS), error rate under load, and the breaking point — plus what saturates first and whether failure is graceful or cliff-edge. SLA-ready numbers, not averages that hide the pain.
Staging or production?
Production-like staging with realistically seeded data is the default. Where staging parity is impossible, carefully rate-limited production verification runs under signed rules of engagement with an emergency stop on both sides.
Our staging data is tiny — does the test still mean anything?
Not without fixing that — which is part of Phase 2: we seed data at realistic scale (rows, users, catalog size) before any ramp, because a 12-row database lies to everyone. The seeding scripts are part of the rig you keep.
Can our engineers re-run the tests later?
Yes — that's the design goal. Scripts are parameterized, documented and live in your repo with your dashboards; re-running before any launch is a command, not a contract.
Know your ceiling before launch day finds it.
Scope your load test in 60 seconds — or bring your launch date to a 30-minute call and leave with a workload model sketch and one number.
