performance.service · rig armed

Performance Testing Services

Load, stress and scalability testing that finds your breaking point before launch day does — JMeter, k6 and Gatling rigs, p95 targets per journey, and bottleneck guidance your engineers can act on the same week.

▶ Find the breaking point
p95 / p99 targetsbreaking-point analysisscripts owned by youfix guidance included

what it is

Your launch is a load test. The only question is who runs it first.

Performance testing measures how your system behaves under real traffic — and where it breaks: load testing at expected volume, stress testing beyond it to find the ceiling, and endurance runs that surface the leaks only hours of sustained pressure reveal. The output isn't a graph for a slide — it's p95 targets per journey, a measured breaking point, and a ranked list of bottlenecks with fixes.

p95 · p99latency targets per user journey — numbers, not vibes
the ceilingyour measured breaking point, found in staging — not on launch day
ranked fixesevery bottleneck named, owned, and paired with guidance

is this you?

Check what's true. The page will be honest back.

Tick anything that sounds familiar — we'll tell you honestly whether a load test is your next move.

how we work

Four phases. An artifact at the end of each.

PHASE · DAYS 1–3

Model

We profile your real traffic — journeys, ratios, peaks, geography — and define the targets that matter: expected load, surge load, and the p95 each journey must hold.

→ artifact: workload model + target table your team signs off
PHASE · WEEK 1

Rig

JMeter/k6 scripts built per journey with realistic data at scale, parameterized and versioned in your repo — plus a staging environment check for production parity.

→ artifact: re-runnable test rig, owned by you
PHASE · WEEK 2

Ramp

Baseline at expected load, stress past it to the breaking point, soak for endurance. Live dashboards while it runs — your engineers can watch the saturation happen.

→ artifact: baseline report — p95/p99 per journey + measured breaking point
PHASE · WEEK 2–3

Harden

Bottlenecks ranked by impact with concrete guidance — pool sizes, query plans, cache strategy, autoscaling thresholds. Fix, then we re-run and verify the new ceiling.

→ artifact: ranked fix list + verified re-test of the new headroom

see it run

Find the breaking point. Live.

A 15-second simulation of a stress ramp to 10,000 users — including the moment that justifies the entire engagement: the ceiling.

qacraft@perf — k6 stress rampIDLE
virtual users0
p95 latency
throughput
error rate0.0%
▶ press run — ramp to 10,000 users, one ceiling
simulation · your real baseline reports exactly like this

the rig

Open-source rigs. Production-grade discipline.

JMJMeter k6Grafana k6 GtGatling GfGrafana PrPrometheus AWAWS

Scripts live in your repo, dashboards in your Grafana — so the baseline isn't a one-time PDF, it's a capability your team re-runs before every launch, forever.

deliverables

Numbers you can plan launches against.

Workload model

Your real traffic, modeled: journey mix, peak ratios, growth scenarios — signed off before anything runs.

Re-runnable test rig

JMeter/k6 scripts in your repo, parameterized, documented — a permanent capability, not a one-off.

Baseline report

p95/p99 per journey at expected and surge load — the numbers your SLAs should be written from.

Breaking-point analysis

The measured ceiling, what saturates first, and how it fails — graceful or cliff-edge.

Ranked bottleneck fixes

Pool sizes, query plans, cache strategy, scaling thresholds — concrete, owned, prioritized by impact.

Verified re-test

After your fixes, we re-ramp and certify the new headroom — in writing.

proof

Ceilings found in staging, not in headlines.

SAMPLE DATA — replace with verified client story
3,200 → 11,500 usersverified ceiling before seasonal peak · Ecommerce, US

A pre-Black-Friday stress ramp found connection-pool saturation at 3,200 concurrent users — far below the projected peak. Two ranked fixes later, the re-test certified 11,500 users of headroom and the peak weekend passed without a single slowdown.

"The QACraft team continued to improve in both functional QA and processes. I'd highly recommend them."
Thomas Bennett · United States · ongoing quality engagement

pricing logic

Fixed price for the baseline. A program for the fast-moving.

fixed-price

Load-Test Baseline

Scoped against journeys and target load — one quote, 2–3 weeks, rig + report + verified re-test.

the numbers your launch plan is missing
continuous

Performance Program

Baselines twice yearly inside a quality pod, regression-checked in CI — headroom as a habit, not an event.

included in Growth Pod engagements

The re-test is included. After your fixes we re-ramp and certify the new ceiling at no additional cost — the engagement ends with verified headroom, not a list of problems.

straight answers

Asked on every performance call. Answered here.

How much does performance testing cost?

A load-test baseline is fixed-price against your journey count and target load — typically 2–3 weeks with one quote, verified re-test included. Continuous programs fold twice-yearly baselines into a quality pod. A 30-minute scoping call gets you one exact number.

Load vs stress vs endurance — what do we actually need?

Load testing answers "are we fine at expected traffic"; stress finds the ceiling; endurance finds what breaks over hours. Pre-launch, you need all three once — the workload model in Phase 1 tells us the proportions, and the rig runs them all.

What metrics do you report?

p95/p99 latency per journey, throughput (RPS), error rate under load, and the breaking point — plus what saturates first and whether failure is graceful or cliff-edge. SLA-ready numbers, not averages that hide the pain.

Staging or production?

Production-like staging with realistically seeded data is the default. Where staging parity is impossible, carefully rate-limited production verification runs under signed rules of engagement with an emergency stop on both sides.

Our staging data is tiny — does the test still mean anything?

Not without fixing that — which is part of Phase 2: we seed data at realistic scale (rows, users, catalog size) before any ramp, because a 12-row database lies to everyone. The seeding scripts are part of the rig you keep.

Can our engineers re-run the tests later?

Yes — that's the design goal. Scripts are parameterized, documented and live in your repo with your dashboards; re-running before any launch is a command, not a contract.

Know your ceiling before launch day finds it.

Scope your load test in 60 seconds — or bring your launch date to a 30-minute call and leave with a workload model sketch and one number.

Book a Scoping Call