Performance Testing — Load, Stress & Benchmarks
Performance bugs don't show up in functional tests. They show up when 10,000 users arrive at once. Here's how to test for load, stress, and durability before your users do.

Functional testing asks: does it work? Performance testing asks: does it work under pressure? Both matter. Only one is usually skipped.
Performance bugs are the silent killers of good products. The app works beautifully in dev with 10 users. Then marketing drives 10,000 users, and the site crashes.
The three types of performance testing
Load testing — Simulates expected peak traffic. Confirms the system handles it without degradation.
Stress testing — Pushes beyond expected peak to find the breaking point.
Soak testing — Runs moderate traffic for hours/days to find leaks and slow failures.
Metrics that actually matter
Throughput — Requests per second
Response time (P50, P95, P99) — Not averages
Error rate — % of requests that fail under load
Concurrency — Simultaneous users handled
Resource usage — CPU, memory, DB connections at load
Recovery time — Return to baseline after traffic drops
Skip "average response time" — it hides the P99 experience.
Tooling
k6 — Modern, developer-friendly. Our recommendation for most teams.
JMeter — Mature, feature-rich. Steeper learning curve.
Lighthouse — Frontend performance (different category).
Locust — Python-based.
Gatling — Scala-based.
How to run performance tests
- Define your target: expected peak traffic + response time SLO
- Build realistic scenarios (full user journeys, not single endpoints)
- Baseline at low traffic
- Load test at expected peak
- Stress test beyond peak
- Soak test for durability
- Monitor everything
What we typically find
- Database queries that scale linearly
- Missing indexes causing slow queries
- Connection pool exhaustion
- Memory leaks that surface after hours
- Third-party APIs that throttle unexpectedly
- N+1 query patterns in ORMs
None of these show up in functional testing.
When to run performance tests
Before major launches — Non-negotiable
- After infrastructure changes
- After scaling milestones
- Continuously in production (synthetic monitoring)
Common mistakes
- Testing single endpoints instead of full journeys
- Skipping soak testing
- Not testing from realistic locations
- Ignoring third-party API limits
- Testing in dev instead of production-like environments
Key takeaways
- Load, stress, and soak testing are all required
- Track P95 and P99, not averages
- k6 is the best starting point for most teams
- Test full user journeys, not single endpoints
- Run before launches, after infra changes, and continuously
Further reading
About the author
Senior QA Engineer →Senior QA Engineer · Quality Assurance Labs



