How many requests per second can one Postgres handle?

Short answer: it is CPU per query times the number of cores. We measured one Postgres on a 2-vCPU cloud VM with a simple indexed read: it served about 6,000 queries per second at 100 % CPU, roughly 0.33 ms of CPU per query.

The number to plan with is lower. With random (Poisson) arrivals, p99 stayed under 7 ms up to about 75 % CPU (4,800 queries/s), doubled by 88 %, and at 100 % requests waited for seconds.

"Can one Postgres take our launch?" has no universal answer, because a query can cost 50 µs or 50 ms of CPU. Blog numbers range from a few hundred to millions of queries per second, and they rarely say what the query was. So here is one measurement we made ourselves, with the full setup, and a method to get your number.

The setup

The results

p99 latency against throughput for one Postgres on 2 vCPU: evenly spaced arrivals stay under 6 ms up to 5,760 queries/s; random arrivals climb from about 3.5 ms at 3,840 to 21 ms at 5,760; saturation near 6,000. 0 2,000 4,000 6,000 queries/s 0 10 20 ms CPU 100 % random (Poisson) arrivals evenly spaced arrivals
p99 of the HTTP request (app + one query) per 60-second step. The last step (6,080 offered) is off the chart: about 1.6–2.0 s.
Measured 2026-09-30 on Google Cloud n2-standard-2. Latencies are of the whole HTTP request and include the app and two network hops. At the last step the offered load exceeded capacity; latency there grows for as long as the overload lasts.
Offered req/sDB CPU (random)p99 randomDB CPU (even)p99 even
6409 %1.4 ms8 %1.1 ms
1,92028 %1.6 ms20 %0.9 ms
3,20049 %2.5 ms35 %1.1 ms
3,84060 %3.5 ms44 %1.3 ms
4,48070 %5.7 ms58 %1.6 ms
4,80075 %6.9 ms65 %2.0 ms
5,12080 %7.3 ms73 %3.3 ms
5,44088 %12.6 ms83 %4.2 ms
5,76094 %21.0 ms91 %6.0 ms
6,08099 %1,966 ms99 %1,572 ms

Three things stand out:

  1. Throughput is linear until the CPU is full. Every step up to 5,760 requests/s was served completely; at 6,080 offered only about 5,950–6,000 got through (a second run on the same machine type: 5,890).
  2. The tail climbs long before 100 %. With random arrivals p99 doubled between 70 % and 88 % CPU and nearly doubled again by 94 %. That is queueing: when requests arrive in clumps, the second and third in a clump wait for the first.
  3. Burstiness matters as much as the average. The same load with evenly spaced arrivals had a p99 2.5–3.5× lower from 60 % CPU on. Real user traffic is closer to random than to a metronome, and a retry storm or a cron job on the hour is worse than random.

How to estimate your own number

  1. Find the CPU cost of your typical query. Run two steady load levels (for example 20 % and 50 % of what you expect) and take (busy cores at level 2 − busy cores at level 1) / (queries/s at level 2 − queries/s at level 1). Our run: about 0.29 ms below 80 % CPU and 0.33 ms at saturation.
  2. Capacity ≈ vCPU ÷ CPU per query. Here 2 ÷ 0.33 ms ≈ 6,000/s. Use the CPU per query measured near saturation, and measure on the machine type you will run: two hyperthreads of one core often do less than two full cores.
  3. Plan for 60–70 % of it if you care about p99, less if your traffic is bursty.
  4. Check what the number leaves out (next section). Writes, disk reads and lock contention each lower it, often by a lot.

What this measurement does not tell you

Why we measured this

Stackrig simulates architectures in the browser and shows where they break, and a Postgres primary is usually the first thing that does. We run these measurements to calibrate the simulation against real machines, and we publish how far the model is off: it got throughput and CPU right, but near a database's capacity its p99 was too optimistic in these runs, mostly because the calibration design left out the wait for a connection in the application's pool. The engine models pool limits since v0.2; re-checking it against these runs is in progress. The details are on How accurate is Stackrig?

On the start page, the live demo runs the Web app with database design: pull the traffic to 3× and watch the database become the bottleneck.

Open “Web app with database” in the playground Watch the live demo Get early access

More