How many requests per second can one Postgres handle?
Short answer: it is CPU per query times the number of cores. We measured one Postgres on a 2-vCPU cloud VM with a simple indexed read: it served about 6,000 queries per second at 100 % CPU, roughly 0.33 ms of CPU per query.
The number to plan with is lower. With random (Poisson) arrivals, p99 stayed under 7 ms up to about 75 % CPU (4,800 queries/s), doubled by 88 %, and at 100 % requests waited for seconds.
"Can one Postgres take our launch?" has no universal answer, because a query can cost 50 µs or 50 ms of CPU. Blog numbers range from a few hundred to millions of queries per second, and they rarely say what the query was. So here is one measurement we made ourselves, with the full setup, and a method to get your number.
The setup
- Database VM: Google Cloud
n2-standard-2(2 vCPU = 1 physical core with 2 hyperthreads, 8 GB),europe-west3, Debian 12, PostgreSQL 15 from the Debian packages,shared_buffers = 1GB, everything else default. - Data: one table with 1,000,000 rows and a covering index; the whole table fits in memory, so reads never touch the disk.
- Query: one aggregate per request over about 1,000 rows of one bucket:
SELECT count(*), sum(v) FROM items WHERE bucket = $1. Read-only. - Application: a Node.js service on its own
n2-standard-2VM, 2 workers, a connection pool of 16 to Postgres, one query per HTTP request. - Load: an open-loop generator on a third VM (
n2-standard-4), 256 connections, 10 steps of 60 seconds from 640 to 6,080 requests/s; latency counted from the scheduled send time. Two runs: evenly spaced arrivals and random (Poisson) arrivals. - CPU per query: the slope of the database VM's busy CPU over throughput between load steps, not busy CPU divided by queries at one step. The ratio at low load charges the idle background work to the queries and overstated the cost in our first runs by up to 2×.
The results
| Offered req/s | DB CPU (random) | p99 random | DB CPU (even) | p99 even |
|---|---|---|---|---|
| 640 | 9 % | 1.4 ms | 8 % | 1.1 ms |
| 1,920 | 28 % | 1.6 ms | 20 % | 0.9 ms |
| 3,200 | 49 % | 2.5 ms | 35 % | 1.1 ms |
| 3,840 | 60 % | 3.5 ms | 44 % | 1.3 ms |
| 4,480 | 70 % | 5.7 ms | 58 % | 1.6 ms |
| 4,800 | 75 % | 6.9 ms | 65 % | 2.0 ms |
| 5,120 | 80 % | 7.3 ms | 73 % | 3.3 ms |
| 5,440 | 88 % | 12.6 ms | 83 % | 4.2 ms |
| 5,760 | 94 % | 21.0 ms | 91 % | 6.0 ms |
| 6,080 | 99 % | 1,966 ms | 99 % | 1,572 ms |
Three things stand out:
- Throughput is linear until the CPU is full. Every step up to 5,760 requests/s was served completely; at 6,080 offered only about 5,950–6,000 got through (a second run on the same machine type: 5,890).
- The tail climbs long before 100 %. With random arrivals p99 doubled between 70 % and 88 % CPU and nearly doubled again by 94 %. That is queueing: when requests arrive in clumps, the second and third in a clump wait for the first.
- Burstiness matters as much as the average. The same load with evenly spaced arrivals had a p99 2.5–3.5× lower from 60 % CPU on. Real user traffic is closer to random than to a metronome, and a retry storm or a cron job on the hour is worse than random.
How to estimate your own number
- Find the CPU cost of your typical query. Run two steady load levels (for example 20 % and 50 % of
what you expect) and take
(busy cores at level 2 − busy cores at level 1) / (queries/s at level 2 − queries/s at level 1). Our run: about 0.29 ms below 80 % CPU and 0.33 ms at saturation. - Capacity ≈ vCPU ÷ CPU per query. Here 2 ÷ 0.33 ms ≈ 6,000/s. Use the CPU per query measured near saturation, and measure on the machine type you will run: two hyperthreads of one core often do less than two full cores.
- Plan for 60–70 % of it if you care about p99, less if your traffic is bursty.
- Check what the number leaves out (next section). Writes, disk reads and lock contention each lower it, often by a lot.
What this measurement does not tell you
- Writes. Our query is read-only. Every committed write waits for the WAL to reach disk; durable write throughput depends on the disk and on group commit and is usually far lower than reads.
- Data larger than memory. The table fits in
shared_buffers. Once reads go to disk, the disk's IOPS and latency set the limit instead of the CPU. - Heavier queries. Joins, sorts and scans over more rows cost more CPU per query; the method above still works, the number changes.
- Connections. We used a pool of 16. Hundreds of direct connections cost memory and context switches; put a pooler in front.
- Managed databases and other machine types. This is a self-managed VM. Cloud SQL, RDS or Aurora add their own overhead and limits; bigger VMs scale with cores until locks or the WAL become the limit.
- One workload, two runs. Treat it as one data point with its setup, not a benchmark of Postgres.
Why we measured this
Stackrig simulates architectures in the browser and shows where they break, and a Postgres primary is usually the first thing that does. We run these measurements to calibrate the simulation against real machines, and we publish how far the model is off: it got throughput and CPU right, but near a database's capacity its p99 was too optimistic in these runs, mostly because the calibration design left out the wait for a connection in the application's pool. The engine models pool limits since v0.2; re-checking it against these runs is in progress. The details are on How accurate is Stackrig?
On the start page, the live demo runs the Web app with database design: pull the traffic to 3× and watch the database become the bottleneck.
Open “Web app with database” in the playground Watch the live demo Get early access