"Duplicate submission protection" and "API rate limiting" both sound small — and both are genuinely hard to get right under concurrency. Our form endpoint surfaced three concurrency problems under load testing, each rooted in PostgreSQL's transaction semantics. Here are the wrong pattern, the root cause, the fix, and the concurrency tests we settled on.
Trap 1: idempotent retries kept reading an old snapshot
The endpoint guards duplicates with a "submission key". The natural pattern is check-then-write:
existing = search([('key', '=', key)])
if existing:
return existing.reference # retry: hand back the existing reference
record = create(...) # first time: file it
Single-threaded, perfectly correct. But under repeatable read, querying again inside the same transaction still shows the snapshot from the transaction's start. The sequence: request A files a record but hasn't committed; request B (a retry) cannot see A in its snapshot, so B files too — and the unique constraint finally rejects it. The error code is right, but the "check first" optimisation did nothing and the failure path got exercised anyway.
Two fixes:
- Push correctness into the database: a unique constraint on the key; check-then-write becomes only the fast path, and on conflict you read back and return the existing record — that is the actual idempotency guarantee;
- To read fresh data, switch connections: when you need to see what someone just committed, re-read on a separate connection or new transaction — inside the RR transaction, every re-query shows the same old snapshot.
Trap 2: serialization failures on the rate-limit counter upsert
Rate limiting uses a counter table, incrementing by "key + window" with INSERT ... ON CONFLICT:
INSERT INTO rate_bucket (key, window_start, count)
VALUES (%s, %s, 1)
ON CONFLICT (key, window_start)
DO UPDATE SET count = rate_bucket.count + 1;
Single-threaded it is fine; under concurrency the database throws serialization failures — in repeatable read, two transactions updating the same row can force the later one to roll back. Three layers of handling:
- Commit the counter separately: rate-limit ticks must not ride inside the business transaction, or a business rollback "refunds" consumed quota;
- Retry a bounded number of times: serialization failures are transient; a little retry converges;
- Define the stance: prefer letting one request through over failing every request — counter errors default to allow, with an alert recorded.
Trap 3: concurrent same-key requests — whoever wins, the answer is the same
Even with a unique constraint, one question remains: when two same-key requests arrive nearly simultaneously, what does the "loser" get?
Our answer: 200 with the existing reference, leaning directly on the constraint — on insert conflict, read back and return the existing record. Who arrived first never enters the outcome.
How to write the concurrency tests
All three traps share a property: single-threaded tests never catch them. We froze the scenarios into tests:
| Scenario | How | Expected |
|---|---|---|
| Same key, same content, doubled | Two connections submitting at once | 201 + 200, same reference |
| Same key, different content, doubled | Two connections at once | 201 + 409, explicit conflict |
| Rate-limit boundary | Concurrent requests past the threshold | Threshold admitted, rest 429 |
| Retry reads fresh data | B retries right after A commits | B sees A's record; no duplicate |
The key: actually use two database connections. Simulating concurrency inside one transaction tests nothing.
Takeaways
The pattern is plain: concurrency correctness belongs to database constraints and transaction semantics; the application shapes only the experience. Check-then-write, in-process locks, in-transaction re-reads — all "cheaper" looking — can be wrong under repeatable read. One genuinely concurrent test per constraint is far cheaper than fixing production data later.
