Upgrading PostgreSQL from 16 to 18 is a copy-paste from the docs; the time actually goes into aligning the old and new environments down to the details. We went through it gating every step — verify, then advance — and kept the record.
Context and principles
The target is a containerised business database (application and database in separate containers). Two principles up front: no trial and error on production — anything unproven runs in a scratch environment first; and the rollback path comes before the upgrade path — confirm how to get back before you go forward.
Gate one: the data directory moved
The new major version's official images changed the data directory layout. A mount written the old way leaves the container with an empty or unwritable data directory — showing up as failed initialisation or data that "vanished". Nothing is corrupted; the mount simply does not land where the new version expects it.
The check is direct: compare the old and new image layouts and confirm mount point, data directory and ownership agree. The lesson — read the target image's changelog before a major upgrade, and treat mounts and ownership as objects of review — not "the old compose file will do".
Gate two: where environment variables come from
The upgrade also hit a classic: environment variables baked into the image at build time. An entrypoint written as "default if unset" gets bypassed by a variable that is never unset, and the service connects to the wrong host. The symptom looks like a networking problem, so networking is where you look first.
The trail (documented in a separate postmortem): print the runtime environment → compare item by item against expectations → locate the override. The cure: the entrypoint exports everything explicitly, plus a startup self-check that fails loudly when a key dependency is unreachable instead of degrading silently.
Gate three: a backup only counts once it restores
The most important prerequisite: a backup is not "a file appeared" — it is "verified restorable". Our checklist:
| Step | Verification | Pass criterion |
|---|---|---|
| Full export | List the archive contents | Complete schema, no errors |
| Restore in scratch | An actual restore | Service boots, no warnings |
| Business spot checks | Key table counts, latest documents, attachment integrity | Counts match the source |
| Rollback path | Simulated rollback in scratch | Return to the old version's running state |
"The service starts" is not enough — a service starting and business data being intact are two different things.
Trial upgrade and the change window
With a verified backup in hand, the real upgrade begins: first a full-scale trial upgrade in the scratch environment, then smoke tests of the application (login, key documents, reports). Once the trial passes, the production upgrade inside the window is largely table-driven:
- Stop writes (window opens; the app flips to read-only);
- Final incremental backup, verified;
- Upgrade the data directory (timed throughout);
- Boot the new version and run the smoke tests;
- Resume writes and watch key metrics through the observation window (connections, slow queries, lock waits).
Three reusable conclusions
- The order is not optional: read the changelog → review mounts and ownership → compare the runtime environment → verify the backup restores → trial upgrade → business-level checks → cut over. None of the steps is hard; wrong order means rework.
- The work of a major upgrade happens before the upgrade: the execution itself is short — the preparation decides whether it takes ten minutes or ten hours.
- Feed the checklists back: all three trap classes (mounts, environment variables, backup verification) went into the deployment checklist — so the next upgrade does not have to rediscover them.
