跳到主要内容
新品云 ERP 正式上线,几分钟开通专属实例

技术分享

Field notes from a PostgreSQL major upgrade: fix the mount, then the data

Upgrading PostgreSQL from 16 to 18 is a copy-paste from the docs; the time actually goes into aligning the old and new environments down to the details. We went through it gating every step — verify, then advance — and kept the record.

Context and principles

The target is a containerised business database (application and database in separate containers). Two principles up front: no trial and error on production — anything unproven runs in a scratch environment first; and the rollback path comes before the upgrade path — confirm how to get back before you go forward.

Gate one: the data directory moved

The new major version's official images changed the data directory layout. A mount written the old way leaves the container with an empty or unwritable data directory — showing up as failed initialisation or data that "vanished". Nothing is corrupted; the mount simply does not land where the new version expects it.

The check is direct: compare the old and new image layouts and confirm mount point, data directory and ownership agree. The lesson — read the target image's changelog before a major upgrade, and treat mounts and ownership as objects of review — not "the old compose file will do".

Gate two: where environment variables come from

The upgrade also hit a classic: environment variables baked into the image at build time. An entrypoint written as "default if unset" gets bypassed by a variable that is never unset, and the service connects to the wrong host. The symptom looks like a networking problem, so networking is where you look first.

The trail (documented in a separate postmortem): print the runtime environment → compare item by item against expectations → locate the override. The cure: the entrypoint exports everything explicitly, plus a startup self-check that fails loudly when a key dependency is unreachable instead of degrading silently.

Gate three: a backup only counts once it restores

The most important prerequisite: a backup is not "a file appeared" — it is "verified restorable". Our checklist:

Step Verification Pass criterion
Full export List the archive contents Complete schema, no errors
Restore in scratch An actual restore Service boots, no warnings
Business spot checks Key table counts, latest documents, attachment integrity Counts match the source
Rollback path Simulated rollback in scratch Return to the old version's running state

"The service starts" is not enough — a service starting and business data being intact are two different things.

Trial upgrade and the change window

With a verified backup in hand, the real upgrade begins: first a full-scale trial upgrade in the scratch environment, then smoke tests of the application (login, key documents, reports). Once the trial passes, the production upgrade inside the window is largely table-driven:

  1. Stop writes (window opens; the app flips to read-only);
  2. Final incremental backup, verified;
  3. Upgrade the data directory (timed throughout);
  4. Boot the new version and run the smoke tests;
  5. Resume writes and watch key metrics through the observation window (connections, slow queries, lock waits).

Three reusable conclusions

  • The order is not optional: read the changelog → review mounts and ownership → compare the runtime environment → verify the backup restores → trial upgrade → business-level checks → cut over. None of the steps is hard; wrong order means rework.
  • The work of a major upgrade happens before the upgrade: the execution itself is short — the preparation decides whether it takes ten minutes or ten hours.
  • Feed the checklists back: all three trap classes (mounts, environment variables, backup verification) went into the deployment checklist — so the next upgrade does not have to rediscover them.

相关产品与 ERP 实践文章在宏斋博客。 宏斋云ERP · Blog

返回列表