PostgreSQL Large Object Cleanup
This page applies to installations running on PostgreSQL only. Installations on MariaDB or MSSQL are not affected and need no action.
Why It Is Needed
On PostgreSQL, some of the data Dossier Organizer stores — collection metadata, document details
and historic versions — is not kept inside the table row itself but in a separate area of the
database. PostgreSQL does not automatically free that data when the row it belonged to is deleted
or changed, and the routine VACUUM maintenance does not clean it up either.
This leftover data accumulates over time. Deletions are not the only contributor — updating a collection also leaves the previous copy behind — so the amount grows during regular use, and independently of whether the other housekeeping jobs are enabled.
The Large Object Cleanup Job removes this leftover data. It uses vacuumlo, a standard PostgreSQL
maintenance tool, and only ever removes data that no longer belongs to any row.
Enabling this job is recommended for every PostgreSQL installation.
Configuration via Helm
The job runs as a scheduled Kubernetes job and is disabled by default. Enable it under
organizer.db.vacuumlo:
organizer:
db:
vacuumlo:
enabled: false # Set to true to enable the cleanup job
schedule: "0 5 * * *" # Cron expression (UTC); daily at 05:00
image: "postgres:18" # Image providing the vacuumlo tool
connectionUri: "" # Only needed in the special case described below
activeDeadlineSeconds: 1800 # Maximum runtime of a single run, in seconds
- enabled: Enables or disables the job
- schedule: Standard Kubernetes cron expression
- image: The container image providing the
vacuumlotool. The default works with every supported PostgreSQL server version, so it does not need to match your server. If you useglobal.imageRegistryfor a mirrored or offline installation, it applies here as well. - connectionUri: Normally leave this empty — the job derives the connection from your existing
database URL. Set it only if your database URL contains options that the PostgreSQL tools do not
understand, most commonly
ssl=true, which has to be written assslmode=requirehere. Do not include credentials here; the job takes them from the existing database secret. - activeDeadlineSeconds: Maximum time a single run may take, in seconds. The default is 30 minutes.
Example
To enable the cleanup and run it every night at 04:30:
organizer:
db:
vacuumlo:
enabled: true
schedule: "30 4 * * *"