Skip to content

Pod crash due to full disk triggers failing 'Expected empty archive' check on recovery #40

Description

@lasseborly

When the primary CNPG pod (cloudnative-pg-cluster-1) runs out of disk space, it crashes. After increasing the PVC size and restarting the pod, the operator fails to resume WAL archiving.

The barman-cloud-check-wal-archive tool returns an exit code 1 with the error Expected empty archive. It is misidentifying the cluster's own existing historical WAL files in s3://openwebui-backup-bucket/ as a conflicting deployment.

Because archiving remains blocked after recovery, local WAL files saturate the newly allocated disk space rapidly, causing a secondary, faster crash loop.

Current Workaround (Unacceptable long-term)

We are forced to manually alter the S3 destination path (e.g., appending /v2/) to provide an empty directory so the recovery pre-flight check passes. This creates substantial manual maintenance overhead and breaks backup retention linearity.

Expected Behavior

Upon recovery from a disk-full crash, the primary instance should recognize its own system ID/timeline and safely append WAL files to the existing bucket path instead of failing with an Expected empty archive error.

Suggested Solutions to Evaluate

  • Operator Upgrade: Investigate if this recovery false-positive check is resolved in newer releases of the CNPG operator.
  • Archiver Flag Override: Expose a mechanism to bypass the Expected empty archive restriction on crash-recovery loops.
  • Automated Volume Expansion: Implement threshold-based PVC expansion via Prometheus/Alertmanager to prevent the disk from ever hitting 100% saturation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

Projects

  • Status
    Løst

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions