Skip to content

Research: production compression replay and published evidence #34

Description

@awksedgreep

Parent: #28. Depends on the harness, sweep, and candidate profiles.

Summary

Validate candidate profiles under production-shaped replay and publish reproducible evidence suitable for operator guidance and technical articles.

Replay scenarios

  • Sustained, burst, trickle, and backfill ingestion.
  • Concurrent narrow/wide queries and field discovery.
  • Quiet periods available for settling and continuous load without idle time.
  • Retention, WAL checkpoints, backups, restart, SIGTERM, and kill-9 during optimization.
  • CPU throttling, constrained memory, slow storage, and supported ARM/x86 classes.

Published evidence

  • Corpus/generator identity and privacy methodology.
  • Exact hardware, CPU features, memory, thread count, commit, toolchain, and commands.
  • Foreground latency/throughput, settled ratio, time to settle, optimizer CPU/memory/write amplification, query effects, and operational file accounting.
  • Raw machine-readable measurements plus a human-readable comparison.
  • Clear statements of tradeoffs and workloads where another profile wins.

Acceptance criteria

  • Every advertised profile claim is reproducible from public or generated inputs.
  • Reports distinguish logical raw bytes, data payload, index, WAL, and whole-file size.
  • Regression thresholds protect the selected profiles in CI or a documented periodic benchmark gate.
  • Documentation and blog-ready material are generated from the same evidence rather than manually copied figures.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions