Skip to content

Commit 4a33d53

Browse files
JSv4claude
andauthored
Chunked (resumable) uploads to get past the 100MB proxy cap (#1840)
* Add chunked (resumable) uploads to get past the 100MB proxy cap Upstream proxies (Cloudflare) cap a single proxied request body at 100MB, which blocked large document and zip imports through the REST endpoints. This adds a generic chunked-upload mechanism: the client slices a file into sub-100MB parts, uploads each independently, and the server reassembles them before handing the whole file to the existing import services (so there is one import code path per kind, chunked or not). Backend: - New models ChunkedUploadSession + ChunkedUploadPart. Parts persist through Django storage (not local disk) so any web/worker process can reassemble a session. Initial migration 0001_initial. - REST endpoints: POST /api/imports/chunked/start/, PUT|POST /api/imports/chunked/<id>/parts/<index>/, POST /api/imports/chunked/<id>/complete/, and GET .../<id>/ for resume. complete returns the same response shape as the matching single-request endpoint. Part PUTs use a looser throttle scope (document_import_chunks). - Service layer start_chunked_upload / store_chunk / complete_chunked_upload with streaming reassembly (peak memory bounded to one 8MB block), per-user IDOR isolation, arithmetic + integrity validation, and fast-fail permission gating at start. All four import kinds supported (single document + the three zip flows). - Hourly purge_stale_chunked_uploads GC; new CHUNKED_UPLOAD_* settings. Frontend: - importHttp transparently routes files above CHUNK_THRESHOLD_BYTES (50MB) through the chunked protocol; call sites and response handling unchanged. - Raised the artificial 100MB dropzone cap (2GB single document; bulk-zip dropzone uses the 500MB cap via new FileDropZone props). Tests: backend round-trip/validation/IDOR/GC coverage (test_document_imports_chunked.py) and a frontend chunked-transport suite. CHANGELOG updated. * Make chunked-upload temp-file cleanup non-silent Address code-quality "empty except" findings on the best-effort temp-file unlink in the chunked-upload assembler/completer. Extract a shared _safe_unlink helper that swallows OSError (cleanup failure is non-fatal — the file lingers in the OS temp dir and the original exception must not be masked) but records it at debug level instead of an unexplained pass. * Fix linter findings in chunked upload (pyupgrade + mypy) - pyupgrade: drop the now-redundant quotes on the _chunk_part_path forward-ref annotation (safe under `from __future__ import annotations`). - mypy: restore corpus narrowing after the _resolve_corpus_for_edit refactor by branching on `corpus is None` (the helper returns Optional[Corpus], so the old `corpus_error is not None` check left `corpus` typed Optional at the later `.import_content`/`.id` uses). - mypy: widen the zip service `zip_source` params from `UploadedFile | bytes` to `File | bytes` so the chunked completer can hand them a plain django File (UploadedFile remains a valid subtype). * Consolidate File import with the .base import line (consistency) * Address review: fix chunked-upload concurrency races, public normalise_optional, COMPLETED-session retention - store_chunk: lock the session row (select_for_update) around the check-then-create so concurrent same-index uploads can't race the uniq_chunk_part_per_session constraint into a 500. - complete_chunked_upload: claim the session via an atomic PENDING->ASSEMBLING compare-and-swap UPDATE; a 0-row result (double-complete) is refused with 409 instead of assembling/importing twice. - Document the DOCUMENT-kind memory behaviour: it buffers the whole file (bounded by MAX_DOCUMENT_IMPORT_SIZE_BYTES) because import_content takes bytes; the per-block streaming guarantee holds only for the ZIP kinds. - Rename _normalise_optional -> normalise_optional (public) so views no longer import a private helper across module boundaries; update its test. - purge_stale_chunked_uploads: also purge COMPLETED sessions older than CHUNKED_UPLOAD_COMPLETED_RETENTION_DAYS (default 30; 0 disables) so the audit-trail rows don't grow unbounded. New setting + test. - Add clarifying comment to test_non_zip_bytes_rejected_for_zip_kind. * Address PR #1840 review: chunked-upload resilience and GC race Resolve the Medium and Low items from the latest review: - Frontend uploads parts with bounded concurrency (default 4) instead of strictly sequentially, with per-part exponential-backoff retry on transient/5xx/network errors (4xx fail fast), and an optional onProgress callback on every public import helper. - purge_stale_chunked_uploads no longer purges ASSEMBLING sessions inside the normal stale window (could delete parts under a live complete reassembly); they are reclaimed after a longer grace window (CHUNKED_UPLOAD_ASSEMBLING_ GRACE_HOURS, default 6h) so crashed mid-assembly workers are still cleaned up. - store_chunk reports session info from the freshly-locked row. - ChunkedUploadCompleteView declares only [JSONParser] (no request body). - Celery purge task exposes completed_retention_days. - Service cap helpers use direct settings access instead of dead getattr fallbacks; cross-reference comment between frontend chunk size and backend part cap. - Test imports normalise_optional from services (its real home). - Single-document whole-file-in-RAM reassembly tracked as follow-up (#1843). New tests: ASSEMBLING grace-window GC; frontend concurrency cap, per-part retry, 4xx fail-fast, and progress reporting. * Address review: chunked-upload GC batching, configurable knobs, admin, nit - purge_stale_chunked_uploads streams each queryset via .iterator(chunk_size=100) instead of materializing every stale session into memory (bounds peak memory when the GC runs against a large backlog). - Move CHUNKED_UPLOAD_ASSEMBLING_GRACE_HOURS and CHUNK_ASSEMBLY_BLOCK_SIZE to env-backed settings in config/settings/base.py for operator tunability, consistent with the other CHUNKED_UPLOAD_* knobs; services.py reads them from settings (module symbols preserved for the existing test import + comments). - Register read-only ChunkedUploadSession / ChunkedUploadPart admins so operators can inspect stale/FAILED sessions without dropping into the shell. - importHttp.ts: drop redundant '?? undefined' — withOptional/appendIfDefined already guard null. --------- Co-authored-by: Claude <noreply@anthropic.com>
1 parent 0de93d2 commit 4a33d53

17 files changed

Lines changed: 2755 additions & 174 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 65 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,71 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
77

88
## [Unreleased]
99

10+
### Added
11+
12+
- **Chunked (resumable) uploads for large files** — work around the 100 MB
13+
per-request body ceiling that upstream proxies (Cloudflare) impose on the
14+
document-import REST endpoints. The client slices a file into sub-100 MB
15+
parts, uploads each independently, and the server reassembles them before
16+
handing the whole file to the **existing** import services (so there is one
17+
import code path per kind, chunked or not).
18+
- **Backend models** (`opencontractserver/document_imports/models.py`):
19+
`ChunkedUploadSession` + `ChunkedUploadPart`. Parts persist through Django
20+
storage (not local disk) so any web/worker process can reassemble a
21+
session. Initial migration `0001_initial`.
22+
- **REST endpoints** (`opencontractserver/document_imports/views.py`,
23+
`urls.py`): `POST /api/imports/chunked/start/`,
24+
`PUT|POST /api/imports/chunked/<id>/parts/<index>/`,
25+
`POST /api/imports/chunked/<id>/complete/`, and
26+
`GET /api/imports/chunked/<id>/` (progress, for resuming). `complete`
27+
returns the **same** response shape as the matching single-request
28+
endpoint. Part PUTs use a looser throttle scope (`document_import_chunks`,
29+
5000/hour) than the whole-file `document_imports` scope (120/hour).
30+
- **Service layer** (`opencontractserver/document_imports/services.py`):
31+
`start_chunked_upload` / `store_chunk` / `complete_chunked_upload` plus a
32+
streaming reassembler (peak memory bounded to one 8 MB block regardless of
33+
file size) and `purge_stale_chunked_uploads`. All four import kinds are
34+
supported (single document + the three zip flows). Per-user session
35+
isolation (IDOR-safe 404s), arithmetic + integrity validation, and
36+
fast-fail permission gating at `start`.
37+
- **Periodic cleanup** (`opencontractserver/document_imports/tasks.py`,
38+
`config/settings/base.py` `CELERY_BEAT_SCHEDULE`): hourly
39+
`purge_stale_chunked_uploads` GCs abandoned sessions and their stored parts
40+
after `CHUNKED_UPLOAD_STALE_HOURS` (default 24h).
41+
- **New settings** (`config/settings/base.py`):
42+
`CHUNKED_UPLOAD_PART_MAX_BYTES` (default 90 MB, must stay below the proxy
43+
cap), `CHUNKED_UPLOAD_MAX_PARTS`, `CHUNKED_UPLOAD_STALE_HOURS`, and the
44+
`document_import_chunks` throttle rate.
45+
- **Frontend** (`frontend/src/utils/importHttp.ts`): the four public import
46+
helpers transparently route files above `UPLOAD.CHUNK_THRESHOLD_BYTES`
47+
(50 MB) through the chunked protocol — call sites and response handling are
48+
unchanged. The artificial 100 MB dropzone cap was raised
49+
(`constants.ts`: `MAX_FILE_SIZE_BYTES` → 2 GB for single documents;
50+
bulk-zip dropzone now uses the 500 MB `MAX_IMPORT_ZIP_BYTES` cap via new
51+
`FileDropZone` `maxSizeBytes`/`maxSizeDisplay` props).
52+
- **Resilient frontend transport** (`frontend/src/utils/importHttp.ts`,
53+
`constants.ts`): parts now upload with bounded concurrency
54+
(`UPLOAD.CHUNK_CONCURRENCY`, default 4) instead of strictly sequentially,
55+
each part retries with exponential backoff on transient/5xx/network
56+
failures (`CHUNK_MAX_ATTEMPTS`, `CHUNK_RETRY_BASE_DELAY_MS`) while 4xx
57+
client errors fail fast, and an optional `onProgress(fraction)` callback
58+
on every public import helper drives a progress bar for large uploads.
59+
- **GC race hardening** (`opencontractserver/document_imports/services.py`):
60+
`purge_stale_chunked_uploads` no longer purges `ASSEMBLING` sessions inside
61+
the normal stale window (which could delete parts out from under a live
62+
`complete` reassembly); they are reclaimed only after a longer grace window
63+
(`CHUNKED_UPLOAD_ASSEMBLING_GRACE_HOURS`, default 6h) so a crashed
64+
mid-assembly worker is still cleaned up. The Celery task now also accepts
65+
`completed_retention_days` so an operator can override retention at enqueue
66+
time. The single-document reassembly's whole-file-in-RAM tradeoff is
67+
tracked as a streaming follow-up (issue #1843).
68+
- **Tests**: `opencontractserver/tests/test_document_imports_chunked.py`
69+
(round-trip byte-exact reassembly, validation, IDOR isolation, integrity
70+
checks, zip kinds, stale-session GC, ASSEMBLING grace window) and the
71+
`importHttp chunked transport` suite in
72+
`frontend/src/utils/__tests__/importHttp.test.ts` (concurrency cap,
73+
per-part retry, 4xx fail-fast, progress reporting).
74+
1075
### Security
1176

1277
- **SmartLabelListMutation IDOR fix** (`config/graphql/smart_label_mutations.py`)

‎config/settings/base.py‎

Lines changed: 47 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -792,6 +792,11 @@
792792
"schedule": MEMORY_CURATION_CHECK_INTERVAL_SECONDS,
793793
"options": {"queue": "celery"},
794794
},
795+
"chunked-uploads-purge-stale": {
796+
"task": "opencontractserver.document_imports.tasks.purge_stale_chunked_uploads",
797+
"schedule": 3600.0, # hourly
798+
"options": {"queue": "celery"},
799+
},
795800
}
796801

797802
# Worker Upload Processing
@@ -820,6 +825,43 @@
820825
)
821826
)
822827

828+
# Chunked (resumable) upload limits
829+
# ------------------------------------------------------------------------------
830+
# Back the /api/imports/chunked/* endpoints, which slice a large file into
831+
# sub-ceiling parts to get past the 100 MB per-request body cap on upstream
832+
# proxies (Cloudflare). The assembled-file total is still bounded by
833+
# MAX_DOCUMENT_IMPORT_SIZE_BYTES above.
834+
#
835+
# Maximum size (bytes) of a single uploaded part. MUST stay below the smallest
836+
# upstream proxy body limit (Cloudflare: 100 MB). Default: 90 MB.
837+
CHUNKED_UPLOAD_PART_MAX_BYTES = int(
838+
env("CHUNKED_UPLOAD_PART_MAX_BYTES", default=str(90 * 1024 * 1024))
839+
)
840+
# Hard cap on the number of parts in one session (bounds metadata / abuse).
841+
CHUNKED_UPLOAD_MAX_PARTS = int(env("CHUNKED_UPLOAD_MAX_PARTS", default="100000"))
842+
# Hours of inactivity before an unfinished session and its stored parts are
843+
# eligible for garbage collection by ``purge_stale_chunked_uploads``.
844+
CHUNKED_UPLOAD_STALE_HOURS = int(env("CHUNKED_UPLOAD_STALE_HOURS", default="24"))
845+
# Days a COMPLETED session row is retained as an audit trail before
846+
# ``purge_stale_chunked_uploads`` removes it (its parts were already deleted on
847+
# completion, so this only reclaims small metadata rows). Prevents the table
848+
# from growing unboundedly. Set to 0 to keep COMPLETED rows forever.
849+
CHUNKED_UPLOAD_COMPLETED_RETENTION_DAYS = int(
850+
env("CHUNKED_UPLOAD_COMPLETED_RETENTION_DAYS", default="30")
851+
)
852+
# Grace window (hours) before an ``ASSEMBLING`` session is treated as a crashed
853+
# worker and made eligible for GC. Deliberately far larger than any real
854+
# reassembly so the staleness GC can never delete parts out from under a live
855+
# assembly. See ``document_imports.services.purge_stale_chunked_uploads``.
856+
CHUNKED_UPLOAD_ASSEMBLING_GRACE_HOURS = int(
857+
env("CHUNKED_UPLOAD_ASSEMBLING_GRACE_HOURS", default="6")
858+
)
859+
# Block size (bytes) used when streaming stored parts into the reassembled temp
860+
# file. Bounds peak assembly memory to O(block), independent of file size.
861+
CHUNK_ASSEMBLY_BLOCK_SIZE = int(
862+
env("CHUNK_ASSEMBLY_BLOCK_SIZE", default=str(8 * 1024 * 1024))
863+
)
864+
823865
# Maximum metadata JSON size (in bytes) accepted by the worker upload endpoint.
824866
# Default: 500 MB. Set to 0 to disable the limit.
825867
MAX_WORKER_METADATA_SIZE_BYTES = int(
@@ -895,6 +937,11 @@
895937
"annotation_images": "200/hour", # Image retrieval endpoint, authenticated (higher bandwidth)
896938
"annotation_images_anon": "200/hour", # Image retrieval endpoint, anonymous
897939
"document_imports": "120/hour", # Multipart document import endpoints
940+
# Chunked-upload part PUTs: one large file fans out into many part
941+
# requests, so this scope is far looser than ``document_imports``
942+
# (which is sized for whole-file imports). ``start``/``complete`` stay
943+
# on the strict ``document_imports`` scope.
944+
"document_import_chunks": "5000/hour",
898945
},
899946
}
900947

‎frontend/src/assets/configurations/constants.ts‎

Lines changed: 49 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -137,17 +137,60 @@ export const DEBOUNCE = {
137137

138138
// Upload constraints
139139
export const UPLOAD = {
140-
/** Maximum file size in bytes (100MB) */
141-
MAX_FILE_SIZE_BYTES: 100 * 1024 * 1024,
142-
/** Maximum file size display string */
143-
MAX_FILE_SIZE_DISPLAY: "100MB",
140+
/**
141+
* Maximum single-document size in bytes (2GB). Files larger than
142+
* ``CHUNK_THRESHOLD_BYTES`` are uploaded in chunks (see ``importHttp``),
143+
* so this is no longer bounded by the 100MB upstream proxy (Cloudflare)
144+
* request cap. The backend ``MAX_DOCUMENT_IMPORT_SIZE_BYTES`` is the
145+
* authoritative ceiling and returns 413 above it.
146+
*/
147+
MAX_FILE_SIZE_BYTES: 2 * 1024 * 1024 * 1024,
148+
/** Maximum single-document size display string */
149+
MAX_FILE_SIZE_DISPLAY: "2GB",
150+
/**
151+
* Slice size (50MB) for chunked uploads. Must stay below the smallest
152+
* upstream proxy body limit (Cloudflare caps proxied requests at 100MB);
153+
* 50MB leaves ~2x headroom for multipart framing. Mirrors the backend
154+
* ``CHUNKED_UPLOAD_PART_MAX_BYTES`` guard (90MB).
155+
*
156+
* NOTE: this MUST stay below the backend ``CHUNKED_UPLOAD_PART_MAX_BYTES``
157+
* (``config/settings/base.py``); a part larger than that backend cap is
158+
* rejected with a 413 at ``store_chunk``. Do not widen this past the
159+
* backend value without bumping it too.
160+
*/
161+
CHUNK_SIZE_BYTES: 50 * 1024 * 1024,
162+
/**
163+
* Files larger than this are uploaded via the chunked endpoints instead
164+
* of a single request. Equal to one chunk, so anything that would risk
165+
* the proxy cap is chunked while small files keep the single-shot path.
166+
*/
167+
CHUNK_THRESHOLD_BYTES: 50 * 1024 * 1024,
168+
/**
169+
* How many parts to upload concurrently. The backend serialises writes
170+
* per ``(session, index)`` and supports idempotent re-upload, so parallel
171+
* part PUTs are safe; 4 keeps a fat pipe busy without overwhelming the
172+
* browser's per-host connection pool.
173+
*/
174+
CHUNK_CONCURRENCY: 4,
175+
/**
176+
* Maximum attempts per part before the whole upload aborts. The backend
177+
* persists parts and accepts idempotent re-upload, so a transient network
178+
* blip on one part of a long upload can be retried instead of discarding
179+
* the whole transfer.
180+
*/
181+
CHUNK_MAX_ATTEMPTS: 3,
182+
/**
183+
* Base delay (ms) for exponential backoff between part retries:
184+
* attempt N waits ``CHUNK_RETRY_BASE_DELAY_MS * 2**(N-1)``.
185+
*/
186+
CHUNK_RETRY_BASE_DELAY_MS: 500,
144187
/** Progress percentage shown while bulk upload is in flight (before completion) */
145188
BULK_PROGRESS_INITIAL: 50,
146189
/** Maximum number of corpuses to show in the inline selector preview */
147190
CORPUS_PREVIEW_LIMIT: 5,
148-
/** Maximum corpus-import ZIP size in bytes (500MB) */
191+
/** Maximum corpus-import / bulk ZIP size in bytes (500MB) */
149192
MAX_IMPORT_ZIP_BYTES: 500 * 1024 * 1024,
150-
/** Maximum corpus-import ZIP size display string */
193+
/** Maximum corpus-import / bulk ZIP size display string */
151194
MAX_IMPORT_ZIP_DISPLAY: "500MB",
152195
} as const;
153196

‎frontend/src/components/widgets/modals/UploadModal/UploadModal.tsx‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -457,6 +457,8 @@ export const UploadModal: React.FC<UploadModalProps> = ({
457457
selectedFile={zipFile}
458458
disabled={isUploading}
459459
acceptedFileTypes={acceptedFileTypes}
460+
maxSizeBytes={UPLOAD.MAX_IMPORT_ZIP_BYTES}
461+
maxSizeDisplay={UPLOAD.MAX_IMPORT_ZIP_DISPLAY}
460462
onFilesSelected={handleFilesSelected}
461463
onFileRejected={handleFileRejected}
462464
/>

‎frontend/src/components/widgets/modals/UploadModal/components/FileDropZone.tsx‎

Lines changed: 16 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -23,6 +23,15 @@ interface FileDropZoneProps {
2323
hasFiles?: boolean;
2424
/** Accepted file types for single mode (from backend). Falls back to PDF-only. */
2525
acceptedFileTypes?: AcceptedFileType[];
26+
/**
27+
* Maximum accepted file size in bytes. Defaults to the single-document
28+
* cap; callers pass the ZIP cap for bulk mode. Files above the cap are
29+
* still chunked on upload, but the dropzone bounds what a user can pick
30+
* so the backend's per-flow ceiling isn't exceeded.
31+
*/
32+
maxSizeBytes?: number;
33+
/** Human-readable form of ``maxSizeBytes`` (e.g. "2GB"). */
34+
maxSizeDisplay?: string;
2635
onFilesSelected: (files: File[]) => void;
2736
onFileRejected?: (rejections: FileRejection[]) => void;
2837
}
@@ -39,6 +48,8 @@ export const FileDropZone: React.FC<FileDropZoneProps> = ({
3948
selectedFile,
4049
hasFiles = false,
4150
acceptedFileTypes,
51+
maxSizeBytes = UPLOAD.MAX_FILE_SIZE_BYTES,
52+
maxSizeDisplay = UPLOAD.MAX_FILE_SIZE_DISPLAY,
4253
onFilesSelected,
4354
onFileRejected,
4455
}) => {
@@ -93,7 +104,7 @@ export const FileDropZone: React.FC<FileDropZoneProps> = ({
93104
accept: acceptConfig,
94105
multiple: mode === "single",
95106
disabled: disabled || (mode === "single" && hasFiles),
96-
maxSize: UPLOAD.MAX_FILE_SIZE_BYTES,
107+
maxSize: maxSizeBytes,
97108
});
98109

99110
const handleBrowseClick = useCallback(
@@ -115,13 +126,13 @@ export const FileDropZone: React.FC<FileDropZoneProps> = ({
115126
const oversizedFiles: FileRejection[] = [];
116127

117128
for (const file of files) {
118-
if (file.size > UPLOAD.MAX_FILE_SIZE_BYTES) {
129+
if (file.size > maxSizeBytes) {
119130
oversizedFiles.push({
120131
file,
121132
errors: [
122133
{
123134
code: "file-too-large",
124-
message: `File exceeds ${UPLOAD.MAX_FILE_SIZE_DISPLAY} limit`,
135+
message: `File exceeds ${maxSizeDisplay} limit`,
125136
},
126137
],
127138
});
@@ -223,8 +234,8 @@ export const FileDropZone: React.FC<FileDropZoneProps> = ({
223234
</div>
224235
<div className="secondary-text">
225236
{mode === "bulk"
226-
? `ZIP should contain documents (max ${UPLOAD.MAX_FILE_SIZE_DISPLAY})`
227-
: `Supported: ${acceptedLabels} · Max ${UPLOAD.MAX_FILE_SIZE_DISPLAY} per file`}
237+
? `ZIP should contain documents (max ${maxSizeDisplay})`
238+
: `Supported: ${acceptedLabels} · Max ${maxSizeDisplay} per file`}
228239
</div>
229240
</DropZoneText>
230241
<Button

0 commit comments

Comments
 (0)