-
Notifications
You must be signed in to change notification settings - Fork 2
Expand file tree
/
Copy pathlambda_function.py
More file actions
2463 lines (2141 loc) · 125 KB
/
Copy pathlambda_function.py
File metadata and controls
2463 lines (2141 loc) · 125 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
"""
CDR Lambda — Content Disarmament and Reconstruction
Triggered by EventBridge on S3 ObjectCreated events.
Supported formats:
Office — .docx .xlsx .pptx and all ZIP/OOXML variants (removes macros, OLE, external links)
PDF — strips JS, OpenAction, Launch, form submit actions, embedded files
Images — re-encodes via Pillow to purge EXIF / ICC exploits
Processing flow (see ``handler``):
download → size/ZIP-structure validation → fail-closed routing by extension →
format-specific CDR (``cdr_office`` / ``cdr_pdf`` / ``cdr_image``) → upload to
SANITISED_BUCKET → publish result + delete source. Rejected/errored/unsupported files
are quarantined instead. Unknown extensions FAIL CLOSED — they are never labelled
sanitised.
Design rules that must not be weakened:
* Side effects (SNS publish, source delete, metrics) are fault-isolated — they can never
turn a successful CDR into an EventBridge retry.
* ZIP entries are read through ``_read_zip_entry_safe`` (chunked counter, never trusts
the central-directory ``file_size``) to defend against decompression bombs.
* CDR drops/neutralises content in place; it never re-serialises through an Office
library, so anything not explicitly touched is preserved.
Environment variables:
SANITISED_BUCKET destination bucket for clean files (required)
QUARANTINE_BUCKET destination for rejected/errored files (optional)
RESULT_TOPIC_ARN SNS topic for CDR result metadata (optional)
CDR_MAX_FILE_BYTES pre-download size limit in bytes (default 104857600 = 100 MB)
CDR_MAX_ENTRY_BYTES per-ZIP-entry decompression limit (default 209715200 = 200 MB)
CDR_MAX_IMAGE_PIXELS decompression-bomb pixel cap for cdr_image (default 40000000 = 40 MP)
"""
import decimal
import html
import io
import json
import logging
import os
import posixpath
import re
import struct
import warnings
import xml.etree.ElementTree as ET
import zipfile
from datetime import datetime, timezone
from collections.abc import Iterable
from typing import Optional
import boto3
import openpyxl
import pikepdf
import pyxlsb
from botocore.config import Config
from PIL import Image, ImageSequence, UnidentifiedImageError
logger = logging.getLogger()
logger.setLevel(logging.INFO)
# Explicit client config rather than botocore's defaults. On a 300 s Lambda under high
# utilisation the defaults are actively harmful: a 60 s connect/read timeout with legacy
# retries lets a single hung S3 socket consume most of the invocation budget, and every
# second spent blocked holds one of the 20 reserved-concurrency slots. Bounded timeouts
# turn a stalled call into a fast, retryable failure. "standard" retry mode adds jitter
# and honours throttling responses, which matters when many concurrent invocations hit
# the same prefix. tcp_keepalive keeps warm-container sockets from being silently dropped.
_BOTO_CONFIG = Config(
retries={"max_attempts": 3, "mode": "standard"},
connect_timeout=5,
read_timeout=30,
tcp_keepalive=True,
max_pool_connections=10,
)
s3 = boto3.client("s3", config=_BOTO_CONFIG)
sns = boto3.client("sns", config=_BOTO_CONFIG)
cw = boto3.client("cloudwatch", config=_BOTO_CONFIG)
SANITISED_BUCKET = os.environ["SANITISED_BUCKET"]
QUARANTINE_BUCKET = os.environ.get("QUARANTINE_BUCKET", "")
RESULT_TOPIC_ARN = os.environ.get("RESULT_TOPIC_ARN", "")
_MAX_FILE_BYTES = int(os.environ.get("CDR_MAX_FILE_BYTES", str(100 * 1024 * 1024)))
_MAX_ENTRY_BYTES = int(os.environ.get("CDR_MAX_ENTRY_BYTES", str(200 * 1024 * 1024)))
# Aggregate decompression budget for one package; _MAX_ENTRY_BYTES bounds only a single
# entry, so many just-under-cap entries otherwise expand without limit (pitfall #46).
# Sized under the 1024 MB MemorySize rather than above it. Corpus max: 757 KB.
_MAX_TOTAL_ENTRY_BYTES = int(os.environ.get("CDR_MAX_TOTAL_BYTES", str(512 * 1024 * 1024)))
# Corpus max: 12 entries. Real OOXML packages run to a few thousand at the extreme.
_MAX_ZIP_ENTRIES = int(os.environ.get("CDR_MAX_ZIP_ENTRIES", "20000"))
# Image.MAX_IMAGE_PIXELS bounds a single frame; cdr_image materialises every frame at
# once, so total pixels and frame count need their own caps (pitfall #46).
_MAX_TOTAL_IMAGE_PIXELS = int(os.environ.get("CDR_MAX_TOTAL_IMAGE_PIXELS", str(80_000_000)))
_MAX_IMAGE_FRAMES = int(os.environ.get("CDR_MAX_IMAGE_FRAMES", "2000"))
# Explicit decompression-bomb cap, sized to this Lambda's 1024 MB default memory budget
# rather than relying on Pillow's own undocumented default (89.5M soft-warn / 179M
# hard-error) — the soft-warn band between those two lets a large-but-under-the-hard-cap
# image decode+re-encode silently, which at several bytes/pixel/buffer can approach the
# Lambda's memory ceiling well before Pillow's own error would ever fire.
Image.MAX_IMAGE_PIXELS = int(os.environ.get("CDR_MAX_IMAGE_PIXELS", str(40_000_000)))
# Pillow decoder allowlist (pitfall #47). IMAGE_EXTS gates which *extensions* route into
# cdr_image, but Image.open() picks its decoder by sniffing CONTENT, so without this the
# extension constrains only the save format and every plugin in Pillow's registry (43 in
# the pinned build — FLI, PSD, DDS, SGI, XPM, JPEG2000, WMF, EPS…) is reachable by naming
# a file .png. That is a native-code attack surface, not a disarm bypass: output is still
# re-encoded, but the attacker chooses which C decoder parses their bytes, which is the
# reachability precondition for a libtiff/FLI/DDS memory-corruption CVE. Deliberately the
# 7-format set, NOT the single declared extension — extension/content mismatch among these
# seven is common in legitimate files, and rejecting real business documents is a
# production incident. Keep in sync with IMAGE_EXTS.
_PILLOW_FORMATS: list[str] = ["JPEG", "PNG", "GIF", "BMP", "TIFF", "WEBP"]
# ── OOXML (Office) dangerous relationship types ────────────────────────────────
STRIP_REL_TYPES: set[str] = {
# vbaProject ships under two interchangeable rel-type URIs in the wild: the Microsoft
# office/2006 form AND the openxmlformats officeDocument/2006 form (LibreOffice and
# python-docx-authored macro docs emit the latter). Both must be stripped, or the rel
# dangles at the removed vbaProject.bin part and python-docx/Word reject the doc.
"http://schemas.microsoft.com/office/2006/relationships/vbaProject",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/vbaProject",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/oleObject",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/externalLink",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/externalLinkPath",
"http://schemas.microsoft.com/office/2006/relationships/attachedToolbars",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/queryTable",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/connections",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/control",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/attachedTemplate",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/subDocument",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/frame",
# altChunk imports an arbitrary HTML/MHTML/RTF chunk that the field/macro scrub never
# inspects (it only scans the host document.xml) — active/remote-content smuggling.
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/aFChunk",
# Embedded package parts (the relationship form of an embedded OLE/package object).
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/package",
# Office Web Add-in (task pane / web extension) relationships.
"http://schemas.microsoft.com/office/2011/relationships/webextensiontaskpanes",
"http://schemas.microsoft.com/office/2011/relationships/webextension",
# customXml parts are stripped by STRIP_ZIP_ENTRIES ("customXml/"). Their relationships
# MUST be dropped here too, or the surviving rel dangles at the deleted part and breaks
# strict OPC consumers (python-docx, Word) — corrupting otherwise-legitimate documents.
# customXml is a data-island carrier (custom doc properties / content-control databinding),
# safe to remove wholesale. Both the 2006 type and the 2010 customXmlProps variant.
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/customXml",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/customXmlProps",
# ── Dangling-rel siblings of parts dropped by STRIP_ZIP_ENTRIES ───────────────
# Each of these points at a part the byte-strip removes; if the rel survives, the
# reconstructed package dangles and strict OPC consumers (python-docx, Word) reject
# it — the same failure mode as vbaProject/customXml above. (Matching is by LOCAL
# NAME — see STRIP_REL_LOCALNAMES — so namespace variants are covered automatically;
# these full URIs are kept for documentation and the membership tests.)
"http://schemas.microsoft.com/office/2006/relationships/xlMacrosheet", # xl/macrosheets/
"http://schemas.microsoft.com/office/2006/relationships/xlIntlMacrosheet",
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/tags", # ppt/tags/
"http://schemas.microsoft.com/office/2006/relationships/activeXControlBinary", # <app>/activeX*.bin
# Excel long-path (>218 char) external-link / OLE-link rels — [MS-OI29500] §12.4.
# Note the deliberately mixed year segments (2019/04 vs 2009/04) per the spec.
"http://schemas.microsoft.com/office/2019/04/relationships/externalLinkLongPath",
"http://schemas.microsoft.com/office/2019/04/relationships/oleObjectLinkLongPath",
"http://schemas.microsoft.com/office/2019/04/relationships/xlExternalLinkLongPath/xlStartup",
"http://schemas.microsoft.com/office/2019/04/relationships/xlExternalLinkLongPath/xlAlternateStartup",
"http://schemas.microsoft.com/office/2009/04/relationships/xlExternalLinkLongPath/xlPathMissing",
"http://schemas.microsoft.com/office/2009/04/relationships/xlExternalLinkLongPath/xlLibrary",
}
def _rel_local(rel_type: str) -> str:
"""Reduce an OPC relationship-type URI to its lowercased local name (the final path
segment). Relationship Types are namespaced URIs, but the SAME logical relationship
ships under multiple interchangeable namespaces in the wild — transitional
(``schemas.openxmlformats.org/officeDocument/2006/…``), Microsoft
(``schemas.microsoft.com/office/{2006,2011,2019/04}/…``), and ISO/IEC-29500 Strict
(``purl.oclc.org/ooxml/officeDocument/…``, emitted after a Word "Strict Open XML"
save). Matching on the namespace-independent local name catches all of them with one
entry instead of enumerating every URI per type — and closes the dual-namespace bug
class (vbaProject under both ms-office and openxmlformats was the original instance).
Lowercased so Word's ``aFChunk`` and the ISO-standard ``afChunk`` both normalise to
one value (altChunk is an active-content import vector — both spellings must strip).
"""
return rel_type.rsplit("/", 1)[-1].lower()
# Local names matched by _strip_rels (namespace-independent — see _rel_local). Derived
# from the documented full-URI set above, so adding a URI there is enough; the extra
# literals below are local names with no single canonical URI in STRIP_REL_TYPES.
STRIP_REL_LOCALNAMES: set[str] = {_rel_local(t) for t in STRIP_REL_TYPES}
# External hyperlink relationship type. NOT stripped (deleting the rel would dangle the
# document's r:id reference) — instead its Target is rewritten to inert in _strip_rels
# (UNC paths leak NTLM creds; arbitrary URLs enable phishing/SSRF). Matched by local name
# so the transitional/strict namespace variants are all covered.
HYPERLINK_REL_TYPE = (
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/hyperlink"
)
HYPERLINK_REL_LOCALNAME = _rel_local(HYPERLINK_REL_TYPE)
STRIP_ZIP_ENTRIES: set[str] = {
"word/vbaProject.bin",
"xl/vbaProject.bin",
"ppt/vbaProject.bin",
"word/activeX",
"xl/activeX",
"ppt/activeX",
"customXml/",
"word/attachedToolbars/",
"word/externalLinks/",
"xl/externalLinks/",
"xl/macrosheets/",
"xl/queryTables/",
"xl/connections.xml",
"ppt/tags/",
# Embedded OLE/package objects (oleObject*.bin, renamed .exe/.lnk/.hta, or a nested
# macro doc). The oleObject relationship is already stripped, but the payload bytes
# remain extractable unless the part itself is dropped.
"word/embeddings/",
"xl/embeddings/",
"ppt/embeddings/",
# Office Web Add-ins (task panes / content add-ins live under webextensions/). A
# webextension auto-loads remote code from an attacker-controlled SourceLocation
# without any VBA — remote-content execution gated only by tenant/admin policy and
# consent prompts. Drop the parts (taskpanes.xml + the webextension*.xml definitions).
"word/webextensions/",
"xl/webextensions/",
"ppt/webextensions/",
# PostScript/EPS image parts (typically under <app>/media/). EPS is a Turing-complete
# interpreter language and a historic RCE surface — drop the payload bytes; the
# matching [Content_Types].xml declaration is stripped in _sanitise_content_types.
".eps",
".ps",
# altChunk payload parts. The aFChunk RELATIONSHIP is dropped (STRIP_REL_TYPES) and
# the <w:altChunk> element is neutralised in _strip_xml_macros, but the imported
# chunk itself — an HTML/MHTML part such as word/afchunk.htm carrying <script>,
# remote references, or an MHTML-smuggled payload — stayed in the package and rode
# into SANITISED_BUCKET intact. It is never inspected by the XML scrub (which only
# ever sees the host part), so anything extracting the package still reaches it.
# Strip the PART as well as the reference, per the strip-part-and-rel rule.
".htm",
".html",
".xhtml",
".mht",
".mhtml",
}
OFFICE_EXTS: set[str] = {
"docx", "docm", "dotx", "dotm",
"xlsx", "xlsm", "xltx", "xltm", "xlam", "xlsb",
"pptx", "pptm", "potx", "potm", "ppsx", "ppsm", "ppam",
}
LEGACY_EXTS: set[str] = {"doc", "xls", "ppt"}
IMAGE_EXTS: set[str] = {"jpg", "jpeg", "png", "gif", "bmp", "tiff", "webp"}
# Formats DELIBERATELY rejected — never given a CDR handler. These carry active content
# as loose, parser-divergent structure (embedded/linked OLE, remote-template refs, control
# words a forgiving consumer acts on) rather than as cleanly-excisable content. A
# reconstruction pass can only defend the grammar IT parses; the attacker targets the
# grammar the *consumer* parses. RTF history bears this out: CVE-2017-0199 (remote OLE
# template), CVE-2017-11882 / CVE-2018-0802 (Equation Editor), CVE-2023-21716 (heap
# corruption in the font table). They are quarantined fail-closed by a dedicated handler
# block that runs BEFORE ZIP validation and the unknown-extension gate — listed explicitly
# so the rejection is a documented, order-independent decision, not an accidental gap a
# future contributor "fixes" by adding a handler. See pitfall #38.
FAIL_CLOSED_EXTS: set[str] = {"rtf"}
EXT_REMAP: dict[str, str] = {
"docm": "docx", "dotm": "dotx",
"xlsm": "xlsx", "xltm": "xltx", "xlam": "xlsx", "xlsb": "xlsx",
"pptm": "pptx", "potm": "potx", "ppsm": "ppsx", "ppam": "pptx",
}
# Includes both container-level types and real OPC part-level (*.main+xml) types.
MACRO_CONTENT_TYPE_REMAP: dict[str, str] = {
"application/vnd.ms-word.document.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.wordprocessingml.document",
"application/vnd.ms-word.document.macroEnabled.main+xml":
"application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml",
"application/vnd.ms-word.template.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.wordprocessingml.template",
"application/vnd.ms-word.template.macroEnabledTemplate.main+xml":
"application/vnd.openxmlformats-officedocument.wordprocessingml.template.main+xml",
"application/vnd.ms-excel.sheet.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
"application/vnd.ms-excel.sheet.macroEnabled.main+xml":
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml",
"application/vnd.ms-excel.sheet.binary.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
"application/vnd.ms-excel.template.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.spreadsheetml.template",
"application/vnd.ms-excel.template.macroEnabled.main+xml":
"application/vnd.openxmlformats-officedocument.spreadsheetml.template.main+xml",
"application/vnd.ms-excel.addin.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
"application/vnd.ms-powerpoint.presentation.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.presentationml.presentation",
"application/vnd.ms-powerpoint.presentation.macroEnabled.main+xml":
"application/vnd.openxmlformats-officedocument.presentationml.presentation.main+xml",
"application/vnd.ms-powerpoint.template.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.presentationml.template",
"application/vnd.ms-powerpoint.template.macroEnabled.main+xml":
"application/vnd.openxmlformats-officedocument.presentationml.template.main+xml",
"application/vnd.ms-powerpoint.slideshow.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.presentationml.slideshow",
"application/vnd.ms-powerpoint.slideshow.macroEnabled.main+xml":
"application/vnd.openxmlformats-officedocument.presentationml.slideshow.main+xml",
"application/vnd.ms-powerpoint.addin.macroEnabled.12":
"application/vnd.openxmlformats-officedocument.presentationml.presentation",
}
# PostScript/EPS content types (lowercased). EPS is a Turing-complete interpreter
# language and a historic RCE surface; a valid OOXML package can legitimately declare
# such a part, so we strip both the [Content_Types].xml declaration and the part bytes.
POSTSCRIPT_CONTENT_TYPES: set[str] = {
"application/postscript",
"application/eps",
"application/x-eps",
"image/eps",
"image/x-eps",
}
def _is_postscript_ct(ct: str) -> bool:
"""True if ``ct`` is a PostScript/EPS content type. Normalises case, surrounding
whitespace, and any RFC-2045 parameter suffix (``;charset=…``) before comparison so
a declaration like ``application/postscript; charset=utf-8`` cannot evade the set."""
return ct.split(";", 1)[0].strip().lower() in POSTSCRIPT_CONTENT_TYPES
_ZIP_MAGIC = b"\x50\x4b\x03\x04"
_SAFE_COMPRESS_METHODS = {0, 8} # stored, deflate
# ── Pure CDR decision core ──────────────────────────────────────────────────────
# Disposition values returned by cdr_dispatch in result["status"]:
# "sanitised" — content disarmed; result["data"] is the clean bytes
# "rejected" — ZIP structural anomaly; hard reject (source must be deleted)
# "unsupported-format" — legacy OLE, fail-closed carrier (RTF), or unknown extension
# cdr_dispatch performs NO I/O (no S3/SNS/CloudWatch). It is the single decision path
# shared by the Lambda handler and any local consumer (e.g. the FastAPI service in
# app.py), so the two can never drift apart on a security decision. Side-effect signals
# the caller may need to act on are returned, not emitted:
# result["metric"] — "zip-anomaly" | "passthrough" | None (caller emits to CloudWatch)
# result["delete_source"] — bool: whether the source object should be deleted on this path
def cdr_dispatch(data: bytes, ext: str, *, max_file_bytes: Optional[int] = None) -> dict:
"""Pure CDR gate: decide and (if applicable) disarm ``data`` for extension ``ext``.
Mirrors the routing in ``handler`` exactly — size guard, legacy-OLE reject,
fail-closed carrier reject, ZIP structural validation, unknown-extension fail-close,
then format-specific CDR — but performs no I/O. Returns a dict:
{
"status": "sanitised" | "rejected" | "unsupported-format",
"data": bytes | None, # clean output (only when sanitised)
"original_ext": str,
"sanitised_ext": str,
"cdr_mode": str | None,
"report": dict | None,
"reason": str | None, # why rejected/unsupported
"metric": str | None, # "zip-anomaly" | "passthrough"
"delete_source": bool,
}
Raises whatever the format CDR functions raise (e.g. on a corrupt PDF) — the caller
decides quarantine-and-reraise vs. surface-the-error, matching the Lambda error path.
"""
ext = ext.lower()
sanitised_ext = EXT_REMAP.get(ext, ext)
limit = _MAX_FILE_BYTES if max_file_bytes is None else max_file_bytes
def _result(status, **kw):
base = {
"status": status, "data": None, "original_ext": ext,
"sanitised_ext": sanitised_ext, "cdr_mode": None, "report": None,
"reason": None, "metric": None, "delete_source": False,
}
base.update(kw)
return base
# ── Pre-CDR size guard ────────────────────────────────────────────────────
if len(data) > limit:
return _result("rejected", reason="file too large", delete_source=False)
# ── Legacy binary (OLE) formats — unsupported, fail closed ────────────────
if ext in LEGACY_EXTS:
return _result("unsupported-format",
reason="OLE binary format not supported",
metric="passthrough", delete_source=True)
# ── Deliberately-rejected carriers (RTF) — FAIL CLOSED (highest priority) ──
if ext in FAIL_CLOSED_EXTS:
return _result("unsupported-format",
reason=f"format rejected by design: {ext}",
metric="passthrough", delete_source=True)
# ── ZIP structural validation (Office only) ───────────────────────────────
if ext in OFFICE_EXTS:
valid, zip_anomalies = _validate_zip_structure(data)
if not valid:
return _result("rejected", reason=zip_anomalies[0],
metric="zip-anomaly", delete_source=True)
# ── Unknown extension — FAIL CLOSED ───────────────────────────────────────
if ext not in OFFICE_EXTS and ext != "pdf" and ext not in IMAGE_EXTS:
return _result("unsupported-format",
reason=f"unsupported extension: {ext}",
metric="passthrough", delete_source=True)
# A CdrReject is a deterministic verdict — hard-reject rather than let it escape as
# an error and burn the EventBridge retry budget on input that fails identically.
try:
if ext in OFFICE_EXTS:
clean_bytes, report = cdr_office(data, ext)
elif ext == "pdf":
clean_bytes, report = cdr_pdf(data)
else: # one of IMAGE_EXTS, per the fail-closed guard above
clean_bytes, report = cdr_image(data, ext)
except CdrReject as exc:
return _result("rejected", reason=str(exc),
metric="zip-anomaly", delete_source=True)
# An empty payload must never be uploaded labelled "sanitised".
if not clean_bytes:
raise ValueError(
f"CDR produced no output for ext={ext} — refusing to label empty content sanitised"
)
return _result("sanitised", data=clean_bytes, report=report,
cdr_mode=report.get("cdr_mode", "full"), delete_source=True)
# ── Entry point ────────────────────────────────────────────────────────────────
def handler(event: dict, context) -> dict:
"""Lambda entry point — receives an EventBridge S3 ObjectCreated event."""
detail = event.get("detail", {})
try:
bucket = detail["bucket"]["name"]
key = detail["object"]["key"]
except KeyError as exc:
logger.error("Malformed EventBridge event — missing field %s: %s", exc, json.dumps(event)[:512])
raise ValueError(f"Malformed event: missing {exc}") from exc
size = detail.get("object", {}).get("size", 0)
logger.info("CDR start: bucket=%s key=%s size=%d", bucket, key, size)
# Validate key for path traversal (S3 keys are literal but defend against confused callers)
if ".." in key.split("/"):
logger.warning("Key contains path traversal segments, rejecting: %s", key)
_publish_result_safe(bucket, key, "rejected", {"reason": "invalid key"})
return {"status": "rejected", "reason": "invalid key"}
# ── Pre-download size guard ───────────────────────────────────────────────
if size > _MAX_FILE_BYTES:
logger.warning("File too large: key=%s size=%d max=%d", key, size, _MAX_FILE_BYTES)
if QUARANTINE_BUCKET:
try:
s3.copy_object(
CopySource={"Bucket": bucket, "Key": key},
Bucket=QUARANTINE_BUCKET,
Key=f"oversized/{key}",
TaggingDirective="REPLACE",
Tagging="cdr-status=rejected&cdr-reason=file-too-large",
)
except Exception as exc:
logger.warning("Could not copy oversized file to quarantine: key=%s error=%s", key, exc)
_publish_result_safe(bucket, key, "rejected", {"reason": "file too large", "size": size})
return {"status": "rejected", "reason": "file too large"}
# ── Download ──────────────────────────────────────────────────────────────
try:
file_bytes, content_type = _download(bucket, key)
except Exception as exc:
_classify_download_error(bucket, key, exc)
raise
# Take the extension from the BASENAME. Splitting the whole key on "." reads a dot in
# a directory name as an extension ("reports.v2/summary" -> ext "v2/summary"), which
# fails closed but logs and alarms under a nonsense extension.
basename = posixpath.basename(key)
ext = basename.rsplit(".", 1)[-1].lower() if "." in basename else ""
# ── Pure CDR decision core ────────────────────────────────────────────────
# cdr_dispatch makes every routing/disarm decision with NO I/O; the handler maps
# its result onto S3/SNS/CloudWatch. The CDR functions can still raise (corrupt PDF
# etc.) — that is the error path below.
try:
decision = cdr_dispatch(file_bytes, ext)
except Exception as exc:
logger.exception("CDR processing failed: key=%s error=%s", key, exc)
if QUARANTINE_BUCKET:
try:
_upload(QUARANTINE_BUCKET, f"error/{key}", file_bytes, content_type,
{"cdr-status": "error", "cdr-error": str(exc)[:256]})
except Exception as q_exc:
logger.warning("Quarantine upload failed: key=%s error=%s", key, q_exc)
_publish_result_safe(bucket, key, "error", {"error": str(exc), "original_ext": ext})
raise
sanitised_ext = decision["sanitised_ext"]
# Emit the side-effect metric the decision asked for (CloudWatch — caller's job).
if decision["metric"] == "zip-anomaly":
_emit_zip_anomaly_metric()
elif decision["metric"] == "passthrough":
_emit_passthrough_metric(ext)
# ── Reject / unsupported routing ──────────────────────────────────────────
if decision["status"] != "sanitised":
reason = decision["reason"]
status = decision["status"]
prefix = "rejected" if status == "rejected" else "unsupported"
tag_status = "rejected" if status == "rejected" else "unsupported-format"
logger.warning("CDR %s: key=%s ext=%s reason=%s", status, key, ext, reason)
quarantine_failed = False
if QUARANTINE_BUCKET:
try:
_upload(QUARANTINE_BUCKET, f"{prefix}/{key}", file_bytes, content_type,
{"cdr-status": tag_status, "cdr-reason": str(reason)[:256],
"cdr-original-ext": ext, "cdr-timestamp": _now()})
except Exception as q_exc:
logger.warning("Quarantine upload failed: key=%s error=%s", key, q_exc)
quarantine_failed = True
_publish_result_safe(bucket, key, status, {"reason": reason, "original_ext": ext})
# If quarantine was supposed to hold the only remaining copy of this file and the
# upload failed, deleting the source would destroy it entirely — never delete in
# that case (pitfall: never destroy the only copy of unprocessable input).
if decision["delete_source"] and not quarantine_failed:
_delete_source_safe(bucket, key)
return {"status": status, "reason": reason}
# ── Sanitised output ──────────────────────────────────────────────────────
clean_bytes = decision["data"]
report = decision["report"]
dest_key = _sanitised_key(key, sanitised_ext)
cdr_mode = decision["cdr_mode"]
sanitised_content_type = _content_type_for_ext(sanitised_ext, content_type)
removal_count = len(report.get("removed", []))
_upload(
SANITISED_BUCKET, dest_key, clean_bytes, sanitised_content_type,
{
"cdr-status": "sanitised",
"cdr-source": f"s3://{bucket}/{key}",
"cdr-timestamp": _now(),
"cdr-removals": str(removal_count),
"cdr-original-ext": ext,
# A structural anomaly is a hard reject, so anything reaching this upload had
# none by construction. (Previously computed from a local that was always [].)
"cdr-zip-anomaly": "false",
"cdr-mode": cdr_mode,
},
)
logger.info("CDR complete: key=%s dest=%s ext=%s removals=%d mode=%s",
key, dest_key, ext, removal_count, cdr_mode)
result_payload = {
"original_ext": ext,
"sanitised_ext": sanitised_ext,
"cdr_mode": cdr_mode,
"zip_anomalies": [],
"report": report,
}
_publish_result_safe(bucket, key, "sanitised", result_payload)
_delete_source_safe(bucket, key)
return {
"status": "sanitised",
"destination": f"s3://{SANITISED_BUCKET}/{dest_key}",
"report": result_payload,
}
# ── Office CDR ─────────────────────────────────────────────────────────────────
def _is_xml_ct(ct: str) -> bool:
"""True when a content type's media subtype is XML. ``+xml`` covers the OOXML family
(…wordprocessingml.document.main+xml, …drawing+xml, …); ``text/xml`` and
``application/xml`` cover the plain forms.
Any ``;`` parameter is discarded before the suffix test. OPC declares a bare media
type, so ``…main+xml; charset=utf-8`` is malformed and both LibreOffice and
python-docx refuse the package outright — the parameter is stripped for robustness,
not to close a live bypass (see pitfall #58). Without it the trailing parameter moves
``+xml`` off the end of the string and the part is never classified as XML."""
ct = ct.split(";", 1)[0].strip().lower()
return ct.endswith("+xml") or ct in ("text/xml", "application/xml")
def _declared_parts(ct_xml: bytes, entry_names: Iterable[str],
matches_ct) -> set[str]:
"""Return the normalised part names (lower, leading '/' stripped) that
``[Content_Types].xml`` declares with a content type satisfying ``matches_ct``.
OPC offers **two** declaration mechanisms and both are authoritative:
* ``Override`` binds a content type to one exact PartName, and
* ``Default`` binds it to *every* part whose extension matches.
Resolving only ``Override`` (pitfall #54) left the ``Default`` half open: a package
declaring ``<Default Extension="dat" ContentType="…wordprocessingml.document.main+xml"/>``
with the officeDocument relationship pointing at ``word/doc.dat`` needs no Override at
all, and python-docx opened the sanitised output and still saw a live DDEAUTO payload
(pitfall #55). ``Default`` is resolved against the real ZIP entry names because it binds
by extension rather than by name.
Returns an empty set on unparseable XML — the suffix rules remain the backstop."""
parts: set[str] = set()
_reject_xml_doctype(ct_xml, "[Content_Types].xml")
try:
root = ET.fromstring(ct_xml)
except ET.ParseError:
return parts
ns = "http://schemas.openxmlformats.org/package/2006/content-types"
default_exts: set[str] = set()
for child in root:
if not matches_ct(child.get("ContentType", "")):
continue
if child.tag == f"{{{ns}}}Override":
pn = child.get("PartName", "").replace("\\", "/").lstrip("/").lower()
if pn:
parts.add(pn)
elif child.tag == f"{{{ns}}}Default":
ext = child.get("Extension", "").strip().lstrip(".").lower()
if ext:
default_exts.add(ext)
if default_exts:
for name in entry_names:
norm = name.replace("\\", "/").lstrip("/").lower()
_, _, ext = norm.rpartition(".")
if ext and ext in default_exts:
# Both spellings: callers match this set against the plain lowered name in
# one place and against the posixpath-canonical form in another, and a
# part named "./word/doc.dat" must be caught either way.
parts.add(norm)
parts.add(posixpath.normpath(norm).lstrip("/"))
return parts
def _postscript_override_parts(ct_xml: bytes,
entry_names: Iterable[str] = ()) -> set[str]:
"""Part names declared with a PostScript/EPS content type in ``[Content_Types].xml``.
The declaration is the authoritative signal for the part-byte strip: an EPS payload can
be declared PostScript while stored as ``word/media/image1.png``, which the ``.eps``/
``.ps`` suffix rule in STRIP_ZIP_ENTRIES would miss. Covers both declaration forms —
see ``_declared_parts``."""
return _declared_parts(ct_xml, entry_names, _is_postscript_ct)
def _xml_override_parts(ct_xml: bytes, entry_names: Iterable[str] = ()) -> set[str]:
"""Part names ``[Content_Types].xml`` declares as XML, whatever their filename suffix.
A real OOXML consumer locates the main document part by following the package
relationship and reads it as XML because the *content type* says so — the ``.xml``
suffix is a convention, not the binding. Covers both declaration forms — see
``_declared_parts``."""
return _declared_parts(ct_xml, entry_names, _is_xml_ct)
# Relationship types (local name) whose Target is a sheet part. `xlBinaryIndex` is the
# type xlsb actually uses for its BIFF12 sheet binaries; the rest are the OOXML sheet
# types, included so a package that binds a .bin under a conventional sheet rel is caught
# too. Matched on local name because the same type ships under both the
# officeDocument/2006 and microsoft.com/office namespaces (mirrors `_strip_rels`).
# Cell-value prefixes Excel treats as the start of a formula. '=' is the only one openpyxl
# turns into a live <f> element in xlsx; the other three matter on CSV export (pitfall #50).
_FORMULA_INJECTION_PREFIXES: tuple[str, ...] = ("=", "+", "-", "@")
_WORKSHEET_REL_TYPES: frozenset[str] = frozenset({
"xlbinaryindex", "worksheet", "chartsheet", "dialogsheet", "macrosheet",
})
def _xlsb_worksheet_parts(src: zipfile.ZipFile, budget: "_DecompressionBudget") -> set[str]:
"""Return the canonical part names an xlsb's workbook relationships bind to a sheet.
OPC part names are arbitrary — a real parser (pyxlsb included) follows the ``Target``
in ``xl/_rels/workbook.bin.rels``, not the conventional ``xl/worksheets/sheetN.bin``
path. Matching on the conventional name let a relocated sheet binary skip conversion
(pitfall #49), so resolve the rels instead.
Deliberately matched on the relationship type's LOCAL NAME (mirrors the rel-type
matching in ``_strip_rels``): the same worksheet type ships under both the
``officeDocument/2006`` and ``microsoft.com/office`` namespaces, and a full-URI compare
misses one of them. Targets are resolved relative to the rels file's owning directory
and canonicalised the same way the ``STRIP_ZIP_ENTRIES`` match is, so ``../`` and
``//`` cannot dodge the comparison.
A ``CdrReject`` from ``_read_zip_entry_safe`` (decompression budget trip) is allowed to
propagate: swallowing it would yield an empty set and fall through to the pass-through
ZIP path, which is exactly the bypass this function exists to close.
"""
parts: set[str] = set()
for item in src.infolist():
name = item.filename.replace("\\", "/").lower()
# Any *.rels under a _rels/ directory — not just xl/_rels/workbook.bin.rels, since
# the workbook part itself is named by the package root rels.
if not (name.endswith(".rels") and "_rels/" in name):
continue
raw = _read_zip_entry_safe(src, item, budget)
try:
_reject_xml_doctype(raw, item.filename)
root = ET.fromstring(raw)
except (ET.ParseError, CdrReject):
continue
base = posixpath.dirname(posixpath.dirname(name)) # strip "_rels/<file>.rels"
for child in root:
rel_type = child.get("Type", "").rsplit("/", 1)[-1].lower()
if rel_type not in _WORKSHEET_REL_TYPES:
continue
target = child.get("Target", "").replace("\\", "/")
if not target or "://" in target:
continue # external target — never a local part
resolved = target.lstrip("/") if target.startswith("/") else posixpath.join(base, target)
canonical = posixpath.normpath(resolved).lstrip("/").lower()
if canonical.endswith(".bin"):
parts.add(canonical)
return parts
def cdr_office(data: bytes, ext: str) -> tuple[bytes, dict]:
"""Disarm an OOXML (ZIP) Office file by rebuilding the archive entry-by-entry.
Drops dangerous parts (VBA, ActiveX, embeddings, external links, …), scrubs the
surviving `.rels`, `[Content_Types].xml`, and XML parts, and dispatches xlsb worksheet
binaries to ``cdr_xlsb``. Never re-serialises through an Office library, so anything
not explicitly touched is preserved byte-for-byte.
Returns ``(clean_bytes, report)`` where report is ``{"format", "removed", "cdr_mode"}``.
"""
removed: list[str] = []
out_buf = io.BytesIO()
budget = _DecompressionBudget()
with zipfile.ZipFile(io.BytesIO(data), "r") as src:
# Pre-pass: resolve which parts [Content_Types].xml declares as PostScript, and
# which it declares as XML, via either OPC declaration form (Override on an exact
# PartName, or Default on an extension). infolist() order is not guaranteed, so
# this must run before the main loop to catch such parts wherever they appear.
# The entry names are passed in because Default binds by extension, not by name.
ps_parts: set[str] = set()
xml_parts: set[str] = set()
entry_names = [i.filename for i in src.infolist()]
for item in src.infolist():
if item.filename.replace("\\", "/").lower() == "[content_types].xml":
try:
ct_raw = _read_zip_entry_safe(src, item, budget)
ps_parts = _postscript_override_parts(ct_raw, entry_names)
xml_parts = _xml_override_parts(ct_raw, entry_names)
except ValueError:
# Bomb-guard trip; suffix rules remain the backstop.
ps_parts = set()
xml_parts = set()
break
# Pitfall #49: resolve xlsb worksheet parts from the workbook relationships, not
# from a hardcoded "xl/worksheets/sheet*.bin" name. OPC part names are arbitrary —
# the rel Target is what a real parser follows — so an xlsb whose sheet binary is
# relocated to e.g. "xl/binparts/data1.bin" skipped conversion entirely and its
# BIFF12 records (FORMULA / DDE / OLEOBJECT / DEFINEDNAME) passed through
# byte-for-byte into the sanitised bucket. Resolved *before* the main loop because
# infolist() order is not guaranteed and the rels part may follow the sheet.
# A budget trip inside raises CdrReject, which propagates — deliberately NOT caught
# and converted into "no sheets found", since that would fall through to the ZIP
# path and pass BIFF12 bytes through unconverted. Fail closed.
xlsb_sheet_parts: set[str] = (
_xlsb_worksheet_parts(src, budget) if ext == "xlsb" else set()
)
with zipfile.ZipFile(out_buf, "w", compression=zipfile.ZIP_DEFLATED) as dst:
for item in src.infolist():
# Normalise backslashes: spec mandates forward slashes
name_lower = item.filename.replace("\\", "/").lower()
# Canonical form for the STRIP_ZIP_ENTRIES match only: collapses "./",
# "//", leading "/", and ".." segments the way OPC part-name / Word's path
# resolution does, so "./word/vbaProject.bin" or "word//vbaProject.bin"
# can't dodge the literal-prefix/suffix check while still resolving to the
# same part at open time.
name_canonical = posixpath.normpath(name_lower).lstrip("/")
# 1. Drop entries matching dangerous file/prefix patterns
skip = False
for pattern in STRIP_ZIP_ENTRIES:
if pattern.endswith("/"):
if name_canonical.startswith(pattern.lower()):
removed.append(item.filename)
skip = True
break
else:
if name_canonical.endswith(pattern.lower()):
removed.append(item.filename)
skip = True
break
if skip:
continue
# 1b. Drop parts declared PostScript/EPS by an Override (content-type
# -driven, regardless of file extension — closes the Override bypass).
if name_lower in ps_parts:
removed.append(item.filename)
continue
# 2. Drop ActiveX control bins/xmls
if re.search(r"activex/activex\d*\.(bin|xml)$", name_lower):
removed.append(item.filename)
continue
# 3. Decompression bomb guard — read in chunks to catch falsified
# file_size, and charge the read against the package-wide budget so a
# package of individually-legal entries cannot decompress to tens of GB.
raw = _read_zip_entry_safe(src, item, budget)
# 4. Sanitise [Content_Types].xml
if name_lower == "[content_types].xml":
raw, ct_removed = _sanitise_content_types(raw)
removed.extend(ct_removed)
# 5. Scrub relationship files
elif name_lower.endswith(".rels"):
raw, rels_removed = _strip_rels(raw)
removed.extend(rels_removed)
# 6. Scrub macro attributes from XML content. `.vml` is included: a VML
# drawing part is XML that carries shape definitions with the same
# action attributes (onClick/onAction) and field-code carriers the
# scrub targets, and it is the classic home of the OLE-object shape.
# Matching only ".xml" left every vmlDrawing*.vml part unscrubbed.
# The Override-declared set is consulted alongside the suffix because
# OPC binds a content type to an exact PartName irrespective of the
# filename: a document part stored as "word/document.bin" and declared
# wordprocessingml via an Override is still read as XML by every real
# consumer, so a suffix-only dispatch left it entirely unscrubbed
# (pitfall #54). Checked before the xlsb branch cannot apply, since an
# xlsb sheet binary is never declared with an XML content type.
elif (name_lower.endswith(".xml") or name_lower.endswith(".vml")
or name_canonical in xml_parts):
raw, xml_removed = _strip_xml_macros(raw, item.filename)
removed.extend(xml_removed)
# 7. xlsb worksheet binary — convert the whole file via cdr_xlsb()
# rather than attempting ZIP-level surgery on BIFF12 records. Only
# worksheet binaries (xl/worksheets/sheet*.bin) trigger conversion;
# other .bin parts (e.g. xl/workbook.bin metadata) must not divert a
# VBA-only/metadata-only xlsb away from the normal ZIP CDR path.
elif ext == "xlsb" and name_canonical in xlsb_sheet_parts:
return cdr_xlsb(data)
dst.writestr(item, raw)
return out_buf.getvalue(), {"format": ext, "removed": removed, "cdr_mode": "full"}
def cdr_xlsb(data: bytes) -> tuple[bytes, dict]:
"""Convert an xlsb file to clean xlsx by extracting cell values via pyxlsb
and re-serialising with openpyxl. Strips all BIFF12 formula records, DDE
references, external links, VBA, and metadata — only plain cell values survive.
Security notes:
- String values starting with '=' are forced to plain text to prevent openpyxl
from serialising them as live formulas in the output xlsx.
- Each xlsb ZIP entry is read through _read_zip_entry_safe before pyxlsb processes
the file, enforcing the same decompression-bomb limit as the standard ZIP CDR path.
"""
removed = ["BIFF12 binary sheet(s) converted to xlsx (all formulas, DDE, and active content stripped)"]
# Enforce per-entry decompression limit on the xlsb ZIP before handing to pyxlsb.
# pyxlsb reads entries internally, bypassing _read_zip_entry_safe — pre-reading
# all entries here ensures a decompression bomb cannot OOM the Lambda.
_budget = _DecompressionBudget()
with zipfile.ZipFile(io.BytesIO(data), "r") as _zf:
for _item in _zf.infolist():
# Raises CdrReject on the per-entry cap or the aggregate package budget.
_read_zip_entry_safe(_zf, _item, _budget)
wb_out = openpyxl.Workbook(write_only=False)
wb_out.remove(wb_out.active) # remove default empty sheet
# Pitfall #49: malformed/hostile BIFF12 makes pyxlsb and openpyxl raise plain
# KeyError/ValueError/IndexError — an unresolvable sheet rel target, a sheet name with
# a character openpyxl forbids, an out-of-range column index. Those are deterministic
# verdicts about the bytes, so translate them into CdrReject; left as generic errors
# they re-raise and burn the full EventBridge retry budget on input that fails
# identically every attempt, ending in the DLQ instead of quarantine.
try:
_convert_xlsb_sheets(data, wb_out)
except CdrReject:
raise
except (KeyError, ValueError, IndexError, StopIteration, struct.error) as exc:
raise CdrReject(f"unparseable xlsb workbook: {type(exc).__name__}") from exc
out_buf = io.BytesIO()
wb_out.save(out_buf)
return out_buf.getvalue(), {
"format": "xlsb",
"converted_to": "xlsx",
"removed": removed,
"cdr_mode": "full",
}
def _convert_xlsb_sheets(data: bytes, wb_out: "openpyxl.Workbook") -> None:
"""Read every sheet's cell values via pyxlsb and append them to ``wb_out``.
Split out of ``cdr_xlsb`` so the whole pyxlsb/openpyxl traversal sits behind one
deterministic-reject boundary (pitfall #49)."""
with pyxlsb.open_workbook(io.BytesIO(data)) as wb_in:
for sheet_name in wb_in.sheets:
ws_out = wb_out.create_sheet(title=sheet_name)
with wb_in.get_sheet(sheet_name) as ws_in:
for row in ws_in.rows(sparse=True):
if not row:
continue
max_col = max(cell.c for cell in row)
out_row: list = [None] * (max_col + 1)
for cell in row:
v = cell.v
# Force formula-injection prefixes to plain text (pitfall #50).
# Only '=' makes openpyxl serialise a live <f> element, so that is
# the whole xlsx threat and the rest are inert *here*. They are not
# inert downstream: '+', '-' and '@' are the standard CSV-injection
# prefixes, and Excel evaluates all four when a sanitised sheet is
# exported to CSV and reopened — so a '@SUM'/'+DDE' payload survives
# CDR as text and goes live one export later. Leading whitespace is
# stripped before the check because Excel tolerates it ahead of the
# prefix. Prefixing with an apostrophe is the standard force-text
# convention and is idempotent for our purposes.
if isinstance(v, str) and v.lstrip("\t\r\n ").startswith(
_FORMULA_INJECTION_PREFIXES
):
v = "'" + v
out_row[cell.c] = v
ws_out.append(out_row)
class _DecompressionBudget:
"""Running total of decompressed bytes across every entry of ONE package.
The per-entry cap alone bounds no aggregate: a package whose entries are each under
_MAX_ENTRY_BYTES can still decompress to tens of GB in total (see
_MAX_TOTAL_ENTRY_BYTES). One budget instance is threaded through every read of a
single archive so the total is enforced, not just the maximum.
"""
__slots__ = ("used", "limit")
def __init__(self, limit: int = None):
self.used = 0
self.limit = _MAX_TOTAL_ENTRY_BYTES if limit is None else limit
def consume(self, n: int, entry_name: str) -> None:
self.used += n
if self.used > self.limit:
raise CdrReject(
f"package exceeds total decompression budget {self.limit} bytes "
f"(while reading '{entry_name}')"
)
def _read_zip_entry_safe(src: zipfile.ZipFile, item: zipfile.ZipInfo,
budget: Optional["_DecompressionBudget"] = None) -> bytes:
"""Read a ZIP entry in chunks, raising if the actual decompressed size exceeds the limit.
Does not trust item.file_size which can be falsified in the central directory.
``budget``, when supplied, additionally enforces the aggregate across all entries of
the package — checked per chunk, so an oversized entry is abandoned mid-read rather
than after it has already been buffered.
"""
buf = bytearray() # extend-in-place: a chunk list + b"".join() doubles peak memory,
total = 0 # 400 MB for one max-size entry inside a 1024 MB container.
with src.open(item) as f:
while True:
chunk = f.read(65536)
if not chunk:
break
total += len(chunk)
if total > _MAX_ENTRY_BYTES:
raise CdrReject(
f"ZIP entry '{item.filename}' exceeds decompression limit {_MAX_ENTRY_BYTES}"
)
if budget is not None:
budget.consume(len(chunk), item.filename)
buf.extend(chunk)
return bytes(buf)
def _strip_rels(data: bytes) -> tuple[bytes, list[str]]:
"""Scrub an OPC ``.rels`` part: drop relationships of dangerous types (VBA, OLE,
external links, altChunk, …) and rewrite external hyperlink Targets to inert while
keeping the rel (so the document's ``r:id`` references don't dangle).
Returns ``(rels_bytes, removed)``; unparseable XML is returned unchanged.
"""
removed: list[str] = []
_reject_xml_doctype(data, "(.rels part)")
try:
root = ET.fromstring(data)