Repository navigation
fix: recover PDF text after inline images - #1889
Yufeng He (he-yufeng) wants to merge 3 commits into
Conversation
|
Rebased onto current Focused validation after the rebase: Result: |
b75be1c to
678aa48
Compare
|
Gentle nudge — freshly rebased with green checks. Would appreciate a look when someone has a minute. |
|
Still open on main: |
|
I checked this with a generated PDF containing an inline image in a compressed content stream: the decoded stream contains Could the detection inspect decoded page content streams, and include a valid compressed-PDF fixture in the regression? The current fixture is uncompressed and both extractors are mocked, so it cannot catch this case. |
|
Confirmed and fixed in 7cdf6e8, exactly the gap you described. Detection now walks On the decoded-stream alternative you suggested: I went with zlib over |
|
Thanks, the Flate detection case is now covered; your focused suite passes (10 passed, 2 skipped). I found one binary-data edge case in the new decode step: I prepared a small follow-up on top of 7cdf6e8: cdb9e9f (patch). Feel free to cherry-pick it. Passing the untrimmed bytes to zlib preserves the checksum; zlib already ignores the PDF newline delimiter after its end-of-stream marker. The four regression cases build complete PDFs with a page tree and cross-reference table, cover both checksum endings and LF/CRLF delimiters, and first verify that pdfminer decodes the page stream correctly. All four fail at the detection assertion before the fix and pass afterward. No PyMuPDF dependency is needed for these tests. Validation on Python 3.12: focused PDF suite 14 passed, 2 skipped; full core suite 311 passed, 33 skipped; Black 23.7.0 and |
|
Sharp catch, thank you. Verified your analysis before landing it: zlib.decompress does ignore the trailing PDF delimiter, and a stream ending in 0x0a exists and fails after the rstrip. Your fix is in 4554569, converter plus your four regression cases unchanged (the padding values land the Adler-32 on both byte endings, nice trick). GitHub's .patch endpoint refuses to serve patches containing PDF bytes, so I applied the two files by hand from the API view and credited you in the commit message; full pdf suite passes 50 passed / 2 skipped here. |
…ection Review on 678aa48: the BI/ID/EI operators live in a page's decoded content stream, so the raw-byte search misses them whenever the stream is Flate compressed, and neither recovery nor the warning fired. Detection now walks stream..endstream blocks and zlib-decodes the Flate ones before giving up. The regression runs the warn path against both an uncompressed fixture and a compressed one, plus a direct detection check on both.
…ction cagdasyurekli's follow-up on 7cdf6e8: rstrip(b"\r\n") also cuts a compressed stream whose Adler-32 checksum happens to end in 0x0a or 0x0d, and the truncated bytes then fail to decompress so detection silently misses. Pass the bytes to zlib untrimmed; it already ignores the PDF newline delimiter after its end-of-stream marker. His four regression cases cover both checksum endings and LF/CRLF delimiters, verified against pdfminer's own decode. Applied by hand since GitHub's .patch endpoint filters PDF bytes; authorship credit to cagdasyurekli.
4554569 to
284264f
Compare
|
Rebased onto current main and resolved the conflicts. The repo's test consolidation (#2579) had absorbed The converter logic is unchanged in spirit: same inline-image detection (now reading through Flate streams), same PyMuPDF recovery comparison, same warning when PyMuPDF is absent. Local verification on this head: |
Summary
Addresses #1870.
Test plan
python -m pytest packages/markitdown/tests/test_pdf_memory.py -q -k "inline_image"python -m pytest packages/markitdown/tests/test_pdf_memory.py packages/markitdown/tests/test_pdf_tables.py -qpython -m py_compile packages/markitdown/src/markitdown/converters/_pdf_converter.py packages/markitdown/tests/test_pdf_memory.pypython -m mypy --ignore-missing-imports packages/markitdown/src/markitdown/converters/_pdf_converter.py packages/markitdown/tests/test_pdf_memory.pypython -m pip install -e "packages/markitdown[pymupdf]"git diff --check