Skip to content

Add MHTML (.mhtml/.mht) converter - #2676

Open
Bicheng (Kenneth) (BichengWang) wants to merge 1 commit into
microsoft:mainfrom
BichengWang:feat/mhtml-converter
Open

Bicheng (Kenneth) (BichengWang) wants to merge 1 commit into
microsoft:mainfrom
BichengWang:feat/mhtml-converter

Conversation

@BichengWang

Copy link
Copy Markdown

Closes #228.

MHTML files currently fall through to the plain-text converter, so the output is the raw MIME container (headers, boundaries, base64 images).

This adds MhtmlConverter. It parses the container with the stdlib email package, picks the page's HTML part, and hands it to HtmlConverter. No new dependencies.

  • Picks the part named by the start parameter if there is one, otherwise the first HTML part (text/html or application/xhtml+xml).
  • Decodes quoted-printable and base64 bodies, and uses the part's charset (UTF-8 if none is declared).
  • Falls back to the message Subject as the title when the page has no <title>.
  • Embedded images and stylesheets are ignored. cid: image references stay as they are in the HTML.
  • Matches .mhtml, .mht, and the application/x-mimearchive type Magika reports, so extensionless files work too.
  • Deliberately does not claim message/rfc822, so .eml handling is not affected.

Tests are in tests/test_mhtml.py. The full suite passes locally.

Pull the page's HTML part out of the MIME container with the stdlib email
parser and convert it with HtmlConverter. Handles quoted-printable/base64
bodies, declared charsets, XHTML roots and the multipart/related start
parameter. Embedded resources are ignored.

Closes microsoft#228
@BichengWang

Copy link
Copy Markdown
Author

Hello members, is there any one can help additional review? Thank you!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FEATURE REQUEST] MHTML Support

1 participant