Repository navigation
fix: prefer cp1252 when the charset detector ties - #2679
Open
Muhammad Usman Mateen (usmanmateen) wants to merge 1 commit into
Open
Muhammad Usman Mateen (usmanmateen) wants to merge 1 commit into
Muhammad Usman Mateen (usmanmateen) wants to merge 1 commit into
Conversation
Western European text in cp1252 often scores exactly like cp1250 or a non-Latin code page in charset_normalizer, and best() then picks by name. "Café crème" came out as "Cafﻠ crﻟme" and "crème" as "crčme" in plain text, CSV and ZIP members. Add best_charset_match(): on a tie in both chaos and coherence, take cp1252, unless the data has S or Z with caron, which point to cp1250. Use it for the stream charset guess and the plain text and CSV fallbacks.
Author
|
@microsoft-github-policy-service agree |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #2678
Western European text in Windows-1252 often scores exactly like
cp1250(or a non-Latin code page) incharset_normalizer, andbest()then returns whichever comes first. "Café crème" came out as "Cafﻠ crﻟme" and "crème" as "crčme".Changes
best_charset_match()in_charset_utils.py. Whencp1252ties with the best match in both chaos and coherence, it returnscp1252, unless the data contains S or Z with caron (0x8A 0x8E 0x9A 0x9E). Those bytes are the same in cp1250 and cp1252 and point to Croatian or Slovene text. In every other case it returnsbest()as before._get_stream_info_guessesand the fallbacks inPlainTextConverterandCsvConverteruse it instead offrom_bytes(...).best().OutlookMsgConverterunchanged. It only runs when a declared code page fails to decode, and its tests patchfrom_bytesin that module. I'm happy to switch it too if you prefer.The trade-off is in #2678: Croatian, Slovene or Bosnian text without š or ž is byte-for-byte as ambiguous as French, so a short sample like that can now decode as cp1252.
Tests
test_western_european_cp1252_text_is_decoded(French, Spanish, Portuguese, Danish) andtest_western_european_cp1252_csv_is_decoded: fail onmain, pass here.test_central_european_cp1250_text_is_still_decoded(Polish, Czech, Hungarian, Croatian): pass onmainand here.packages/markitdown: 1104 passed, 12 skipped.test_speech_transcriptionfails both onmainand here, because it needs the speech recognition service.black23.7.0 (the pre-commit version) leaves the changed files unchanged.