Skip to content

fix: prefer cp1252 when the charset detector ties - #2679

Open
Muhammad Usman Mateen (usmanmateen) wants to merge 1 commit into
microsoft:mainfrom
usmanmateen:fix/cp1252-charset-tie
Open

Muhammad Usman Mateen (usmanmateen) wants to merge 1 commit into
microsoft:mainfrom
usmanmateen:fix/cp1252-charset-tie

Conversation

@usmanmateen

Copy link
Copy Markdown

Fixes #2678

Western European text in Windows-1252 often scores exactly like cp1250 (or a non-Latin code page) in charset_normalizer, and best() then returns whichever comes first. "Café crème" came out as "Cafﻠ crﻟme" and "crème" as "crčme".

Changes

  • New best_charset_match() in _charset_utils.py. When cp1252 ties with the best match in both chaos and coherence, it returns cp1252, unless the data contains S or Z with caron (0x8A 0x8E 0x9A 0x9E). Those bytes are the same in cp1250 and cp1252 and point to Croatian or Slovene text. In every other case it returns best() as before.
  • _get_stream_info_guesses and the fallbacks in PlainTextConverter and CsvConverter use it instead of from_bytes(...).best().
  • I left the fallback in OutlookMsgConverter unchanged. It only runs when a declared code page fails to decode, and its tests patch from_bytes in that module. I'm happy to switch it too if you prefer.

The trade-off is in #2678: Croatian, Slovene or Bosnian text without š or ž is byte-for-byte as ambiguous as French, so a short sample like that can now decode as cp1252.

Tests

  • test_western_european_cp1252_text_is_decoded (French, Spanish, Portuguese, Danish) and test_western_european_cp1252_csv_is_decoded: fail on main, pass here.
  • test_central_european_cp1250_text_is_still_decoded (Polish, Czech, Hungarian, Croatian): pass on main and here.
  • Full suite in packages/markitdown: 1104 passed, 12 skipped. test_speech_transcription fails both on main and here, because it needs the speech recognition service.
  • black 23.7.0 (the pre-commit version) leaves the changed files unchanged.

Western European text in cp1252 often scores exactly like cp1250 or a
non-Latin code page in charset_normalizer, and best() then picks by name.
"Café crème" came out as "Cafﻠ crﻟme" and "crème" as "crčme" in plain
text, CSV and ZIP members.

Add best_charset_match(): on a tie in both chaos and coherence, take
cp1252, unless the data has S or Z with caron, which point to cp1250. Use
it for the stream charset guess and the plain text and CSV fallbacks.
@usmanmateen

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Windows-1252 text is decoded with the wrong charset ("Café crème" becomes "Cafﻠ crﻟme")

1 participant