Skip to content

Add SRT (.srt) subtitle converter - #2677

Open
Bicheng (Kenneth) (BichengWang) wants to merge 1 commit into
microsoft:mainfrom
BichengWang:feat/srt-converter
Open

Bicheng (Kenneth) (BichengWang) wants to merge 1 commit into
microsoft:mainfrom
BichengWang:feat/srt-converter

Conversation

@BichengWang

Copy link
Copy Markdown

SRT files currently go through the plain-text converter, so the output keeps cue numbers, millisecond timing arrows and styling tags, which is noisy for an LLM reader.

This adds SrtConverter, which turns each cue into one line prefixed with its start time:

[00:00:01] Hello there. Second line.

[00:01:05] Café time
  • Cue numbers, end times, milliseconds and <i>, <b>, <u>, <font> and {\an8} styling are dropped. Other angle brackets (<Bob>, x<y) are kept.
  • A cue runs from its timing line to the next one, so a blank line inside a cue or a missing separator between cues doesn't lose or merge text.
  • Handles BOMs and CR, LF and CRLF line endings. If the detected charset fails on the full file (for example ascii guessed from the first 64 KiB), it falls back to a whole-file guess.
  • A .srt with no cues raises, so the plain-text converter still handles it.
  • Matches .srt, text/srt, text/x-srt and application/x-subrip.

No new dependencies. Tests are in tests/test_srt.py, and the full suite passes locally.

Convert .srt files to a transcript with one '[HH:MM:SS] text' line per cue,
dropping cue numbers, millisecond timings and styling tags (<i>, <b>, <u>,
<font>, {\an8}). A cue runs from its timing line to the next one, so blank
lines inside a cue or a missing separator do not lose text. Text that only
looks like markup (<Bob>, x<y) is kept, and a .srt with no cues is left to the
plain-text converter.
@BichengWang

Copy link
Copy Markdown
Author

Hello members, is there any one can help additional review? Thank you!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant