All Tools View Categories About Contact Privacy

Subtitle Language Detector

Detect the writing scripts used in your subtitles.

Runs entirely in your browser — no data is uploaded or stored.

Use before translating or styling so you pick the right font and RTL handling.

About this tool & how to use it
  • Nine script families: The detector recognises Latin, Cyrillic, Greek, Arabic, Hebrew, CJK, Hangul, Devanagari and Thai characters. Each is tested independently, so a file can be reported as containing several at once.
  • Per-cue counts, not just a flag: The report tells you how many cues contain each script instead of only saying yes or no. A file showing 480 Latin cues and 12 Cyrillic cues is obviously a translation that was left unfinished.
  • Single-field workflow: The only control is the SRT Content textarea, so there is nothing to configure and nothing to get wrong. Paste, press Run, read the answer.
  • No translation, no guessing: The tool never attempts to identify a language by vocabulary or machine translation; it only reports the writing systems present. That makes the result deterministic and repeatable.
  • Mixed-script visibility: Cues that combine two writing systems, such as an Arabic line quoting an English brand name, are counted under both scripts so bidirectional problems surface early.
  • Total cue count included: The report opens with how many cues were parsed, which doubles as a sanity check that your SRT was read correctly in the first place.
  • Plain-text output: Results download with a .txt extension, ready to paste into a QC log, an email to a vendor or a delivery checklist.
  • Entirely offline: All analysis happens in the page itself, so embargoed scripts and unreleased episode subtitles never leave your computer.
  1. Open your subtitle file in any text editor, select all of it, and paste it into the SRT Content box on this page.
  2. If you would rather see how the detector behaves first, click Load sample to drop a short example subtitle into the box and work from that.
  3. Press Run to analyse the file; the tool parses the cues and tests each one against the built-in script ranges.
  4. Read the report in the result panel, which lists the total number of cues followed by one line per detected script and the number of cues it appeared in.
  5. Check the numbers against what you expected: a single dominant script normally means a clean monolingual file, while two large counts usually mean source and target text are both still present.
  6. Click Copy to put the report on your clipboard, or Download to save it as a .txt file alongside the subtitle it describes.
  7. Use the answer to choose your next step, such as picking a font that covers the script or switching on right-to-left handling before you style the file.

Example 1 - a half-finished translation. A vendor returns what should be a fully Russian file, but a spot check feels wrong.

Detected scripts across 412 cues:
  Cyrillic: 371 cue(s)
  Latin: 44 cue(s)

Forty-four cues still hold English text, and because a few contain both scripts the counts exceed the cue total. That is a rejected delivery caught in seconds.

Example 2 - confirming right-to-left handling. An unlabelled file arrives before a burn-in job.

Detected scripts across 96 cues:
  Arabic: 96 cue(s)
  Latin: 9 cue(s)

Every cue is Arabic and nine also carry Latin characters, typically numbers or brand names, so pick an Arabic-capable font and apply right-to-left formatting first.

About Subtitle Language Detector

The Subtitle Language Detector takes an SRT pasted into the SRT Content box and reports which writing scripts actually appear inside it. It parses the file into cues, inspects the characters of each one, and returns a plain-text report naming every recognised script with the number of cues containing it, ready to copy or save as a .txt file.

This matters most to people who receive subtitle files rather than write them. Localisation vendors, dubbing studios and e-learning teams routinely inherit an SRT whose filename promises one language and whose contents deliver another, or a file where an outsourced translator left half the cues untranslated. A script check before translation, styling or burn-in catches those mix-ups while they are still cheap to fix.

Detection is purely character-range based, so nothing is guessed and nothing is sent to a translation service. Each cue is tested against Unicode ranges for Latin, Cyrillic, Greek, Arabic, Hebrew, CJK, Hangul, Devanagari and Thai, and a cue containing two scripts is counted under both, which is exactly how a partial translation shows up. The Latin test looks for A-Z and a-z, and the CJK test covers the unified ideograph block, so heavily accented or kana-only lines may not register. If nothing matches, the report says so rather than guessing.

Everything runs in your browser, so confidential pre-release scripts are never uploaded, and a feature-length file with thousands of cues is analysed the moment you press Run.

Features

  • Nine script families: The detector recognises Latin, Cyrillic, Greek, Arabic, Hebrew, CJK, Hangul, Devanagari and Thai characters. Each is tested independently, so a file can be reported as containing several at once.
  • Per-cue counts, not just a flag: The report tells you how many cues contain each script instead of only saying yes or no. A file showing 480 Latin cues and 12 Cyrillic cues is obviously a translation that was left unfinished.
  • Single-field workflow: The only control is the SRT Content textarea, so there is nothing to configure and nothing to get wrong. Paste, press Run, read the answer.
  • No translation, no guessing: The tool never attempts to identify a language by vocabulary or machine translation; it only reports the writing systems present. That makes the result deterministic and repeatable.
  • Mixed-script visibility: Cues that combine two writing systems, such as an Arabic line quoting an English brand name, are counted under both scripts so bidirectional problems surface early.
  • Total cue count included: The report opens with how many cues were parsed, which doubles as a sanity check that your SRT was read correctly in the first place.
  • Plain-text output: Results download with a .txt extension, ready to paste into a QC log, an email to a vendor or a delivery checklist.
  • Entirely offline: All analysis happens in the page itself, so embargoed scripts and unreleased episode subtitles never leave your computer.

How to Use

  1. Open your subtitle file in any text editor, select all of it, and paste it into the SRT Content box on this page.
  2. If you would rather see how the detector behaves first, click Load sample to drop a short example subtitle into the box and work from that.
  3. Press Run to analyse the file; the tool parses the cues and tests each one against the built-in script ranges.
  4. Read the report in the result panel, which lists the total number of cues followed by one line per detected script and the number of cues it appeared in.
  5. Check the numbers against what you expected: a single dominant script normally means a clean monolingual file, while two large counts usually mean source and target text are both still present.
  6. Click Copy to put the report on your clipboard, or Download to save it as a .txt file alongside the subtitle it describes.
  7. Use the answer to choose your next step, such as picking a font that covers the script or switching on right-to-left handling before you style the file.

Examples

Example 1 - a half-finished translation. A vendor returns what should be a fully Russian file, but a spot check feels wrong.

Detected scripts across 412 cues:
  Cyrillic: 371 cue(s)
  Latin: 44 cue(s)

Forty-four cues still hold English text, and because a few contain both scripts the counts exceed the cue total. That is a rejected delivery caught in seconds.

Example 2 - confirming right-to-left handling. An unlabelled file arrives before a burn-in job.

Detected scripts across 96 cues:
  Arabic: 96 cue(s)
  Latin: 9 cue(s)

Every cue is Arabic and nine also carry Latin characters, typically numbers or brand names, so pick an Arabic-capable font and apply right-to-left formatting first.

Benefits

  • Catch bad deliveries before your client does: Spotting untranslated cues at intake means you can send the file back to the vendor instead of explaining a failed QC to the broadcaster.
  • Choose the right font the first time: Knowing the file contains Devanagari or Thai before styling stops the classic run of empty boxes and missing glyphs in a rendered video.
  • Plan bidirectional work up front: Seeing Arabic or Hebrew in the report tells you the job needs right-to-left treatment, so mixed numbers and Latin words are handled properly rather than patched later.
  • Answer a question in seconds: A file that would take several minutes to skim by eye is characterised in one click, which adds up quickly across a delivery of dozens of language versions.
  • Keep confidential material private: Because detection happens entirely in your browser, embargoed scripts and unreleased episodes never touch a third-party server.
  • Produce evidence you can share: The downloadable text report gives you something concrete to attach to a QC record or a vendor email instead of an impression.

Frequently Asked Questions

Does it translate?
No. It only detects the script (alphabet) by analyzing characters, not the spoken language.
Which scripts are recognized?
Latin, Cyrillic, Greek, Arabic, Hebrew, CJK, Hangul, Devanagari, Thai and more via Unicode ranges.
Why detect script?
To pick the right font, RTL handling and translation pipeline before processing.
Is my file uploaded?
No, detection runs locally.