CharacterCodes.net

Homoglyph & Confusable Character Checker

Some characters look identical but are different Unicode code points. A Cyrillic а looks just like a Latin a, and a zero-width space is invisible. Attackers use these to fake domain names, usernames and filenames, and they also cause hard-to-find bugs in code. Paste text below to find look-alike characters, words that mix scripts, and hidden characters, or compare two strings to see if they only look the same.

If this site has been useful, we’d love your support! Consider buying us a coffee to keep things going strong!
Try an example:

How the checker works

Many Unicode characters are drawn almost identically to one another, and some have no visible shape at all. The checker looks at every character in your text, then judges it in the context of the word it sits in, because the same letter can be harmless in one place and dangerous in another.

  1. Classify each characterIt is matched against a list of known look-alikes, Unicode compatibility forms, special spaces and invisible characters.
  2. Group into wordsLetters, digits and marks are grouped into words, and the Unicode script of each character (Latin, Cyrillic, Greek and so on) is detected.
  3. Rate the riskA look-alike is rated by its context: mixed scripts and hidden characters score highest, genuine foreign-language text scores lowest.

Risk levels

High riskTreat as a likely spoof or hidden trick

A strong sign that the text is not what it appears to be. This level is used for:

  • A look-alike character inside a word that mixes scripts, such as a Cyrillic letter among Latin letters.
  • Invisible characters: zero-width space, soft hyphen, word joiner, byte order mark, filler characters and Unicode tag characters.
  • Bidirectional controls such as the right-to-left override, which can make a filename or line of code display differently from how it is stored.

Why it matters: this is how fake domains, impersonated usernames, disguised file extensions and "Trojan Source" code attacks work. Do not trust the text until you have checked it.

Examplepаypal.comCyrillic а (U+0430) inside a Latin word
SuspiciousWorth a closer look

Unusual, and sometimes intentional, but not proof of a problem on its own. This level is used for:

  • A word of three or more letters made up entirely of look-alikes from a single non-Latin script, which can spell a Latin word without a single Latin letter.
  • Styled forms that imitate ASCII letters and digits: fullwidth, mathematical bold or italic, circled and Roman numeral characters.
  • Punctuation and spaces that imitate ASCII, such as the Greek question mark (which looks like a semicolon), the fraction slash, hyphen variants and ideographic or thin spaces.
  • Zero-width joiners and non-joiners outside emoji sequences. These are legitimate in some scripts, but are also used to hide text.

Why it matters: a look-alike semicolon in source code can break a build, and styled letters can slip past search, filters and username rules.

ExamplePayPalFullwidth letters that resemble “PaY”
InfoProbably legitimate

The character resembles a Latin letter, but the surrounding text is ordinary writing in one script. This level is used for:

  • Letters in normal Russian, Ukrainian, Greek, Armenian and other single-script words that happen to look like Latin letters.
  • The no-break space, which is very common in typeset text.

Why it matters: it usually does not. These characters are highlighted so you can see them, and the cleaner leaves them unchanged so genuine text is never damaged.

ExampleПриветRussian “hello”, where р resembles Latin p

What it detects

Cross-script look-alikes

Cyrillic, Greek, Armenian, Cherokee, Lisu and other letters that resemble Latin letters, digits and punctuation, based on Unicode's confusables data (UTS #39).

Styled and compatibility forms

Fullwidth, mathematical, circled, small-form and Roman numeral characters, found through Unicode compatibility normalization (NFKC).

Mixed-script words

Words containing letters from more than one script. Normal combinations such as Han with Hiragana and Katakana (Japanese) or Han with Hangul (Korean) are not flagged.

Invisible characters

Zero-width characters, bidirectional controls, soft hyphens, filler characters and tag characters. Emoji joiners and flag sequences are recognised and ignored.

Limitations

The built-in list covers the most common look-alikes and is not a complete copy of the Unicode confusables file. Look-alikes within plain ASCII, such as 0 and O or l, I and 1, and letter pairs such as rn and m, are not reported. How similar two characters look also depends on the font. A clean result is therefore a good sign, but not a guarantee that text is safe.

If this site has been useful, we’d love your support! Consider buying us a coffee to keep things going strong!