Unicode inspector

Type or paste a character or a text and get a row for each character: its code point, UTF-8 bytes, HTML entity, name, type and alphabet.

  • Free
  • No sign-up
  • Runs in your browser
unicode-inspector

100% private — your text is analysed in your browser and never sent to any server.

How it works

Type or paste a text

A single mysterious symbol or a whole paragraph.

Read the table

Code point, UTF-8 bytes, HTML entity, name, type and alphabet of each character.

Check the summary

Visible characters, code points, UTF-16 units and UTF-8 bytes.

What is this character, exactly?

Every character on a screen is identified by a number, its Unicode code point, and that number is written to a file or sent over a network as one or more bytes. When a symbol looks wrong, a length does not match, a database complains about an encoding or you need to write a character in code, you want to see that identity. This inspector shows, for each character of your text, everything you usually need to know.

What each column tells you

  • Code point: the number of the character in the Unicode standard, written as U+00E9 for é. It is what you use in escapes such as \u00E9 in programming languages.
  • UTF-8: the bytes the character takes in a UTF-8 file, in hexadecimal. Plain letters use one byte, accented letters two, most symbols three and emoji four.
  • HTML: the numeric entity that you can paste into a web page.
  • Name: the official name for common characters, such as LATIN SMALL LETTER E WITH ACUTE. For less common characters, the type and alphabet still tell you what it is.
  • Type and alphabet: whether it is a letter, digit, punctuation mark, symbol, space, control or invisible format character, and which script it belongs to.

Why lengths do not match

The summary explains many "off by one" surprises. A text can have a different number of visible characters, code points, UTF-16 units and bytes. An emoji is one visible symbol but two UTF-16 units and four bytes, and a family emoji is several code points joined together. A letter with an accent can be one code point or two, depending on how it was typed. Knowing which count a program uses explains why a field with a 10-character limit rejects text that looks shorter.

Spot the impostors

Invisible characters, such as zero-width spaces, appear in the table with a coloured tag instead of a blank cell, so you can see them. Words that mix alphabets, such as a Latin word with a Cyrillic letter that looks identical, are flagged with a warning. Both are common in copied text and in phishing attempts.

What it does not do

Character names are built into the tool only for common characters, so a rare symbol may show a type and alphabet without a name. The table lists the first 300 characters of a long text, though the summary counts all of them. For a list of only the hidden characters and a clean copy, use the invisible character detector.

Frequently asked questions

How do I find the Unicode code of a character?
Type or paste the character. Its code point, written as U+ followed by hexadecimal digits, appears in the table.
What is the difference between a code point and a byte?
The code point is the character's number in the Unicode standard. Bytes are how that number is stored in a file; UTF-8 uses one to four bytes per character.
Why does an emoji count as two characters in some programs?
Some programs count UTF-16 units, and emoji outside the basic range use two. The summary shows the visible, code point, UTF-16 and byte counts side by side.
How can I tell if a character is invisible?
Invisible characters show as a coloured tag in the table, with their code and name, instead of an empty cell.
Can it detect lookalike letters?
Yes. Words that mix alphabets, such as a Cyrillic letter inside a Latin word, trigger a warning.
Is my text uploaded anywhere?
No. It is analysed in your browser.