Five useful counts
- Characters: extended grapheme clusters, which approximate user-perceived characters.
- Words: segments that the browser labels word-like for the selected language.
- Code points: Unicode values in the string. One emoji can use several.
- Lines: explicit newline-separated lines, including a final empty line after a trailing newline. Visual wrapping does not count.
- UTF-8 bytes: the byte length of the text export.
Why not split on spaces?
Languages such as Chinese, Japanese and Thai do not separate every word with spaces. TypeMay uses Intl.Segmenter for language-aware segmentation. Browser versions and their language data can produce different counts. These values are not a guarantee of any publisher’s word-count policy.
Emoji, accents and limits
A combined accent and its base may count as one character but two code points. Some emoji are a whole sequence. If the browser lacks Intl.Segmenter, word and grapheme counts show a dash instead of a misleading space-based estimate. Code-point and byte counts still work.
The input limit uses UTF-16 units, a JavaScript string measure that is different from all five displayed counts. The current cap is 100,000 units. Counts describe the current preview, so cleanup may change them.
For a character-by-character view, use the Unicode inspector.