Word count vs character count: which one to trust, and when

Updated 5 min read

You paste a text into two counters and get two different numbers. Neither is necessarily wrong, because “word”, “character” and “length” each have more than one definition, and the limits you are trying to meet are defined by whoever wrote the platform.

Characters: spaces, newlines and what counts as one

Three figures get called “character count”:

FigureSpacesLine breaks
Characters (with spaces)countedusually not counted
Characters without spacesnot countednot counted
UTF-16 code units, what string.length returns in JavaScriptcountedcounted as one or two

For plain English text the first two differ by the number of spaces, and which you need depends on the rule you are meeting. Essay portals and the usual meta description limit use characters with spaces. Some translation quotes use characters without spaces. Check which one before trusting a number.

The third row is the source of most mismatches. Programmers read "hello".length and assume it equals the number of characters a person sees. For ordinary letters that holds. For other text it does not.

Emoji and accents: one character to a reader, several to a program

A family emoji such as 👨‍👩‍👧 is three emoji joined by invisible joiner characters. Its JavaScript .length is 8 and its UTF-8 size is 18 bytes, and a reader sees one symbol. A letter with an accent can be stored as one code point or as a base letter plus a combining mark, and the two look identical on screen.

A counter can count what people see by grouping code points into user-perceived characters (graphemes). The word and character counter does this, so an emoji or an accented letter counts as one. Spreadsheet length formulas and many scripts count code units, which is why a cell with three emoji can report a number far above three.

Whenever a platform counts differently from what you see, its rule wins. X, for instance, weights Chinese characters as two, and a counter that shows one per visible character will read lower than the platform does. Leave a margin on any limit.

Words: spaces are not always the boundary

The simplest word counter splits on spaces. That works for a sentence like this one and fails in several places:

  • don't should be one word, and a naive split on punctuation makes it two.
  • A hyphenated phrase like state-of-the-art is one word to most readers and four to a segmenter that splits at hyphens. The tool joins a hyphen with a word on each side back into one.
  • Chinese and Japanese have no spaces. 你好,世界 has no space-separated words at all.

For Chinese and Japanese the tool uses the browser’s built-in word segmenter, which splits by dictionary. That is an approximation: different engines and versions can segment the same sentence slightly differently, so the word figure for CJK text can move between browsers. This is why the tool reports Chinese, Japanese and Korean characters as a separate, exact number beside the words. For Chinese, the character figure is the one to quote.

Chinese counts: which number do editors mean?

In Chinese publishing “字数” usually means characters, and the details differ by who is asking: some requests count punctuation and some do not, some count digits and Latin letters individually. Ask before you deliver a text to a specification.

In the tool the total character count includes punctuation, digits and letters. The CJK count covers only Han, Hiragana, Katakana and Hangul characters, so full-width punctuation like , and 。 is in the first number and left out of the second. Subtract one from the other and you have a quick check on how much of the text is punctuation and Latin script.

Word processors apply their own rules too. Microsoft Word’s word count treats East Asian characters differently from space-separated words, so a mixed Chinese and English document can read differently there than in a browser tool. When the number matters, use the counter your recipient uses.

SMS: 160, 70 and what happens when you cross them

A single SMS carries 160 characters in the GSM-7 alphabet, or 70 if the message needs UCS-2. One character outside GSM-7 is enough to switch the whole message to UCS-2. Chinese characters trigger it, and so do most emoji and typographic quotes. Longer messages are split into concatenated segments of 153 GSM-7 characters or 67 UCS-2 characters each, because the joining header takes space.

Two further points:

  • Some GSM-7 characters, including €, [, ], { and }, take two positions.
  • An emoji in a UCS-2 message takes two of the 70 positions, because it is stored as a pair of code units.

The tool’s limit meter has a 160 option and a custom field, and it counts one per visible character. It does not detect which encoding your message will use. For a Chinese message set the custom limit to 70; for a message with a stray emoji or curly apostrophe, assume the shorter limit applies.

Title, description and tweet limits

LimitWhat it isCaveat
Title tag, about 60A guideline for search resultsTruncation depends on pixel width, so 60 is an estimate
Meta description, about 160A display conventionSearch engines may show a different snippet entirely
X post, 280A hard limitWeighted: Chinese characters count as two, and links have a fixed cost
SMS segment, 160 / 70A billing and delivery unitDepends on encoding, as above

The tool offers the 60, 160 and 280 presets and a custom one. The first two are rules of thumb, so treat a meter reading just under the line as a rough pass, and use a pixel-width preview for titles that matter. The 280 for X is a count the platform computes itself, and its own composer is the final judge.

Reading time is an average

The tool estimates reading time at 238 words per minute for spaced languages and 400 characters per minute for Chinese, Japanese and Korean, and speaking time at 130 words and 220 characters per minute. A mixed text adds both. Treat these as planning figures: a dense technical page is read slower, and a skimmed list faster.

Which number to use

  • Meeting a platform limit: use that platform’s own counter, and treat any other as an estimate.
  • Quoting a Chinese or Japanese length: use the CJK character count.
  • Writing for a spaced language: words for length, characters with spaces for fitting fields.
  • Storing text in a database column: bytes in UTF-8, because column limits are often defined in bytes and a Chinese character takes three.

Open the tool: Word & Character Counter

Back to guides

More guides