Keyword density: what the number can and cannot tell you

Updated 5 min read

Many SEO plugins show a green light when a keyword reaches some percentage of the text, and many writers ask which percentage to aim for. Google has said that there is no ideal keyword density, so a target number would be invented. Density does have a use, which is spotting repetition you did not intend.

How the number is calculated

Density is the number of times a term appears divided by the total number of terms in the text. A 1,000-word article that uses “email marketing” 8 times has a phrase density of 0.8%. One that uses the single word “email” 15 times has a word density of 1.5%.

Two details change the result and are rarely stated:

  • What counts as a term. If the denominator is all words, the figure is lower than if stopwords such as “the” and “of” are removed first. Two tools can disagree on the same text without either being wrong.
  • Overlap. A phrase is made of words that are also counted individually. In the sentence “email marketing tools help email marketing teams”, “email” appears twice as a word and “email marketing” appears twice as a phrase, both over the same words. Percentages from different tables therefore cannot be added up.

The keyword density checker divides every row by the total terms in the whole text, whether or not stopwords are hidden from the tables. Hiding them changes what you see, and the denominator stays put, so a density you read from a filtered table matches the unfiltered one.

There is no target, so stop looking for one

A number such as “2%” circulates because it is easy to write in a checklist. Search engines do not publish a density figure, and they have said repeatedly that no ideal density exists. What Google does document is the opposite problem: its spam policies describe keyword stuffing as filling a page with keywords or numbers in an attempt to manipulate rankings, such as repeating the same words so often that the text sounds unnatural.

That gives density a defensive role. A very high figure for one phrase is a prompt to reread the text, and a normal figure does not tell you the page will rank. Natural writing about a subject repeats its central term, as well as synonyms, related terms and pronouns, and that mix is what a density table cannot measure.

What a density table is good for

Finding a phrase you leaned on. After editing a draft in pieces, one phrase can end up in every paragraph. The two-word and three-word tables put it at the top, with the count beside it. Reading the sentences around the top rows usually shows which ones can say “it”, “this approach” or a shorter version.

Checking a draft against the brief. If a page is meant to be about “invoice templates” and that phrase does not appear in the top rows, while “free download” does, the draft has drifted. This catches an off-topic first draft more reliably than a percentage.

Looking at a competitor page or your own live page. Paste the page source and turn on HTML extraction. The tool parses the markup, removes script, style, template and SVG content, and counts the visible text. It tells you how many characters of tags and script it removed, so you can see the extraction worked. That makes the figures reflect what a reader sees, though the extracted text still includes navigation and footer text if they sit in the same HTML.

Spotting boilerplate. Repeated phrases such as “add to cart”, cookie notices or repeated footer lines climb the table on long pages and dilute the real topic words. Seeing them counted explains why a density figure looks low.

Stopwords change the picture

Without removing stopwords, the top of the table is “the”, “of”, “and”, “to”. Turning stopword removal on hides those rows, along with two-word phrases containing one and three-word phrases that start or end with one. The setting is useful, and it has a side effect: “to be or not to be” disappears entirely, and so does a brand name that happens to be a stopword. If a term you expect is missing, switch the setting off before concluding the text lacks it.

The tool also offers a filter for terms used at least twice. A phrase used once is not repetition, and hiding singletons leaves a shorter list to read.

Chinese text: character groups, not words

Chinese is written without spaces, and splitting it into words correctly needs a dictionary and a segmentation model. The tool does not pretend otherwise. For Chinese it counts overlapping groups of two and three characters, so a run such as 搜索引擎优化 produces 搜索, 索引, 引擎, 擎优, 优化 and the three-character groups beside them. Some of those groups are real words and some are accidents of adjacency, such as 索引 inside 搜索引擎.

Three consequences follow:

  • The count for a group shows how often those characters sit next to each other, not how often a word was used.
  • The total in the denominator is English words plus Chinese characters, so a Chinese density is a share of characters, which is a different quantity from a share of words.
  • Groups starting with common function characters such as 的 or 了 are skipped when stopword removal is on, which makes the table more informative and is a heuristic, not a linguistic judgment.

For Chinese pages the table still does its main job: a group that appears far more often than the rest is a cue to read that paragraph again.

Readability numbers

The tool also reports sentence count and average sentence length. A Flesch Reading Ease score is shown only when at least 80% of the terms are English words and the text has enough words, because the formula relies on syllable counts and space-separated words. Syllables are estimated from vowel groups, so the score is a rough band, and it is hidden for Chinese. For Chinese the average number of characters per sentence is the figure to look at.

A short workflow

  1. Paste the text, or the page source with extraction on.
  2. Read the top five rows of the word, two-word and three-word tables. Does the list describe the topic?
  3. If one phrase has a far higher count than anything else, find its occurrences and rewrite some of them.
  4. Turn stopword removal off once, to confirm nothing important was hidden.
  5. Stop. Ranking depends on the page being useful and matching what searchers want, and a density table cannot tell you that.

Open the tool: Keyword Density Checker & Word Frequency Analysis

Back to guides

More guides