Vocabulary Diversity Analyzer guide
Measure unique words, type-token ratio, root type-token ratio, single-use words, lexical density, and average word length.
What this tool does
The Vocabulary Diversity Analyzer tokenizes text into lowercase words and builds a frequency table. Total words are tokens, while unique words are distinct normalized forms. Type-token ratio divides unique words by total words and expresses the result as a percentage. Root type-token ratio divides unique words by the square root of total words to reduce, but not eliminate, sensitivity to passage length.
The tool also counts hapax words, meaning forms that occur exactly once. Lexical words are tokens not included in the utility’s common English function-word list, and lexical density is their share of all tokens. Average word length counts Unicode characters in each normalized token.
These are descriptive mathematical signals, not a model of meaning. Related forms such as run, runs, and running count separately. Synonyms remain separate, while one repeated spelling counts as the same type even when it has different meanings.
How to use it
- Paste an English passage containing at least ten words.
- Select Analyze vocabulary.
- Review total and unique word counts.
- Compare type-token and root type-token ratios only with passages of reasonably similar genre and length.
- Use hapax, density, and word-length results to guide a manual reading of the text.
For draft comparison, analyze equivalent sections rather than one paragraph against an entire document. Type-token ratio normally falls as a passage grows because common words repeat.
Benefits
- Reports eight transparent vocabulary metrics
- Shows raw counts alongside normalized ratios
- Includes a length-moderated root TTR measure
- Identifies words used exactly once
- Processes unpublished text locally
Responsible interpretation
Vocabulary variety is not automatically clarity or quality. Technical instructions may repeat exact terms to prevent ambiguity. Accessible writing often uses familiar words deliberately. Creative writing can use a wide range of words while still being confusing, and short passages can receive very high type-token ratios by chance.
The English common-word list affects lexical density. Other languages, domain terms, contractions, names, and tokenization choices can make the result less representative. Do not use vocabulary metrics to assess intelligence, diagnose a condition, identify an author, grade a person automatically, or make educational and employment decisions.
Practical notice: These statistics describe surface word patterns. They are not educational, psychological, authorship, or professional assessments.