Readability
The readability of written text -- a book, letter, newspaper article etc. -- refers to the ease with which it can be understood by readers. Quoting from Wikipedia (article retrieved Jan. 16, 2023): “In natural language, the readability of text depends on its content (the complexity of its vocabulary and syntax) and its presentation (such as typographic aspects that affect legibility, like font size, line height, character spacing, and line length).”
While readability is a complex topic and this is not the place for a primer, Word Stats has several built-in readability metrics meant to give an impression of a text’s readability. The current version includes most metrics mentioned in the above-referenced Wikipedia article.
The metrics are numerical formulas that use some of the statistics that Word Stats collects during analysis, such as average word or sentence length, or the average number of words with fewer than 3 syllables (‘simple’ words) or at least 3 syllables (‘complex’ words). One metric indicates reading ease on a scale of 0 to 100 (with 100 indicating the highest ease of reading); the others all estimate the U.S. grade level or number of years of education needed to read the text:

A brief description of each metric follows; all quotes are from the linked Wikipedia articles. Because the metrics were developed for U.S. English, they are only calculated for analyses in that language and locale, and are not available when a secondary language is selected. In the calculations, numbers are treated as single-syllable words.
Flesch Reading Ease
Published in 1948, the Flesch Reading Ease metric estimates readability on a scale from 0 to 100, where 0 to 50 indicate that the text is difficult to read and aims at professionals (0) to college freshmen (50) while 50 to 100 indicate the text is easy to read with 12th grade (50) to 4th grade (100) education. The formula is
where
Flesch-Kincaid Grade Level
A 1975 effort sponsored by the U.S. Navy to turn the Flesch Reading Ease metric into a grade-level score resulted in the Flesch–Kincaid formula, which is “one of the most popular and heavily tested formulas. It correlates 0.91 with comprehension as measured by reading tests.” The formula for this metric is
with ASL and ASW defined as above. Scores larger than 12 indicate “the number of years of education generally required to understand this text”.
Gunning Fog Grade Level
The Gunning Fog metric, published in 1952, “correlates 0.91 with comprehension as measured by reading tests.” The formula is:
where ASL is as defined above and
Word Stats currently ignores the restrictions that i) proper nouns, familiar jargon, or compound words should be excluded, and that ii) common suffixes (such as -es, -ed, or -ing) should not be included as a syllable.
SMOG Grade Level
Published in 1969, the SMOG formula – SMOG being an acronym for Simple Measure of Gobbledygook – correlates 0.985 “with the grades of readers who had 100% comprehension of test materials.” It mainly aims at healthcare, with a “2010 study published in the Journal of the Royal College of Physicians of Edinburgh [stating] that "SMOG should be the preferred measure of readability when evaluating consumer-oriented healthcare material."” The formula is:
where
The formula is meant to be applied to a sample of three 10-sentence words; the above definition allows application to either such a sample or to an entire text.
FORCAST Grade Level
“In 1973, a study commissioned by the US military of the reading skills required for different military jobs produced the FORCAST formula. Unlike most other formulas, it uses only a vocabulary element, making it useful for texts without complete sentences.” Its formula is
where
The formula is meant to be applied to a 150-word sample; the above definition allows application to either such a sample or to an entire text.
ARI Grade Level
The Automated Readability Index (ARI), published in 1967, “was designed for real-time monitoring of readability on electric typewriters.” Its formula is:
where
Linsear Write Grade Level
While “purportedly developed for the United States Air Force to help them calculate the readability of their technical manuals,” the Linsear Write metric, published in 1977, “is specifically designed to calculate the United States grade level of a text sample.” Its formula is:
where
The formula is meant to be applied to a 100-word sample; the above definition allows application to either such a sample or to an entire text.
Coleman–Liau Grade Level
Published in 1975, the Coleman-Liau metric “was designed to be easily calculated mechanically from samples of hard-copy text.” Its formula is:
where
The formula is meant to be applied to a 100-word sample; the above definition allows application to either such a sample or to an entire text.
List-based metrics
Two additional readability metrics were considered for inclusion in Word Stats: the New Dale-Chall Grade Level and the Spache Grade Level. Both of these rely on a fixed list of words that fourth-grade American students can reliably understand. They were not included for the following two reasons:
The word lists are fairly old. The 3,000-word Dale-Chall list was released in 1948 and updated in 1995, while the 1,063-word Spache list was created in 1953 and updated in 1978. As such, neither list contains a number of words that most modern-day fourth graders (9-10 years of age) would readily know, like computer, app, smartphone, laptop, download, video game, etc.
Both word lists contain only the basic forms of verbs and nouns. Words such as regular plurals of nouns, regular past tense forms, progressive forms of verbs etc. are considered reliably understandable, but need to be manually added. Extending the original lists is not obvious though. For example, some words can both be a noun and a verb, and a fourth grader may understand one but not the other. The word ‘jar’ for example has the associated plural ‘jars’ when considered a noun, but the past tense ‘jarred’ and the present participle ‘jarring’ when considered a verb; the latter is also an adjective. Should these all be added?
Future versions of Word Stats may contain these two metrics with only the basic word lists, with tentative extensions of these lists, and/or with optional user-defined lists.