Benchmarking

Word Stats was benchmarked using the first five paragraphs of The Great Gatsby, copied in Appendix A. For analysis, these were stored in a plain text file in UTF-8 format; each paragraph was a single line of text, and paragraphs were separated by blank lines. A few small modifications, indicated with boldfaced font, were made to the original text: two words were converted to numbers (three→3, fifty-one→1851) to test the handling of numbers, and a mid-sentence abbreviation (U.S.) was inserted to test sentence segmentation (the splitting of a body of text into sentences). For the analysis, the primary language was set to American English and all hyphenation minima to 2 .

The table below compares the Word Stats results with those of eight free online text analysis or wordcount tools; the tests were performed on May 6, 2023. In addition, an AI analysis by ChatGPT was added; see Appendix B for details. Finally, results obtained from Microsoft Word were added, even though it provides only limited statistics; Appendix C has a few comments.

To obtain the ground truth, a separate, manual analysis was done with the aid of other tools, notably, Notepad++, Microsoft Excel and Microsoft Calculator. Results matching the ground truth are indicated in green; differences, even when very small, are indicated in orange. Numbers between parentheses were derived by combining results. The header legend follows below the table.

Header legend: 1=Ground truth; 2=`Word Stats`; 3=[readabilityformulas.com](https://readabilityformulas.com/freetests/six-readability-formulas.php); 4=[usingenglish.com](https://www.usingenglish.com/resources/text-statistics/); 5=[lexicool.com](https://www.lexicool.com/text_analyzer.asp); 6=[online-utility.org](https://www.online-utility.org/text/analyzer.jsp); 7=[webtools.services](https://www.webtools.services/text-analyzer); 8=[lumoslearning.com](https://www.lumoslearning.com/llwp/free-text-complexity-analysis.html); 9=[wordcounter.net](https://wordcounter.net/); 10=[wordcount.com](https://wordcount.com/); 11=[ChatGPT 4](https://openai.com/index/gpt-4/); 12=Microsoft Word 365

The main discrepancy in the table relates to the syllable count. This is probably caused by the absence of generally accepted unique hyphenation rules for English. Because the syllable count is used in the calculation of readability scores, we took a closer look by comparing Word Stats’ results to those from several free online hyphenation tools that provide the entire hyphenated text.

Tool Syllables
Word Stats 839
hyphenation.one 824
hyphenator.net 824
hyphenation24.com   830
juiciobrennan.com 863
plazoo.com 778

The results from the first three online tools are close to those of Word Stats. Manual comparison of the hyphenated texts revealed that differences mainly (but not exclusively) stem from improper processing of em dashes (—); e.g., the sentence fragment “temperament—it was” produces syllables tem, pera, ment—it and was rather than tem, pera, ment, it and was, resulting in a missed syllable. The differences with the last two tools may stem from those tools appearing to use hyphenation minima of 1 and 3 letters, respectively, resulting in more viz. fewer syllables. Based on these findings, we consider the Word Stats syllable count trustworthy.

Appendix A. Text used for benchmarking

In my younger and more vulnerable years my father gave me some advice that I’ve been turning over in my mind ever since.

“Whenever you feel like criticizing anyone,” he told me, “just remember that all the people in this world haven’t had the advantages that you’ve had.”

He didn’t say any more, but we’ve always been unusually communicative in a reserved way, and I understood that he meant a great deal more than that. In consequence, I’m inclined to reserve all judgements, a habit that has opened up many curious natures to me and also made me the victim of not a few veteran bores. The abnormal mind is quick to detect and attach itself to this quality when it appears in a normal person, and so it came about that in college I was unjustly accused of being a politician, because I was privy to the secret griefs of wild, unknown men. Most of the confidences were unsought—frequently I have feigned sleep, preoccupation, or a hostile levity when I realized by some unmistakable sign that an intimate revelation was quivering on the horizon; for the intimate revelations of young men, or at least the terms in which they express them, are usually plagiaristic and marred by obvious suppressions. Reserving judgements is a matter of infinite hope. I am still a little afraid of missing something if I forget that, as my father snobbishly suggested, and I snobbishly repeat, a sense of the fundamental decencies is parceled out unequally at birth.

And, after boasting this way of my tolerance, I come to the admission that it has a limit. Conduct may be founded on the hard rock or the wet marshes, but after a certain point I don’t care what it’s founded on. When I came back from the East last autumn I felt that I wanted the world to be in uniform and at a sort of moral attention forever; I wanted no more riotous excursions with privileged glimpses into the human heart. Only Gatsby, the man who gives his name to this book, was exempt from my reaction—Gatsby, who represented everything for which I have an unaffected scorn. If personality is an unbroken series of successful gestures, then there was something gorgeous about him, some heightened sensitivity to the promises of life, as if he were related to one of those intricate machines that register earthquakes ten thousand miles away. This responsiveness had nothing to do with that flabby impressionability which is dignified under the name of the “creative temperament”—it was an extraordinary gift for hope, a romantic readiness such as I have never found in any other person and which it is not likely I shall ever find again. No—Gatsby turned out all right at the end; it is what preyed on Gatsby, what foul dust floated in the wake of his dreams that temporarily closed out my interest in the abortive sorrows and short-winded elations of men.

My family have been prominent, well-to-do people in this Middle Western U.S. city for 3 generations. The Carraways are something of a clan, and we have a tradition that we’re descended from the Dukes of Buccleuch, but the actual founder of my line was my grandfather’s brother, who came here in 1851, sent a substitute to the Civil War, and started the wholesale hardware business that my father carries on today.

Appendix B. ChatGPT analysis

The file with the first five paragraphs of The Great Gatsby that was used for the other analyses was, on October 30, 2024, uploaded to OpenAI's ChatGPT 4.0 which was subsequently asked:

"In the UTF-8 text file just uploaded, what are the number of lines, paragraphs, sentences, words, unique words, numbers, syllables, bytes, characters, letters, digits and spaces, and what is the average number of characters per word, average number of syllables per word, average number of words per sentence, and average number of sentences per paragraph?"

ChatGPT proceeded with an analysis whose results are included in the table above. While these results from a general purpose platform are quite impressive, they are not 100% correct. Therefore, as with all other tools, use AI platforms with care.

Appendix C. Microsoft Word analysis

The file with the first five paragraphs of The Great Gatsby that was used for the other analyses was opened in Microsoft Word 365. Statistics obtained by double-clicking on the word count at the lower left were added to the table above. They are all correct except for the line count. The latter differs for obvious reasons: Word Stats counts the number of strings separated by an end-of-line character, whereas Word counts the number of lines visible on the screen, which depends on the selected font. The number of spaces was calculated by subtracting the number of characters without spaces from the total number of characters.

As Microsoft Word's word count is the de facto metric used by many writers, we looked at it a bit more closely by analyzing the full text of The Great Gatsby, which was used to obtain the results reported in the user manual. The file was manually separated into tokens that were subsequently cleaned, as per Word Stats' algorithm. This yielded 48,406 tokens. Further inspection showed that these contained 10 numbers. Word Stats' analysis of this text reported 48,386 words and 10 numbers. Comparing the ground truth, Word Stats' words + numbers and Microsoft Word's word count, we get:

Source  Count  Difference
Ground truth   48,406   n/a
Word Stats   48,396   -10 (0.021%)
Microsoft Word   48,489   +83 (0.171%)

Closer inspection showed that the text contains 16 tokens that are either times (9:00) or ordinal numbers (158th). Word Stats does currently not recognize these as either a word or number and adds them to the unique non-words list. If they were counted as words, the difference with ground truth would be +6 (0.012%).

In conclusion, while both Microsoft Word's and Word Stats' word count are close to the ground truth, the former deviates more from it than the latter does. Moreover, Word Stats' word count can be further refined.