Choosing an English novel is also a matching problem: a book should be challenging without becoming inaccessible. This research explored a practical way to estimate reading difficulty from the text itself.

27Novels in the experiment
FleschReadability baseline
2021Conference acceptance

The research question

The study asked whether linguistic features could be combined to approximate the difficulty of English novels, using Lexile levels as a reference.

The starting point was the Flesch Reading Ease formula, which relies on sentence length and syllables per word. We explored whether alternative feature combinations could better reflect the difficulty faced by English learners.

From text to features

The experiment used a small set of 27 English novels. Text processing extracted sentence, syllable, and vocabulary measures, followed by feature selection and normalisation.

Candidate feature combinations were compared with the reference difficulty levels. Beyond the conventional Flesch inputs, the research considered vocabulary categories associated with different stages of English education.

What emerged

Three feature combinations outperformed the original feature strategy within this experiment. CET-4 vocabulary counts appeared in each of the selected combinations, alongside sentence length, syllable length, or vocabulary from earlier school stages.

The result suggested a useful direction for personalised reading tools: difficulty can depend on the learner’s vocabulary context, as well as general properties of a text.

A small study, with clear limits

The sample was small, and the mapping between computed scores and reference levels was fixed. The manuscript treats the result as a feasible feature-construction approach with room for further optimisation.

A larger and more diverse dataset, together with reader-level validation, would be needed before treating it as a general-purpose difficulty assessment.

Publication record

K. Jiang, Z. Tang et al. “An optimization of feature construction strategy for difficulty calculation algorithm of English novels.” CSTNED, accepted 2021.

The title and acceptance status are recorded in my archived CV. A final proceedings link and DOI have not been verified, so this page retains the accepted-paper status.