Semantic Error Prediction: Estimating Word Production Complexity

Abstract

Estimating word complexity is a well-established task in computer-assisted language learning. So far, however, complexity estimation has been largely limited to comprehension. This neglects words that are easy to comprehend, but hard to produce. We introduce semantic error prediction (SEP) as a novel task that assesses the production complexity of content words. Given the corrected version of a learner-produced text, a system has to predict which content words replace tokens from the original text. We present and analyse one example of such a semantic error prediction dataset, which we generate from an error correction dataset. As neural baselines, we use BERT, RoBERTa, and LLAMA2 embeddings for SEP. We show that our models can already improve downstream applications, such as predicting essay vocabulary scores.

Other Versions

No versions found

Links

PhilArchive

External links

  • This entry has no external links. Add one.
Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

Similar books and articles

Analytics

Added to PP
2024-10-19

Downloads
583 (#96,557)

6 months
98 (#122,640)

Historical graph of downloads
How can I increase my downloads?

Author's Profile

David Strohmaier
Cambridge University

Citations of this work

No citations found.

Add more citations