All-words Word Sense Disambiguation for Russian Using Automatically Generated Text Collection

The limited amount of the sense annotated data is a big challenge for the word sense disambiguation task. As a solution to this problem, we propose an algorithm of automatic generation and labelling of the training collections based on the monosemous relatives concept. In this article we explore the limits of this algorithm: we employ it to harvest training collections for all ambiguous nouns, verbs and adjectives presented in RuWordNet thesaurus and then evaluate the quality of the obtained collections. We demonstrate that our approach can create high-quality labelled collections with almost full-coverage of the RuWordNet polysemous words. Furthermore, we show that our method can be applied to the Word-in-Context task.

eISSN:: 1314-4081
Language:: English

Publication timeframe:: 4 times per year
Journal Subjects:: Computer Sciences, Information Technology

Journal RSS Feed

All-words Word Sense Disambiguation for Russian Using Automatically Generated Text Collection

Published Online: Dec 10, 2020

Page range: 90 - 107

Received: Oct 15, 2020

Accepted: Oct 29, 2020

DOI: https://doi.org/10.2478/cait-2020-0049

Keywords
Word Sense Disambiguation, Word-in-Context task, automatic annotation of training collections, monosemous relatives, Russian dataset, RuWordNet thesaurus

© 2020 Bolshina Angelina et al., published by Sciendo

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

All-words Word Sense Disambiguation for Russian Using Automatically Generated Text Collection

Published Online: Dec 10, 2020

Page range: 90 - 107

Received: Oct 15, 2020

Accepted: Oct 29, 2020

DOI: https://doi.org/10.2478/cait-2020-0049

KeywordsWord Sense Disambiguation, Word-in-Context task, automatic annotation of training collections, monosemous relatives, Russian dataset, RuWordNet thesaurus

© 2020 Bolshina Angelina et al., published by Sciendo

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Keywords
Word Sense Disambiguation, Word-in-Context task, automatic annotation of training collections, monosemous relatives, Russian dataset, RuWordNet thesaurus