Publications

Scientific publications

A.N. Kirillov, N.B. Krizhanovskaya, A.A. Krizhanovsky.
WSD algorithm based on a new method of vector-word contexts proximity calculation via epsilon-filtration
// Труды КарНЦ РАН. No 7. Сер. Математическое моделирование и информационные технологии. 2018. C. 149-163
Keywords: synonym; synset; corpus linguistics; word2vec; Wikisource; WSD; RusVectores; Wiktionary
The problem of word sense disambiguation (WSD) is considered in the article. Set of synonyms (synsets) and sentences with these synonyms are taken. It is necessary to automatically select the meaning of the word in the sentence.

1285 sentences were tagged by experts, namely, one of the dictionary meanings was selected by experts for target words.

To solve the WSD problem, an algorithm based on a new method of vector-word contexts proximity calculation is proposed. A preliminary epsilon-filtering of words is performed, both in the sentence and in the set of synonyms, in order to achieve higher accuracy.

An extensive program of experiments was carried out. Four algorithms are implemented, including the new algorithm. Experiments have shown that in some cases the new algorithm produces better results.

The developed software and the tagged corpus have an open license and are available online. Wiktionary and Wikisource are used. A brief description of this work can be viewed as slides. A video lecture in Russian about this research is available online.
Indexed at RISC, Google Scholar
Last modified: December 11, 2018