Gender Asymmetry of Visegrád Group Languages as Reflected by Word Embeddings

Today, word embeddings have become a standard method in natural language processing, largely due to the availability of large language corpora. The models effectively reflect the semantic relationships between words without any additional linguistic input. Recently, more emphasis has been placed on interpreting the seemingly discriminatory results of some queries, with the goal of de-biasing language models.

However, if we consider the vector space to be a reasonably valid model of a linguistic semantic space, does not the asymmetry and subsequent discrimination in word embeddings reflect the (average) discriminatory tendencies inherent in the language? This article explores word embedding models for the Visegrád group languages and we apply basic vector arithmetic to demonstrate the basic language asymmetry present in the models.

It is well known that in English models, vector transfers result in eerily accurate predictions when swapping genders (the famous king – man + woman = queen), but these transfers also result in rather uncomplimentary roles for certain occupations (doctor – man + woman = nurse, or computer programmer – man + woman = homemaker). The article explores similar transfers in models of V4 languages – Slovak, Czech, Polish, and Hungarian. With Hungarian gender neutrality, Polish strong generic masculine, and close parallels between Slovak and Czech, we hope to uncover interesting similarities and differences in gender asymmetry in these languages, based on real language data.

Sprache:: Englisch

Zeitrahmen der Veröffentlichung:: 2 Hefte pro Jahr
Fachgebiete der Zeitschrift:: Linguistik und Semiotik, Theorien und Fachgebiete, Linguistik, andere

Zeitschrift RSS Feed

Gender Asymmetry of Visegrád Group Languages as Reflected by Word Embeddings

Radovan Garabík

Jana Wachtarczyková

Online veröffentlicht: 27. März 2023

Seitenbereich: 354 - 379

DOI: https://doi.org/10.2478/jazcas-2023-0013

Schlüsselwörterword embeddings, discrimination, NLP, grammatical gender, gender stereotypes, generic masculine, gender symmetry, gender asymmetry

© 2022 Radovan Garabík et al., published by Sciendo

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Schlüsselwörter
word embeddings, discrimination, NLP, grammatical gender, gender stereotypes, generic masculine, gender symmetry, gender asymmetry