Slovene and Croatian word embeddings in terms of gender occupational analogies


  • Matej Ulčar University of Ljubljana, Faculty of Computer and Information Science, Slovenia
  • Anka Supej Jožef Stefan Institute, Ljubljana, Slovenia
  • Marko Robnik-Šikonja University of Ljubljana, Faculty of Computer and Information Science, Slovenia
  • Senja Pollak Jožef Stefan Institute, Ljubljana, Slovenia



word embeddings, gender bias, word analogy task, occupations, natural language processing


In recent years, the use of deep neural networks and dense vector embeddings for text representation have led to excellent results in the field of computational understanding of natural language. It has also been shown that word embeddings often capture gender, racial and other types of bias. The article focuses on evaluating Slovene and Croatian word embeddings in terms of gender bias using word analogy calculations. We compiled a list of masculine and feminine nouns for occupations in Slovene and evaluated the gender bias of fastText, word2vec and ELMo embeddings with different configurations and different approaches to analogy calculations. The lowest occupational gender bias was observed with the fastText embeddings. Similarly, we compared different fastText embeddings on Croatian occupational analogies.


Download data is not yet available.


How to Cite

Ulčar, M., Supej, A., Robnik-Šikonja, M., & Pollak, S. (2021). Slovene and Croatian word embeddings in terms of gender occupational analogies. Slovenščina 2.0: Empirical, Applied and Interdisciplinary Research, 9(1), 26–59. (Original work published July 1, 2021)