Skip to content

Nikola Ljubesic

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

44

Venues

9

Active years

2008–2026

Best venue rank

A*

Where they publish

Papers

44 indexed papers, newest first.

YearVenueTitleAuthors
2026LRECParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian.Nikola Ljubesic, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungersek
2026LRECThe Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora.Taja Kuzman Pungersek, Peter Rupnik, Vit Suchomel, Nikola Ljubesic
2026LRECROG: A Multi-Layer Manually Annotated Corpus of Spoken Slovenian.Kaja Dobrovoljc Zor, Darinka Verdonik, Jaka Cibej, Peter Rupnik, Nikola Ljubesic
2025COLINGProceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects.Yves Scherrer, Tommi Jauhiainen, Nikola Ljubesic, Preslav Nakov, Jrg Tiedemann, Marcos Zampieri
2025ECIROverview of Touch 2025: Argumentation Systems - Extended Abstract.Johannes Kiesel, agri ltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Frbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matys Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein
2025InterspeechIdentifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models.Nikola Ljubesic, Ivan Porupski, Peter Rupnik
2024COLINGA Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection.Filip Dobranic, Bojan Evkoski, Nikola Ljubesic
2024COLINGCLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation.Nikola Ljubesic, Taja Kuzman
2024COLINGThe ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings.Michal Mochtak, Peter Rupnik, Nikola Ljubesic
2024COLINGDo Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages.Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljubesic, Miquel Espl-Gomis, Gema Ramrez-Snchez, Antonio Toral
2024COLINGGos 2: A New Reference Corpus of Spoken Slovenian.Darinka Verdonik, Kaja Dobrovoljc, Tomaz Erjavec, Nikola Ljubesic
2024ECIROverview of Touch 2024: Argumentation Systems.Johannes Kiesel, agri ltekin, Maximilian Heinrich, Maik Frbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matys Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Mnstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein
2024NAACLUniversal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark.Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Brje Karlsson, Peiqin Lin, Nikola Ljubesic, Lester James V. Miranda, Barbara Plank, Arij Riabi, Yuval Pinter
2023EAMTMaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages.Marta Ban, Malina Chichirau, Miquel Espl-Gomis, Mikel L. Forcada, Aarn Galiano Jimnez, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramrez-Snchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Jaume Zaragoza-Bernabeu
2022EAMTMaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages.Marta Ban, Miquel Espl-Gomis, Mikel L. Forcada, Cristian Garca-Romero, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramrez-Snchez, Peter Rupnik, Vt Suchomel, Antonio Toral, Tobias van der Werff, Jaume Zaragoza
2022LRECThe GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild.Taja Kuzman, Peter Rupnik, Nikola Ljubesic
2021RANLPCultural Topic Modelling over Novel Wikipedia Corpora for South-Slavic Languages.Filip Markoski, Elena Markoska, Nikola Ljubesic, Eftim Zdravevski, Ljupco Kocarev
2020LRECCoSimLex: A Resource for Evaluating Graded Word Similarity in Context.Carlos Santos Armendariz, Matthew Purver, Matej Ulcar, Senja Pollak, Nikola Ljubesic, Mark Granroth-Wilding
2020LRECGigafida 2.0: The Reference Corpus of Written Standard Slovene.Simon Krek, Spela Arhar Holdt, Tomaz Erjavec, Jaka Cibej, Andraz Repar, Polona Gantar, Nikola Ljubesic, Iztok Kosem, Kaja Dobrovoljc
2018ACLBleaching Text: Abstract Features for Cross-lingual Gender Prediction.Rob van der Goot, Nikola Ljubesic, Ian Matroos, Malvina Nissim, Barbara Plank
2016COLINGTweetGeo - A Tool for Collecting, Processing and Analysing Geo-encoded Linguistic Data.Nikola Ljubesic, Tanja Samardzic, Curdin Derungs
2016EAMTCollaborative Development of a Rule-Based Machine Translator between Croatian and Serbian.Filip Klubicka, Gema Ramrez-Snchez, Nikola Ljubesic
2016EAMTDealing with Data Sparseness in SMT with Factured Models and Morphological Expansion: a Case Study on Croatian.Vctor M. Snchez-Cartagena, Nikola Ljubesic, Filip Klubicka
2016EAMTAbu-MaTran: automatic building of machine translation.Antonio Toral, Sergio Ortiz-Rojas, Mikel L. Forcada, Nikola Ljubesic, Prokopis Prokopidis
2016LRECCorpus vs. Lexicon Supervision in Morphosyntactic Tagging: the Case of Slovene.Nikola Ljubesic, Tomaz Erjavec
2016LRECCorpus-Based Diacritic Restoration for South Slavic Languages.Nikola Ljubesic, Tomaz Erjavec, Darja Fiser
2016LRECProducing Monolingual and Parallel Web Corpora at the Same Time - SpiderLing and Bitextor's Love Affair.Nikola Ljubesic, Miquel Espl-Gomis, Antonio Toral, Sergio Ortiz-Rojas, Filip Klubicka
2016LRECNew Inflectional Lexicons and Training Corpora for Improved Morphosyntactic Annotation of Croatian and Serbian.Nikola Ljubesic, Filip Klubicka, Zeljko Agic, Ivo-Pavao Jazbec
2016LRECCroatian Error-Annotated Corpus of Non-Professional Written Language.Vanja Stefanec, Nikola Ljubesic, Jelena Kuvac Kraljevic
2015EAMTAbu-MaTran: Automatic building of Machine Translation.Antonio Toral, Flammie A. Pirinen, Andy Way, Gema Ramrez-Snchez, Sergio Ortiz-Rojas, Raphael Rubino, Miquel Espl-Gomis, Mikel L. Forcada, Vassilis Papavassiliou, Prokopis Prokopidis, Nikola Ljubesic
2015RANLPPredicting Inflectional Paradigms and Lemmata of Unknown Words for Semi-automatic Expansion of Morphological Lexicons.Nikola Ljubesic, Miquel Espl-Gomis, Filip Klubicka, Nives Mikelic Preradovic
2015RANLPPredicting the Level of Text Standardness in User-generated Content.Nikola Ljubesic, Darja Fiser, Tomaz Erjavec, Jaka Cibej, Dafne Marko, Senja Pollak, Iza Skrjanec
2014CICLINGStandardizing Tweets with Character-Level Machine Translation.Nikola Ljubesic, Tomaz Erjavec, Darja Fiser
2014LRECThe SETimes.HR Linguistically Annotated Corpus of Croatian.Zeljko Agic, Nikola Ljubesic
2014LRECComparing two acquisition systems for automatically building an English-Croatian parallel corpus from multilingual websites.Miquel Espl-Gomis, Filip Klubicka, Nikola Ljubesic, Sergio Ortiz-Rojas, Vassilis Papavassiliou, Prokopis Prokopidis
2014LRECTweetCaT: a tool for building Twitter corpora of smaller languages.Nikola Ljubesic, Darja Fiser, Tomaz Erjavec
2014LRECcaWaC - A web corpus of Catalan and its application to language modeling and machine translation.Nikola Ljubesic, Antonio Toral
2014LRECQuality Estimation for Synthetic Parallel Data Generation.Raphal Rubino, Antonio Toral, Nikola Ljubesic, Gema Ramrez-Snchez
2012COLINGEfficient Discrimination Between Closely Related Languages.Jrg Tiedemann, Nikola Ljubesic
2012LRECAddressing polysemy in bilingual lexicon extraction from comparable corpora.Darja Fiser, Nikola Ljubesic, Ozren Kubelka
2011RANLPBilingual lexicon extraction from comparable corpora for closely related languages.Darja Fiser, Nikola Ljubesic
2010LRECTowards Sentiment Analysis of Financial Texts in Croatian.Zeljko Agic, Nikola Ljubesic, Marko Tadic
2010LRECBuilding a Gold Standard for Event Detection in Croatian.Nikola Ljubesic, Tomislava Lauc, Damir Boras
2008LRECGenerating a Morphological Lexicon of Organization Entity Names.Nikola Ljubesic, Tomislava Lauc, Damir Boras