Nikola Ljubesic
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
44
Venues
9
Active years
2008–2026
Best venue rank
A*
Where they publish
Papers
44 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2026 | LREC | ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian. | Nikola Ljubesic, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungersek |
| 2026 | LREC | The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora. | Taja Kuzman Pungersek, Peter Rupnik, Vit Suchomel, Nikola Ljubesic |
| 2026 | LREC | ROG: A Multi-Layer Manually Annotated Corpus of Spoken Slovenian. | Kaja Dobrovoljc Zor, Darinka Verdonik, Jaka Cibej, Peter Rupnik, Nikola Ljubesic |
| 2025 | COLING | Proceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects. | Yves Scherrer, Tommi Jauhiainen, Nikola Ljubesic, Preslav Nakov, Jrg Tiedemann, Marcos Zampieri |
| 2025 | ECIR | Overview of Touch 2025: Argumentation Systems - Extended Abstract. | Johannes Kiesel, agri ltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Frbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matys Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein |
| 2025 | Interspeech | Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models. | Nikola Ljubesic, Ivan Porupski, Peter Rupnik |
| 2024 | COLING | A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection. | Filip Dobranic, Bojan Evkoski, Nikola Ljubesic |
| 2024 | COLING | CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation. | Nikola Ljubesic, Taja Kuzman |
| 2024 | COLING | The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings. | Michal Mochtak, Peter Rupnik, Nikola Ljubesic |
| 2024 | COLING | Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages. | Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljubesic, Miquel Espl-Gomis, Gema Ramrez-Snchez, Antonio Toral |
| 2024 | COLING | Gos 2: A New Reference Corpus of Spoken Slovenian. | Darinka Verdonik, Kaja Dobrovoljc, Tomaz Erjavec, Nikola Ljubesic |
| 2024 | ECIR | Overview of Touch 2024: Argumentation Systems. | Johannes Kiesel, agri ltekin, Maximilian Heinrich, Maik Frbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matys Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Mnstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein |
| 2024 | NAACL | Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark. | Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Brje Karlsson, Peiqin Lin, Nikola Ljubesic, Lester James V. Miranda, Barbara Plank, Arij Riabi, Yuval Pinter |
| 2023 | EAMT | MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages. | Marta Ban, Malina Chichirau, Miquel Espl-Gomis, Mikel L. Forcada, Aarn Galiano Jimnez, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramrez-Snchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Jaume Zaragoza-Bernabeu |
| 2022 | EAMT | MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages. | Marta Ban, Miquel Espl-Gomis, Mikel L. Forcada, Cristian Garca-Romero, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramrez-Snchez, Peter Rupnik, Vt Suchomel, Antonio Toral, Tobias van der Werff, Jaume Zaragoza |
| 2022 | LREC | The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild. | Taja Kuzman, Peter Rupnik, Nikola Ljubesic |
| 2021 | RANLP | Cultural Topic Modelling over Novel Wikipedia Corpora for South-Slavic Languages. | Filip Markoski, Elena Markoska, Nikola Ljubesic, Eftim Zdravevski, Ljupco Kocarev |
| 2020 | LREC | CoSimLex: A Resource for Evaluating Graded Word Similarity in Context. | Carlos Santos Armendariz, Matthew Purver, Matej Ulcar, Senja Pollak, Nikola Ljubesic, Mark Granroth-Wilding |
| 2020 | LREC | Gigafida 2.0: The Reference Corpus of Written Standard Slovene. | Simon Krek, Spela Arhar Holdt, Tomaz Erjavec, Jaka Cibej, Andraz Repar, Polona Gantar, Nikola Ljubesic, Iztok Kosem, Kaja Dobrovoljc |
| 2018 | ACL | Bleaching Text: Abstract Features for Cross-lingual Gender Prediction. | Rob van der Goot, Nikola Ljubesic, Ian Matroos, Malvina Nissim, Barbara Plank |
| 2016 | COLING | TweetGeo - A Tool for Collecting, Processing and Analysing Geo-encoded Linguistic Data. | Nikola Ljubesic, Tanja Samardzic, Curdin Derungs |
| 2016 | EAMT | Collaborative Development of a Rule-Based Machine Translator between Croatian and Serbian. | Filip Klubicka, Gema Ramrez-Snchez, Nikola Ljubesic |
| 2016 | EAMT | Dealing with Data Sparseness in SMT with Factured Models and Morphological Expansion: a Case Study on Croatian. | Vctor M. Snchez-Cartagena, Nikola Ljubesic, Filip Klubicka |
| 2016 | EAMT | Abu-MaTran: automatic building of machine translation. | Antonio Toral, Sergio Ortiz-Rojas, Mikel L. Forcada, Nikola Ljubesic, Prokopis Prokopidis |
| 2016 | LREC | Corpus vs. Lexicon Supervision in Morphosyntactic Tagging: the Case of Slovene. | Nikola Ljubesic, Tomaz Erjavec |
| 2016 | LREC | Corpus-Based Diacritic Restoration for South Slavic Languages. | Nikola Ljubesic, Tomaz Erjavec, Darja Fiser |
| 2016 | LREC | Producing Monolingual and Parallel Web Corpora at the Same Time - SpiderLing and Bitextor's Love Affair. | Nikola Ljubesic, Miquel Espl-Gomis, Antonio Toral, Sergio Ortiz-Rojas, Filip Klubicka |
| 2016 | LREC | New Inflectional Lexicons and Training Corpora for Improved Morphosyntactic Annotation of Croatian and Serbian. | Nikola Ljubesic, Filip Klubicka, Zeljko Agic, Ivo-Pavao Jazbec |
| 2016 | LREC | Croatian Error-Annotated Corpus of Non-Professional Written Language. | Vanja Stefanec, Nikola Ljubesic, Jelena Kuvac Kraljevic |
| 2015 | EAMT | Abu-MaTran: Automatic building of Machine Translation. | Antonio Toral, Flammie A. Pirinen, Andy Way, Gema Ramrez-Snchez, Sergio Ortiz-Rojas, Raphael Rubino, Miquel Espl-Gomis, Mikel L. Forcada, Vassilis Papavassiliou, Prokopis Prokopidis, Nikola Ljubesic |
| 2015 | RANLP | Predicting Inflectional Paradigms and Lemmata of Unknown Words for Semi-automatic Expansion of Morphological Lexicons. | Nikola Ljubesic, Miquel Espl-Gomis, Filip Klubicka, Nives Mikelic Preradovic |
| 2015 | RANLP | Predicting the Level of Text Standardness in User-generated Content. | Nikola Ljubesic, Darja Fiser, Tomaz Erjavec, Jaka Cibej, Dafne Marko, Senja Pollak, Iza Skrjanec |
| 2014 | CICLING | Standardizing Tweets with Character-Level Machine Translation. | Nikola Ljubesic, Tomaz Erjavec, Darja Fiser |
| 2014 | LREC | The SETimes.HR Linguistically Annotated Corpus of Croatian. | Zeljko Agic, Nikola Ljubesic |
| 2014 | LREC | Comparing two acquisition systems for automatically building an English-Croatian parallel corpus from multilingual websites. | Miquel Espl-Gomis, Filip Klubicka, Nikola Ljubesic, Sergio Ortiz-Rojas, Vassilis Papavassiliou, Prokopis Prokopidis |
| 2014 | LREC | TweetCaT: a tool for building Twitter corpora of smaller languages. | Nikola Ljubesic, Darja Fiser, Tomaz Erjavec |
| 2014 | LREC | caWaC - A web corpus of Catalan and its application to language modeling and machine translation. | Nikola Ljubesic, Antonio Toral |
| 2014 | LREC | Quality Estimation for Synthetic Parallel Data Generation. | Raphal Rubino, Antonio Toral, Nikola Ljubesic, Gema Ramrez-Snchez |
| 2012 | COLING | Efficient Discrimination Between Closely Related Languages. | Jrg Tiedemann, Nikola Ljubesic |
| 2012 | LREC | Addressing polysemy in bilingual lexicon extraction from comparable corpora. | Darja Fiser, Nikola Ljubesic, Ozren Kubelka |
| 2011 | RANLP | Bilingual lexicon extraction from comparable corpora for closely related languages. | Darja Fiser, Nikola Ljubesic |
| 2010 | LREC | Towards Sentiment Analysis of Financial Texts in Croatian. | Zeljko Agic, Nikola Ljubesic, Marko Tadic |
| 2010 | LREC | Building a Gold Standard for Event Detection in Croatian. | Nikola Ljubesic, Tomislava Lauc, Damir Boras |
| 2008 | LREC | Generating a Morphological Lexicon of Organization Entity Names. | Nikola Ljubesic, Tomislava Lauc, Damir Boras |