Skip to content

Peter Rupnik

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

9

Venues

4

Active years

2022–2026

Best venue rank

B

Where they publish

Papers

9 indexed papers, newest first.

YearVenueTitleAuthors
2026LRECParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian.Nikola Ljubesic, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungersek
2026LRECThe Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora.Taja Kuzman Pungersek, Peter Rupnik, Vit Suchomel, Nikola Ljubesic
2026LRECROG: A Multi-Layer Manually Annotated Corpus of Spoken Slovenian.Kaja Dobrovoljc Zor, Darinka Verdonik, Jaka Cibej, Peter Rupnik, Nikola Ljubesic
2025InterspeechIdentifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models.Nikola Ljubesic, Ivan Porupski, Peter Rupnik
2024COLINGThe ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings.Michal Mochtak, Peter Rupnik, Nikola Ljubesic
2024COLINGDo Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages.Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljubesic, Miquel Espl-Gomis, Gema Ramrez-Snchez, Antonio Toral
2023EAMTMaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages.Marta Ban, Malina Chichirau, Miquel Espl-Gomis, Mikel L. Forcada, Aarn Galiano Jimnez, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramrez-Snchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Jaume Zaragoza-Bernabeu
2022EAMTMaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages.Marta Ban, Miquel Espl-Gomis, Mikel L. Forcada, Cristian Garca-Romero, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramrez-Snchez, Peter Rupnik, Vt Suchomel, Antonio Toral, Tobias van der Werff, Jaume Zaragoza
2022LRECThe GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild.Taja Kuzman, Peter Rupnik, Nikola Ljubesic