| 2024 | JCoLA: Japanese Corpus of Linguistic Acceptability. | Taiga Someya, Yushi Sugimoto, Yohei Oseki |
| 2024 | At the Crossroad of Cuneiform and NLP: Challenges for Fine-grained Part-of-speech Tagging. | Gustav Ryberg Smidt, Els Lefever, Katrien de Graef |
| 2024 | Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks. | Jakub Smd, Pavel Pribn, Ondrej Prazk, Pavel Krl |
| 2024 | CuSINeS: Curriculum-driven Structure Induced Negative Sampling for Statutory Article Retrieval. | T. Y. S. S. Santosh, Kristina Kaiser, Matthias Grabmair |
| 2024 | New Datasets for Automatic Detection of Textual Entailment and of Contradictions between Sentences in French. | Maximos Skandalis, Richard Moot, Christian Retor, Simon Robillard |
| 2024 | A Typology of Errors for User Utterances in Chatbots. | Anu Singh, Esm Manandise |
| 2024 | EROS: Entity-Driven Controlled Policy Document Summarization. | Joykirat Singh, Sehban Fazili, Rohan Jain, Md. Shad Akhtar |
| 2024 | Generating Clarification Questions for Disambiguating Contracts. | Anmol Singhal, Chirag Jain, Preethu Rose Anish, Arkajyoti Chakraborty, Smita Ghaisas |
| 2024 | Reconstruction of Cuneiform Literary Texts as Text Matching. | Fabian Simonjetz, Jussi Laasonen, Yunus Cobanoglu, Alexander Fraser, Enrique Jimnez |
| 2024 | PPORTAL_ner: An Annotated Corpus of Portuguese Literary Entities. | Mariana O. Silva, Mirella M. Moro |
| 2024 | Generating Multiple-choice Questions for Medical Question Answering with Distractors and Cue-masking. | Damien Sileo, Kanimozhi Uma, Marie-Francine Moens |
| 2024 | tasksource: A Large Collection of NLP tasks with a Structured Dataset Preprocessing Framework. | Damien Sileo |
| 2024 | The Low Saxon LSDC Dataset at Universal Dependencies. | Janine Siewert, Jack Rueter |
| 2024 | SPICED: News Similarity Detection Dataset with Multiple Topics and Complexity Levels. | Elena Shushkevich, Long Thanh Mai, Manuel V. Loureiro, Steven Derby, Tri Kurniawan Wijaya |
| 2024 | The Effects of Pretraining in Video-Guided Machine Translation. | Ammon Shurtz, Lawry Sorenson, Stephen D. Richardson |
| 2024 | Continual Reinforcement Learning for Controlled Text Generation. | Velizar Shulev, Khalil Sima'an |
| 2024 | Negation Triplet Extraction with Syntactic Dependency and Semantic Consistency. | Yuchen Shi, Deqing Yang, Jingping Liu, Yanghua Xiao, Zongyu Wang, Huimin Xu |
| 2024 | Generative Multimodal Entity Linking. | Senbao Shi, Zhenran Xu, Baotian Hu, Min Zhang |
| 2024 | Analyzing the Dynamics of Climate Change Discourse on Twitter: A New Annotated Corpus and Multi-Aspect Classification. | Shuvam Shiwakoti, Surendrabikram Thapa, Kritesh Rauniyar, Akshyat Shah, Aashish Bhandari, Usman Naseem |
| 2024 | Deconstructing In-Context Learning: Understanding Prompts via Corruption. | Namrata Shivagunde, Vladislav Lialin, Sherin Muckatira, Anna Rumshisky |
| 2024 | QA-based Event Start-Points Ordering for Clinical Temporal Relation Annotation. | Seiji Shimizu, Lis Pereira, Shuntaro Yada, Eiji Aramaki |
| 2024 | Phonotactic Complexity across Dialects. | Ryan Soh-Eun Shim, Kalvin Chang, David R. Mortensen |
| 2024 | Find-the-Common: A Benchmark for Explaining Visual Patterns from Images. | Yuting Shi, Naoya Inoue, Houjing Wei, Yufeng Zhao, Tao Jin |
| 2024 | UzbekVerbDetection: Rule-based Detection of Verbs in Uzbek Texts. | Maksud Sharipov, Elmurod Kuriyozov, Ollabergan Yuldashov, Ogabek Sobirov |
| 2024 | CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation. | Yujie Shao, Xinrong Yao, Xingwei Qu, Chenghua Lin, Shi Wang, Wenhao Huang, Ge Zhang, Jie Fu |