| 2025 | OpenMent: A Dataset of Mentor-Mentee Interactions in Google Summer of Code. | Erfan Raoofian, Fatemeh H. Fard, Ifeoma Adaji, Gema Rodrguez-Prez |
| 2025 | Investigating the Understandability of Review Comments on Code Change Requests. | Md Shamimur Rahman, Zadia Codabux, Chanchal K. Roy |
| 2025 | Revisiting Defects4J for Fault Localization in Diverse Development Scenarios. | Md Nakhla Rafi, An Ran Chen, Tse-Hsun Peter Chen, Shaohua Wang |
| 2025 | Understanding Software Vulnerabilities in the Maven Ecosystem: Patterns, Timelines, and Risks. | Md. Fazle Rabbi, Rajshakhar Paul, Arifa Islam Champa, Minhaz F. Zibran |
| 2025 | Chasing the Clock: How Fast Are Vulnerabilities Fixed in the Maven Ecosystem? | Md. Fazle Rabbi, Arifa Islam Champa, Rajshakhar Paul, Minhaz F. Zibran |
| 2025 | EvoChain: A Framework for Tracking and Visualizing Smart Contract Evolution. | Ilham A. Qasse, Mohammad Hamdaqa, Bjrn r Jnsson |
| 2025 | HaPy-Bug - Human Annotated Python Bug Resolution Dataset. | Piotr Przymus, Mikolaj Fejzer, Jakub Narebski, Radoslaw Wozniak, Lukasz Halada, Aleksander Kazecki, Mykhailo Molchanov, Krzysztof Stencel |
| 2025 | Out of Sight, Still at Risk: The Lifecycle of Transitive Vulnerabilities in Maven. | Piotr Przymus, Mikolaj Fejzer, Jakub Narebski, Krzysztof Rykaczewski, Krzysztof Stencel |
| 2025 | Wolves in the Repository: A Software Engineering Analysis of the XZ Utils Supply Chain Attack. | Piotr Przymus, Thomas Durieux |
| 2025 | Human-In-The-Loop Software Development Agents: Challenges and Future Directions. | Jirat Pasuksmit, Wannita Takerngsaksiri, Patanamon Thongtanunam, Chakkrit Tantithamthavorn, Ruixiong Zhang, Shiyan Wang, Fan Jiang, Jing Li, Evan Cook, Kun Chen, Ming Wu |
| 2025 | CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance. | Kunal Suresh Pai, Premkumar T. Devanbu, Toufique Ahmed |
| 2025 | Characterizing Packages for Vulnerability Prediction. | Saviour Owolabi, Francesco Rosati, Ahmad Abdellatif, Lorenzo De Carli |
| 2025 | Software Composition Analysis and Supply Chain Security in Apache Projects: an Empirical Study. | Sabato Nocera, Sira Vegas, Giuseppe Scanniello, Natalia Juristo |
| 2025 | Are the Majority of Public Computational Notebooks Pathologically Non-Executable? | Tien Nguyen, Waris Gill, Muhammad Ali Gulzar |
| 2025 | How Effective are LLMs for Data Science Coding? A Controlled Experiment. | Nathalia Nascimento, Everton Guimares, Sai Sanjna Chintakunta, Santhosh Anitha Boominathan |
| 2025 | Jupyter Notebook Activity Dataset. | Tomoki Nakamaru, Tomomasa Matsunaga, Tetsuro Yamazaki |
| 2025 | Decoding Dependency Risks: A Quantitative Study of Vulnerabilities in the Maven Ecosystem. | Costain Nachuma, Md Mosharaf Hossan, Asif Kamal Turzo, Minhaz F. Zibran |
| 2025 | GHALogs: Large-Scale Dataset of GitHub Actions Runs. | Florent Moriconi, Thomas Durieux, Jean-Rmy Falleri, Raphal Troncy, Aurlien Francillon |
| 2025 | E2EGit: A Dataset of End-to-End Web Tests in Open Source Projects. | Sergio Di Meglio, Luigi Libero Lucio Starace, Valeria Pontillo, Ruben Opdebeeck, Coen De Roover, Sergio Di Martino |
| 2025 | Does Functional Package Management Enable Reproducible Builds at Scale? Yes. | Julien Malka, Stefano Zacchiroli, Tho Zimmermann |
| 2025 | ICVul: A Well-labeled C/C++ Vulnerability Dataset with Comprehensive Metadata and VCCs. | Chaomeng Lu, Tianyu Li, Toon Dehaene, Bert Lagaisse |
| 2025 | Too Noisy To Learn: Enhancing Data Quality for Code Review Comment Generation. | Chunhua Liu, Hong Yi Lin, Patanamon Thongtanunam |
| 2025 | From Industrial Practices to Academia: Uncovering the Gap in Vulnerability Research and Practice. | Zhuang Liu, Xing Hu, Jiayuan Zhou, Xin Xia |
| 2025 | DataTD: A Dataset of Java Projects Including Test Doubles. | Mengzhen Li, Mattia Fazzini |
| 2025 | Analyzing Dependency Clusters and Security Risks in the Maven Central Repository. | George Lake, Minhaz F. Zibran |