| 2026 | SIGIR | ERASE - A Real-World Aligned Benchmark for Unlearning in Recommender Systems. | Pierre Sicco Lubitzsch, Maarten de Rijke, Sebastian Schelter |
| 2025 | EDBT | A Deep Dive Into Cross-Dataset Entity Matching with Large and Small Language Models. | Zeyu Zhang, Paul Groth, Iacer Calixto, Sebastian Schelter |
| 2025 | ICDE | Towards Regaining Control Over Messy Machine Learning Pipelines. | Stefan Grafberger, Hao Chen, Olga Ovcharenko, Sebastian Schelter |
| 2025 | ICDE | Navigating Data Errors in Machine Learning Pipelines: Identify, Debug, and Learn. | Bojan Karlas, Babak Salimi, Sebastian Schelter |
| 2025 | ICML | scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell Data. | Olga Ovcharenko, Florian Barkmann, Philip Toma, Imant Daunhawer, Julia E. Vogt, Sebastian Schelter, Valentina Boeva |
| 2025 | RecSys | Scalable Data Debugging for Neighborhood-based Recommendation with Data Shapley Values. | Barrie Kersbergen, Olivier Sprangers, Bojan Karlas, Maarten de Rijke, Sebastian Schelter |
| 2025 | SIGMOD | Navigating Data Errors in Machine Learning Pipelines: Identify, Debug, and Learn. | Bojan Karlas, Babak Salimi, Sebastian Schelter |
| 2024 | ICDE | Etude - Evaluating the Inference Latency of Session-Based Recommendation Models at Scale. | Barrie Kersbergen, Olivier Sprangers, Frank Kootte, Shubha Guha, Maarten de Rijke, Sebastian Schelter |
| 2024 | ICDE | Directions Towards Efficient and Automated Data Wrangling with Large Language Models. | Zeyu Zhang, Paul Groth, Iacer Calixto, Sebastian Schelter |
| 2024 | ICLR | Data Debugging with Shapley Importance over Machine Learning Pipelines. | Bojan Karlas, David Dao, Matteo Interlandi, Sebastian Schelter, Wentao Wu, Ce Zhang |
| 2023 | CHIIR | How to Make an Outlier? Studying the Effect of Presentational Features on the Outlierness of Items in Product Search Results. | Fatemeh Sarvi, Mohammad Aliannejadi, Sebastian Schelter, Maarten de Rijke |
| 2023 | CIDR | Reconstructing and Querying ML Pipeline Intermediates. | Sebastian Schelter |
| 2023 | ICDE | Automated Data Cleaning Can Hurt Fairness in Machine Learning-based Decision Making. | Shubha Guha, Falaah Arif Khan, Julia Stoyanovich, Sebastian Schelter |
| 2023 | WWW | Provenance Tracking for End-to-End Machine Learning Pipelines. | Stefan Grafberger, Paul Groth, Sebastian Schelter |
| 2023 | SIGIR | On the Impact of Outlier Bias on User Clicks. | Fatemeh Sarvi, Ali Vardasbi, Mohammad Aliannejadi, Sebastian Schelter, Maarten de Rijke |
| 2023 | SIGIR | Forget Me Now: Fast and Exact Unlearning in Neighborhood-based Recommendation. | Sebastian Schelter, Mozhdeh Ariannezhad, Maarten de Rijke |
| 2023 | SIGMOD | Proactively Screening Machine Learning Pipelines with ARGUSEYES. | Sebastian Schelter, Stefan Grafberger, Shubha Guha, Bojan Karlas, Ce Zhang |
| 2023 | WSDM | A Personalized Neighborhood-based Model for Within-basket Recommendation in Grocery Shopping. | Mozhdeh Ariannezhad, Ming Li, Sebastian Schelter, Maarten de Rijke |
| 2022 | CIDR | Screening Native Machine Learning Pipelines with ArgusEyes. | Sebastian Schelter, Stefan Grafberger, Shubha Guha, Olivier Sprangers, Bojan Karlas, Ce Zhang |
| 2022 | ICDE | GitSchemas: A Dataset for Automating Relational Data Preparation Tasks. | Till Dhmen, Madelon Hulsebos, Christian Beecks, Sebastian Schelter |
| 2022 | WWW | Understanding Financial Information Seeking Behavior from User Interactions with Company Filings. | Mozhdeh Ariannezhad, Mohamed Yahya, Edgar Meij, Sebastian Schelter, Maarten de Rijke |
| 2022 | SIGIR | ReCANet: A Repeat Consumption-Aware Neural Network for Next Basket Recommendation in Grocery Shopping. | Mozhdeh Ariannezhad, Sami Jullien, Ming Li, Min Fang, Sebastian Schelter, Maarten de Rijke |
| 2022 | SIGMOD | Serenade - Low-Latency Session-Based Recommendation in e-Commerce at Scale. | Barrie Kersbergen, Olivier Sprangers, Sebastian Schelter |
| 2022 | WSDM | Understanding and Mitigating the Effect of Outliers in Fair Ranking. | Fatemeh Sarvi, Maria Heuss, Mohammad Aliannejadi, Sebastian Schelter, Maarten de Rijke |
| 2021 | CIDR | Lightweight Inspection of Data Preprocessing in Native Machine Learning Pipelines. | Stefan Grafberger, Julia Stoyanovich, Sebastian Schelter |
| 2021 | CIKM | Understanding Multi-channel Customer Behavior in Retail. | Mozhdeh Ariannezhad, Sami Jullien, Pim Nauts, Min Fang, Sebastian Schelter, Maarten de Rijke |
| 2021 | EDBT | Automating Data Quality Validation for Dynamic Data Ingestion. | Sergey Redyuk, Zoi Kaoudi, Volker Markl, Sebastian Schelter |
| 2021 | EDBT | JENGA - A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models. | Sebastian Schelter, Tammo Rukat, Felix Biessmann |
| 2021 | ICDE | Learnings from a Retail Recommendation System on Billions of Interactions at bol.com. | Barrie Kersbergen, Sebastian Schelter |
| 2021 | KDD | Probabilistic Gradient Boosting Machines for Large-Scale Probabilistic Regression. | Olivier Sprangers, Sebastian Schelter, Maarten de Rijke |
| 2021 | SIGMOD | MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines. | Stefan Grafberger, Shubha Guha, Julia Stoyanovich, Sebastian Schelter |
| 2021 | SIGMOD | HedgeCut: Maintaining Randomised Trees for Low-Latency Machine Unlearning. | Sebastian Schelter, Stefan Grafberger, Ted Dunning |
| 2020 | CIDR | "Amnesia" - Machine Learning Models That Can Forget User Data Very Fast. | Sebastian Schelter |
| 2020 | DAC | Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost. | Ji Zhang, Yuanzhang Wang, Yangtao Wang, Ke Zhou, Sebastian Schelter, Ping Huang, Bin Cheng, Yongguang Ji |
| 2020 | EDBT | Zooming Out on an Evolving Graph. | Amir Aghasadeghi, Vera Zaychik Moffitt, Sebastian Schelter, Julia Stoyanovich |
| 2020 | EDBT | Towards Unsupervised Data Quality Validation on Dynamic Data. | Sergey Redyuk, Volker Markl, Sebastian Schelter |
| 2020 | EDBT | FairPrep: Promoting Data to a First-Class Citizen in Studies on Fairness-Enhancing Interventions. | Sebastian Schelter, Yuxuan He, Jatin Khilnani, Julia Stoyanovich |
| 2020 | RecSys | Three Challenges in Building Industrial-Scale Recommender Systems. | Sebastian Schelter |
| 2020 | SIGMOD | Elastic Machine Learning Algorithms in Amazon SageMaker. | Edo Liberty, Zohar S. Karnin, Bing Xiang, Laurence Rouesnel, Baris Coskun, Ramesh Nallapati, Julio Delgado, Amir Sadoughi, Yury Astashonok, Piali Das, Can Balioglu, Saswata Chakravarty, Madhav Jha, Philip Gautier, David Arpin, Tim Januschowski, Valentin Flunkert, Yuyang Wang, Jan Gasthaus, Lorenzo Stella, Syama Sundar Rangapuram, David Salinas, Sebastian Schelter, Alex Smola |
| 2020 | SIGMOD | Learning to Validate the Predictions of Black Box Classifiers on Unseen Data. | Sebastian Schelter, Tammo Rukat, Felix Biemann |
| 2020 | USENIX | HDDse: Enabling High-Dimensional Disk State Embedding for Generic Failure Detection System of Heterogeneous Disks in Large Data Centers. | Ji Zhang, Ping Huang, Ke Zhou, Ming Xie, Sebastian Schelter |
| 2019 | ICDE | Differential Data Quality Verification on Partitioned Data. | Sebastian Schelter, Stefan Grafberger, Philipp Schmidt, Tammo Rukat, Mario Kieling, Andrey Taptunov, Felix Biemann, Dustin Lange |
| 2019 | SIGMOD | Learning to Validate the Predictions of Black Box Machine Learning Models on Unseen Data. | Sergey Redyuk, Sebastian Schelter, Tammo Rukat, Volker Markl, Felix Biemann |
| 2019 | SIGMOD | Unit Testing Data with Deequ. | Sebastian Schelter, Felix Biemann, Dustin Lange, Tammo Rukat, Philipp Schmidt, Stephan Seufert, Pierre Brunelle, Andrey Taptunov |
| 2019 | SIGMOD | DEEM 2019: Workshop on Data Management for End-to-End Machine Learning. | Sebastian Schelter, Neoklis Polyzotis, Manasi Vartak, Stephan Seufert |
| 2019 | SSDBM | Efficient Incremental Cooccurrence Analysis for Item-Based Collaborative Filtering. | Sebastian Schelter, Ufuk Celebi, Ted Dunning |
| 2018 | CIKM | "Deep" Learning for Missing Value Imputationin Tables with Non-Numerical Data. | Felix Biemann, David Salinas, Sebastian Schelter, Philipp Schmidt, Dustin Lange |
| 2017 | BTW | Gilbert: Declarative Sparse Linear Algebra on Massively Parallel Dataflow Systems. | Till Rohrmann, Sebastian Schelter, Tilmann Rabl, Volker Markl |
| 2016 | IC2E | Apache Flink: Stream Analytics at Scale. | Asterios Katsifodimos, Sebastian Schelter |
| 2016 | ICDM | Structural Patterns in the Rise of Germany's New Right on Facebook. | Sebastian Schelter, Felix Biemann, Malisa Zobel, Nedelina Teneva |
| 2016 | ICWSM | Tracking the Trackers: A Large-Scale Analysis of Embedded Web Trackers. | Sebastian Schelter, Jrme Kunegis |
| 2015 | ICDE | Efficient sample generation for scalable meta learning. | Sebastian Schelter, Juan Soto, Volker Markl, Douglas Burdick, Berthold Reinwald, Alexandre V. Evfimievski |
| 2015 | SIGMOD | Optimistic Recovery for Iterative Dataflows in Action. | Sergey Dudoladov, Chen Xu, Sebastian Schelter, Asterios Katsifodimos, Stephan Ewen, Kostas Tzoumas, Volker Markl |
| 2014 | SIGMOD | Scaling data mining in massively parallel dataflow systems. | Sebastian Schelter |
| 2013 | CIKM | "All roads lead to Rome": optimistic recovery for distributed iterative data processing. | Sebastian Schelter, Stephan Ewen, Kostas Tzoumas, Volker Markl |
| 2013 | RecSys | Distributed matrix factorization with mapreduce using a series of broadcast-joins. | Sebastian Schelter, Christoph Boden, Martin Schenck, Alexander Alexandrov, Volker Markl |
| 2013 | SIGMOD | Iterative parallel data processing with stratosphere: an inside look. | Stephan Ewen, Sebastian Schelter, Kostas Tzoumas, Daniel Warneke, Volker Markl |
| 2012 | RecSys | Scalable similarity-based neighborhood methods with MapReduce. | Sebastian Schelter, Christoph Boden, Volker Markl |