| 2024 | Spatially Augmented Speech Bubble to Character Association via Comic Multi-task Learning. | Grkan Soykan, Deniz Yuret, Tevfik Metin Sezgin |
| 2024 | A Comprehensive Gold Standard and Benchmark for Comics Text Detection and Recognition. | Grkan Soykan, Deniz Yuret, Tevfik Metin Sezgin |
| 2024 | CICA: Content-Injected Contrastive Alignment for Zero-Shot Document Image Classification. | Sankalp Sinha, Muhammad Saif Ullah Khan, Talha Uddin Sheikh, Didier Stricker, Muhammad Zeshan Afzal |
| 2024 | The Learnable Typewriter: A Generative Approach to Text Analysis. | Ioannis Siglidis, Nicolas Gonthier, Julien Gaubil, Tom Monnier, Mathieu Aubry |
| 2024 | WikiDT: Visual-Based Table Recognition and Question Answering Dataset. | Hui Shi, Yusheng Xie, Luis Goncalves, Sicun Gao, Jishen Zhao |
| 2024 | Cross-Domain Image Conversion by CycleDM. | Sho Shimotsumagari, Shumpei Takezaki, Daichi Haraguchi, Seiichi Uchida |
| 2024 | Towards End-to-End Semi-supervised Table Detection with Semantic Aligned Matching Transformer. | Tahira Shehzadi, Shalini Sarode, Didier Stricker, Muhammad Zeshan Afzal |
| 2024 | A Hybrid Approach for Document Layout Analysis in Document Images. | Tahira Shehzadi, Didier Stricker, Muhammad Zeshan Afzal |
| 2024 | SlideCraft: Synthetic Slides Generation for Robust Slide Analysis. | Travis Seng, Axel Carlier, Thomas Forgione, Vincent Charvillat, Wei Tsang Ooi |
| 2024 | Are Layout Analysis and OCR Still Useful for Document Information Extraction Using Foundation Models? | Anna Scius-Bertrand, Atefeh Fakhari, Lars Vgtlin, Daniel Ribeiro Cabral, Andreas Fischer |
| 2024 | DocXplain: A Novel Model-Agnostic Explainability Method for Document Image Classification. | Saifullah Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed |
| 2024 | Source-Free Domain Adaptation for Optical Music Recognition. | Adrian Rosello, Eliseo Fuentes-Martnez, Mara Alfaro-Contreras, David Rizo, Jorge Calvo-Zaragoza |
| 2024 | Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription. | Antonio Ros-Vila, Jorge Calvo-Zaragoza, Thierry Paquet |
| 2024 | Toward Accessible Comics for Blind and Low Vision Readers. | Christophe Rigaud, Jean-Christophe Burie, Samuel Petit |
| 2024 | StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation. | Nauman Riaz, Saifullah Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed |
| 2024 | Enhancing CRNN HTR Architectures with Transformer Blocks. | George Retsinas, Konstantina Nikolaidou, Giorgos Sfikas |
| 2024 | The KuiSCIMA Dataset for Optical Music Recognition of Ancient Chinese Suzipu Notation. | Tristan Repolusk, Eduardo E. Veas |
| 2024 | A Multimodal Framework For Structuring Legal Documents. | Thibaud Real, Pauline Chavallard |
| 2024 | Self-supervised Vision Transformers for Writer Retrieval. | Tim Raven, Arthur Matei, Gernot A. Fink |
| 2024 | Binarizing Documents by Leveraging both Space and Frequency. | Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Rita Cucchiara |
| 2024 | Weakly Supervised Training for Hologram Verification in Identity Documents. | Glen Pouliquen, Guillaume Chiron, Joseph Chazalon, Thierry Graud, Ahmad Montaser Awal |
| 2024 | ClusterTabNet: Supervised Clustering Method for Table Detection and Table Structure Recognition. | Marek Polewczyk, Marco Spinaci |
| 2024 | LayeredDoc: Domain Adaptive Document Restoration with a Layer Separation Approach. | Maria Pilligua, Nil Biescas, Javier Vazquez-Corral, Josep Llads, Ernest Valveny, Sanket Biswas |
| 2024 | Typographic Text Generation with Off-the-Shelf Diffusion Model. | KhayTze Peong, Seiichi Uchida, Daichi Haraguchi |
| 2024 | SAGHOG: Self-supervised Autoencoder for Generating HOG Features for Writer Retrieval. | Marco Peer, Florian Kleber, Robert Sablatnig |