17 papers · ranked by Valyu relevance
Vedangi Wagh, Snehal Khandve, Isha Joshi, Apurva Wani + 2 more
'Geetanjali Kale' 'Raviraj Joshi'] > Abstract. The amount of information stored in the form of documents on the internet has been increasing rapidly. Thus it has become a necessity to organize and maintain these documents in an optimum manner. Text classification algorithms study the complex relationships between words…
Sadia Zaman Mishu, S M Rafiuddin
The demand for text classification is growing significantly in web searching, data mining, web ranking, recommendation systems, and so many other fields of information and technology. This paper illustrates the text classification process on different dataset using some standard supervised machine learning techniques.…
Jordy Van Landeghem, Sanket Biswas, Matthew B. Blaschko, Marie‐Francine Moens
'Marie‐Francine Moens'] This paper highlights the need to bring document classification benchmarking closer to real-world applications, both in the nature of data tested (X: multi-channel, multipaged, multi-industry; Y : class distributions and label set variety) and in classification tasks considered (f: multipage…
Ciaran Cooney, Joana Cavadas, Liam Madigan, Bradley Savage + 2 more
'Rachel Heyburn' "Mairead O'Cuinn"] We propose end-to-end document classification and key information extraction (KIE) for automating document processing in forms. Through accurate document classification we harness known information from templates to enhance KIE from forms. We use text and layout encoding with a…
Narayanan Arvind
In the shipping industry, document classification plays a crucial role in ensuring that the necessary documents are properly identified and processed for customs clearance. OCR technology is being used to automate the process of document classification, which involves identifying important documents such as Commercial…
Fnu Mohbat, Mohammed J. Zaki, Catherine Finegan-Dollak, Ashish Verma
The robustness of a model for real-world deployment is decided by how well it performs on unseen data and distinguishes between in-domain and out-of-domain samples. Visual document classiers have shown impressive performance on in-distribution test sets. However, they tend to have a hard time correctly classifying and…
Sumaia Mohammed Al-Ghuribi, Shahrul Azman Mohd Noah
The rapid growth of the internet has increased the number of online texts. This led to the rapid growth of the number of online texts in the Arabic language. The enormous amount of text must be organized into classes to make the analysis process and text retrieval easier. Text classification is, therefore, a key…
Amin Sadri, M Maruf Hossain
Contribution: In this paper, we introduce Coordinate Matrix Machine (CM 2 ), a structureaware document intelligence model designed to bridge the gap between semantic content and spatial layout. Unlike traditional NLP approaches that prioritize sequential text, CM 2 is particularly effective for low-data…
Zhijian Li, Stefan Larson, Kevin Leach
Rapid document classification is critical in several time-sensitive applications like digital forensics and large-scale media classification. Traditional approaches that rely on heavy-duty deep learning models fall short due to high inference times over vast input datasets and computational resources associated with…
Lei Kang, Mohamed Ali Souibgui, Fei Yang, Lluís Gómez + 2 more
'Ernest Valveny' 'Dìmosthenis Karatzas'] Abstract. Document understanding models have recently demonstrated remarkable performance by leveraging extensive collections of user documents. However, since documents often contain large amounts of personal data, their usage can pose a threat to user privacy and weaken the…
Stefanie Schwaar, Franziska Diez, Michael Trebing, Nils Witznick
In German public administration, there are 45 different offices to which incoming messages need to be distributed. Since these messages are often unstructured, the system has to be based at least partly on message content. For public service no data are given so far and no pretrained model is available. The data we…
Stefan Larson, Gordon Lim, Kevin Leach
The RVL-CDIP benchmark is widely used for measuring performance on the task of document classification. Despite its widespread use, we reveal several undesirable characteristics of the RVL-CDIP benchmark. These include (1) substantial amounts of label noise, which we estimate to be 8.1% (ranging between 1.6% to 16.9%…
Yoshinari Fujinuma, Siddharth Varia, Nishant Sankaran, Srikar Appalaraju + 2 more
'Srikar Appalaraju' 'Bonan Min' 'Yogarshi Vyas'] Document image classification is different from plain-text document classification and consists of classifying a document by understanding the content and structure of documents such as forms, emails, and other such documents. We show that the only existing dataset for…
Jawid Ahmad Baktash, Mursal Dawodi
Today text classification becomes critical task for concerned individuals for numerous purposes. Hence, several researches have been conducted to develop automatic text classification for national and international languages. However, the need for an automatic text categorization system for local languages is felt. The…
Tanmoy Mondal, Abhijit Das, Zuheng Ming
In this work, we adhere to explore a Multi-Tasking learning (MTL) based network to perform document attribute classification such as the font type, font size, font emphasis and scanning resolution classification of a document image. To accomplish these tasks, we operate on either segmented word level or on uniformed…
Taylor Archibald, Tony Martinez
Form Classification Authors: ['Taylor Archibald' 'Tony Martinez'] Abstract. Efficient categorization of historical documents is crucial for fields such as genealogy, legal research, and historical scholarship, where manual classification is impractical for large collections due to its laborintensive and error-prone…
Mohsen Ahmadi, Matin Khajavi, Abbas Varmaghani, Ali Ala + 2 more
Detection with Robust and Context-Aware Text Classification Authors: ['Mohsen Ahmadi' 'Matin Khajavi' 'Abbas Varmaghani' 'Ali Ala' 'Kasra Danesh' 'Danial Javaheri'] Abstract—This study evaluates the effectiveness of different feature extraction techniques and classification algorithms in detecting spam messages within…