26 papers · ranked by Valyu relevance
Loris Nanni, Michelangelo Paci, Sheryl Brahnam, Alessandra Lumini + 2 more
Convolutional neural networks (CNNs) have gained prominence in the research literature on image classification over the last decade. One shortcoming of CNNs, however, is their lack of generalizability and tendency to overfit when presented with small training sets. Augmentation directly confronts this problem by…
João Fonseca, Fernando Bação
In the Machine Learning research community, there is a consensus regarding the relationship between model complexity and the required amount of data and computation power. In real world applications, these computational requirements are not always available, motivating research on regularization methods. In addition…
Khaled Alomar, Halil Ibrahim Aysel, Xiaohao Cai, Jérôme Gilles + 1 more
'Luminiţa Moraru'] In the past decade, deep neural networks, particularly convolutional neural networks, have revolutionised computer vision. However, all deep learning models may require a large amount of data so as to achieve satisfying results. Unfortunately, the availability of sufficient amounts of data for…
Kumar, Saorj, Prince Asiamah, Oluwatoyin Jolaoso + 1 more
—Convolutional Neural Networks (CNNs) serve as the workhorse of deep learning, finding applications in various fields that rely on images. Given sufficient data, they exhibit the capacity to learn a wide range of concepts across diverse settings. However, a notable limitation of CNNs is their susceptibility to…
Z. Wang, Pengfei Wang, Kunpeng Liu, Pengyang Wang + 5 more
'Chang‐Tien Lu' 'Charų C. Aggarwal' 'Jian Pei' 'Yuanchun Zhou'] ZAITIAN WANG∗ and PENGFEI WANG∗ , Computer Network Information Center, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Chinese Academy of Sciences, China KUNPENG LIU, Portland State University, USA PENGYANG WANG, University of…
Melissa Valaee, Shahram Shirani
With a 17.9 million annual mortality rate, cardiovascular disease is the leading global cause of death. As such, early detection and disease diagnosis are critical for effective treatment and symptom management. Cardiac auscultation, the process of listening to the heartbeat, often provides the first indication of…
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song
Due to progression in cell-cycle or duration of storage, classification of morphological changes in human blood cells is important for correct and effective clinical decisions. Automated classification systems help avoid subjective outcomes and are more efficient. Deep learning and more specifically Convolutional…
Domagoj Pluščec, Jan Šnajder
—Data scarcity is a problem that occurs in languages and tasks where we do not have large amounts of labeled data but want to use state-of-the-art models. Such models are often deep learning models that require a significant amount of data to train. Acquiring data for various machine learning problems is accompanied by…
Wei Wang, Zhaowei Shang, Chengxing Li
Data augmentation is an effective technique for automatically expanding training data in deep learning. Brain-inspired methods are approaches that draw inspiration from the functionality and structure of the human brain and apply these mechanisms and principles to artificial intelligence and computer science. When…
Ji-Young Yoon, Gahgene Gweon, Yun Joo Yoo, Peida Zhan
Over recent decades, machine learning, an integral subfield of artificial intelligence, has revolutionized diverse sectors, enabling data-driven decisions with minimal human intervention. In particular, the field of educational assessment emerges as a promising area for machine learning applications, where students can…
Alhassan Mumuni, Fuseini Mumuni
performance comparison with classical data augmentation methods Authors: ['Alhassan Mumuni' 'Fuseini Mumuni'] Abstract—Data augmentation is arguably the most important regularization technique commonly used to improve generalization performance of machine learning models. It primarily involves the application of…
Dan Liu, Samer El Kababji, Nicholas Mitsakakis, Lisa Pilgram + 5 more
Background Small datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution and is often used for imaging and time series data, but there are no evaluations on its…
Teerath Kumar, Muhammad Turab, Kislay Raj, Alessandra Mileo + 2 more
'Rob Brennan' 'Malika Bendechache'] Abstract—Deep learning algorithms have demonstrated remarkable performance in various computer vision tasks, however, limited labeled data can lead to overfitting problems, hindering the network's performance on unseen data. To address this issue, various generalization techniques…
Mason Minot, Sai T. Reddy
Machine learning-guided protein engineering is a rapidly advancing field. Despite major experimental and computational advances however, collecting protein genotype (sequence) and phenotype (function) data remains time and resource intensive. As a result, the quality and quantity of training data is often a limiting…
Yiyang Yu, Shivani Muthukumar, Peter K Koo
Deep neural networks (DNNs) have been widely applied to predict the molecular functions of regulatory regions in the non-coding genome. DNNs are data hungry and thus require many training examples to fit data well. However, functional genomics experiments typically generate limited amounts of data, constrained by the…
Suorong Yang, Hongchao Yang, Suhan Guo, Furao Shen + 1 more
—Data augmentation is widely utilized as an effective technique to enhance the generalization performance of deep models. However, data augmentation may inevitably introduce distribution shifts and noises, which significantly constrain the potential and deteriorate the performance of deep networks. To this end, we…
Jack Cole, Heng Yang, Krasimira Tsaneva-Atanasova, Ke Li
Genomic Language Models (GLMs) suffer from the inherent problem of data scarcity, due to the cost, time and complexity of wet-lab experiments. Data augmentation offers a solution; however traditional methods may unintentionally affect the underlying structure or function. By combining evolutionary signals with the RNA…
Authors not listed
Background: Obstructive sleep apnea (OSA) is growing increasingly prevalent in many countries as obesity rises. Sufficient, effective treatment of OSA entails high social and financial costs for healthcare. Objective: For treatment purposes, predicting OSA patients’ visit expenses for the coming year is crucial.…
Nikita Janakarajan, Mara Graziani, Maria Rodriguez Martinez
Working with transcriptomic data is challenging in deep learning applications due to its high dimensionality and low patient numbers. Deep learning models tend to overfit this data and do not generalize well on out-of-distribution samples and new cohorts. Data augmentation strategies help alleviate this problem by…
Manuel Domínguez-Rodrigo, Gabriel Cifuentes-Alcobendas, Marina Vegara-Riquelme, Enrique Baquedano
Recent critiques of the reliability of deep learning (DL) for taphonomic analysis of bone surface modifications (BSM), such as that presented by 20 based on a selection of earlier published studies, have raised concerns about the efficacy of the method. Their critique, however, overlooked fundamental principles…
Authors not listed
Data augmentation can alleviate the limitations of small molecular datasets for generative deep learning, by ‘artificially inflating’ the number of instances available for training. SMILES enumeration – whereby multiple valid SMILES strings are used to represent the same molecules – has resulted particularly beneficial…
Authors not listed
Predicting reaction yields in synthetic chemistry remains a significant challenge. This study systematically evaluates the impact of tokenization, molecular representation, pre-training data, and adversarial training on a BERT-based model for yield prediction of Buchwald-Hartwig and Suzuki-Miyaura coupling reactions…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
High throughput screening (HTS) is one of the leading techniques for hit identification in drug discovery and comprises of multiple phases, one primary and one or more confirmatory screens which result in multi-fidelity data. Noisy primary screening data are available on a large number of compounds and higher quality…
Authors not listed
Active learning (AL) can significantly accelerate drug discovery by iteratively selecting informative molecules, reducing experimental workload. However, existing AL studies typically assume access to large datasets, an unrealistic scenario for most academic labs. Here, we investigate AL strategies tailored…
Authors not listed
Many successful machine learning models for molecular property prediction rely on Lewis structure representations, commonly encoded as SMILES strings. However, a key limitation arises with molecules exhibiting resonance, where multiple valid Lewis structures represent the same species. This causes inconsistent…
Authors not listed
The scarcity and expense of fatigue data limits optimal design of components and constrains companies to a few well qualified materials when safety-critical applications are concerned. This research investigates different strategies to improve extraction of structured information from unstructured scientific…