25 papers · ranked by Valyu relevance
Connor Shorten, Taghi M. Khoshgoftaar, Borko Furht
Natural Language Processing (NLP) is one of the most captivating applications of Deep Learning. In this survey, we consider how the Data Augmentation training strategy can aid in its development. We begin with the major motifs of Data Augmentation summarized into strengthening local decision boundaries, brute force…
Kumar, Saorj, Prince Asiamah, Oluwatoyin Jolaoso + 1 more
—Convolutional Neural Networks (CNNs) serve as the workhorse of deep learning, finding applications in various fields that rely on images. Given sufficient data, they exhibit the capacity to learn a wide range of concepts across diverse settings. However, a notable limitation of CNNs is their susceptibility to…
Khaled Alomar, Halil Ibrahim Aysel, Xiaohao Cai, Jérôme Gilles + 1 more
'Luminiţa Moraru'] In the past decade, deep neural networks, particularly convolutional neural networks, have revolutionised computer vision. However, all deep learning models may require a large amount of data so as to achieve satisfying results. Unfortunately, the availability of sufficient amounts of data for…
Nurdan Ayse Saran, Murat Saran, Fatih Nar, Sebastian Ventura
In the last decade, deep learning has been applied in a wide range of problems with tremendous success. This success mainly comes from large data availability, increased computational power, and theoretical improvements in the training phase. As the dataset grows, the real world is better represented, making it…
Z. Wang, Pengfei Wang, Kunpeng Liu, Pengyang Wang + 5 more
'Chang‐Tien Lu' 'Charų C. Aggarwal' 'Jian Pei' 'Yuanchun Zhou'] ZAITIAN WANG∗ and PENGFEI WANG∗ , Computer Network Information Center, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Chinese Academy of Sciences, China KUNPENG LIU, Portland State University, USA PENGYANG WANG, University of…
Melissa Valaee, Shahram Shirani
With a 17.9 million annual mortality rate, cardiovascular disease is the leading global cause of death. As such, early detection and disease diagnosis are critical for effective treatment and symptom management. Cardiac auscultation, the process of listening to the heartbeat, often provides the first indication of…
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song
Due to progression in cell-cycle or duration of storage, classification of morphological changes in human blood cells is important for correct and effective clinical decisions. Automated classification systems help avoid subjective outcomes and are more efficient. Deep learning and more specifically Convolutional…
Teerath Kumar, Muhammad Turab, Kislay Raj, Alessandra Mileo + 2 more
'Rob Brennan' 'Malika Bendechache'] Abstract—Deep learning algorithms have demonstrated remarkable performance in various computer vision tasks, however, limited labeled data can lead to overfitting problems, hindering the network's performance on unseen data. To address this issue, various generalization techniques…
Suorong Yang, Weikang Xiao, Mengcheng Zhang, Suhan Guo + 2 more
'Furao Shen'] Deep learning has achieved remarkable results in many computer vision tasks. Deep neural networks typically rely on large amounts of training data to avoid overfitting. However, labeled data for realworld applications may be limited. By improving the quantity and diversity of training data, data…
Wei Wang, Zhaowei Shang, Chengxing Li
Data augmentation is an effective technique for automatically expanding training data in deep learning. Brain-inspired methods are approaches that draw inspiration from the functionality and structure of the human brain and apply these mechanisms and principles to artificial intelligence and computer science. When…
Ji-Young Yoon, Gahgene Gweon, Yun Joo Yoo, Peida Zhan
Over recent decades, machine learning, an integral subfield of artificial intelligence, has revolutionized diverse sectors, enabling data-driven decisions with minimal human intervention. In particular, the field of educational assessment emerges as a promising area for machine learning applications, where students can…
João Fonseca, Fernando Bação
In the Machine Learning research community, there is a consensus regarding the relationship between model complexity and the required amount of data and computation power. In real world applications, these computational requirements are not always available, motivating research on regularization methods. In addition…
Dan Liu, Samer El Kababji, Nicholas Mitsakakis, Lisa Pilgram + 5 more
Background Small datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution and is often used for imaging and time series data, but there are no evaluations on its…
Yiyang Yu, Shivani Muthukumar, Peter K Koo
Deep neural networks (DNNs) have been widely applied to predict the molecular functions of regulatory regions in the non-coding genome. DNNs are data hungry and thus require many training examples to fit data well. However, functional genomics experiments typically generate limited amounts of data, constrained by the…
Min Oh, Liqing Zhang
Predictive models trained on sequencing profiles often fail to achieve expected performance when externally validated on unseen profiles. While many factors such as batch effects, small data sets, and technical errors contribute to the gap between source and unseen data distributions, it is a challenging problem to…
Jack Cole, Heng Yang, Krasimira Tsaneva-Atanasova, Ke Li
Genomic Language Models (GLMs) suffer from the inherent problem of data scarcity, due to the cost, time and complexity of wet-lab experiments. Data augmentation offers a solution; however traditional methods may unintentionally affect the underlying structure or function. By combining evolutionary signals with the RNA…
Mason Minot, Sai T. Reddy
Machine learning-guided protein engineering is a rapidly advancing field. Despite major experimental and computational advances however, collecting protein genotype (sequence) and phenotype (function) data remains time and resource intensive. As a result, the quality and quantity of training data is often a limiting…
Nikita Janakarajan, Mara Graziani, Maria Rodriguez Martinez
Working with transcriptomic data is challenging in deep learning applications due to its high dimensionality and low patient numbers. Deep learning models tend to overfit this data and do not generalize well on out-of-distribution samples and new cohorts. Data augmentation strategies help alleviate this problem by…
Manuel Domínguez-Rodrigo, Gabriel Cifuentes-Alcobendas, Marina Vegara-Riquelme, Enrique Baquedano
Recent critiques of the reliability of deep learning (DL) for taphonomic analysis of bone surface modifications (BSM), such as that presented by 20 based on a selection of earlier published studies, have raised concerns about the efficacy of the method. Their critique, however, overlooked fundamental principles…
Authors not listed
Data augmentation can alleviate the limitations of small molecular datasets for generative deep learning, by ‘artificially inflating’ the number of instances available for training. SMILES enumeration – whereby multiple valid SMILES strings are used to represent the same molecules – has resulted particularly beneficial…
Authors not listed
Predicting reaction yields in synthetic chemistry remains a significant challenge. This study systematically evaluates the impact of tokenization, molecular representation, pre-training data, and adversarial training on a BERT-based model for yield prediction of Buchwald-Hartwig and Suzuki-Miyaura coupling reactions…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
High throughput screening (HTS) is one of the leading techniques for hit identification in drug discovery and comprises of multiple phases, one primary and one or more confirmatory screens which result in multi-fidelity data. Noisy primary screening data are available on a large number of compounds and higher quality…
Angela Lopez-del Rio, Sergio Picart, Alexandre Perera-Lluna
In silico analysis of biological activity data has become an essential technique in pharmaceutical development. Specifically, the so-called proteochemometric models aim to share information between targets in machine learning ligand-target activity prediction models. However, bioactivity datasets used in…
Authors not listed
Many successful machine learning models for molecular property prediction rely on Lewis structure representations, commonly encoded as SMILES strings. However, a key limitation arises with molecules exhibiting resonance, where multiple valid Lewis structures represent the same species. This causes inconsistent…
Authors not listed
The scarcity and expense of fatigue data limits optimal design of components and constrains companies to a few well qualified materials when safety-critical applications are concerned. This research investigates different strategies to improve extraction of structured information from unstructured scientific…