24 papers · ranked by Valyu relevance
Rhett N. D’souza, Po-Yao Huang, Fang-Cheng Yeh
Deep neural networks have gained immense popularity in the Big Data problem; however, the availability of training samples can be relatively limited in certain application domains, particularly medical imaging, and consequently leading to overfitting problems. This “Small Data” challenge may need a mindset that is…
Fleming Kretschmer, Jan Seipp, Marcus Ludwig, Gunnar W. Klau + 1 more
Small molecule machine learning aims to predict chemical, biochemical, or biological properties from molecular structures, with applications such as toxicity prediction, ligand binding, and pharmacokinetics. A recent trend is developing end-to-end models that avoid explicit domain knowledge. These models assume no…
Max Robinson, Jennifer Hadlock, Jiyang Yu, Alireza Khatamian + 5 more
We present a locality-sensitive hashing strategy for summarizing semi-structured data (e.g., in JSON or XML formats) into ‘data fingerprints’: highly compressed representations which cannot recreate details in the data, yet simplify and greatly accelerate the comparison and clustering of semi-structured data by…
Jakub Galgonek, Jiří Vondrášek
The Resource Description Framework (RDF), together with well-defined ontologies, significantly increases data interoperability and usability. The SPARQL query language was introduced to retrieve requested RDF data and to explore links between them. Among other useful features, SPARQL supports federated queries that…
Derek van Tilborg, Helena Brinkmann, Emanuele Criscuolo, Luke Rossen + 2 more
Deep learning is becoming increasingly relevant in drug discovery, from de novo design to protein structure prediction and synthesis planning. However, it is often challenged by the small data regimes typical of certain drug discovery tasks. In such scenarios, deep learning approaches – which are notoriously…
Yanwei Huang, Yan Miao, Di Weng, Adam Perer + 1 more
To evaluate the performance of StructVizor's data preprocessing pipeline, we conducted a technical assessment using a substantial set of semi-structured datasets. To the best of our knowledge, we haven't found any open-source benchmark datasets for parsing general semi-structured data except PADS1 , which was used in…
Authors not listed
Early-stage drug discovery often suffers from data scarcity and out-of-distribution (OOD) shifts, which constrain the reliability of predictive models. While deep learning has advanced representation learning from molecular and biological data, tabular modeling remains indispensable, particularly in small-sample and…
Kohulan Rajan, Achim Zielesny, Christoph Steinbeck
Naming chemical compounds systematically is a complex task governed by a set of rules established by the International Union of Pure and Applied Chemistry (IUPAC). These rules are universal and widely accepted by chemists worldwide, but their complexity makes it challenging for individuals to consistently apply them…
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis + 3 more
'Luis Ibáñez' 'Emilia Kacprzak' 'Paul Groth'] Abstract Generating value from data requires the ability to find, access and make sense of datasets. There are many efforts underway to encourage data sharing and reuse, from scientific publishers asking authors to submit data alongside manuscripts to data marketplaces…
Authors not listed
Computational methods for predictive modeling have been increasingly utilized in the early stages of drug discovery to supplement high-throughput screening. The advent of highly efficient and complex machine learning architectures necessitates new methods of collating the plethora of topological, geometrical, and…
Pablo Pareja-Tobes, Raquel Tobes, Marina Manrique, Eduardo Pareja + 1 more
Next Generation Sequencing and other high-throughput technologies have brought a revolution to the bioinformatics landscape, by offering sheer amounts of data about previously unaccessible domains in a cheap and scalable way. However, fast, reproducible, and cost-effective data analysis at such scale remains elusive. A…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Brenda Farrell, Jason Bengtson, Robert Hoehndorf
In the past scientists reported summaries of their findings; they did not provide their original data collections. Many stakeholders (e.g., funding agencies) are now requesting that such data be made publicly available. This mandate is being adopted to facilitate further discovery, and to mitigate waste and deficits in…
Lucina Hackman, Pauline Mack, Hervé Ménard
Data underpinning science have become one of the most precious assets in research, and while the principles of FAIR (Findable, Accessible, Interoperable and Reusable) have been put forward as a guide to how to approach data handling, data sharing and long-term storage still remain a challenge for many research areas…
Jingyi Shi, Mingna Zheng, Lixia Yao, Yaorong Ge
Background The right dataset is essential to obtain the right insights in data science; therefore, it is important for data scientists to have a good understanding of the availability of relevant datasets as well as the content, structure, and existing analyses of these datasets. While a number of efforts are underway…
Andreea Grigoriu, Amrapali Zaveri, Gerhard Weiss, Michel Dumontier
Background The amount of available data, which can facilitate answering scientific research questions, is growing. However, the different formats of published data are expanding as well, creating a serious challenge when multiple datasets need to be integrated for answering a question. Results This paper presents a…
Deepshikha Singh, Shashank Jatav, Shefali Lathwal, Arman Kazmi + 1 more
Public biomedical repositories contain extensive, valuable datasets, yet identifying datasets that precisely match specific study requirements remains inefficient. Conventional keyword- and schema-based systems frequently fall short when queries encompass multiple biological and experimental facets. To address this, we…
Yihan Gao, Silu Huang, Aditya Parameswaran
Organizations routinely accumulate semi-structured log datasets generated as the output of code; these datasets remain unused and uninterpreted, and occupy wasted space—this phenomenon has been colloquially referred to as "data lake" problem. One approach to leverage these semi-structured datasets is to convert them…
Kristina Edfeldt, Aled M. Edwards, Ola Engkvist, Judith Günther + 20 more
'Matthew Hartley' 'David G. Hulcoop' 'Andrew R. Leach' 'Brian D. Marsden' 'Amelie Menge' 'Leonie Misquitta' 'Susanne Müller' 'Dafydd R. Owen' 'Kristof T. Schütt' 'Nicholas Skelton' 'Andreas Steffen' 'Alexander Tropsha' 'Erik Vernet' 'Yanli Wang' 'James Wellnitz' 'Timothy M. Willson' 'Djork-Arné Clevert' 'Benjamin…
Leigh Dodds
The FAIR principles need to be applied in context. To do that, we need to understand both the needs of data users and the characteristics of the data to be shared. This Opinion introduces ten different dataset archetypes that can be used to inform plans for how data are to be accessed, used, and shared.
El Kindi Rezig, Michael Cafarella, Vijay Gadepally
AI application developers typically begin with a dataset of interest and a vision of the end analytic or insight they wish to gain from the data at hand. Although these are two very important components of an AI workflow, one often spends the first few weeks (sometimes months) in the phase we refer to as data…
Maxat Kulmanov, Senay Kafkas, Andreas Karwath, Alexander Malic + 3 more
Recent developments in machine learning have lead to a rise of large number of methods for extracting features from structured data. The features are represented as a vectors and may encode for some semantic aspects of data. They can be used in a machine learning models for different tasks or to compute similarities…
Authors not listed
Background: Chemical reactions form intricate, highly connected networks whose exploration is essential for discovering more efficient and sustainable synthetic routes. As reaction data from literature, patents, and high‑throughput experimentation continue to surge, so does the need for tools that can collectively…