16 papers · ranked by Valyu relevance
Peter Sestoft
Research relies on ever larger amounts of data from experiments, automated production equipment, questionnaries, times series such as weather records, and so on. A major task in science is to combine, process and analyse such data to obtain evidence of patterns and correlations. Most research data are on digital form…
Rebekah Duke, Vinayak Bhat, Chad Risko
As buzzwords like “big data,” “machine learning,” and “high-throughput” expand through chemistry, chemists need to consider more than ever their data storage, data management, and data accessibility, whether in their own laboratories or with the broader community. While it is commonplace for chemists to use…
Mateusz Jundzill, Riccardo Spott, Mara Lohde, Martin Hölzer + 2 more
'Adrian Viehweger' 'Christian Brandt'] Title: Abstract With the rapidly growing amount of biological data, powerful but also flexible data management and visualization systems are of increasingly crucial importance. The COVID-19 pandemic has more than highlighted this need and the challenges scientists are facing.…
Shahid Ullah, Wajeeha Rahman, Farhan Ullah, Gulzar Ahmad + 2 more
'Muhmmad Ijaz' 'Tianshun Gao'] Background: The achievement of the human genome project provides a basis for the systematic study of the human genome from evolutionary history to disease-specific medicine. With the explosive growth of biological data, a growing number of biological databases are being established to…
Mikołaj Danielewski, Marlena Szalata, Jan Krzysztof Nowak, Jarosław Walkowiak + 3 more
With the development of genome sequencing technologies, the amount of data produced has greatly increased in the last two decades. The abundance of digital sequence information (DSI) has provided research opportunities, improved our understanding of the genome, and led to the discovery of new solutions in industry and…
Paul L. Soto
Data collection and analysis are central to scientific research, including in applied and basic behavior analysis. A substantial amount of attention has been given to how to rigorously collect and analyze data. Less attention has been paid to storing and maintaining research data, which becomes a critical step in the…
Mingrui Liu, Zelin Ye, Haiyu Liu, Pengzhen Ma + 6 more
This review provides a representative overview of public biomedical databases and their use in biomedical research. These resources are categorized into four major types according to their dominant data content: public health databases, clinical databases, comprehensive cohort databases, and omics databases. For each…
Delphine Steinbach, Michael Alaux, Joelle Amselem, Nathalie Choisne + 13 more
'Sophie Durand' 'Raphaël Flores' 'Aminah-Olivia Keliet' 'Erik Kimmel' 'Nicolas Lapalu' 'Isabelle Luyten' 'Célia Michotey' 'Nacer Mohellibi' 'Cyril Pommier' 'Sébastien Reboux' 'Dorothée Valdenaire' 'Daphné Verdelet' 'Hadi Quesneville'] Data integration is a key challenge for modern bioinformatics. It aims to provide…
A. C. Gorakshakar, K. Ghosh
Bioinformatics is a relatively new discipline. It is a field of science in which Computer science, Mathematics, Molecular biology and Information technology merges to form a single discipline. Database development, sequence alignment, protein structure prediction, RNA folding, evolutionary tree construction are some of…
Leigh Dodds
The FAIR principles need to be applied in context. To do that, we need to understand both the needs of data users and the characteristics of the data to be shared. This Opinion introduces ten different dataset archetypes that can be used to inform plans for how data are to be accessed, used, and shared.
Xiao-Qin Xia, Michael McClelland, Yipeng Wang
Background With advances in high-throughput genomics and proteomics, it is challenging for biologists to deal with large data files and to map their data to annotations in public databases. Results We developed TabSQL, a MySQL-based application tool, for viewing, filtering and querying data files with large numbers of…
Arthur M. Lesk, Anna Tramontano
Computational molecular biology is a relatively new specialty that has arisen in response to the very large amount and quality of data currently being produced, including gene and protein sequences (“one-dimensional” information) and nucleic acid and protein structures (“three-dimensional” information). Many important…
Ted D Wade
We review traits of reusable clinical data and offer a typology of clinical repositories with a range of known examples. Sources of clinical data suitable for research can be classified into types reflecting the data’s institutional origin, original purpose, level of integration and governance. Primary data nearly…
Sergio Lifschitz, Edward H. Haeusler, Marcos Catanho, Antonio B. de Miranda + 6 more
'Antonio B. de Miranda' 'Elvismary Molina de Armas' 'Alexandre Heine' 'Sergio G. M. P. Moreira' 'Cristian Tristão' 'Pietro Pinoli' 'Anna Bernasconi'] DNA sequencers output a large set of very long biological data strings that we should persist in databases rather than basic text file systems. Many different data models…
Sohrab P Shah, Yong Huang, Tao Xu, Macaire MS Yuen + 2 more
Background We present a biological data warehouse called Atlas that locally stores and integrates biological sequences, molecular interactions, homology information, functional annotations of genes, and biological ontologies. The goal of the system is to provide data, as well as a software infrastructure for…
Hasan M Jamil
Background One of the many unique features of biological databases is that the mere existence of a ground data item is not always a precondition for a query response. It may be argued that from a biologist's standpoint, queries are not always best posed using a structured language. By this we mean that approximate and…