13 papers · ranked by Valyu relevance
Heidi J. Imker
In the early 1990s the life sciences quickly adopted online databases to facilitate wide-spread dissemination and use of scientific data. From 1991, the journal Nucleic Acids Research has published an annual Database Issue dedicated to articles describing molecular biology databases. Analysis of these articles reveals…
Heidi J. Imker
Online resources enable unfettered access to and analysis of scientific data and are considered crucial for the advancement of modern science. Despite the clear power of online data resources, including web-available databases, proliferation can be problematic due to challenges in sustainability and long-term…
Dimitri Yatsenko, Jacob Reimer, Alexander S. Ecker, Edgar Y. Walker + 6 more
The rise of big data in modern research poses serious challenges for data management: Large and intricate datasets from diverse instrumentation must be precisely aligned, annotated, and processed in a variety of ways to extract new insights. While high levels of data integrity are expected, research teams have diverse…
Bohdan B. Khomtchouk, Kasra A. Vand, Thor Wahlestedt, Kelly Khomtchouk + 2 more
We propose a search engine and file retrieval system for all bioinformatics databases worldwide. PubData searches biomedical data in a user-friendly fashion similar to how PubMed searches biomedical literature. PubData is built on novel network programming, natural language processing, and artificial intelligence…
Magali Ruffier, Andreas Kähäri, Monika Komorowska, Stephen Keenan + 10 more
The Ensembl software resources are a stable infrastructure to store, access and manipulate genome assemblies and their functional annotations. The Ensembl “Core” database and Application Programming Interface (API) was our first major piece of software infrastructure and remains at the centre of all of our genome…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Po-Ju Yao, Ren-Hua Chung
Computer simulations are routinely conducted to evaluate new statistical methods, to compare the properties among different methods, and to mimic the real data in genetic epidemiology studies. Conducting simulation studies can become a complicated task as several challenges can occur, such as the selection of an…
Ana Claudia Sima, Tarcisio Mendes de Farias, Erich Zbinden, Maria Anisimova + 5 more
Data integration promises to be one of the main catalysts in enabling new insights to be drawn from the wealth of biological data available publicly. However, the heterogeneity of the different data sources, both at the syntactic and the semantic level, still poses significant challenges for achieving interoperability…
Alexander M. Waldrop, John B. Cheadle, Kira Bradford, Alexander Preiss + 14 more
As the number of public data resources continues to proliferate, identifying relevant datasets across heterogenous repositories is becoming critical to answering scientific questions. To help researchers navigate this data landscape, we developed Dug: a semantic search tool for biomedical datasets utilizing…
Mauricio de Alvarenga Mudadu, Adhemar Zerlotini
Genome projects and multiomics experiments generate huge volumes of data that must be stored, mined and transformed into useful knowledge. All this information is supposed to be accessible and, if possible, browsable afterwards. Computational biologists have been dealing with this scenario for over a decade and have…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Gerardo Lagunes-García, Alejandro Rodríguez-González, Lucía Prieto-Santamaría, Eduardo P. García del Valle + 2 more
Within the global endeavour of improving population health, one major challenge is the increasingly high cost associated with drug development. Drug repositioning, i.e. finding new uses for existing drugs, is a promising alternative; yet, its effectiveness has hitherto been hindered by our limited knowledge about…
Marco Falda
FuzzyVariantExplorer is a web interface and a fuzzy search system for exploring richly annotated genomes in a flexible way. The application allows combining vague constraints in expressive logical queries to retrieve graded sets of genes and their associated variants. Results can be further refined in a visual way by…