19 papers · ranked by Valyu relevance
Xiao-Qin Xia, Michael McClelland, Yipeng Wang
Background With advances in high-throughput genomics and proteomics, it is challenging for biologists to deal with large data files and to map their data to annotations in public databases. Results We developed TabSQL, a MySQL-based application tool, for viewing, filtering and querying data files with large numbers of…
M. Nissan
Database audit and transaction logs are fundamental to forensic investigations, but they are vulnerable to tampering by privileged attackers. Malicious insiders or external threats with administrative access can alter, purge, or temporarily disable logging mechanisms, creating significant blind spots and rendering…
Andriy Miranskyy, Zainab Al-Zanbouri, D. Godwin, Ayşe Bener
Conclusions: Our findings provide insights to both practitioners and researchers. Database administrators may use them to select a fast, green release of the MySQL database engine. MySQL database-engine developers may use the software metric to assess products' greenness and performance. Researchers may use our…
Jordi Vilaplana, Francesc Solsona, Ivan Teixido, Anabel Usié + 3 more
'Hiren Karathia' 'Rui Alves' 'Jordi Mateo'] Our group developed two biological applications, Biblio-MetReS and Homol-MetReS, accessing the same database of organisms with annotated genes. Biblio-MetReS is a data-mining application that facilitates the reconstruction of molecular networks based on automated text-mining…
Yu Liang, Hong Hu
In this paper, we propose ParserFuzz, a novel fuzzing framework that automatically extracts grammar rules from DBMSs' built-in syntax definition files for SQL query generation. Without any input corpus, ParserFuzz can generate diverse query statements to saturate the grammar features of the tested DBMSs, which grammar…
Florent Chuffart, Gaël Yvert
Laboratory stocks are the hardware of research. They must be stored and managed with mimimum loss of material and information. Plasmids, oligonucleotides and strains are regularly exchanged between collaborators within and between laboratories. Managing and sharing information about every item is crucial for retrieval…
Lukas Iffländer, Alexandra Dmitrienko, Christoph Hagen, M. Jobst + 1 more
'Samuel Kounev'] Ransomware is an emerging threat which imposed a $ 5 billion loss in 2017 and is predicted to hit 11.5 billion in 2019. While initially targeting PC (client) platforms, ransomware recently made the leap to server-side databases – starting in January 2017 with the MongoDB Apocalypse attack, followed by…
Manuel Rigger, Zhendong Su
Relational databases are used ubiquitously. They are managed by database management systems (DBMS), which allow inserting, modifying, and querying data using a domainspecific language called Structured Query Language (SQL). Popular DBMS have been extensively tested by fuzzers, which have been successful in finding…
Kok Swee Sim, Sze Siang Chong, Chih Ping Tso, Mohsen Esmaeili Nia + 2 more
patients Authors: ['Kok Swee Sim' 'Sze Siang Chong' 'Chih Ping Tso' 'Mohsen Esmaeili Nia' 'Aun Kee Chong' 'Siti Fathimah Abbas'] Data analysis based on breast cancer risk factors such as age, race, breastfeeding, hormone replacement therapy, family history, and obesity was conducted on breast cancer patients using a…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Zhijie He, Cong Wang, Xudong Guo, Heyun Sun + 6 more
PE/PPE proteins, highly abundant in the Mycobacterium genome, play a vital role in virulence and immune modulation. Understanding their functions is key to comprehending the internal mechanisms of Mycobacterium. However, a lack of dedicated resources has limited research into PE/PPE proteins. Addressing this gap, we…
Carol A. Soderlund
De novo transcriptome sequencing and analysis provides a way for researchers of non-model organisms to explore the differences between various conditions and species. The results are typically not definitive but will lead to new hypotheses to study. Therefore, it is important that the results be reproducible…
Lubos Molcan
Physiological processes oscillate in time. Circadian oscillations, over approximately 24-h, are very important and among the most studied. To evaluate the presence and significance of 24-h oscillations, physiological time distributed data (TDD) are often set to a cosinor model using a wide range of irregularly updated…
Tiqing Liu, Linda Hwang, Stephen K Burley, Carmen I Nitsche + 3 more
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry…
Shicai Wang, Ioannis Pandis, Chao Wu, Sijin He + 4 more
'Ibrahim Emam' 'Florian Guitton' 'Yike Guo'] Background High-throughput transcriptomic data generated by microarray experiments is the most abundant and frequently stored kind of data currently used in translational medicine studies. Although microarray data is supported in data warehouses such as tranSMART, when…
Tina Sharma, Rakesh Kumar, Anshu Bhardwaj
Ab-AMR is a comprehensive repository of drug resistance mechanisms in Acinetobacter baumannii. The current version of Ab-AMR provides a drug resistance profile of 788 genomes. In order to ensure that the datasets in Ab-AMR have relevance both to the research and clinical community, standards of defining MIC…
Shawn Yates, Martin Lagüe, Ron Knox, Richard Cuthbert + 5 more
With the advent of next-generation marker platforms and phenomics in crop breeding programs, the volume of both the genotypic and phenotypic data produced has increased exponentially. Often the data remain underutilized if not properly collated, managed and accessed. Effective management of the data is paramount to…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…