23 papers · ranked by Valyu relevance
Dilpreet Singh, Chandan K Reddy
The primary purpose of this paper is to provide an in-depth analysis of different platforms available for performing big data analytics. This paper surveys different hardware platforms available for big data analytics and assesses the advantages and drawbacks of each of these platforms based on various metrics such as…
Christian Lovis, Jianbo Lei, Weihua Meng, Miye Wang + 7 more
Background With the advent of data-intensive science, a full integration of big data science and health care will bring a cross-field revolution to the medical community in China. The concept big data represents not only a technology but also a resource and a method. Big data are regarded as an important strategic…
Raghavendra Kune, Pramodkumar Konugurthi, Arun Agarwal, Raghavendra Rao Chillarige + 1 more
'Raghavendra Rao Chillarige' 'Rajkumar Buyya'] Abstract- Advances in information technology and its widespread growth in several areas of business, engineering, medical and scientific studies are resulting in information/data explosion. Knowledge discovery and decision making from such rapidly growing voluminous data…
Michał Zasadziński, Michael H. Theodoulou, Markus Thurner, Kshitij Ranganath
'Kshitij Ranganath'] Abstract—Data Analytics provides core business reporting needs in many software companies, acts as a source of truth for key information, and enables building advanced solutions, e.g., predictive models, machine learning, real-time recommendations, to grow the business. Typically, companies ingest…
Patricia L. Mabry, Xiaoran Yan, Valentin Pentchev, Robert Van Rennes + 2 more
Big bibliographic datasets hold promise for revolutionizing the scientific enterprise when combined with state-of-the-science computational capabilities. Yet, hosting proprietary and open big bibliographic datasets poses significant difficulties for libraries, both large and small. Libraries face significant barriers…
Yao Wu, Henan Guan
—Various tools, softwares and systems are proposed and implemented to tackle the challenges in big data on different emphases, e.g., data analysis, data transaction, data query, data storage, data visualization, data privacy. In this paper, we propose datar, a new prospective and unified framework for Big Data…
Authors not listed
The Digital Catalysis Platform (DigCat) is a pioneering integration of big data and AI tailored for catalysis materials research. It encompasses over 400,000 experimental performance data for electro-, thermo-, and photocatalysts, alongside more than 300,000 catalyst structures. DigCat provides dynamic data…
Ji-Long Liu, Li-Guang Yang, Qing-Yu Xiao, Zhao-Qiang Li + 7 more
The advent of high throughput sequencing has ushered life science and clinical research into the era of big data, posing significant challenges for reproducibility due to the complexity of data integration and analysis. Although the FAIR principles advocate for the transparent and reliable sharing of scientific data…
Anthony Mammoliti, Petr Smirnov, Minoru Nakano, Zhaleh Safikhani + 3 more
Reproducibility is essential to Open Science, as there is limited relevance for finding that cannot be reproduced by independent research groups, regardless of its validity. It is therefore crucial for scientists to describe their experiments in sufficient detail so they can be reproduced, challenged, and built upon.…
Saraswati Koppad, Annappa B, Georgios V Gkoutos, Animesh Acharjee
Analytics Authors: ['Saraswati Koppad' 'Annappa B' 'Georgios V Gkoutos' 'Animesh Acharjee'] High-throughput experiments enable researchers to explore complex multifactorial diseases through large-scale analysis of omics data. Challenges for such high-dimensional data sets include storage, analyses, and sharing. Recent…
Nathan C. Sheffield, Vivien R. Bonazzi, Philip E. Bourne, Tony Burdett + 4 more
'Tony Burdett' 'Timothy Clark' 'Robert L. Grossman' 'Ola Spjuth' 'Andrew D. Yates'] The biomedical research community is investing heavily in biomedical cloud platforms. Cloud computing holds great promise for addressing challenges with big data and ensuring reproducibility in biology. However, despite their…
Arthur W Toga, Ivo D Dinov
Background The promise of Big Biomedical Data may be offset by the enormous challenges in handling, analyzing, and sharing it. In this paper, we provide a framework for developing practical and reasonable data sharing policies that incorporate the sociological, financial, technical and scientific requirements of a…
Todor Ivanov, Nikolaos Korfiatis, Roberto V. Zicari
The well-known 3V architectural paradigm for Big Data introduced by Laney (2011) provides a simplified framework for defining the architecture of a big data platform to be deployed in various scenarios tackling processing of massive datasets. While additional components such as Variability and Veracity have been…
Aleem Akhtar
With the increase in amount of Big Data being generated each year, tools and technologies developed and used for the purpose of storing, processing and analyzing Big Data has also improved. Open-Source software has been an important factor in the success and innovation in the field of Big Data while Apache Software…
Ivan Merelli, Horacio Pérez-Sánchez, Sandra Gesing, Daniele D'Agostino
"Daniele D'Agostino"] The explosion of the data both in the biomedical research and in the healthcare systems demands urgent solutions. In particular, the research in omics sciences is moving from a hypothesis-driven to a data-driven approach. Healthcare is additionally always asking for a tighter integration with…
Alex M. Ascension, Marcos J. Araúzo-Bravo
Big Data analysis is a discipline with a growing number of areas where huge amounts of data is extracted and analyzed. Parallelization in Python integrates Message Passing Interface via mpi4py module. Since mpi4py does not support parallelization of objects greater than 2^31^ bytes, we developed BigMPI4py, a Python…
Ravi Madduri, Kyle Chard, Mike D’Arcy, Segun C. Jung + 14 more
Big biomedical data create exciting opportunities for discovery, but make it difficult to capture analyses and outputs in forms that are findable, accessible, interoperable, and reusable (FAIR). In response, we describe tools that make it easy to capture, and assign identifiers to, data and code throughout the data…
David Mayer, Seth Russell, Melissa P. Wilson, Michael G. Kahn + 1 more
One of the challenges of teaching applied data science courses is managing individual students’ local computing environment. This is especially challenging when teaching massively open online courses (MOOCs) where students come from across the globe and have a variety of access to and types of computing systems. There…
Yasin El Abiead, Michael Strobel, Thomas Payne, Eoin Fahy + 13 more
Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, we've developed pan-repository universal identifiers and harmonized cross-repository metadata. This novel ecosystem…
Authors not listed
Machine learning models are transforming data-driven research across scientific disciplines, yet their deployment as accessible and reliable web services remains a significant challenge. We introduce the NERDD framework, a scalable, maintainable, and secure microservices platform designed to support the sustainable…
Monika Vogler, Jonas Busk, Hamidreza Hajiyani, Peter Bjørn Jørgensen + 9 more
The future of materials science is borderless, cooperative, and distributed across the globe. This necessitates flexible, reconfigurable software defined research workflows, which we herein demonstrate by integrating multiple disciplines and modalities. Our brokering approach to research orchestration exposes entire…
Rebecca Brunk, Kriti Shukla, Bryant Hutson, Yue Wang + 7 more
Genomic sequencing and other big biological data is unquestionably of paramount value, however the success in recruiting highly skilled individuals with diverse backgrounds has been limited. A main reason for this deficiency could be due to the lack of educational resources and early exposure to the field. With the…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…