27 papers · ranked by Valyu relevance
Nawsher Khan, Ibrar Yaqoob, Ibrahim Abaker Targio Hashem, Zakira Inayat + 4 more
'Zakira Inayat' 'Waleed Kamaleldin Mahmoud Ali' 'Muhammad Alam' 'Muhammad Shiraz' 'Abdullah Gani'] Big Data has gained much attention from the academia and the IT industry. In the digital and computing world, information is generated and collected at a rate that rapidly exceeds the boundary range. Currently, over 2…
Jonathan S. Ward, Adam Barker
The term big data has become ubiquitous. Owing to a shared origin between academia, industry and the media there is no single unified definition, and various stakeholders provide diverse and often contradictory definitions. The lack of a consistent definition introduces ambiguity and hampers discourse relating to big…
Gunther Eysenbach, Luca Toldo, Junfeng Gao, Weiqi Wang + 1 more
Background In the past few decades, medically related data collection saw a huge increase, referred to as big data. These huge datasets bring challenges in storage, processing, and analysis. In clinical medicine, big data is expected to play an important role in identifying causality of patient symptoms, in predicting…
Scott Monteith, Tasha Glenn, John Geddes, Michael Bauer
Big data are coming to the study of bipolar disorder and all of psychiatry. Data are coming from providers and payers (including EMR, imaging, insurance claims and pharmacy data), from omics (genomic, proteomic, and metabolomic data), and from patients and non-providers (data from smart phone and Internet activities…
Luca Clissa, Mario Lassnig, Lorenzo Rinaldi
The contemporary surge in data production is fueled by diverse factors, with contributions from numerous stakeholders across various sectors. Comparing the volumes at play among different big data entities is challenging due to the scarcity of publicly available data. This survey aims to offer a comprehensive…
Branimir K. Hackenberger
Simply put, Hadoop is an open-source storage and a data storage framework for large data sets, which stores data via a so-called distributed Hadoop file system, while MapReduce takes care of the processing. MapReduce is, in other words, a software model that enables the processing of massive data stored in Hadoop.…
Kevin Taylor-Sakyi
—Steve Jobs, one of the greatest visionaries of our time was quoted in 1996 saying "a lot of times, people don't know what they want until you show it to them"[38] indicating he advocated products to be developed based on human intuition rather than research. With the advancements of mobile devices, social networks and…
M. Zanin, D. Papo, P. A. Sousa, E. Menasalvas + 3 more
The increasing power of computer technology does not dispense with the need to extract meaningful in-formation out of data sets of ever growing size, and indeed typically exacerbates the complexity of this task. To tackle this general problem, two methods have emerged, at chronologically different times, that are now…
Nataliya Shakhovska, Uyrii Bolubash, Oleh Veres
The article deals with the problem which led to Big Data. Big Data information technology is the set of methods and means of processing different types of structured and unstructured dynamic large amounts of data for their analysis and use of decision support. Features of NoSQL databases and categories are described.…
Shakhovska Nataliya, Veres Oleh, Mariia Hirnyak
This article dwells on the basic characteristic features of the Big Data technologies. It is analyzed the existing definition of the "big data" term. The article proposes and describes the elements of the generalized formal model of big data. It is analyzed the peculiarities of the application of the proposed model…
Alex M. Ascension, Marcos J. Araúzo-Bravo
Big Data analysis is a discipline with a growing number of areas where huge amounts of data is extracted and analyzed. Parallelization in Python integrates Message Passing Interface via mpi4py module. Since mpi4py does not support parallelization of objects greater than 2^31^ bytes, we developed BigMPI4py, a Python…
L. Clissa, M. Lassnig, L. Rinaldi
The modern increase in data production is driven by multiple factors, and several stakeholders from various sectors contribute to it. Although drawing a comparison of the sizes at stake for different big data players is hard due to the lack of official data, this report tries to reconstruct the yearly orders of…
Raghavendra Kune, Pramodkumar Konugurthi, Arun Agarwal, Raghavendra Rao Chillarige + 1 more
'Raghavendra Rao Chillarige' 'Rajkumar Buyya'] Abstract- Advances in information technology and its widespread growth in several areas of business, engineering, medical and scientific studies are resulting in information/data explosion. Knowledge discovery and decision making from such rapidly growing voluminous data…
B. Manjulatha, Suresh Pabboju
The term, Big Data, has been authored to refer to the extensive heave of data that can't be managed by traditional data handling methods or techniques. The field of Big Data plays an indispensable role in various fields, such as agriculture, banking, data mining, education, chemistry, finance, cloud computing…
Sean M. Bagshaw, Stuart L. Goldstein, Claudio Ronco, John A. Kellum
''] The world is immersed in “big data”. Big data has brought about radical innovations in the methods used to capture, transfer, store and analyze the vast quantities of data generated every minute of every day. At the same time; however, it has also become far easier and relatively inexpensive to do so. Rapidly…
Leticia Leone Lauricella, Paulo Manuel Pêgo-Fernandes
“Information is the oil of the 21st century, and analytics is the combustion engine” said Peter Sondergaard, senior vice president of Gartner Research. The more information we have, the more likely we are to find correlations that are not obvious to the eye and that can completely change the way we think or act. We are…
Pablo Pareja-Tobes, Raquel Tobes, Marina Manrique, Eduardo Pareja + 1 more
Next Generation Sequencing and other high-throughput technologies have brought a revolution to the bioinformatics landscape, by offering sheer amounts of data about previously unaccessible domains in a cheap and scalable way. However, fast, reproducible, and cost-effective data analysis at such scale remains elusive. A…
Simeone Marino, Yi Zhao, Nina Zhou, Yiwang Zhou + 9 more
Health advances are contingent on continuous development of new methods and approaches to foster data driven discovery in the biomedical and clinical health sciences. Open-science offers hope for tackling some of the challenges associated with Big Data and team-based scientific discovery. Domain-independent…
Matthias Scheffler, Stefan Bauer, Peter Benner, Tristan Bereau + 57 more
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Matthias Scheffler
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Yasin El Abiead, Michael Strobel, Thomas Payne, Eoin Fahy + 13 more
Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, we've developed pan-repository universal identifiers and harmonized cross-repository metadata. This novel ecosystem…
Yue Chang, Wei Du, Ruo Shi, Tang Lei + 4 more
This study explored the theory of medical enterprise management and big data. Based on the Delphi method, two rounds of expert opinions were consulted on the capability of a health care enterprise big data application index system covering 11 dimensions, 46 first-level indicators and 111 second-level indicators. The…
Rebecca Brunk, Kriti Shukla, Bryant Hutson, Yue Wang + 7 more
Genomic sequencing and other big biological data is unquestionably of paramount value, however the success in recruiting highly skilled individuals with diverse backgrounds has been limited. A main reason for this deficiency could be due to the lack of educational resources and early exposure to the field. With the…
Authors not listed
Raman spectroscopy is an increasingly powerful and fast-growing analytical technique across diverse disciplines, from materials science and chemistry to biology and medicine, thanks to advances in Raman instrumentation and greatly supported by the flourishing of chemometrics and artificial intelligence (AI). However…
Ricardo Stefani
The use of data science, artificial intelligence, and big data in the field of chemistry has recently grown to speed up the discovery of new materials, drugs, and synthetic substances and the identification of automated compounds. Machine learning and data science are commonly used in organic chemistry to predict…
Jack D. Huey, Nezar Abdennur
The BigWig and BigBed file formats were originally designed for the visualization of next-generation sequencing data through a genome browser. Due to their versatility, these formats have long since become ubiquitous for the storage of processed sequencing data and regularly serve as the basis for downstream data…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…