23 papers · ranked by Valyu relevance
Helder Prado Santos, Methanias Colaço Júnior, Ricardo Valentim, João Paulo Queiroz dos Santos + 8 more
Context The growth of data in healthcare brings both challenges and opportunities. The term Big Data refers to the handling of large volumes of data using advanced techniques and scalable infrastructure. Distributed processing and parallel computing accelerate data processing and necessitate a robust architecture. A…
Javier Gamboa-Cruzado, Kiara Fernandez-Perez, Herber Laura-Abarca, Cristina Alzamora Rivero + 4 more
Introduction The application of Big Data in healthcare enhances disease management by improving diagnostics and treatments, offering a foundation for positive impacts on the health of diverse populations. Objective This paper aims to analyze the impact of Big Data on healthcare through a comprehensive review of…
Yan Li, Tao Huang, Xiang Li, Ahmed Abdelwahab Ibrahim El-Sayed
This study examines the discrepancy between big data talent training and industry demand. The study analyzed 85 training programs and over 10,000 job postings from two job boards in China (51job and Zhaopin). Using content analysis, social network analysis, and the BERTopic-TOPSIS model, it mined implicit information…
Hugo Morvan, Jonas Agholme, Bjorn Eliasson, Katarina Olofsson + 3 more
Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling…
Yasir Ali Soomro, Yasser Baeshen
Introduction This study examines the impact of Customer Big Data Analytics (CBDA) on customer satisfaction and firm performance in business-to-business (B2B) firms operating in emerging markets, specifically Pakistan. Despite the growing adoption of big data technologies, empirical evidence on their strategic value in…
Sonia Katyal
In this Article, I explore the impending conflict between the protection of civil rights and artificial intelligence (AI). While both areas of law have amassed rich and well-developed areas of scholarly work and doctrinal support, a growing body of scholars are interrogating the intersection between them. This Article…
Evelyn Smith, Emma Harvey, Jacob Goldin, Daniel E. Ho
Big data are the foundation for an increasing share of academic research and AI models deployed in both the public and private sectors, prompting substantial growth over time in reliance on brokered datasets. Brokered property records, which are ubiquitous in studies of gentrification, inequality, and the property tax…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Enrico Seiler, Myrthe Willemsen, Vitor C. Piro, Knut Reinert
A continued decrease in sequencing costs has facilitated the exponential increase in available sequencing data, with public databases like the European Nucleotide Archive (ENA) and Sequence Read Archive (SRA) reaching well in the order of petabases [1, 2]. This has been the incentive to develop more scalable tools for…
Song Rixin, Ramayah, Ma Xinrui, Syed Hamid Hussain Madni
This study explores how small and medium-sized enterprises (SMEs) can leverage big data analytics (BDA) to gain competitive advantage (CA), highlighting the mediating role of data-driven innovation (DDI) and the foundational importance of data governance (DG) and data-driven culture (DDC). Despite the transformative…
Edgar Ribeiro João, Manuel Parra-Royón, Julián Garrido
The unprecedented volume of data from the Square Kilometre Array (SKA) telescopes will require the implementation of robust and solid strategies for efficient data processing and management. In this context, the SKA Regional Centre Network (SRCNet) a collaborative global infrastructure comprising multiple regional…
Uwaise Ibna Islam, Davide Cozzi, Travis Gagie, Rahul Varki + 4 more
Large, phase-resolved haplotype panels—now emerging from efforts such as UK Biobank, TOPMed, All of Us, and the Mexican Biobank—enable fine-grained analyses of admixture, local ancestry, and imputation. The Positional Burrows–Wheeler Transform (PBWT) is a natural index for these data, supporting efficient phase-aware…
Sebastian Bernhard Kloubert, Simon Renders, Oliver Bannach, Thomas Dino Rockel + 2 more
We present Ultra-Content Screening (UCS), a novel, scalable method combining cyclic immunostaining with high-dimensional image-based single-cell proteomics. UCS utilizes fluorescein isothiocyanate-conjugated antibodies and iterative staining-photobleaching cycles to analyze up to 40 markers in up to 100,000 peripheral…
Sayed Hoseini, Christoph Quix, Stefan Decker
Modern organizations generate and consume massive volumes of heterogeneous data at high speed. This requires a continuous development of new techniques for more efficient and reliable data management. Designing appropriate data architectures has therefore become a strategic necessity, as they shape how data is…
Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan
Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids. We target arbitrarily tall data: a fixed feature space may contain arbitrarily many, possibly…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Authors not listed
Solubility is a crucial property of each organic compound, impacting its potential applications in synthetic chemistry, materials science and drug design. Moreover, in technological processes mixtures of solvents are often utilized, making the solubility assessment more complicated. Predicting solubility values in…
Mahbuba Tasmin, Saishradha Mohanty, Sanjana Kulkarni, Maha R. Farhat + 1 more
Foundation models aim to learn useful representations of biological sequences. However, the applicability of these representations for a wide range of tasks, including phenotype prediction and variant discovery, is still in question, in large part due to the relatively small set of benchmark tasks. To this end, we…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Authors not listed
The analysis of metabolic profiles using high resolution mass spectrometry (MS) data gives deep insights into the biological processes. In metabolomics, MS generates a large number of features that represent metabolites. However, identifying specific metabolites from these features can be challenging. One of the major…
Authors not listed
High-throughput experimentation (HTE) accelerates chemical discovery by shortening the lead times for molecule synthesis. The choice of initial reaction conditions directly influences the outcome and length of a screening campaign. But human involvement in plate design and data analysis remains a significant cost…
Daniel Schütte, Thomas J. Y. Kono, Roland F. Schwarz
Single-cell DNA sequencing (scDNA-seq) has emerged as a primary method for studying the evolution of cancer genomes and intra-tumor heterogeneity. However, despite technological advances, scDNA-seq remains noisy and is affected by amplification biases and allelic dropouts. Accurately determining the presence or absence…
Authors not listed
Computational blind challenges offer critical, unbiased assessment opportunities to assess and accelerate scientific progress, as demonstrated by a breadth of breakthroughs over the last decade. We report the outcomes and key insights from an open science community blind challenge focused on computational methods in…