24 papers · ranked by Valyu relevance
Nawsher Khan, Ibrar Yaqoob, Ibrahim Abaker Targio Hashem, Zakira Inayat + 4 more
'Zakira Inayat' 'Waleed Kamaleldin Mahmoud Ali' 'Muhammad Alam' 'Muhammad Shiraz' 'Abdullah Gani'] Big Data has gained much attention from the academia and the IT industry. In the digital and computing world, information is generated and collected at a rate that rapidly exceeds the boundary range. Currently, over 2…
By Huan Chen, Brian Caffo, Genevieve Stein-O’Brien, Jinrui Liu + 3 more
Integrative analysis of multiple data sets has the potential of fully leveraging the vast amount of high throughput biological data being generated. In particular such analysis will be powerful in making inference from publicly available collections of genetic, transcriptomic and epigenetic data sets which are designed…
Mike Phuycharoen, Verena Kaestele, Thomas Williams, Lijing Lin + 3 more
We introduce the Unbiasing Variational Autoencoder (UVAE), a novel computational framework developed for the integration of unpaired biomedical data streams, with a particular focus on clinical flow cytometry. UVAE effectively addresses the challenges of batch effect correction and data alignment by training a…
José L. Torrecilla, Juan Romo
Technology is generating a huge and growing availability of observations of diverse nature. This big data is placing data learning as a central scientific discipline. It includes collection, storage, preprocessing, visualization and, essentially, statistical analysis of enormous batches of data. In this paper, we…
Stefan Petrescu, Floris den Hengst, Alexandru Uta, Jan S. Rellermeyer
'Jan S. Rellermeyer'] Abstract—Due to the complexity and size of modern software systems, the amount of logs generated is tremendous. Hence, it is infeasible to manually investigate these data in a reasonable time, thereby requiring automating log analysis to derive insights about the functioning of the systems.…
Jinfeng Wang, Robert Haining, Tonglin Zhang, Chengdong Xu + 1 more
'Huanjun Liu'] Abstract. Spatial statistics is dominated by spatial autocorrelation (SAC) based Kriging and BHM, and spatial local heterogeneity based hotspots and geographical regression methods, appraised as the first and second laws of Geography (Tobler 1970; Goodchild 2004), respectively. Spatial stratified…
Yipeng Song
ter verkrijging van de graad van doctor aan de Universiteit van Amsterdam op gezag van de Rector Magnificus prof. dr. ir. K.I.J. Maex ten overstaan van een door het College voor Promoties ingestelde commissie, in het openbaar te verdedigen in de Aula der Universiteit op vrijdag 27 september 2019 om 11.00 uur
Todor Ivanov, Nikolaos Korfiatis, Roberto V. Zicari
The well-known 3V architectural paradigm for Big Data introduced by Laney (2011) provides a simplified framework for defining the architecture of a big data platform to be deployed in various scenarios tackling processing of massive datasets. While additional components such as Variability and Veracity have been…
Cindy Cheng, Luca Messerschmidt, Isaac Bravo, Marco Waldbauer + 6 more
Data harmonization is an important method for combining or transforming data. To date however, articles about data harmonization are field-specific and highly technical, making it difficult for researchers to derive general principles for how to engage in and contextualize data harmonization efforts. This commentary…
Authors not listed
Artificial intelligence (AI) is poised to transform heterogeneous catalysis, ushering in a new paradigm for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental, and chemical…
Neal D. Goldstein, Brianne Olivieri-Mui, Igor Burstyn
There has been a proliferation of large-scale electronic health record (EHR) data platforms that pool across multiple healthcare organizations, such as the National Institutes of Health’s All of Us in the federal space and TriNetX and Epic Cosmos in the commercial space. There are unique issues that occur when EHR data…
Ramalingam Shanmugam, Gerald Ledlow, Karan P. Singh
In this paper, heterogeneity is formally defined, and its properties are explored. We define and distinguish observable versus non-observable heterogeneity. It is proposed that heterogeneity among the vulnerable is a significant factor in the contagion impact of COVID-19, as demonstrated with incidence rates on a…
Jingru Zhang, Erjia Cui, Hongzhe Li, Haochang Shou
Wearable devices and digital phenotyping are increasingly used in observational and interventional studies to measure real-time biosignals such as physical activity. However, integrating and comparing data across studies and cohorts remains challenging due to variability in device types, acquisition protocols, and…
Authors not listed
Applications of deep learning (DL) to design nanomaterials are hampered by a lack of suitable data representations and training data. We report efforts to overcome these limitations and leverage DL to optimize the nonlinear optical properties of core-shell upconverting nanoparticles (UCNPs). UCNPs, which have…
Jean-Baptiste Cazier, Liudmila Sergeevna Mainzer, Weihao Ge, Justina Žurauskienė + 1 more
The world is diverse, and this needs to be better recognized and addressed in health research. Health Disparities (HD) are a growing concern, which affects not only the world at a global scale, but individual countries and their own diversity . The spectrum of individual health is moulded not solely by genetics or…
Zachary Madaj, Mao Ding, Carmen Khoo, Ember Tokarski + 7 more
Disease heterogeneity is a persistent challenge in medicine, complicating both research and treatment. Standard analytical pipelines often assume patient populations are homogeneous, overlooking variance patterns that may signal biologically distinct subgroups. Variance heterogeneity (VH)—including skewness, outliers…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
Hal Caswell, Silke F. van Daalen
A heterogeneous population is a mixture of groups differing in vital rates. In such a population, some of the variance in demographic outcomes (e.g., longevity, lifetime reproduction) is due to heterogeneity and some is the result of stochastic demographic processes. Many studies have partitioned variance into its…
Yuxiang Xie, Nanyu Chen, Xiaolin Shi
Online controlled experiments (a.k.a. A/B testing) have been used as the mantra for data-driven decision making on feature changing and product shipping in many Internet companies. However, it is still a great challenge to systematically measure how every code or feature change impacts millions of users with great…
Matthias Scheffler
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Matthias Scheffler, Stefan Bauer, Peter Benner, Tristan Bereau + 57 more
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Mike W.-L. Cheung, Suzanne Jak
Big data is a field that has traditionally been dominated by disciplines such as computer science and business, where mainly data-driven analyses have been performed. Psychology, a discipline in which a strong emphasis is placed on behavioral theories and empirical research, has the potential to contribute greatly to…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…
Scott Monteith, Tasha Glenn, John Geddes, Michael Bauer
Big data are coming to the study of bipolar disorder and all of psychiatry. Data are coming from providers and payers (including EMR, imaging, insurance claims and pharmacy data), from omics (genomic, proteomic, and metabolomic data), and from patients and non-providers (data from smart phone and Internet activities…