25 papers · ranked by Valyu relevance
Federico Ottomano, Giovanni De Felice, Vladimir Gusev, Taylor Sparks
Recent Machine Learning (ML) developments have opened new perspectives on accelerating the discovery of new materials. However, in the field of materials informatics, the performance of ML estimators is heavily limited by the nature of the available training datasets, which are often severely restricted and unbalanced.…
Minh-Duc Nguyen, A. P. Kryukov, Julia Dubenskaya, Е. Е. Коростелева + 3 more
'Igor Bychkov' 'Andrey Mikhailov' 'Alexey Shigarov'] > Abstract. German-Russian Astroparticle Data Life Cycle Initiative is an international project whose aim is to develop a distributed data storage system that aggregates data from the storage systems of different astroparticle experiments. The prototype of such a…
Haikady N. Nagaraja, Shane Sanders, Alan D Hutson
The relationship between social choice aggregation rules and non-parametric statistical tests has been established for several cases. An outstanding, general question at this intersection is whether there exists a non-parametric test that is consistent upon aggregation of data sets (not subject to Yule-Simpson…
Vasiliki Rahimzadeh
The Office of the National Coordinator for Health Information Technology estimates that 96% of all U.S. hospitals use a basic electronic health record, but only 62% are able to exchange health information with outside providers. Barriers to information exchange across EHR systems challenge data aggregation and analysis…
Maria Ulan, Welf Löwe, Morgan Ericsson, Anna Wingkvist
Regression uses supervised machine learning to find a model that combines several independent variables to predict a dependent variable based on ground truth (labeled) data, i.e., tuples of independent and dependent variables (labels). Similarly, aggregation also combines several independent variables to a dependent…
Ana Helena Tavares, Ana Silva, Tiago Freitas, Maria Costa + 3 more
Despite the advances on data analysis methodologies in the last decades, most of the traditional regression methods cannot be directly applied to large-scale data. Although aggregation methods are especially designed to deal with large-scale data, their performance may be strongly reduced in ill-conditioned problems…
Katarzyna Malinowska, Michał Wawrzynowicz, Katarzyna Markowska, Tomasz Chodkiewicz + 2 more
Conservation decision-making requires accurate identification of causes of population changes. Ecologists often rely on analytical protocols that aggregate high-dimensional monitoring data. We hypothesise that compressing data - either spatially, as in conventional time series (TS) analysis, or temporally, as in static…
Hezheng Lyu, Hassan Gharibi, Zhaowei Meng, Bohdana Sokolova + 2 more
Protein-level statistical tests in proteomics aimed at obtaining p-value are conventionally made on protein abundances aggregated from peptide data. This integral approach overlooks peptide-level heterogeneity and ignores important information coded in individual peptide data, while protein p-value can also be obtained…
Chengjie Qin, Florin Rusu
In this paper we introduce the first framework for parallel online aggregation in which the estimation virtually does not incur any overhead on top of the actual execution. We define a generic interface to express any estimation model that abstracts completely the execution details. We design a novel estimator…
Nico M. Franz, Beckett W. Sterner
Growing concerns about the quality of aggregated biodiversity data are lowering trust in large-scale data networks. Aggregators frequently respond to quality concerns by recommending that biologists work with original data providers to correct errors "at the source". We show that this strategy falls systematically…
Eric Simon, Bernd Amann, Rutian Liu, Stéphane Gançarski
We present a comprehensive set of conditions and rules to control the correctness of aggregation queries within an interactive data analysis session. The goal is to extend self-service data preparation and BI tools to automatically detect semantically incorrect aggregate queries on analytic tables and views built by…
Kanat Tangwongsan, Martin Hirzel, Scott Schneider
Sliding-window aggregation is a widely-used approach for extracting insights from the most recent portion of a data stream. The aggregations of interest can usually be expressed as binary operators that are associative but not necessarily commutative nor invertible. Non-invertible operators, however, are difficult to…
Yancheng Shi, Zhenjiang Zhang, Han-Chieh Chao, Bo Shen
With the rapid development of information technology, large-scale personal data, including those collected by sensors or IoT devices, is stored in the cloud or data centers. In some cases, the owners of the cloud or data centers need to publish the data. Therefore, how to make the best use of the data in the risk of…
Katarína Furmanová, Samuel Gratzl, Holger Stitz, Thomas Zichner + 3 more
'Miroslava Jarešová' 'Alexander Lex' 'Marc Streit'] Most tabular data visualization techniques focus on overviews, yet many practical analysis tasks are concerned with investigating individual items of interest. At the same time, relating an item to the rest of a potentially large table is important. In this work we…
Jaclyn Smith, Yao Shi, Michael Benedikt, Milos Nikolic
Targeted diagnosis and treatment options are dependent on insights drawn from multi-modal analysis of large-scale biomedical datasets. Advances in genomics sequencing, image processing, and medical data management have supported data collection and management within medical institutions. These efforts have produced…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
Zhaoyang Zhang, Christopher R. Cotter, Zhe Lyu, Lawrence J. Shimkets + 1 more
Single mutations frequently alter several aspects of cell behavior but it is often not clear whether a particular statistically significant change is biologically significant. To determine which behavioral changes are most important for multicellular self-organization, we devised a new methodology using Myxococcus…
Cindy Cheng, Luca Messerschmidt, Isaac Bravo, Marco Waldbauer + 6 more
Data harmonization is an important method for combining or transforming data. To date however, articles about data harmonization are field-specific and highly technical, making it difficult for researchers to derive general principles for how to engage in and contextualize data harmonization efforts. This commentary…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Jacob Bien, Xiaohan Yan, Léo Simpson, Christian L. Müller
Modern high-throughput sequencing technologies provide low-cost microbiome survey data across all habitats of life at unprecedented scale. At the most granular level, the primary data consist of sparse counts of amplicon sequence variants or operational taxonomic units that are associated with taxonomic and…
Bradley Voytek, Philip E. Bourne
Another benefit of large-scale data availability is that it could uncover sampling bias by allowing researchers to combine data from multiple studies. For example, sampling bias is rampant in psychology, in which 96% of studies published from the top six psychology journals consisted of data collected from people…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Leyla Cabugos, Katheryn Buble, Jenna Daenzer, Sook Jung + 7 more
To improve the FAIRness of agricultural genomic, genetic, and breeding ( [GGB]() ) data, the AgBioData FAIR Scientific Literature Working Group developed a free search tool that helps researchers identify appropriate databases for submitting their data. Existing repository discovery tools lack the specificity needed…
Belinda Boehm, Christopher McNeill, David Huang
Understanding the solution-phase behaviour of organic semiconducting polymers is important for systematically improving the performance of devices based on solution-processed thin films of these molecules. Conventional polymer theory predicts that polymer conformations become more compact as solvent quality decreases…
Aaron Liu, Myeongyeon Lee, Rahul Venkatesh, Jessica Bonsu + 4 more
Polymer-based semiconductors and organic electronics encapsulate a significant research thrust for informatics-driven materials development. However, device measurements are described by a complex array of design and parameter choices, many of which are sparsely reported. For example, the mobility of a polymer-based…