21 papers · ranked by Valyu relevance
Ana Helena Tavares, Ana Silva, Tiago Freitas, Maria Costa + 3 more
Despite the advances on data analysis methodologies in the last decades, most of the traditional regression methods cannot be directly applied to large-scale data. Although aggregation methods are especially designed to deal with large-scale data, their performance may be strongly reduced in ill-conditioned problems…
Hayat Mohammad Khan, Farhana Jabeen, Abid Khan, Muhammad Waqar + 2 more
The Internet of Things (IoT) has transformed multiple industries, providing significant potential for automation, efficiency, and enhanced decision-making. The incorporation of IoT and data analytics in smart grid represents a groundbreaking opportunity for the energy sector, delivering substantial advantages in…
Qiang Qin, Yongjiao Yang, Jiaxin Lin, Hanye Huang + 1 more
The widespread adoption of Internet of Things (IoT) devices has opened new possibilities for data-driven decision making while simultaneously raising serious concerns about the protection of sensitive personal information. This paper presents an integrated privacy-preserving data aggregation framework that…
Hezheng Lyu, Hassan Gharibi, Zhaowei Meng, Bohdana Sokolova + 2 more
Protein-level statistical tests in proteomics aimed at obtaining p-value are conventionally made on protein abundances aggregated from peptide data. This integral approach overlooks peptide-level heterogeneity and ignores important information coded in individual peptide data, while protein p-value can also be obtained…
Xiang Huang, Lei Jin, Kequan Lin, Wenpeng Wu + 2 more
In a cloud-edge-end collaborative system, data generated by terminal devices often contains users’ sensitive information and is constantly generated and changing, leading to potential data privacy leaks in caches. Additionally, due to the inability to promptly capture these dynamic changes and the failure to consider…
Rohul Amin, Md. Mehedi Hasan Rana, Sumya Aktar
Federated learning (FL) enables collaborative clinical model training without centralized data sharing, yet its deployment is hindered by statistical heterogeneity (non-IID data) and inherent class imbalance across healthcare institutions. Conventional aggregation strategies such as FedAvg and FedProx weight client…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
Stephen Kasica, Charles Berret, Tamara Munzner
Data journalists routinely integrate records across multiple independently published sources to support accountability reporting, yet no existing interactive wrangling tool treats the collection of tables -- rather than the single table -- as its primary unit of work. We present OpenRoundup, an open-source…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Kristýna Petrlíková, Tomas Petricek
Vega is a popular declarative language for creating interactive data visualizations. It supports reactive data transformations using its streaming dataflow architecture. Despite its widespread adoption, the exact semantics of Vega is subtle and poorly documented. This leads to incorrect or confusing visualizations and…
Zheng Zhang, Cuong C. Nguyen, Kevin Wells, Gustavo Carneiro
The rapid development of large language models (LLMs) has motivated research on decision-making in multi-agent systems, where multiple agents collaborate to achieve shared objectives. Existing aggregation approaches, such as voting and debate, are largely ad-hoc and lack formal guarantees regarding the informativeness…
Richard Arthur, Virginia DiDomizio, Louis Hoebel
In some complex domains, certain problem-specific decompositions can provide advantages over monolithic designs by enabling comprehension and specification of the design. In this paper we present an intuitive and tractable approach to reasoning over large and complex data sets. Our approach is based on Active Data…
Cydney M. Martell, Kamil K. Gebis, Han My Van, Yulia M. Gutierrez + 3 more
Protein aggregation is an obstacle for engineering effective recombinant proteins for biotechnology and therapeutic applications. Predicting protein aggregation propensity remains challenging due to the complex interplay of sequence, structure, environmental factors, and external stress conditions, particularly for…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Puneet Rawat, R. Prabakaran, Niccolò Cardente, Sandeep Kumar + 2 more
Protein aggregation is central to amyloid-related disorders and remains a major developability challenge for protein therapeutics. Over the past two decades, significant advances have been made to predict aggregation-prone regions (APRs) and estimate aggregation propensity in proteins and peptides. In contrast, the…
Rubén Grillo-Risco, Maksym Kupchyk Tiurin, Carla Perpiñá-Clérigues, Francisco J. Cordero Felipe + 3 more
The growing number of omics datasets in public repositories provides an opportunity to enhance data reusability through data integration; however, complex statistical barriers often hinder the effective combination of independent studies. To address this problem, we present MetaOmixTools, an interactive web-based suite…
Authors not listed
Predicting solution conformation and aggregation of conjugated polymers remains a bottleneck for translating solution processing into controlled film microstructure and for closing the loop in self-driving laboratories. We construct a cleaned, machine-readable dataset of 256 entries that links polymer size…
Xi Zheng, Yinghui Huang, Xiangyu Chang, Ruoxi Jia + 1 more
Rigorous valuation of individual data sources is critical for fair compensation in data markets, informed data acquisition, and transparent development of ML/AI models. Classical Data Shapley (DS) provides a essential axiomatic framework for data valuation but is constrained by its symmetry axiom that assumes…
Authors not listed
The analysis of metabolic profiles using high resolution mass spectrometry (MS) data gives deep insights into the biological processes. In metabolomics, MS generates a large number of features that represent metabolites. However, identifying specific metabolites from these features can be challenging. One of the major…
Authors not listed
mRNA-based technology has emerged as a new class of medicines with a wide range of applications, including viral vaccines, cancer vaccines, and therapeutics for the treatment of metabolic diseases and cardiovascular conditions. Impurities, including double-stranded RNA (dsRNA), mRNA fragments, and mRNA multimers…
Joanna Schroeder, Alan Wang, Kathryn Linehan, Joel Thurston + 1 more
This article describes the use of metadata and standards in the Social Impact Data Commons to expose official statisticians to an innovative project built on actionable and evaluable metadata, which produces a FAIR data system. We begin by introducing the concept of the Data Commons, focusing on its features, and…