21 papers · ranked by Valyu relevance
Arhit Chakrabarti, Yang Ni, Yuchao Jiang, Bani K. Mallick
We consider the problem of clustering nested or hierarchical data, where observations are grouped and there are both group-level and observation-level variables. In our motivating OneK1K dataset, observations consist of single-cell RNA-sequencing (scRNA-seq) data from 982 individuals (groups), totaling 1.27 million…
Harshavardhanan Deekeswar
Serialization formats designed for document interchange impose structural overhead that becomes prohibitive when large language models consume operational data at scale. A modest dataset of 1,000 IoT sensor readings serialized as JSON requires approximately 80,000 tokens - the majority spent on repeated field names…
Mymuna Monem, Ian L. Dryden, Florence George
The method of Principal Nested Spheres (PNS) is a non-linear dimension reduction technique for spherical data. The method is a backwards fitting procedure, starting with fitting a high-dimensional sphere and then successively reducing dimension at each stage. After reviewing the PNS method in detail, we introduce some…
Kyle Cox, Benjamin Kelcey
Bayesian and structural-after-measurement (SAM) approaches have been developed, in part, to address limitations of conventional estimators in the context of structural equation models (SEMs) with latent interactions. Although both approaches have shown promise in a variety of contexts including small-sample studies…
Timothy LaRock, Yanting Zhang, Jean-Gabriel Young, Nicole Eikmeier + 2 more
In contrast to dyadic interactions, higher-order interactions may contain one another, with subgroups naturally embedded within larger groups. These containment patterns arise empirically in ecology, sociology, computer science and the science of science, and have been studied under the names nestedness, simpliciality…
Alejandro Flores Sepúlveda, Jean-Louis Reymond
Visualize Billions of Molecules Authors: Alejandro Flores Sepúlveda, Jean-Louis Reymond Here, we present a visualization and clustering framework enabling the exploration of billion-sized chemical data sets, exemplified with the REAL database of 9.6 billion make-on-demand molecules. We represent molecules as…
Perry A. LaBoone, Antara Anika Piya, Raquel Assis
Gene expression divergence is a major source of phenotypic variation, yet the factors that shift regulatory optima remain incompletely understood. In particular, it is unclear how broad taxonomic differences and local genomic architecture interact to shape the tempo and mode of expression evolution. Here, we analyze…
Chunqing (Tony) Liang, Tajveer Grewal, Asees Singh, Amrit Singh
Multimodal biomedical studies increasingly profile multiple molecular and clinical modalities from the same samples, creating new opportunities for disease prediction and biological discovery. However, benchmarking multimodal integration methods remains difficult because studies often use inconsistent preprocessing…
Angela John, Selvyn Allotey, Till Koebe, Alexandra Tyukavina + 1 more
Afforestation and reforestation are popular strategies for mitigating climate change by enhancing carbon sequestration. However, the effectiveness of these efforts is often self-reported by project developers, or certified through processes with limited external validation. This leads to concerns about data reliability…
Authors not listed
Accurate prediction of redox potentials of iron (Fe) complexes, in tandem with uncertainty quantification, is essential to advance technologies related to electro-deposition and energy storage by enabling reliable modeling, guiding experimental design, and improving the efficiency of material discovery. Since…
Daniel Ohene-Kwofie, Nkosinathi Masilela, Kerry Glover, Samuel Iddi + 6 more
Key Features1. The Multimorbidity in Africa Digital Innovation, Visualisation, and Application (MADIVA) research hub was established to facilitate the study of multimorbidity and associated risk factors, disease clustering, and stratification in African populations. 2. The hub has harmonized and integrated data from…
Xi Shen, Andreas Bjerregaard, Yan Li, Anders Krogh
Biomedical datasets are often heterogeneous and affected by noise or missing values. Deep Generative Decoders (DGD) provide a promising framework for latent representation learning, but their standard training procedure relies on sample-level stochastic gradient descent (SGD), which performs poorly with incomplete…
Qiyang Chen, Guozheng Li, Xingqi Wang, Gerile Aodeng + 2 more
Hierarchical tables are an important structure for organizing data with inherent hierarchical relationships. Existing studies have extensively explored methods for data fact exploration from tabular data. In particular, some studies have directly integrated visual data facts into the original table structure to support…
Mathew Thomson, Jean-David Therrien, Nikho Hizon, Janet Ting-mei Lin + 6 more
Wastewater surveillance (WWS) has quickly emerged as an invaluable tool for public health surveillance, particularly in the wake of the COVID-19 pandemic. Its long-term utility is constrained, however, by fragmented data systems, inconsistent metadata practices, and poor interoperability. The Public Health and…
Ignacio Castillo-Barrios, Melesio Crespo-Sanchez, Hugo G. Reyes-Anastacio, Jose L. Gonzalez-Compean + 7 more
This paper presents Jub, a Life Science and Healthcare Data Platform (LSHDP) based on generic sandboxes that integrate AI tools and cloud storage into big data science services. Jub automatically and transparently creates data science services to transform datasets into massive information products by using a profiling…
Heinrich Lukas Weil, Kevin Schneider, Dominik Brilhaus, Timo Mühlhaus + 1 more
Scientific communication depends on the production of FAIR (findable, accessible, interoperable, reusable) data. Yet when datasets are shared, they remain hard to interpret across domains because annotations are tied to heterogeneous file formats and implicit semantics. We present a format-agnostic method for…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Daniel Schönberger, Zachary G. MacDonald, B. Christian Schmidt, Julian R. Dupuis
Quantifying niche divergence is crucial to understanding the ecological and evolutionary processes underlying range limits, coexistence, speciation, biogeography, and macroevolution. Yet available approaches rely on low-dimensional climate summaries, are vulnerable to multiple biases, or struggle with high-dimensional…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Tobias K. Mildenberger, Federico Maioli, Casper W. Berg
Scientific bottom-trawl surveys provide essential fisheries-independent data for fisheries and ecosystem research. In the Northeast Atlantic, the ICES Database of Trawl Surveys (DATRAS) compiles haul-level information, species- and length-specific catch data, and individual biological observations across multiple…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…