25 papers · ranked by Valyu relevance
Samyajoy Pal, Christian Heumann, Sheetal Kalyani
A model-based clustering method for compositional data is explored in this article. Most methods for compositional data analysis require some kind of transformation. The proposed method builds a mixture model using Dirichlet distribution which works with the unit sum constraint. The mixture model uses a hard EM…
Elliott Gordon-Rodríguez, Thomas P. Quinn, John P. Cunningham
Data augmentation plays a key role in modern machine learning pipelines. While numerous augmentation strategies have been studied in the context of computer vision and natural language processing, less is known for other data modalities. Our work extends the success of data augmentation to compositional data, i.e.…
Greenacre, Michael, Graeve, Martin
In certain fields where compositional data are studied, the compositional components, called parts, can be combined into certain subsets, called amalgamations, which are based on domain knowledge. Furthermore, these subsets can form a natural hierarchy of amalgamations subdividing into sub-amalgamations. The authors, a…
Abel Ruiz-Giralt, Stefano Biagetti, Carla Lancelotti, Óscar Parque + 4 more
Anthropic Activity Markers (AAMs) were formalized over 10 years ago as a toolkit to infer human activities from biological and geochemical signatures preserved in sediments. This paper presents AAMs 2.0, a revised analytical framework that integrates compositional data analysis (CoDA) and geostatistics, as the…
Jeseok Lee, Byungwon Kim, Hongchuan Yu
Through the Human Microbiome Project, research on human-associated microbiomes has been conducted in various fields. New sequencing techniques such as Next Generation Sequencing (NGS) and High-Throughput Sequencing (HTS) have enabled the inclusion of a wide range of features of the microbiome. These advancements have…
Jesse Pasanen, Tuija Leskinen, Kristin Suorsa, Anna Pulakka + 3 more
'Joni Virta' 'Kari Auranen' 'Sari Stenholm'] We utilized compositional data analysis (CoDA) to study changes in the composition of the 24-h movement behaviors during an activity tracker based physical activity intervention. A total of 231 recently retired Finnish retirees were randomized into intervention and control…
Onur Batın Doğan, Fatma Sevinç Kurnaz
Crime Trends for 2022 Authors: ['Onur Batın Doğan' 'Fatma Sevinç Kurnaz'] This article investigates crime patterns across European countries in 2022 using Compositional Data Analysis (CoDA) to address limitations of traditional statistical approaches in dealing with the relative nature of crime data. Recognizing crime…
Andrey Shternshis, Bangzhuo Tong, Carolina Wählby, Dave Zachariah + 2 more
Time-series of compositional data are a common format for many high-throughput studies of biological molecules, analyzing e.g. response to a treatment or with the aim to predict an outcome. However, data from some time points may be missing, which reduces the size of the complete dataset. We propose a method for binary…
Jingjing Ma, Dinesh Kumar Nishad
To improve the prediction accuracy of compositional data time series (CDTSs), the aggregation of compositional data was considered and applied to construct a combination forecasting model. Different from current arithmetic mean based aggregation of compositional data, the aggregation method of compositional data from…
Alberte Sloth Carlsen, Te Chen, Nicholas Luke Cowie, Christian Brinch + 2 more
Isotopic Metabolic Flux Analysis (I-MFA) is a standard approach for estimating intracellular metabolic fluxes. I-MFA infers fluxes by comparing simulated and measured metabolite isotopologue distributions (MIDs) of metabolites from isotope labeling experiments. MIDs represent fractional abundances that strictly sum to…
Eric Grunsky, Michael Greenacre, B A Kjarsgaard
Geochemical data are compositional in nature and are subject to the problems typically associated with data that are restricted to the real non-negative number space with constant-sum constraint, that is, the simplex. Geochemistry can be considered a proxy for mineralogy, comprised of atomically ordered structures that…
Alemu Takele Assefa, Bie Verbist, Koen Van den Berge
In single-cell studies, a common question is whether there is a change in cell composition between conditions. While ideally, one needs absolute cell counts (number of cells per volumetric unit in a sample) to address these questions, current experimentation typically obtains cell counts that only carry relative…
Siyuan Ma, Curtis Huttenhower, Lucas Janson
A major task of microbiome epidemiology is association analysis, where the goal is to identify microbial features related to host health. This is commonly performed by differential abundance (DA) analysis, which, by design, examines each microbe as isolated from the rest of the microbiome. This does not properly…
Stephanie Wankowicz, James Fraser
In their folded state, biomolecules exchange between multiple conformational states, crucial for their function. However, most structural models derived from experiments and computational predictions only encode a single state. To represent biomolecules more accurately, we must move towards modeling and predicting…
Samuel D. Gamboa-Tuz, Marcel Ramos, Eric Franzosa, Curtis Huttenhower + 3 more
Previous benchmarking of differential abundance (DA) analysis methods in microbiome studies have employed synthetic data, simulations, and “real data” examples, but to the best of our knowledge, none have yet employed experimental data with known “ground truth” differential abundance. A key debate in the field centers…
Stephanie Wankowicz, James Fraser
In their folded state, biomolecules exchange between multiple conformational states, crucial for their function. However, most structural models derived from experiments and computational predictions only encode a single state. To represent biomolecules more accurately, we must move towards modeling and predicting…
Maxwell Venetos, Masha Elkin, Connor Delaney, John Hartwig + 1 more
NMR spectroscopy is an important analytical technique in synthetic organic chemistry, but its integration into high-throughput experimentation workflows has been limited by the necessity to manually analyze NMR spectra of new chemical entities. Current efforts to automate the analysis of NMR spectra rely on comparisons…
Michael Hagmann, Michael Staniek, Stefan Riezler
This work investigates whether time series of natural phenomena can be understood as being generated by sequences of latent states which are ordered in systematic and regular ways. We focus on clinical time series and ask whether clinical measurements can be interpreted as being generated by meaningful physiological…
Authors not listed
The vastness of chemical space presents a long-standing challenge for the exploration of new compounds with pre-determined properties. In materials science, crystal structure prediction has become a mature tool for mapping from composition to structure based on global optimisation techniques. Generative artificial…
Authors not listed
Transition metal phosphates (TMPs) are extensively explored for electrochemical and catalytical applications due to their structural versatility and chemical stability. Within this material class, novel high-entropy metal phosphates (HEMPs)—containing multiple transition metals combined into a single-phase…
Eugene Wu
We now relax the safety rules for statistical composition op for cases like Figure 5(e,f) where the two views do not have identical query schemas, but the grouping attributes A 1 gb in Q1 are a strict super set of the grouping attributes A 2 gb in Q2. In these cases, each row in Q2 potentially matches many rows in Q1…
Stephen A. Allegri, Kevin McCoy, Cassie S. Mitchell
Large networks are quintessential to bioinformatics, knowledge graphs, social network analysis, and graph-based learning. CompositeView is a Python-based open-source application that improves interactive complex network visualization and extraction of actionable insight. CompositeView utilizes specifically formatted…
Yu-Chieh Huang, Pierre Tremouilhac, Stefan Kuhn, Pei-Chi Huang + 6 more
A method for data review in chemical sciences with a focus on data for the characterization of synthetic molecules is described. As current procedures for data curation in chemistry rely almost exclusively on manual checking or peer reviewing, a (semi-)automatic procedure for the evaluation of data assigned to…
Hyunsoo Park, Anthony Onwuli, Keith T. Butler, Aron Walsh
The combination of elements from the Periodic Table defines a vast chemical space. Only a small fraction of these combinations yield materials that occur naturally or are accessible synthetically. Here, we enumerate binary, ternary, and quaternary element combinations to produce an extensive library of over 10^10…
Jiaqi Wu
Many comparative analyses operate on rectangular matrices whose columns represent the same variables across observations. Phylogenomic measurements, however, are attached to tree branches. Converting locus-specific trees into a common locus-by-coordinate matrix is straightforward only when loci contain the same taxa…