26 papers · ranked by Valyu relevance
Louise Foley, Dorothea Dumuid, Andrew J. Atkin, Timothy Olds + 1 more
'David Ogilvie'] Background Active travel (walking or cycling for transport) is associated with favourable health outcomes in adults. However, little is known about the concurrent patterns of health behaviour associated with active travel. We used compositional data analysis to explore differences in how people doing…
Samyajoy Pal, Christian Heumann, Sheetal Kalyani
A model-based clustering method for compositional data is explored in this article. Most methods for compositional data analysis require some kind of transformation. The proposed method builds a mixture model using Dirichlet distribution which works with the unit sum constraint. The mixture model uses a hard EM…
Guojun Gan, Emiliano A. Valdez
Compositional data are multivariate observations that carry only relative information between components. Applying standard multivariate statistical methodology directly to analyze compositional data can lead to paradoxes and misinterpretations. Compositional data also frequently appear in insurance, especially with…
Thomas P. Quinn, Ionas Erb, Greg Gloor, Cedric Notredame + 2 more
Next-generation sequencing (NGS) has made it possible to determine the sequence and relative abundance of all nucleotides in a biological or environmental sample. Today, NGS is routinely used to understand many important topics in biology from human disease to microorganism diversity. A cornerstone of NGS is the…
Abel Ruiz-Giralt, Stefano Biagetti, Carla Lancelotti, Óscar Parque + 4 more
Anthropic Activity Markers (AAMs) were formalized over 10 years ago as a toolkit to infer human activities from biological and geochemical signatures preserved in sediments. This paper presents AAMs 2.0, a revised analytical framework that integrates compositional data analysis (CoDA) and geostatistics, as the…
Thomas P. Quinn, Ionas Erb
In the health sciences, many data sets produced by next-generation sequencing (NGS) only contain relative information because of biological and technical factors that limit the total number of nucleotides observed for a given sample. As mutually dependent elements, it is not possible to interpret any component in…
Michail Tsagris
In compositional data, an observation is a vector with non-negative components which sum to a constant, typically 1. Data of this type arise in many areas, such as geology, archaeology, biology, economics and political science among others. The goal of this paper is to extend the taxicab metric and a newly suggested…
Kellyn F Arnold, Laurie Berrie, Peter W G Tennant, Mark S Gilthorpe
Consider three random variables-X, Y and Z-for which X + Y = Z. The relationship among these variables is depicted in the DAG in [dyaa021-F2], which employs the previously introduced notation for deterministic relationships. Although X and Y (the ‘components’) together determine Z (the ‘whole’ or ‘total’), no time flow…
Onur Batın Doğan, Fatma Sevinç Kurnaz
Crime Trends for 2022 Authors: ['Onur Batın Doğan' 'Fatma Sevinç Kurnaz'] This article investigates crime patterns across European countries in 2022 using Compositional Data Analysis (CoDA) to address limitations of traditional statistical approaches in dealing with the relative nature of crime data. Recognizing crime…
Julie Rendlová, Karel Hron, Kamila Fačevicová, Peter Filzmoser
A data table which is arranged according to two factors can often be considered as a compositional table. An example is the number of unemployed people, split according to gender and age classes. Analyzed as compositions, the relevant information would consist of ratios between different cells of such a table. This is…
Alberte Sloth Carlsen, Te Chen, Nicholas Luke Cowie, Christian Brinch + 2 more
Isotopic Metabolic Flux Analysis (I-MFA) is a standard approach for estimating intracellular metabolic fluxes. I-MFA infers fluxes by comparing simulated and measured metabolite isotopologue distributions (MIDs) of metabolites from isotope labeling experiments. MIDs represent fractional abundances that strictly sum to…
Alemu Takele Assefa, Bie Verbist, Koen Van den Berge
In single-cell studies, a common question is whether there is a change in cell composition between conditions. While ideally, one needs absolute cell counts (number of cells per volumetric unit in a sample) to address these questions, current experimentation typically obtains cell counts that only carry relative…
Michael Hagmann, Michael Staniek, Stefan Riezler
This work investigates whether time series of natural phenomena can be understood as being generated by sequences of latent states which are ordered in systematic and regular ways. We focus on clinical time series and ask whether clinical measurements can be interpreted as being generated by meaningful physiological…
Fatih Dikbaş
Correlation remains to be one of the most widely used statistical tools for assessing the strength of relationships between data series. This paper presents a novel compositional correlation method for detecting linear and nonlinear relationships by considering the averages of all parts of all possible compositions of…
Stephanie Wankowicz, James Fraser
In their folded state, biomolecules exchange between multiple conformational states, crucial for their function. However, most structural models derived from experiments and computational predictions only encode a single state. To represent biomolecules more accurately, we must move towards modeling and predicting…
Arun Srinivasan, Lingzhou Xue, Xiang Zhan
A critical task in microbiome data analysis is to explore the association between a scalar response of interest and a large number of microbial taxa that are summarized as compositional data at different taxonomic levels. Motivated by fine-mapping of the microbiome, we propose a two-step compositional knockoff filter…
Matteo Negri, Davide Bergamini, Carlo Baldassi, Riccardo Zecchina + 1 more
Generative processes in biology and other fields often produce data that can be regarded as resulting from a composition of basic features. Here we present an unsupervised method based on autoencoders for inferring these basic features of data. The main novelty in our approach is that the training is based on the…
Samuel D. Gamboa-Tuz, Marcel Ramos, Eric Franzosa, Curtis Huttenhower + 3 more
Previous benchmarking of differential abundance (DA) analysis methods in microbiome studies have employed synthetic data, simulations, and “real data” examples, but to the best of our knowledge, none have yet employed experimental data with known “ground truth” differential abundance. A key debate in the field centers…
Stijn Hawinkel, Frederiek-Maarten Kerckhof, Luc Bijnens, Olivier Thas
Explorative visualization techniques provide a first summary of microbiome read count datasets through dimension reduction. A plethora of dimension reduction methods exists, but many of them focus primarily on sample ordination, failing to elucidate the role of the bacterial species. Moreover, implicit but often…
Stephanie Wankowicz, James Fraser
In their folded state, biomolecules exchange between multiple conformational states, crucial for their function. However, most structural models derived from experiments and computational predictions only encode a single state. To represent biomolecules more accurately, we must move towards modeling and predicting…
Maxwell Venetos, Masha Elkin, Connor Delaney, John Hartwig + 1 more
NMR spectroscopy is an important analytical technique in synthetic organic chemistry, but its integration into high-throughput experimentation workflows has been limited by the necessity to manually analyze NMR spectra of new chemical entities. Current efforts to automate the analysis of NMR spectra rely on comparisons…
Authors not listed
The vastness of chemical space presents a long-standing challenge for the exploration of new compounds with pre-determined properties. In materials science, crystal structure prediction has become a mature tool for mapping from composition to structure based on global optimisation techniques. Generative artificial…
Authors not listed
Transition metal phosphates (TMPs) are extensively explored for electrochemical and catalytical applications due to their structural versatility and chemical stability. Within this material class, novel high-entropy metal phosphates (HEMPs)—containing multiple transition metals combined into a single-phase…
Stephen A. Allegri, Kevin McCoy, Cassie S. Mitchell
Large networks are quintessential to bioinformatics, knowledge graphs, social network analysis, and graph-based learning. CompositeView is a Python-based open-source application that improves interactive complex network visualization and extraction of actionable insight. CompositeView utilizes specifically formatted…
Yu-Chieh Huang, Pierre Tremouilhac, Stefan Kuhn, Pei-Chi Huang + 6 more
A method for data review in chemical sciences with a focus on data for the characterization of synthetic molecules is described. As current procedures for data curation in chemistry rely almost exclusively on manual checking or peer reviewing, a (semi-)automatic procedure for the evaluation of data assigned to…
Stuart J. Chalk
With the move toward global, Internet enabled science there is an inherent need to capture, store, aggregate and search scientific data across a large corpus of heterogeneous data silos. As a result, standards development is needed to create an infrastructure capable of representing the diverse nature of scientific…