21 papers · ranked by Valyu relevance
Giancarlo Jug
The problems of the intermediate-range atomic structure of glasses and of the mechanism for the glass transition are approached from the low-temperature end in terms of a scenario for the atomic organization that justifies the use of an extended tunneling model. The latter is crucial for the explanation of the magnetic…
Katherine Eason, Gift Nyamundanda, Anguraj Sadanandam
To stratify cancer patients for most beneficial therapies, it is a priority to define robust molecular subtypes using clustering methods and “big data”. If each of these methods produces different numbers of clusters for the same data, it is difficult to achieve an optimal solution. Here, we introduce “polyCluster”, a…
Nicolas F. Fernandez, Gregory W. Gundersen, Adeeb Rahman, Mark L. Grimes + 3 more
Most tools developed to visualize hierarchically clustered heatmaps generate static images. Clustergrammer is a web-based visualization tool with interactive features such as: zooming, panning, filtering, reordering, sharing, performing enrichment analysis, and providing dynamic gene annotations. Clustergrammer can be…
Joshua L Phillips, Michael E Colvin, Shawn Newsam
Background Molecular dynamics (MD) simulation is a powerful technique for sampling the meta-stable and transitional conformations of proteins and other biomolecules. Computational data clustering has emerged as a useful, automated technique for extracting conformational states from MD simulation data. Despite extensive…
Jakub Kubečka, Vitus Besel, Ivo Neefjes, Yosef Knattrup + 3 more
Computational modeling of atmospheric molecular clusters requires a comprehensive understanding of their complex configurational spaces, interaction patterns, stabilities against fragmentation, and even dynamic behaviors. To address these needs, we introduce the Jammy Key framework, a collection of automated scripts…
Mohith Manjunath, Yi Zhang, Steve H. Yeo, Omar Sobh + 5 more
Clustering is one of the most common techniques used in data analysis to discover hidden structures by grouping together data points that are similar in some measure into clusters. Although there are many programs available for performing clustering, a single web resource that provides both state-of-the-art clustering…
Mohith Manjunath, Yi Zhang, Yeonsung Kim, Steve H. Yeo + 7 more
'Nathan Russell' 'Christian Followell' 'Colleen Bushell' 'Umberto Ravaioli' 'Jun S. Song' 'Kjiersten Fagnan'] Background Clustering is one of the most common techniques in data analysis and seeks to group together data points that are similar in some measure. Although there are many computer programs available for…
Charles Sing, Jian Qin
Complex coacervation is a phase separation phenomenon, driven by the electrostatic attraction between oppositely-charged macromolecular species. A recent surge of interest in coacervation between polyelectrolytes has been driven by both fundamental advances in experimental characterization of these systems, along with…
Caitlin C. Bannan, David Mobley
Force fields are used in a variety of research fields including computer-aided drug design, biomaterials, and polymer chemistry. However, force fields also continue to limit the accuracy of predictions of physical properties. Current parameterization of these force fields involves a huge amount of human effort -- often…
Raquel Lopez-Rios de Castro, Alejandro Santana-Bonilla, Robert M. Ziolek, Christian D. Lorenz
Molecular dynamics simulations have become an essential tool in the study of soft matter and biological macromolecules. The large amount of high-dimensional data produced by such simulations does not immediately elucidate the atomistic mechanisms that underlie complex materials and molecular processes. Analysis of…
Nuno Fachada, Diogo de Andrade
Synthetic data is essential for assessing clustering techniques, complementing and extending real data, and allowing for more complete coverage of a given problem's space. In turn, synthetic data generators have the potential of creating vast amounts of data—a crucial activity when real-world data is at premium—while…
Linda Dib, Alessandra Carbone
Background Searching for similarities in a set of biological data is intrinsically difficult due to possible data points that should not be clustered, or that should group within several clusters. Under these hypotheses, hierarchical agglomerative clustering is not appropriate. Moreover, if the dataset is not known…
Ji Xu, Guoyin Wang
1 School of Information Science & Technology, Southwest Jiaotong University, Chengdu, China 2Chongqing Key Laboratory of Computational Intelligence, Chongqing University of Posts and Telecommunications, Chongqing, China 3 Institute of Electronic Information Technology, Chongqing Institute of Green and Intelligent…
Adelchi Azzalini, Giovanna Menardi
The traditional approach to the clustering problem, also called 'unsupervised classification' in the machine learning literature, hinges on some notion of distance or dissimilarity between objects. Once one such notion has been adopted among the many existing alternatives, the clustering process aims at grouping…
Andreas Adolfsson, Margareta Ackerman, Naomi C. Brownstein
Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study of clusterability, which evaluates whether data possesses such structure, is an integral part of…
Stijn van Dongen
The result of applying single linkage to the RCL matrix is a hierarchical clustering in the form of a binary tree called the RCL tree. This tree has almost as many internal nodes as there are data elements.^1^ By picking a subset of internal nodes in the tree it is possible to obtain a flat clustering or a simplified…
Authors not listed
Digital polymer chemistry leverages computational methods to design and optimize polymer materials. While there have been advances in using machine learning to accelerate the design of polymers, the field is hampered by the lack of standards, which precludes comparability and makes it difficult to build on top of prior…
Marco Cavallo, Çağatay Demiralp
D4: Represent clustering instances compactly It is important for users to be able to examine different clustering instances fluidly and independently without visual clutter or cognitive overload. The Clustrophile 2 interface employs the "Clustering View" element as the atomic component representing a clustering…
Metin Balaban, Niema Moshiri, Uyen Mai, Siavash Mirarab
Clustering homologous sequences based on their similarity is a problem that appears in many bioinformatics applications. The fact that sequences cluster is ultimately the result of their phylogenetic relationships. Despite this observation and the natural ways in which a tree can define clusters, most applications of…
Mark A. Newell, Dianne Cook, Heike Hofmann, Jean‐Luc Jannink
> A first step in exploring population structure in crop plants and other organisms is to define the number of subpopulations that exist for a given data set. The genetic marker data sets being generated have become increasingly large over time and commonly are of the high-dimension, low sample size (HDLSS) situation.…
David M. Swanson, Tonje Lien, Helga Bergholtz, Therese Sørlie + 1 more
Unsupervised clustering is important in disease subtyping, among having other genomic applications. As genomic data has become more multifaceted, how to cluster across data sources for more precise subtyping is an ever more important area of research. Many of the methods proposed so far, including iCluster and Cluster…