22 papers · ranked by Valyu relevance
Tianhong Huang, Victor Agostinelli, Lizhong Chen
Compactness in deep learning can be critical to a model's viability in low-resource applications, and a common approach to extreme model compression is quantization. We consider Iterative Product Quantization (iPQ) with Quant-Noise (Fan et al., 2020) to be state-of-the-art in this area, but this quantization framework…
Chenyue W. Hu, Hanyang Li, Amina A. Qutub
Background Many common clustering algorithms require a two-step process that limits their efficiency. The algorithms need to be performed repetitively and need to be implemented together with a model selection criterion. These two steps are needed in order to determine both the number of clusters present in the data…
Yuan Yin, Masanao Yajima, Joshua D. Campbell
Assays such as CITE-seq can measure the abundance of cell surface proteins on individual cells using antibody derived tags (ADTs). However, many ADTs have high levels of background noise that can obfuscate down-stream analyses. Using an exploratory analysis of PBMC datasets, we find that some droplets that were…
Y. A. Joarder, Mosabbir Ahmed
Big Data is a massive volume of both structured and unstructured data that is too large and it also difficult to process using traditional techniques. Clustering algorithms have developed as a powerful learning tool that can exactly analyze the volume of data that produced by modern applications. Clustering in data…
Frank Nielsen, Richard Nock
The k-means clustering problem asks to partition the data into k clusters so as to minimize the sum of the squared Euclidean distances of the data points to their closest cluster center. Finding the optimal k-means clustering of a d-dimensional data set is NP-hard in general and many heuristics have been designed for…
Lisa Strasser, Tomos E. Morgan, Felipe Guapo, Florian Füssl + 3 more
Adeno-associated virus (AAV)-based cell and gene therapy is a rapidly developing field, requiring analytical methods for detailed product characterization. One important quality attribute of AAV products that requires monitoring is the amounts of residual empty capsids following downstream processing. Traditionally…
P. L. Krapivsky
The Riviera model mimics a densifying settlement along the coastline. In the lattice version, houses are built sequentially in empty sites with the constraint that every newly built house has at least one empty neighboring site. The distribution of clusters of adjacent houses does not obey a closed set of evolutionary…
Jing Li, Jiahui Yang, Hui Cai, Chi Jiang + 6 more
'Zimeng Lu' 'Lingzhi Li' 'Guanqun Sun' 'Chun Sing Lai'] With the aging of the social population structure, the number of empty-nesters is also increasing. Therefore, it is necessary to manage empty-nesters with data mining technology. This paper proposed an empty-nest power user identification and power consumption…
Luke Zappia, Alicia Oshlack
Clustering techniques are widely used in the analysis of large datasets to group together samples with similar properties. For example, clustering is often used in the field of single-cell RNA-sequencing in order to identify different cell types present in a tissue sample. There are many algorithms for performing…
Lynin Sokhonn, Yun-Soo Park, Mun-Kyu Lee, Ilsun You
Hierarchical clustering is a widely used data analysis technique. Typically, tools for this method operate on data in its original, readable form, raising privacy concerns when a clustering task involving sensitive data that must remain confidential is outsourced to an external server. To address this issue, we…
Luke Zappia, Alicia Oshlack
Clustering techniques are widely used in the analysis of large data sets to group together samples with similar properties. For example, clustering is often used in the field of single-cell RNA-sequencing in order to identify different cell types present in a tissue sample. There are many algorithms for performing…
Dylan Molinié, Kurosh Madani, Véronique Amarger, Grigore Stamatescu + 2 more
For two centuries, the industrial sector has never stopped evolving. Since the dawn of the Fourth Industrial Revolution, commonly known as Industry 4.0, deep and accurate understandings of systems have become essential for real-time monitoring, prediction, and maintenance. In this paper, we propose a machine learning…
Robert A. Kłopotek, Mieczysław A. Kłopotek
This paper investigates the validity of Kleinberg's axioms for clustering functions with respect to the quite popular clustering algorithm called k-means.We suggest that the reason why this algorithm does not fit Kleinberg's axiomatic system stems from missing match between informal intuitions and formal formulations…
Stijn van Dongen
Formulating RCL and generally considering input ensembles of partitions necessitates various comparisons of sets and partitions, quantification of these comparisons, as well a simple trait associated with binary trees. The required terminology and definitions are gathered in this section. The input ensemble ℰ is a set…
Joao C. Marques, Michael B. Orger
How to partition a data set into a set of distinct clusters is a ubiquitous and challenging problem. The fact that data sets vary widely in features such as cluster shape, cluster number, density distribution, background noise, outliers and degree of overlap, makes it difficult to find a single algorithm that can be…
Authors not listed
We present burbuja (Baring Unseen Regions of Bubbles Using Joint-Density Analysis), an automated software tool for detecting and characterizing gas bubbles and other local voids in molecular structures and trajectories containing explicit aqueous solvent. We describe the burbuja algorithm and demonstrate its accuracy…
Jie Yang, Chin‐Teng Lin
Grouping similar objects is a fundamental tool of scientific analysis, ubiquitous in disciplines from biology and chemistry to astronomy and pattern recognition. Inspired by the torque balance that exists in gravitational interactions when galaxies merge, we propose a novel clustering method based on two natural…
Dietger Van den Enyden, Rohan Pokratath, Jikson Pulparayil Mathew, Eline Goossens + 2 more
Metal oxo clusters of the type \ce{M6O4(OH)4(RCOO)12} (M = Zr of Hf) are valuable building blocks for material science. Here, we develop them as smallest conceivable nanocrystal prototypes. We synthesize a series of zirconium and hafnium oxo clusters with ligands that are typically used to stabilize oxide nanocrystals…
Haide Wu, Morten Engsvang, Yosef Knattrup, Jakub Kubečka + 1 more
The nucleation process leading to the formation of new atmospheric particles plays a crucial role in aerosol research. Quantum chemical (QC) calculations can be used to model the early stages of aerosol formation, where atmospheric vapor molecules interact and form stable molecular clusters. However, QC calculations…
Nicoline Frederiks, Danika Heaney, John Kreinbihl, Christopher Johnson
Iodine containing clusters are expected to be central to new particle formation (NPF) events in polar and mid-latitude coastal regions. Iodine oxoacids and iodine oxides are observed in newly formed clusters, and in more polluted mid-latitude settings, theoretical studies suggest ammonia may increase growth rates.…
Jakub Kubečka, Vitus Besel, Ivo Neefjes, Yosef Knattrup + 3 more
Computational modeling of atmospheric molecular clusters requires a comprehensive understanding of their complex configurational spaces, interaction patterns, stabilities against fragmentation, and even dynamic behaviors. To address these needs, we introduce the Jammy Key framework, a collection of automated scripts…
Sudhakar Jonnalagadda, Rajagopalan Srinivasan
Background Clustering techniques are routinely used in gene expression data analysis to organize the massive data. Clustering techniques arrange a large number of genes or assays into a few clusters while maximizing the intra-cluster similarity and inter-cluster separation. While clustering of genes facilitates…