Search · four archives
Search · four archives
16 papers · ranked by Valyu relevance
Pere-Pau Vázquez
The analysis of research paper collections is an interesting topic that can give insights on whether a research area is stalled in the same problems, or there is a great amount of novelty every year. Previous research has addressed similar tasks by the analysis of keywords or reference lists, with different degrees of…
Andrew R. Cohen, Paul Vitányi
Normalized compression distance (NCD) is a parameter-free, feature-free, alignment-free, similarity measure between a pair of finite objects based on compression. However, it is not sufficient for all applications. We propose an NCD of finite nonempty multisets (a.k.a. multiples) of finite objects that is also a…
Rebecca Schuller Borbely
Normalized Compression Distance (NCD) is a popular tool that uses compression algorithms to cluster and classify data in a wide range of applications. Existing discussions of NCD's theoretical merit rely on certain theoretical properties of compression algorithms. However, we demonstrate that many popular compression…
Mariana S. Ramos, Susana Brás, Paula Pinto, Luísa Castro + 1 more
'Agnese Sbrollini'] Intrapartum asphyxia is responsible for approximately 900 000 deaths per year worldwide. These numbers show the urgency of investing in the quality of fetal health care. The heart rate signal is a complex signal and sometimes behaves unpredictably. Thus, it becomes relevant to study approaches that…
Diogo Pratas, Raquel M. Silva, Armando J. Pinho
An efficient DNA compressor furnishes an approximation to measure and compare information quantities present in, between and across DNA sequences, regardless of the characteristics of the sources. In this paper, we compare directly two information measures, the Normalized Compression Distance (NCD) and the Normalized…
John Hurwitz, Charles Nicholas, Edward Raff
Compression and Classification Authors: ['John Hurwitz' 'Charles Nicholas' 'Edward Raff'] It is generally well understood that predictive classification and compression are intrinsically related concepts in information theory. Indeed, many deep learning methods are explained as learning a kind of compression, and that…
Dinu Coltuc, Mihai Datcu, Daniela Coltuc
This paper investigates the usefulness of the normalized compression distance (NCD) for image similarity detection. Instead of the direct NCD between images, the paper considers the correlation between NCD based feature vectors extracted for each image. The vectors are derived by computing the NCD between the original…
Ana Granados
Universidad Autonoma de Madrid ´ Escuela Polit´ecnica Superior ANALYSIS AND STUDY ON TEXT REPRESENTATION TO IMPROVE THE ACCURACY OF THE NORMALIZED COMPRESSION DISTANCE Thesis Ana Granados Fontecha Madrid, 2012 ADVISORS: David Camacho Fern´andez Francisco de Borja Rodr´ıguez Ortiz Tribunal nombrado por el Mgfco. y…
Rudi L. Cilibrasi, Paul M.B. Vitányi
We analyze the whole genome phylogeny and taxonomy of the SARS-CoV-2 virus using compression. This is a new fast alignment-free method called the “normalized compression distance” (NCD) method. It discovers all effective similarities based on Kolmogorov complexity. The latter being incomputable we approximate it by a…
Andrew R. Cohen, Paul Vitányi
Normalized Google distance (NGD) is a relative semantic distance based on the World Wide Web (or any other large electronic database, for instance Wikipedia) and a search engine that returns aggregate page counts. The earlier NGD between pairs of search terms (including phrases) is not sufficient for all applications.…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Bryan A. Dawkins, Trang T. Le, Brett A. McKinney
The performance of nearest-neighbor feature selection and prediction methods depends on the metric for computing neighborhoods and the distribution properties of the underlying data. The effects of the distribution and metric, as well as the presence of correlation and interactions, are reflected in the expected…
Axel Trefzer, Alexandros Stamatakis
Bayesian Markov-Chain Monte Carlo (MCMC) methods for phylogenetic tree inference, that is, inference of the evolutionary history of distinct species using their molecular sequence data, typically generate large sets of phylogenetic trees. The trees generated by the MCMC procedure are samples of the posterior…
Bryan A. Dawkins, Trang T. Le, Brett A. McKinney, Alan D Hutson
The performance of nearest-neighbor feature selection and prediction methods depends on the metric for computing neighborhoods and the distribution properties of the underlying data. Recent work to improve nearest-neighbor feature selection algorithms has focused on new neighborhood estimation methods and distance…
Muzaffar Rafique, Aykut Erbas
Polyelectrolyte gels can generate electric potentials under mechanical deformation. While the underlying mechanism of such response is often attributed to the changes in counterion-condensation levels or alterations in the ionic conditions in the pervaded volume of the hydrogel, exact molecular origins are still…
Yueying He, Yue Xue, Jingyao Wang, Yupeng Huang + 3 more
High-throughput chromosome conformation capture (Hi-C) technique profiles the genomic structure in a genome-wide fashion. The reproducibility and consistency of Hi-C data are essential in characterizing dynamics of genomic structures. We developed a diffusion-based method, C_T_G (Hi-C To Geometry), to deal with the…