7 papers · ranked by Valyu relevance
Timothy C. Au
One advantage of decision tree based methods like random forests is their ability to natively handle categorical predictors without having to first transform them (e.g., by using feature engineering techniques). However, in this paper, we show how this capability can lead to an inherent "absent levels" problem for…
Helen L. Smith, Patrick J. Biggs, Nigel P. French, Adam N. H. Smith + 1 more
To date, there remains no satisfactory solution for absent levels in random forest models. Absent levels are levels of a predictor variable encountered during prediction for which no explicit rule exists. Imposing an order on nominal predictors allows absent levels to be integrated and used for prediction. The ordering…
Sara P. Garcia, Armando J. Pinho, João M. O. S. Rodrigues, Carlos A. C. Bastos + 2 more
'Carlos A. C. Bastos' 'Paulo J. S. G. Ferreira' 'Christian Schönbach'] Minimal absent words have been computed in genomes of organisms from all domains of life. Here, we explore different sets of minimal absent words in the genomes of 22 organisms (one archaeota, thirteen bacteria and eight eukaryotes). We investigate…
Armando J Pinho, Paulo JSG Ferreira, Sara P Garcia, João MOS Rodrigues
'João MOS Rodrigues'] Background The problem of finding the shortest absent words in DNA data has been recently addressed, and algorithms for its solution have been described. It has been noted that longer absent words might also be of interest, but the existing algorithms only provide generic absent words by trivially…
Sara P. Garcia, Armando J. Pinho, Zhanjiang Liu
Minimal absent words have been computed in genomes of organisms from all domains of life. Here, we aim to contribute to the catalogue of human genomic variation by investigating the variation in number and content of minimal absent words within a species, using four human genome assemblies. We compare the reference…
Eugene F Schuster, Eric Blanc, Linda Partridge, Janet M Thornton
Correction of non-specific binding for both PM and MM probes using probe-sequence models can partially remove the probe-sequence bias in Affymetrix microarray experiments and result in better performance of the MAS 5.0 algorithm.
Bora Canbula, R. Bulur, Deniz Canbula, H. Babacan
Collective effects in the level density are not well understood, and including these effects as enhancement factors to the level density does not produce sufficiently consistent predictions of observables. Therefore, collective effects are investigated in the level density parameter instead of treating them as a final…