21 papers · ranked by Valyu relevance
Prasad Krishnan, Lakshmi Natarajan, V. Lalitha, Siu-Wai Ho + 2 more
'Lawrence Ong' 'Kenneth Shum'] The problem of data exchange between multiple nodes with storage and communication capabilities models several current multi-user communication problems like Coded Caching, Data Shuffling, Coded Computing, etc. The goal in such problems is to design communication schemes which accomplish…
Andrea Bittau, Úlfar Erlingsson, Petros Maniatis, Ilya Mironov + 6 more
'Ananth Raghunathan' 'David Lie' 'Mitch Rudominer' 'Ushasree Kode' 'Julien Tinnes' 'Bernhard Seefeld'] The large-scale monitoring of computer users' software activities has become commonplace, e.g., for application telemetry, error reporting, or demographic profiling. This paper describes a principled systems…
Mohamed Attia, Ravi Tandon
—Data shuffling is one of the fundamental building blocks for distributed learning algorithms, that increases the statistical gain for each step of the learning process. In each iteration, different shuffled data points are assigned by a central node to a distributed set of workers to perform local computa tions, which…
Adel Elmahdy, Soheil Mohajer
We consider the data shuffling problem in a distributed learning system, in which a master node is connected to a set of worker nodes, via a shared link, in order to communicate a set of files to the worker nodes. The master node has access to a database of files. In every shuffling iteration, each worker node…
Kai Wan, Daniela Tuninetti, Mingyue Ji, Giuseppe Caire + 1 more
'Pablo Piantanida'] Abstract—Data shuffling of training data among different computing nodes (workers) has been identified as a core element to improve the statistical performance of modern large-scale machine learning algorithms. Data shuffling is often considered as one of the most significant bottlenecks in such…
Carola Engler, Ramona Gruetzner, Romy Kandzia, Sylvestre Marillonnet + 1 more
'Jean Peccoud'] We have developed a protocol to assemble in one step and one tube at least nine separate DNA fragments together into an acceptor vector, with 90% of recombinant clones obtained containing the desired construct. This protocol is based on the use of type IIs restriction enzymes and is performed by simply…
Manuel Penschuck
Shuffling is the process of rearranging a sequence of elements into a random order such that any permutation occurs with equal probability. It is an important building block in a plethora of techniques used in virtually all scientific areas. Consequently considerable work has been devoted to the design and…
Minghui Jiang, James Anderson, Joel Gillespie, Martin Mayne
Background Randomly shuffled sequences are routinely used in sequence analysis to evaluate the statistical significance of a biological sequence. In many cases, biologists need sophisticated shuffling tools that preserve not only the counts of distinct letters but also higher-order statistics such as doublet counts…
Karin M. Cox, Daisuke Kase, Taieb Znati, Robert S. Turner
Oscillations figure prominently as neurological disease hallmarks and neuromodulation targets. To detect oscillations in a neuron’s spiking, one might attempt to seek peaks in the spike train’s power spectral density (PSD) which exceed a flat baseline. Yet for a non-oscillating neuron, the PSD is not flat: The recovery…
Marios Fanourakis
An important feature of data collection frameworks, in which voluntary participants are involved, is that of privacy. Besides data encryption, which protects the data from third parties in case the communication channel is compromised, there are schemes to obfuscate the data and thus provide some anonymity in the data…
Frederic von Wegner, Gesine Hermann, Inken Tödt, Inga Karin Todtenhaupt + 1 more
Higher-order syntax properties of EEG microstate sequences offer insight into the transition dynamics of functional brain networks. We here define higher-order syntax as microstate sequence properties that are not explained by the first-order transition matrix, and we postulate three requirements that surrogate data…
Ke Zeng, Anthony Fodor
In microbiome research, differential abundance analysis aids in identifying significant differences in microbial taxa across two or more conditions. Statistical approaches used for this purpose include classical tests such as the t-test and Wilcoxon test, as well as methods designed to account for the compositional…
Murakami, Takao, Sei Yuichi, Eriguchi + 1 more
—Shuffle DP (Differential Privacy) protocols provide high accuracy and privacy by introducing a shuffler who randomly shuffles data in a distributed system. However, most shuffle DP protocols are vulnerable to two attacks: collusion attacks by the data collector and users and data poisoning attacks. A recent study…
Amy Vennos, Alan Michaels, Yong Deng
This paper models a translation for base-2 pseudorandom number generators (PRNGs) to mixed-radix uses such as card shuffling. In particular, we explore a shuffler algorithm that relies on a sequence of uniformly distributed random inputs from a mixed-radix domain to implement a Fisher-Yates shuffle that calls for…
Thomas Lynn, Julio Ottino, Richard Lueptow, Paul Umbanhowar
Cut-and-shuffle mixing is an instructive candidate system with which to assess the potential of machine learning (ML) as an approach to solve difficult mixing problems. We focus on a specific subset of cut-and-shuffle systems, the one-dimensional interval exchange transform. This class of mixing operations is well…
Tian Yang, Zhuangcheng Zhen, Yuhai Tu, Qi Ouyang + 1 more
Protein complexes are critical for cellular functions, and subunit exchange within these complexes is increasingly recognized as a key regulatory mechanism. In the cyanobacterial circadian clock, subunits shuffling of the core clock protein KaiC is thought to synchronize the clock, though the underlying mechanism…
Carl Veller, Nancy Kleckner, Martin A. Nowak
Comparative studies in evolutionary genetics rely critically on evaluation of the total amount of genetic shuffling that occurs during gamete production. However, such studies have been ham-pered by the fact that there has been no direct measure of this quantity. Existing measures consider crossing over by simply…
Nadella Sunil, G Narsimha
The emergence of digital information has raised many issues over the release of sensitive personal information in data mining activities due to the rapid expansion of the digital information. Privacy-Preserving Data Mining (PPDM) has the goal of allowing a significant analysis of data, as well as safeguarding…
Pijun Hou, Yuepeng Wang, Ziming Shi, Pan Zheng
While encrypting information with color images, most encryption schemes treat color images as three different grayscale planes and encrypt each plane individually. These algorithms produce more duplicated operations and are less efficient because they do not properly account for the link between the various planes of…
Gavin C Conant, Andreas Wagner
The incidence of gene shuffling is estimated in conserved genes in 10 organisms from the three domains of life. Successful gene shuffling is found to be very rare among such conserved genes. This suggests that gene shuffling may not be a major force in reshaping the core genomes of eukaryotes.
Authors not listed
The discoverability and reusability of data is critical for machine learning to drive new discovery in the chemical sciences, and the ‘FAIR Guiding Principles for scientific data management and stewardship’ provide a measurable set of guidelines that can be used to ensure the accessibility of reusable data. We…