22 papers · ranked by Valyu relevance
Georg Meisl, Catherine K Xu, Jonathan D Taylor, Thomas C T Michaels + 7 more
Fibrillar protein aggregates are a hallmark of the pathology of a range of human disorders, from prion diseases to dementias. Yet, the same aggregated structures that are formed in disease are also encountered in several functional contexts. The fundamental properties that determine whether these protein assembly…
Haikady N. Nagaraja, Shane Sanders, Alan D Hutson
The relationship between social choice aggregation rules and non-parametric statistical tests has been established for several cases. An outstanding, general question at this intersection is whether there exists a non-parametric test that is consistent upon aggregation of data sets (not subject to Yule-Simpson…
Ana Helena Tavares, Ana Silva, Tiago Freitas, Maria Costa + 3 more
Despite the advances on data analysis methodologies in the last decades, most of the traditional regression methods cannot be directly applied to large-scale data. Although aggregation methods are especially designed to deal with large-scale data, their performance may be strongly reduced in ill-conditioned problems…
Ponnuswamy Sadayappan, Bradford L. Chamberlain, Guido Juckeland, Hatem Ltaief + 15 more
'Hatem Ltaief' 'Richard L. Graham' 'Lion Levi' 'Devendar Burredy' 'Gil Bloch' 'Gilad Shainer' 'David Cho' 'George Elias' 'Daniel Klein' 'Joshua Ladd' 'Ophir Maor' 'Ami Marelli' 'Valentin Petrov' 'Evyatar Romlet' 'Yong Qin' 'Ido Zemah'] This paper describes the new hardware-based streaming-aggregation capability added…
Authors not listed
When designing compound AI systems, a common approach is to query multiple copies of the same model and aggregate the responses to produce a synthesized output. Given the homogeneity of these models, this raises the question of whether aggregation unlocks access to a greater set of outputs than querying a single model.…
Pritom Saha Akash, Wei-Cheng Lai, Po-Wen Lin
In the current world, OLAP (Online Analytical Processing) is used intensively by modern organizations to perform ad hoc analysis of data, providing insight for better decision making. Thus, the performance for OLAP is crucial; however, it is costly to support OLAP for a large data-set. An approximate query process…
Omer Markovitch, Juntian Wu, Otto Sijbren
Copying information is vital for life's propagation. Current life forms maintain a low error rate in replication using complex machinery to prevent and correct errors. However, primitive life had to deal with higher error rates, limiting its ability to evolve. Discovering mechanisms to reduce errors would alleviate…
Robert P. Goldman, Robert Moseley, Nicholas Roehner, Bree Cummins + 31 more
We describe an experimental campaign that replicated the performance assessment of logic gates engineered into cells of S. cerevisiae by Gander, et al. The experimental campaign used a novel high throughput experimentation framework developed under DARPA’s Synergistic Discovery and Design (SD2) program: a remote…
Chengjie Qin, Florin Rusu
In this paper we introduce the first framework for parallel online aggregation in which the estimation virtually does not incur any overhead on top of the actual execution. We define a generic interface to express any estimation model that abstracts completely the execution details. We design a novel estimator…
Christian L. Staudt, Michael Hamann, Alexander Gutfraind, Ilya Safro + 1 more
networks Authors: Christian L. Staudt, Michael Hamann, Alexander Gutfraind, Ilya Safro, Henning Meyerhenke Research on generative models plays a central role in the emerging field of network science, studying how statistical patterns found in real networks could be generated by formal rules. Output from these…
Dimitra Tsavachidou
Sequencing at single-nucleotide resolution using nanopore devices is performed with reported error rates 10.5–20.7% (2). Since errors occur randomly during sequencing, repeating the sequencing procedure for the same DNA strands several times can generate sequencing results based on consensus derived from replicate…
Ingo Mueller, Andrea Arteaga, Torsten Hoefler, Gustavo Alonso
—Industry-grade database systems are expected to produce the same result if the same query is repeatedly run on the same input. However, the numerous sources of non-determinism in modern systems make reproducible results difficult to achieve. This is particularly true if floating-point numbers are involved, where the…
Fabien C. Y. Benureau, Nicolas P. Rougier
Scientific code is different from production software. Scientific code, by producing results that are then analyzed and interpreted, participates in the elaboration of scientific conclusions. This imposes specific constraints on the code that are often overlooked in practice. We articulate, with a small example, five…
Nicholas Pritchard, A. Wicenec
Computational Workflows Authors: ['Nicholas Pritchard' 'A. Wicenec'] Computational workflow management systems power contemporary data-intensive sciences. The slowly resolving reproducibility crisis presents both a sobering warning and an opportunity to iterate on what science and data processing entails. The Square…
Wenqin Zhuang, Yuao Wang, Guocheng Wang, Li Sun
Federated learning (FL) enables collaborative model training without sharing raw data, but it faces challenges due to client heterogeneity, leading to inefficiency and reduced accuracy. This paper proposes a digital twin (DT)-based dynamic FL aggregation method to address these issues. The framework integrates a DT…
Belinda Boehm, Christopher McNeill, David Huang
Understanding the solution-phase behaviour of organic semiconducting polymers is important for systematically improving the performance of devices based on solution-processed thin films of these molecules. Conventional polymer theory predicts that polymer conformations become more compact as solvent quality decreases…
Fahem Arar, Riad Mokadem, Djamel Eddine Zegour
— Given its intuitive nature, many Cloud providers opt for threshold-based data replication to enable automatic resource scaling. However, setting thresholds effectively needs human intervention to calibrate thresholds for each metric and requires a deep knowledge of current workload trends, which can be challenging to…
Ahmed M. Abdelmoniem, Sameh Abdulah, Walid Atwa
MapReduce Code Authors: ['Ahmed M. Abdelmoniem' 'Sameh Abdulah' 'Walid Atwa'] Data management applications are growing and require more attention, especially in the "big data" era. Thus, supporting such applications with novel and efficient algorithms that achieve higher performance is critical. Array database…
Lakshani Weerarathna, Oliver Weismantel, Tanja Junkers
A fully automated robotic synthesizer for the screening of amphiphilic block copolymer nanoparticle synthesis is presented. To reach this aim, block copolymer solutions are mixed in continuous flow with water, allowing for the automated variation of overall polymer concentration, mixing ratio of the water and organic…
Pelin Icer Baykal, Mike Simonov, Dhrithi Deshpande, Ful Belin Korukoglu + 8 more
Genomic research relies on accurate and reproducible computational analyses of DNA sequencing data to draw reliable biological conclusions. Read mapping, the process of aligning reads to a reference genome, is central to many applications, including variant detection and comparative genomics. While several tools have…
Michał Bałchanowski, Urszula Boryczka, Jiayi Ma
The aim of a recommender system is to suggest to the user certain products or services that most likely will interest them. Within the context of personalized recommender systems, a number of algorithms have been suggested to generate a ranking of items tailored to individual user preferences. However, these algorithms…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…