20 papers · ranked by Valyu relevance
Allan Vikiru, Mfadhili Muiruri, Ismail Lukandu Ateya
Numerous systems run on distributed database systems hosted on cloud architectures. Facebook applies several databases such as MySQL, Apache Hadoop and Apache Cassandra which are distributed across data centres. [1] Due to the amount of data generated by users streaming video, Hulu applied Apache Cassandra to ensure…
Priyanka Dash, Ranjita Rout, Satya Bhusan Pratihari, Sanjay Padhi
Considerable Progress has been made in the last few years in improving the performance of the distributed database systems. The development of Fragment allocation models in Distributed database is becoming difficult due to the complexity of huge number of sites and their communication considerations. Under such…
Hassen Fadoua, Amel Grissa-Touzi
> Abstract . Despite the increasing need for modeling and implementing Distributed Databases (DDB), distributed database management systems are still quite far from helping the designer to directly implement its BDD. Indeed, the fundamental principle of implementation of a DDB is to make the database appear as a…
Ruben Cruz Huacarpuma, Rafael Timoteo de Sousa Junior, Maristela Terto de Holanda, Robson de Oliveira Albuquerque + 6 more
The development of the Internet of Things (IoT) is closely related to a considerable increase in the number and variety of devices connected to the Internet. Sensors have become a regular component of our environment, as well as smart phones and other devices that continuously collect data about our lives even without…
David Bouzaglo, Israel Chasida, Elishai Ezra Tsur
The integration of cloud resources with federated data retrieval has the potential of improving the maintenance, accessibility and performance of specialized databases in the biomedical field. However, such an integrative approach requires technical expertise in cloud computing, usage of a data retrieval engine and…
C. Sunil Kumar, J. Seetha, S.R. Vinotha
Security features must be addressed when escalating a distributed database. The choice between the object oriented and the relational data model, several factors should be considered. The most important of these factors are single and multilevel access controls (MAC), protection and integrity maintenance. While…
Benjamin Tingle, Khanh Tang, Jose Castanon, John Gutierrez + 4 more
Purchasable chemical space has grown rapidly into the tens of billions of molecules providing unprecedented opportunities for ligand discovery, but also straining the tools that might exploit these molecules at scale. We have therefore developed ZINC-22, a database of commercially accessible small molecules derived…
Felipe Castro-Medina, Lisbeth Rodríguez-Mazahua, Asdrúbal López-Chau, Jair Cervantes + 2 more
'Jair Cervantes' 'Giner Alor-Hernández' 'Isaac Machorro-Cano'] Fragmentation is a design technique widely used in multimedia databases, because it produces substantial benefits in reducing response times, causing lower execution costs in each operation performed. Multimedia databases include data whose main…
Changhao Zhu, Junzhe Li, Ziyue Zhong, Cong Yue + 1 more
The success of blockchain technology in cryptocurrencies reveals its potential in the data management field. Recently, there is a trend in the database community to integrate blockchains and traditional databases to obtain security, efficiency, and privacy from the two distinctive but related systems. In this survey…
Dimitri Yatsenko, Jacob Reimer, Alexander S. Ecker, Edgar Y. Walker + 6 more
The rise of big data in modern research poses serious challenges for data management: Large and intricate datasets from diverse instrumentation must be precisely aligned, annotated, and processed in a variety of ways to extract new insights. While high levels of data integrity are expected, research teams have diverse…
Baldeep Singh, Randall Martyr, Thomas Medland, Jamie Astin + 2 more
'Gordon Hunter' 'Jean-Christophe Nebel'] About fifty years ago, the world’s first fully automated system for trading securities was introduced by Instinet in the US. Since then the world of trading has been revolutionised by the introduction of electronic markets and automatic order execution. Nowadays, financial…
Gamze Gürsoy, Charlotte M Brannon, Mark Gerstein
With the advent of precision medicine, pharmacogenomics data is becoming increasingly critical to patient care. These data describe the relationship between a particular variant in the genome and the response to a drug by the patient. As utilizing this kind of data becomes more integral to medical treatment decisions…
Lester Melie-Garcia, Bogdan Draganski, John Ashburner, Ferath Kherif
We propose a Multiple Linear Regression (MLR) methodology for the analysis of distributed and Big Data in the framework of the Medical Informatics Platform (MIP) of the Human Brain Project (HBP). MLR is a very versatile model, and is considered one of the workhorses for estimating dependences between clinical…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
Cláudia Brito, Pedro Ferreira, João Paulo
Breakthroughs in sequencing technologies led to an exponential growth of genomic data, providing unprecedented biological in-sights and new therapeutic applications. However, analyzing such large amounts of sensitive data raises key concerns regarding data privacy, specifically when the information is outsourced to…
Salim Miloudi, Yulin Wang, Wenjia Ding, Joaquín Abellán
Clustering algorithms for multi-database mining (MDM) rely on computing $((n2-n)/2)$ pairwise similarities between n multiple databases to generate and evaluate $(m\in[1,(n2-n)/2])$ candidate clusterings in order to select the ideal partitioning that optimizes a predefined goodness measure. However, when these pairwise…
Haotian Li
Machine learning and deep learning are novel and trending approaches to solving real-world scientific problems. Graph machine learning is dedicated to performing learning methods, such as graph neural networks, on non-Euclidean data such as graphs. Molecules, with their natural graph structures, could be analyzed by…
Authors not listed
Machine learning models are transforming data-driven research across scientific disciplines, yet their deployment as accessible and reliable web services remains a significant challenge. We introduce the NERDD framework, a scalable, maintainable, and secure microservices platform designed to support the sustainable…
Noah Lewis, Harshvardhan Gazula, Sergey M. Plis, Vince D. Calhoun
In this age of big data, large data stores allow researchers to compose robust models that are accurate and informative. In many cases, the data are stored in separate locations requiring data transfer between local sites, which can cause various practical hurdles, such as privacy concerns or heavy network load. This…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…