13 papers · ranked by Valyu relevance
Daniel Engel, Freek Verbeek, Pranav Kumar, Binoy Ravindran
The binary executable format is the standard method for distributing and executing software. Yet, it is also as opaque a representation of software as can be. If the binary format were augmented with metadata that provides security-relevant information, such as which data is intended by the compiler to be executable…
Michael James Bommarito
Deep learning research for binary analysis faces a critical infrastructure gap. Today, existing datasets target single platforms, require specialized tooling, or provide only hand-engineered features incompatible with modern neural architectures; no single dataset supports accessible research and pedagogy on realistic…
Jordan, Herbert, Jezek, Kamil + 4 more
The State Database of a blockchain stores account data and enables authentication. Modern blockchains use fast consensus protocols to avoid forking, improving throughput and finality. However, Ethereum's StateDB was designed for a forking chain that maintains multiple state versions. While newer blockchains adopt…
Andrews Frimpong Adu, Elliot Sarpong Menkah, Peter Amoako-Yirenkyi, Samson Pandam Salifu
De novo genome assembly using de Bruijn graphs (DBGs) typically relies on fixed-length $k$-mers as the nodes of the graph. While this approach is effective, it presents a fundamental trade-off: smaller $k$ values tend to collapse repeats, whereas larger $k$ values can result in fragmentation, particularly in…
Authors not listed
Solubility is a crucial property of each organic compound, impacting its potential applications in synthetic chemistry, materials science and drug design. Moreover, in technological processes mixtures of solvents are often utilized, making the solubility assessment more complicated. Predicting solubility values in…
F. Kirchner, D.C.F. Wieland, S. Irvine, S. Schimek + 11 more
Scientific communities have recognized the importance of well-documented metadata generated during research. However, ensuring that metadata is findable, accessible, interoperable, and reusable (FAIR) remains a significant challenge. To address this, scientific communities are working towards making metadata available…
Ziad Ismaili Alaoui, Detlef Plump
We present an approach to implement binary search trees in the rule-based graph programming language GP 2. (See [[4]] for a brief introduction to GP 2.) Our implementation uses GP 2's rooted graph transformation rules to be fast [[1]] and supports insertion, deletion and query operations. We argue that the worst-case…
Mikhail Karasikov, Harun Mustafa, Daniel Danciu, Oleksandr Kulkov + 4 more
The amount of biological sequencing data available in public repositories is growing rapidly, forming a critical resource for biomedicine. However, making these data efficiently and accurately full-text searchable remains challenging. Here we build on efficient data structures and algorithms for representing large…
Stephen A. Fisher, Josef Hardi, Richard Morgan, Erik Nordgren + 7 more
Since publication of the FAIR Guiding Principles in 2016, the scientific community has increasingly sought to make experimental data findable, accessible, interoperable, and reusable. Operationalizing the FAIR principles in routine scientific workflows remains challenging without a standardized, workable…
Noam Teyssier, Alexander Dobin, Stephan Schiffels
Modern genomics produces billions of sequencing records per run, which are typically stored as gzip-compressed FASTQ files. While this format is widely used, it is not optimal for high-throughput processing due to its reliance on single-threaded decompression and sequential parsing of irregularly sized records. This…
C. Ke, J.J. Koehorst, B. Nijsse, W.T. Scott + 1 more
Molecular profiling using high-throughput ‘omics technologies has tremendously increased our ability to interrogate complex microbial communities at the molecular level. In the context of data reuse, the FAIRification of these extensive datasets is frequently perceived as a secondary administrative task, addressed only…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Maddimsetti Srinivas, Debdoot Sheet
A binary decision tree (BDT) is stochastic and depth-dependent when inference is performed. The lower and upper bounds are derived from the minimum and maximum heights of the leaf nodes. The inherent randomness complicates BDT and random forest (RF) inference processes for fixed-rate streaming data. BDT is reformulated…