21 papers · ranked by Valyu relevance
Chabin, Jacques, Ferrari, Mirian Halfeld + 2 more
We introduce an automated method for structuring textual data into a model-agnostic schema, enabling alignment with any database model. It generates both a schema and its instance. Initially, textual data is represented as semantically enriched syntax trees, which are then refined through iterative tree rewriting and…
K Karthikeyan, Raghuveer Thirukovalluru, David Carlson
Clinical notes contain valuable, context-rich information, but their unstructured format introduces several challenges, including unintended biases (e.g., gender or racial bias), and poor generalization across clinical settings (e.g., models trained on one EHR system may perform poorly on another due to format…
Authors not listed
Computational modeling of enzymes provides molecular-level insight into catalysis, but the preparation of quantum mechanical (QM) calculations starting from experimental structures is a significant bottleneck for high-throughput studies. Automated tools developed to accelerate this process may fail to generalize across…
Shani Alkoby, Ron S. Hirschprung
Introduction Privacy has become a significant concern in the digital world, especially concerning the personal data collected by websites and other service providers on the World Wide Web network. One of the significant approaches to enable the individual to control privacy is the privacy policy document, which…
Igor Martayan, Loup Lobet, Camille Marchet, Charles Paperman
Modern sequencing pipelines routinely produce billions of reads, yet the dominant storage formats (FASTQ and FASTA) are text-based and sequential, making high-throughput parsing a persistent bottleneck in bioinformatics. Their regular, line-oriented structure makes them well-suited to SIMD vectorization, but existing…
Authors not listed
We present an open source collection of scripts and programs for the setup, management and evaluation of calculations with the Vienna ab-initio simulation package (VASP), called utils4VASP. It contains 20 independent Python scripts and Fortran programs, all with a unified and intuitive handling concept based on command…
Rifat Mehreen Amin, Alperen Adatepe, Daniela Fernandes, Daniel Buschek + 1 more
Conversational interfaces powered by large language models (LLMs) are widely used for ideation and analysis, yet their linear structure limits exploration of alternatives and management of long-running interactions. We present CanvasConvo, a conversational interface concept that transforms linear chat into a branching…
James Urban, Roman Joeres, Daniel Bojar, Michael Gromiha
The Universal Input framework presented here addresses several critical challenges in glycoinformatics while offering distinct advantages over existing approaches. While, in some cases, individual converters exist for glycan nomenclatures, no current approach handles as many nomenclatures simultaneously as Universal…
Authors not listed
The ability to generate crystal structures directly from textual descriptions marks a pivotal advancement in materials informatics and underscores the emerging role of large language models (LLMs) in inverse design. In this work, we introduce CrysText, a text-conditioned framework that generates crystal structures in…
Ahmed Fathy Aly Omar Ibrahim, Katarzyna Jeleniewicz, Artur Piekarczuk, Valentino Paolo Berardi + 2 more
Highlights 1. Up to 31% reduction in structural mass was achieved through the use of closed-section profiles. 2. Hybrid configuration (IPE ribs + SHS rings) provided the most efficient structural performance among the three configurations considered, reaching utilization ratios of 0.87 (ribs) and 0.63 (rings). 3.…
Diaa Mohamed Fayed, Aly Aly Fahmy, Mohsen Abdelrazek Rashwan, Wafaa Kamel Fayed
Dictionaries are rich sources of lexical information about words that is required for many applications of natural language processing and human language technology. However, publishers prepare printed dictionaries for human usage not for machine processing. This paper presented a method to structure partly a…
Jane D. Fudyma, Petar Penev, Katerina Estera-Molina, Jordan Hoff + 3 more
Viruses have the potential to influence microbial community structure and elemental cycling in soils, but it remains unclear how these communities are distributed across space and time, which can shape how they respond to environmental change and impact ecosystem processes. Mediterranean grasslands, with their…
Adrian Tkachenko, Sepehr Salem, Ayotomiwa Ezekiel Adeniyi, Zülal Bingöl + 6 more
High-throughput sequencing (HTS) enables population-scale genomics but generates massive datasets, creating bottlenecks in storage, transfer, and analysis. FASTQ, the standard format for over two decades, stores one byte per base and one byte per quality score, leading to inefficient I/O, high storage costs, and…
Gloria Cecchini, Alex Roxin
In neural networks within the brain, the activity of a post-synaptic neuron is determined by the combined influence of many pre-synaptic neurons. This distributed processing enables mechanisms like Hebbian plasticity to associate sensory inputs with specific internal states, as seen in feedforward structures such as…
Akmuhammet Ashyralyyev, Zülal Bingöl, Begüm Filiz Öz, Kaiyuan Zhu + 4 more
Efficient and consistent string processing is critical in the exponentially growing genomic data era. Locally Consistent Parsing (LCP) addresses this need by partitioning an input genome string into short, exactly matching substrings (“cores”), ensuring consistency across partitions. Compared to the popular sketching…
Marcin Zukowski
Huffman encoding has been an enduring technique for 70+ years, ubiquitous in compression algorithms since its invention. In this paper we propose a new approach to Huffman coding, based on a data structure from wavelet trees. The resulting pivot-coded Huffman (PivCo-Huffman) enables high-performance SIMD-friendly…
Authors not listed
Given the recent inclusion of sodium-ion batteries (SIBs) in the energy market, the optimization of their performance becomes a relevant research topic. At the electrode-level, the parameters selected during its manufacturing process influence its microstructure and, consequently, its electrochemical performance. Here…
Jakob Schuster, Kathleen Zeglinski, Lucinda Xiao, Olivia Voulgaris + 5 more
The wide variety of protocols and applications for DNA and RNA sequencing makes flexible tools for read processing an important step in sequence analysis. Beyond trimming and demultiplexing, custom read-level processing is commonly required for data exploration, QC and analysis. Existing tools are often task-specific…
Somashekaracharya G. Bhaskaracharya, Aravind Acharya, Bastian Hagedorn, Vinod Grover
Modern deep learning compilers rely on layout abstractions to manage the complex mapping between logical tensor structures and physical memory arrangements. CuTe layouts and Triton linear layouts are widely adopted industry standards. However, these layout systems operate independently with distinct mathematical…
Mengtao Wang, Yifan Xu, Zaiyang Liu, Hidemitsu Furukawa + 3 more
Voxelizing active composite structures and controlling voxel-level material properties via 4D printing significantly expand design possibilities. However, as the number of voxels increases, the design space grows exponentially, posing significant challenges for predicting structural deformation. Here, a scalable…
Authors not listed
Infrared (IR) spectroscopy provides rich structural information but interpreting spectra at scale remains challenging. Here we introduce j-IR-vis, a vision-based neural model that learns chemically interpretable representations directly from IR spectra for functional-group prediction and downstream molecular…