22 papers · ranked by Valyu relevance
Manuel Krebber, Henrik Barthels, Paolo Bientinesi
—Pattern matching is a powerful tool for symbolic computations. Applications include term rewriting systems, as well as the manipulation of symbolic expressions, abstract syntax trees, and XML and JSON data. It also allows for an intuitive description of algorithms in the form of rewrite rules. We present the open…
Manuel Krebber
Pattern matching is a powerful tool which is part of many functional programming languages as well as computer algebra systems such as Mathematica. Among the existing systems, Mathematica offers the most expressive pattern matching. Unfortunately, no open source alternative has comparable pattern matching capabilities.…
Arshia Ataee Naeini, Amir-Parsa Mobed, Masoud Seddighin, Saeed Seddighin
We study the fully dynamic pattern matching problem where the pattern may contain up to k wildcard symbols, each matching any symbol of the alphabet. Both the text and the pattern are subject to updates (insert, delete, change). We design an algorithm with O(n log 2 n) preprocessing and update/query time O˜(knk/k+1 + k…
Janja Paliska Soldo, Ana Sović Kržić, Damir Seršić
This paper focuses on pattern matching in the DNA sequence. It was inspired by a previously reported method that proposes encoding both pattern and sequence using prime numbers. Although fast, the method is limited to rather small pattern lengths, due to computing precision problem. Our approach successfully deals with…
Florín Manea, Markus L. Schmid
A pattern α (i. e., a string of variables and terminals) matches a word w, if w can be obtained by uniformly replacing the variables of α by terminal words. The respective matching problem, i. e., deciding whether or not a given pattern matches a given word, is generally NP-complete, but can be solved in…
Markus Seiler, Alexander Mehrle, Annemarie Poustka, Stefan Wiemann
Background The identification of patterns in biological sequences is a key challenge in genome analysis and in proteomics. Frequently such patterns are complex and highly variable, especially in protein sequences. They are frequently described using terms of regular expressions (RegEx) because of the user-friendly…
J. A. M. Rexie, Kumudha Raimond, Mythily Murugaaboopathy, D. Brindha + 1 more
'Henock Mulugeta'] An area of medical science, that is, gaining prominence, is DNA sequencing. Genetic mutations responsible for the disease have been detected using DNA sequencing. The research is focusing on pattern identification methodologies for dealing with DNA-sequencing problems relating to various…
Daniel Liu, Sven Rahmann
Next-generation sequencing technologies create large, multiplexed DNA sequences that require preprocessing before any further analysis. Part of this preprocessing includes demultiplexing and trimming sequences. Although there are many existing tools that can handle these preprocessing steps, they cannot be easily…
Nazim Uddin Sheikh, Hasina Rahman, Hamid Al-Qahtani
With the advent of large-scale heterogeneous search engines comes the problem of unified search control resulting in mismatches that could have otherwise avoided. A mechanism is needed to determine exact patterns in web mining and ubiquitous device searching. In this paper we demonstrate the use of an optimized string…
HyunJin Kim, Kang-Il Choi, Sang-Il Choi, Francesco Pappalardo
This paper proposes a memory-efficient bit-split string matching scheme for deep packet inspection (DPI). When the number of target patterns becomes large, the memory requirements of the string matching engine become a critical issue. The proposed string matching scheme reduces the memory requirements using the…
Tim Anderson, Travis J Wheeler
Pattern matching is a key step in a variety of biological sequence analysis pipelines. The FM-index is a compressed data structure for pattern matching, with search run time that is independent of the length of the database text. We present AvxWindowedFMindex (AWFM-index), an open-source, thread-parallel FM-index…
Paolo Ferragina, Bud Mishra
This paper reports an initial design of new data-structures that generalizes the idea of pattern-matching in stringology, from its traditional usage in an (unstructured) set of strings to the arena of a well-structured family of strings. In particular, the object of interest is a family of strings composed of…
Hongyi Xin, Jeremie Kim, Sunny Nahar, Carl Kingsford + 2 more
Approximate String Matching is a pivotal problem in the field of computer science. It serves as an integral component for many string algorithms, most notably, DNA read mapping and alignment. The improved LV algorithm proposes an improved dynamic programming strategy over the banded Smith-Waterman algorithm but suffers…
Rick Beeloo, Ragnar Groot Koerkamp
Approximate string matching (ASM) is the problem of finding all occurrences of a pattern P in a text T while allowing up to k errors. ASM was researched extensively around the 1990s, but with the rise of large-scale datasets, focus shifted towards inexact approaches based on seed-chain-extend. These methods often…
Rick Beeloo, Ragnar Groot Koerkamp
Approximate string matching (ASM) is the problem of finding all occurrences of a pattern in a text while allowing up to k errors. Many modern methods use seed-chain-extend, which is fast in practice, but does not guarantee finding all matches with ≤ k errors. However, applications such as CRISPR off-target detection…
Ping Zeng, Qingping Tan, Xiankai Meng, Zeming Shao + 5 more
'Ying Yan' 'Wei Cao' 'Jianjun Xu' 'Francesco Pappalardo'] In this paper, based on our previous multi-pattern uniform resource locator (URL) binary-matching algorithm called HEM, we propose an improved multi-pattern matching algorithm called MH that is based on hash tables and binary tables. The MH algorithm can be…
Trevor Gokey, David L. Mobley
Molecular mechanics force fields require a chemical perception model to assign parameters to molecules. A recent advancement in force fields is the use of the SMARTS substructure query language as the perception model. Although it is straightforward to write SMARTS patterns to define new force field parameters, it is…
Camila Zanette, Caitlin C. Bannan, Christopher I. Bayly, Josh Fass + 4 more
Molecular mechanics force fields define how the energy and forces of a molecular system are computed from its atomic positions, and enable the study of such systems through computational methods like molecular dynamics and Monte Carlo simulations. Despite progress toward automated force field parameterization…
Martin Priessner, Anna Tomberg, Jon Paul Janet, Richard J. Lewis + 2 more
In the pursuit of improved compound identification and database search tasks, this study explores Heteronuclear Single Quantum Coherence (HSQC) spectra simulation and matching methodologies. HSQC spectra serve as unique molecular fingerprints, enabling a valuable balance of data collection time and information…
Tizian Schulz, Paul Medvedev
Given a sequencing read, the broad goal of read mapping is to find the location(s) in the reference genome that have a “similar sequence”. Traditionally, “similar sequence” was defined as having a high alignment score and read mappers were viewed as heuristic solutions to this well-defined problem. For sketch-based…
Andrew Hoover, Martin Spale, Brian Lahue, Danny Bitton
To solve recurring problems in drug discovery, matched molecular pair (MMP) analysis is used to understand relationships between chemical structure and function. For the MMP analysis of large datasets (>10,000 compounds), available tools lack flexible search and visualization functionality and require computational…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…