24 papers · ranked by Valyu relevance
Song Deng, Dong Yue, Le-chan Yang, Xiong Fu + 2 more
'Jayoung Kim'] For high-dimensional and massive data sets, traditional centralized gene expression programming (GEP) or improved algorithms lead to increased run-time and decreased prediction accuracy. To solve this problem, this paper proposes a new improved algorithm called distributed function mining for gene…
Bing Zhang, Guoyan Huang, Yuqian Wang, Haitao He + 2 more
'Zhong-Ke Gao'] As the quality of crucial entities can directly affect that of software, their identification and protection become an important premise for effective software development, management, maintenance and testing, which thus contribute to improving the software quality and its attack-defending ability. Most…
Aditi T. Merchant, Samuel H. King, Eric Nguyen, Brian L. Hie
Generative genomics models can design increasingly complex biological systems. However, effectively controlling these models to generate novel sequences with desired functions remains a major challenge. Here, we show that Evo, a 7-billion parameter genomic language model, can perform function-guided design that…
Alexey Braver, Dror G. Feitelson
Creating functions is at the center of writing computer programs. But there has been little empirical research on how this is done and what are the considerations that developers use. We design an experiment in which we can compare the decisions made by multiple developers under exactly the same conditions. The…
Zhuang Xiong, Yunfeng Zhang, Xiaodie Chen, Ajia Sha + 6 more
This study utilized 16S rRNA high-throughput sequencing technology to analyze the community structure and function of endophytic bacteria within the roots of three plant species in the vanadium-titanium-magnetite (VTM) mining area. The findings indicated that mining activities of VTM led to a notable decrease in both…
Michael J. Mior
Functional and inclusion dependencies are the most widely used classes of data dependencies in data profiling due to their ability to identify relationships in data such as primary and foreign keys. These relationships are equally important when dealing with nested data formats such as JSON. However, the definition of…
Jing Chen, Aijun Liu, Hongjun Zhang, Shengyi Yang + 3 more
'Ning Zhou' 'Peng Li'] With the rapid development of AI and big data mining technologies, computerized medical decision-making has become increasingly prominent. The aim of high-utility pattern mining (HUPM) is to discover meaningful patterns in medical databases that contribute to maximizing the utility from the…
Priyanka Rahi
The World Wide Web is a popular and interactive medium to distribute information in this scenario. The web is huge, diverse, ever changing, widely disseminated global information service center. We are familiar with terms like e-commerce, e-governance, e-market, e-finance, e-learning, e-banking etc. for an organization…
Patryk Burek, Frank Loebe, Heinrich Herre
Background Gene Ontology (GO) is the largest resource for cataloging gene products. This resource grows steadily and, naturally, this growth raises issues regarding the structure of the ontology. Moreover, modeling and refactoring large ontologies such as GO is generally far from being simple, as a whole as well as…
Jiao Yue, Dongpeng Zhang, Miaomiao Cao, Yukui Li + 4 more
Nine land types in the northern mining area (BKQ) (mining land, smelting land, living area), the old mining area (LKQ) (whole-ore heap, wasteland, grassland), and southern mining area (NKQ) (grassland, shrubs, farmland) of Xikuang Mountain were chosen to explore the composition and functions of soil bacterial…
Robert Starke, Petr Capek, Daniel Morais, Nico Jehmlich + 1 more
Unveiling the relationship between taxonomy and function of the microbiome is crucial to determine its contribution to ecosystem functioning. However, while there is a considerable amount of information on microbial taxonomic diversity, our understanding of its relationship to functional diversity is still scarce. Here…
Adarsh Arun, Zhen Guo, Simon Sung, Alexei Lapkin
Automated prediction of reaction impurities can be useful in facilitating rapid early-stage reaction development, synthesis planning and optimization. Existing reaction predictors are catered towards main product prediction, and are often black-box, making it difficult to troubleshoot erroneous outcomes. This work…
Neelamadhab Padhy
In this paper we have focused a variety of techniques, approaches and different areas of the research which are helpful and marked as the important field of data mining Technologies. As we are aware that many MNC's and large organizations are operated in different places of the different countries. Each place of…
Moises A. Rojas, Gladis Serrano, Jorge Torres, Jaime Ortega + 8 more
Microbial communities inhabiting mining environments harbor a diverse array of bacteria with specialized metabolic capacities adapted to extreme conditions. Here, we utilized comparative genome-resolved metagenomics of a high-quality Illumina-sequenced sample from the Cauquenes copper tailing in central Chile. We…
Philipp Rosenthal, Niels Demke, Frank Mantwill, Oliver Niggemann
The presented approach defines the decomposition problem in terms of a planning problem—a well established field in Artificial Intelligence. For the planning problem, logic-based solvers can be used to find solutions that compute a useful function structure for the design process. Well-known function libraries from…
Laura Kaikkonen, Malcolm R. Clark, Daniel Leduc, Scott D. Nodder + 4 more
Increasing interest in seabed resource use in the ocean is introducing new pressures on deep-sea environments, the ecological impacts of which need to be evaluated carefully. The complexity of these ecosystems and the dearth of comprehensive data pose significant challenges to predicting potential impacts. In this…
Beth N. Orcutt, James Bradley, William J. Brazelton, Emily R. Estes + 7 more
Interest in extracting mineral resources from the seafloor through deep-sea mining has accelerated substantially in the past decade, driven by increasing consumer demand for various metals like copper, zinc, manganese, cobalt and rare earth elements. While there are many on-going discussions and studies evaluating…
Tobias Baum, Steffen Herbold, Kurt Schneider
Data extracted from software repositories is used intensively in Software Engineering research, for example, to predict defects in source code. In our research in this area, with data from open source projects as well as an industrial partner, we noticed several shortcomings of conventional data mining approaches for…
Authors not listed
Iron, the most abundant element on Earth by mass (34.6%), primarily exists as iron minerals due to its inherent reactivity. The study of iron mineral phase transformations under changing environmental conditions remains an important research focus due to its geological, environmental, and industrial significance. Yet…
Laura Kaikkonen, Inari Helle, Kirsi Kostamo, Sakari Kuikka + 4 more
Seabed mining is approaching the commercial mining phase across the world’s oceans. This rapid industrialization of seabed resource use is introducing new pressures to marine environments. The environmental impacts of such pressures should be carefully evaluated prior to permitting new activities, yet observational…
Authors not listed
Curried functions provide a systematic way of transforming multi-argument functions into nested singleargument functions. This transformation allows partial application and supports many central principles of functional programming. Their extension, called curried 𝑘-ary functions, naturally generalizes the familiar…
Kevin Maik Jablonka, Andrew S. Rosen, Aditi S. Krishnapriyan, Berend Smit
The space of all plausible materials for a given application is so large that it cannot be explored using a brute-force approach. This is, in particular, the case for reticular chemistry which provides materials designers with a practically infinite playground on different length scales. One promising approach to guide…
Authors not listed
Artificial intelligence (AI) is reshaping scientific research by accelerating discovery and enabling the analysis of complex data that traditional methods struggle to handle. This review examines over 310,000 journal articles and patents from the CAS Content Collection (2015–2025), with a focus on, biomedical research…
Julian Ivanov, Alan Lipkus, Haitao Chen, Chris Aultman + 3 more
A novel bibliometric methodology based on natural language data processing for identifying emerging topics in science is presented. Along with the usual practice of data collection and preprocessing, our method includes a natural language processing (NLP) technique and an innovative mathematical function data…