25 papers · ranked by Valyu relevance
Stephen R. Piccolo, Michael B. Frampton
When reporting research findings, scientists document the steps they followed so that others can verify and build upon the research. When those steps have been described in sufficient detail that others can retrace the steps and obtain similar results, the research is said to be reproducible. Computers play a vital…
Michael S. Evans, Daniele Fanelli
In this paper I introduce computational techniques to extend qualitative analysis into the study of large textual datasets. I demonstrate these techniques by using probabilistic topic modeling to analyze a broad sample of 14,952 documents published in major American newspapers from 1980 through 2012. I show how…
Sreya Guha
Insights into social phenomenon can be gleaned from trends and patterns in corpora of documents associated with that phenomenon. Recent years have witnessed the use of computational techniques, mostly based on keywords, to analyze large corpora for these purposes. In this paper, we extend these techniques to…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Weerapong Phadungsukanan, Markus Kraft, Joe A Townsend, Peter Murray-Rust
This paper introduces a subdomain chemistry format for storing computational chemistry data called CompChem. It has been developed based on the design, concepts and methodologies of Chemical Markup Language (CML) by adding computational chemistry semantics on top of the CML Schema. The format allows a wide range of ab…
Mohamed Abuella
The management and organization of a large collection of academic documents is an important part of scientific research. This study explores the use of ChatGPT, a large language model from OpenAI, to extract insights from a large collection of academic documents stored in Mendeley. The study found that ChatGPT can be…
S. Erhardt, Mainak Ghosh, Erik Buunk, Michael E. Rose + 1 more
'Dietmar Harhoff'] Abstract—Logic Mill is a scalable and openly accessible software system that identifies semantically similar documents within either one domain-specific corpus or multi-domain corpora. It uses advanced Natural Language Processing (NLP) techniques to generate numerical representations of documents.…
Jens Dörpinghaus, Sebastian Schaaf, Marc Jacobs
Document clustering is widely used in science for data retrieval and organisation. DocClustering is developed to include and use a novel algorithm called PS-Document Clustering that has been first introduced in 2017. This method combines approaches of graph theory with state of the art NLP-technologies. This new…
James Philips, Nasseh Tabrizi
Historical Document Processing is the process of digitizing written material from the past for future use by historians and other scholars. It incorporates algorithms and software tools from various subfields of computer science, including computer vision, document analysis and recognition, natural language processing…
Fotis A. Baltoumas, Sofia Zafeiropoulou, Evangelos Karatzas, Savvas Paragkamian + 7 more
Extracting and processing information from documents is of great importance as lots of experimental results and findings are stored in local files. Therefore, extracting and analysing biomedical terms from such files in an automated way is absolutely necessary. In this article, we present OnTheFly^2.0^, a web…
Samantha Durdy, Cameron J. Hargreaves, Mark Dennison, Benjamin Wagg + 5 more
The discovery of new materials often requires collaboration between experimental and computational chemists. Web based platforms allow more flexibility in this collaboration by giving access to computational tools without the need for access to computational researchers. We present Liverpool Materials Discovery Server…
Stephen R. Piccolo, Zachary E. Ence, Elizabeth C. Anderson, Jeffrey T. Chang + 1 more
Command-line software plays a critical role in biology research. However, processes for installing and executing software differ widely. The Common Workflow Language (CWL) is a community standard that addresses this problem. Using CWL, tool developers can formally describe a tool’s inputs, outputs, and other execution…
Oliver Lee, Malte Gather, Eli Zysman-Colman
We describe a new tool for the efficient management of computational chemistry. Digichem is a program that automates and simplifies nearly the entire computational pipeline, including large-scale batch submission of calculations, analysis and results parsing, the generation of 3D density plots and 2D graphs of…
Sungwook Yoon
Enterprise document management faces a significant challenge: traditional clustering methods focus solely on content similarity while ignoring organizational context, such as priority, workflow status, and temporal relevance. This paper introduces FLACON (Flag-Aware Context-sensitive Clustering), an…
Uttam Mukhopadhyay
MINDS i• a di.tributed •v.tem of cooperating querv engine• that cultomizu document retrieval for each u1er in a dvnamic environment. It improvu it1 performance and adapt1 to changing pattern� of document dutribution bv oburving IJ!Item-ts�er interaction� and modifJ!lng the appropriate certaintv factor•, which act a•…
Marcello Castellano, Giuseppe Mastronardi, Roberto Bellotti, Gianfranco Tarricone
Background A fundamental activity in biomedical research is Knowledge Discovery which has the ability to search through large amounts of biomedical information such as documents and data. High performance computational infrastructures, such as Grid technologies, are emerging as a possible infrastructure to tackle the…
Varun Dogra, Sahil Verma, Kavita, Pushpita Chatterjee + 3 more
'Jaeyoung Choi' 'Muhammad Fazal Ijaz'] With the rapid advancement of information technology, online information has been exponentially growing day by day, especially in the form of text documents such as news events, company reports, reviews on products, stocks-related reports, medical reports, tweets, and so on. Due…
Shadrack Barnabas, Timo Böhme, Stephen Boyer, Matthias Irmer + 5 more
The extraction of chemical information from documents is a demanding task in cheminformatics due to the variety of text and image-based representations of chemistry. The present work describes the extraction of chemical compounds with unique chemical structures from the open access CORE (COnnecting REpositories) and…
Shadrack Barnabas, Timo Böhme, Stephen Boyer, Matthias Irmer + 4 more
The extraction of chemical information from documents is a demanding task in cheminformatics due to the variety of text and image-based representations of chemistry. The present work describes the extraction of chemical compounds with unique chemical structures from the open access CORE (COnnecting REpositories) and…
Ali Mansouri, Youssef Amghar
La gestion de la documentation technique, champ d'étude de notre travail, est une activité très intéressante pour les entreprises. En effet, les entreprises ont besoin de gérer leurs documents de la création jusqu'a l'archivage. Pour cela, le besoin d'élaboration des systèmes permettant la gestion des documents est…
Jan Range, Colin Halupczok, Jens Lohmann, Neil Swainston + 6 more
EnzymeML is an XML–based data exchange format that supports the comprehensive documentation of enzymatic data by describing reaction conditions, time courses of substrate and product concentrations, the kinetic model, and the estimated kinetic constants. EnzymeML is based on the Systems Biology Markup Language, which…
Authors not listed
Artificial intelligence (AI) is reshaping scientific research by accelerating discovery and enabling the analysis of complex data that traditional methods struggle to handle. This review examines over 310,000 journal articles and patents from the CAS Content Collection (2015–2025), with a focus on, biomedical research…
Adam Volanakis, Krawczyk Konrad
There are more than 26 million peer-reviewed biomedical research items according to Medline/PubMed. This breadth of information is indicative of the progress in biomedical sciences on one hand, but an overload for scientists performing literature searches on the other. A major portion of scientific literature search is…
Eugene Krissinel, Andrey A. Lebedev, Ville Uski, Charles B. Ballard + 30 more
'Ronan M. Keegan' 'Oleg Kovalevskiy' 'Robert A. Nicholls' 'Navraj S. Pannu' 'Pavol Skubák' 'John Berrisford' 'Maria Fando' 'Bernhard Lohkamp' 'Marcin Wojdyr' 'Adam J. Simpkin' 'Jens M. H. Thomas' 'Christopher Oliver' 'Clemens Vonrhein' 'Grzegorz Chojnowski' 'Arnaud Basle' 'Andrew Purkiss' 'Michail N. Isupov' 'Stuart…
Lorenzo Magnani
Eco-cognitive computationalism sees computation in context, exploiting the ideas developed in those projects that have originated the recent views on embodied, situated, and distributed cognition. Turing’s original intellectual perspective has already clearly depicted the evolutionary emergence in humans of…