11 papers · ranked by Valyu relevance
Stephen Coshatt, He Yang, Shushan Wu, Jin Ye + 5 more
As machine learning and artificial intelligence are being integrated into cyber-physical systems, it is becoming important for engineers to know and understand these topics. In particular, sensor data is on the rise in these systems and therefore engineers need to understand which models are appropriate to time-series…
Julie R. Pivin-Bachler, Egon L. van den Broek
Title: Summary Machine learning struggles with imbalanced data. Although several mitigation approaches exist, their application depends on the extent of imbalance. To determine the latter, a protocol was developed. Across 428 synthetic and 70 real datasets, 8 imbalance measures were benchmarked and evaluated using…
Authors not listed
TurtleMol is an open-source Python package that aims to help users generate large, complex molec- ular systems. In the current version, users can generate systems by filling volumes defined by basic geometric shapes (e.g. cube, sphere), or by shapes of arbitrary gemoetries defined meshes created in other software (such…
Han Zhang, John Jonides
We present PupEyes, an open-source Python package for preprocessing and visualizing pupil size and fixation data. PupEyes supports data collected from EyeLink and Tobii eye-trackers as well as any generic dataset that conforms to minimal formatting standards. Developed with current best practices, PupEyes provides a…
Authors not listed
We present an open source collection of scripts and programs for the setup, management and evaluation of calculations with the Vienna ab-initio simulation package (VASP), called utils4VASP. It contains 20 independent Python scripts and Fortran programs, all with a unified and intuitive handling concept based on command…
Authors not listed
With the rapid growth of chemical data and information, there is an increasing need for chemistry undergraduates to master Python tools for analyzing large chemical datasets and extracting key or feature information. Currently, more than 100,000 types of metal-organic frameworks (MOFs), as the material recently awarded…
Gino Carmona-Díaz, William Jiménez-Leal, María Alejandra Grisales, Chandra Sripada + 3 more
Analyzing texts such as open-ended responses, headlines, or social media posts is a time- and labor-intensive process highly susceptible to bias. However, large language models (LLMs) are promising tools for text analysis, using either a predefined (top-down) or a data-driven (bottom-up) taxonomy, without sacrificing…
Jaka Kokošar, Ela Praznik, Martin Špendl, Nancy P. Moreno + 4 more
In biomedicine, survival analysis addresses time-to-event data to study outcomes like patient survival and treatment response, and supports biomarker discovery. Yet, teaching this analysis is often hindered by mathematical and programming barriers. We present a structured, hands-on tutorial that goes beyond a typical…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Authors not listed
Computational methods for predictive modeling have been increasingly utilized in the early stages of drug discovery to supplement high-throughput screening. The advent of highly efficient and complex machine learning architectures necessitates new methods of collating the plethora of topological, geometrical, and…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…