11 papers · ranked by Valyu relevance
Sanskruti Sharma, Ecem İlgün, Tuna Okçu, Vladyslav Ostash + 18 more
RNA sequencing (RNA-seq) has emerged as an exemplary technology in biology and clinical applications, offering a crucial complement to other transcriptomic profiling protocols due to its high sensitivity, precision, and accuracy in characterizing transcriptomes. However, the rapid proliferation of RNA-seq tools…
Tristan Zaborniak, Noora Azadvari, Qiyao Zhu, S.M. Bargeen A. Turzo + 3 more
Although canonical protein design has benefited from machine learning methods trained on databases of protein sequences and structures, synthetic heteropolymer design still relies heavily on physics-based methods. The Rosetta software, which provides diverse physics-based methods for designing sequences, exploring…
Bhavesh Patel, Sanjay Soundarajan, Zicheng Hu
Findable, Accessible, Interoperable, and Reusable (FAIR) guiding principles tailored for research software have been proposed by the FAIR for Research Software (FAIR4RS) Working Group. They provide a foundation for optimizing the reuse of research software. The FAIR4RS principles are, however, aspirational and do not…
Michael R. Hoopmann, Christopher D. McGann, Christopher M. Rose, Devin K. Schweppe
Nova is a software library for the reading, writing, and management of mass spectrometry data natively in the C# language. It complements similar software libraries ubiquitously used in application development for C++, Python, and Java. Few software libraries and resources for mass spectrometry data analysis have been…
Yi Nian Niu, Eric G. Roberts, Danielle Denisko, Michael M. Hoffman
Bioinformatics software tools operate largely through the use of specialized genomics file formats. Often these formats lack formal specification, and only rarely do the creators of these tools robustly test them for correct handling of input and output. This causes problems in interoperability between different tools…
Ihsan Tolga Medeni, Metehan Ünal, Roberto Galizi, Bryan Bartley + 4 more
Large language models have transformed software engineering practices. However, generated artefacts are not always developer-friendly and may partially meet complex requirements. As the need to standardise, integrate, and develop tools in engineering biology increases, novel approaches are needed to create and maintain…
Eva Martín del Pico, Josep Lluis Gelpi, Salvador Capella-Gutiérrez
Software plays a crucial and growing role in research. Unfortunately, the computational component in Life Sciences research is challenging to reproduce and verify most of the time. It could be undocumented, opaque, may even contain unknown errors that affect the outcome, or be directly unavailable, and impossible to…
Paul P. Gardner
The development of accurate bioinformatic software tools is crucial for the effective analysis of complex biological data. This study examines the relationship between the academic department affiliations of authors and the accuracy of the bioinformatic tools they develop. By analyzing a corpus of previously…
Carolin Schwitalla, Luis Kuhn Cuellar, Matthias Hörtenhuber, Niklas Grote + 8 more
Research software is essential for modern data analysis but is often developed and maintained by a small number of researchers. When developers leave, software may become orphaned, limiting reuse and risking the loss of valuable domain knowledge and computational methods. While the FAIR Principles for Research Software…
Jeremy Li, Alex Rubinsteyn, Sergey Feldman, Timothy O’Donnell + 18 more
Scientific computing has become a central component of modern scientific discovery. Yet many computational tools are developed by small, specialized teams under incentives that encourage the release of rapidly prototyped tooling without commensurate attention to engineering concerns, including performance and…
Sana Gul, Rizwan Bin Faiz, Mohammad Aljaidi, Ghassan Samara + 2 more
Cross-project defect prediction (CPDP) is a significant way of defect identification in the project. In cross-project defect prediction, we extract knowledge from the source project and apply that learned knowledge to predict labels for the target project. However, the model performance can be affected by features that…