24 papers · ranked by Valyu relevance
Abhinav Parmar, Abhisek Panigrahi, Abhishek Kumar Dwivedi, Abhishek Bhattacharya + 93 more
We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy on Mify-2.5B2 foundation model. Mify-Coder achieves comparable accuracy and safety while outperforming significantly larger baseline models on standard coding and function-calling benchmarks, demonstrating that…
Halimeh Agh, Betül Cimendag, Stefan Wagner
Machine learning systems consist of general-purpose code as well as machine-learning-specific code. While ML-specific code smells have been identified, their connection to project characteristics and their interaction with overall code quality are not well understood. Without this knowledge, quality assurance…
Rong Huang, Su Tao
Automated Machine Learning (AutoML) aims to streamline the end-to-end process of ML models, yet current approaches remain constrained by rigid rule-based frameworks and structured input requirements that create barriers for non-expert users. Despite advances in Large Language Models (LLMs) demonstrating capabilities in…
Agrawal, Shriyansh, Lau, Aidan + 6 more
The prevalence of Large Language Models (LLMs) for generating multilingual text and source code has only increased the imperative for machine-generated content detectors to be accurate and efficient across domains. Current detectors either incur high computational cost or lack sufficient accuracy, often with a…
Nicolas Lacroix, Mireille Blay-Fornarino, Benjamin Benni, Damien Garreau
—Background: Extracting the stages that structure Machine Learning (ML) pipelines from source code is key for gaining a deeper understanding of data science practices. However, the diversity caused by the constant evolution of the ML ecosystem (e.g., algorithms, libraries, datasets) makes this task challenging.…
Jeffery G Klann, Micheal M Mendis, Shyam Visweswaran, Shawn N Murphy + 2 more
We implemented an application programming interface (API) for the i2b2-ML platform to facilitate the creation and execution/application of ML models. The APIs allow the end-users to develop ML models within the i2b2 platform without the need to download data into external environments. The models are stored in i2b2’s…
Wang, Yiran, López, José Antonio Hernández + 4 more
Jupyter notebooks are widely used for machine learning (ML) prototyping. Yet, few debugging tools are designed for ML code in notebooks, partly, due to the lack of benchmarks. We introduce JunoBench, the first benchmark dataset of real-world crashes in Python-based ML notebooks. JunoBench includes 111 curated and…
Hung Q. Vo, Huy Q. Vo, Son T. Ly, Zhihao Wan + 5 more
Conventional tissue image analysis software provides foundational capabilities for cellular analysis, including segmentation, basic morphological feature extraction, and spatial organization analysis. However, these tools often require manual intervention and are not well integrated with code-driven automation…
Cameron S Movassaghi, Amanda Momenzadeh, Jesse G Meyer, Jonathan Wren
We started with a relatively simple algorithm, the random forest, initially described by Leo Breiman in 2001 (). We tested Claude 4 Sonnet, which was unable to accept context as long as the PDF and the prompt. We also tested Gemini Pro, which produced an error related to an edge case where there was no information gain…
Sabine Eichhorn, Franz Niklas Mitze, Fritz Wagner, Inga Marte Charlott Seuthe + 6 more
To verify medical representativeness, the most common ICD-10 codes and OPS codes of each database were selected. The results for each database were first ranked separately and then compared. The primary aim was to compare how many of the codes found in DESTATIS were also found in the ML-Network among the most frequent…
Authors not listed
The integration of machine learning methods is transforming many areas of research by, for instance, accelerating molecular dynamics simulations and enabling improved prediction and optimization of chemical reactions. However, despite this progress, the adoption of data-driven approaches in atomic layer deposition…
Beiji Lu
Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of in-dependently…
Chloe A. Game, Nils Piechaud, Kerry L. Howell
Deep learning (DL) is a powerful tool to extract ecological information from large image datasets efficiently and consistently. However, applying these methods remains challenging, due in part to the complexity of DL workflows and the dynamic nature of available tools. To address this, we created a practical guide and…
Adela Bara, Gabriela Dobrita, Simona-Vasilica Oprea
The purpose of our paper is to develop a unified multi-agent architecture that automates end-to-end machine learning (ML) pipeline generation from datasets and natural-language (NL) goals, improving efficiency, robustness and explainability. A five-agent system is proposed to handle profiling, intent parsing…
Vlastimil Martinek, Andrea Gariboldi, Dimosthenis Tzimotoudis, Mark Galea + 7 more
Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack…
M. Shahbaz Ismail, Sara Shahzad, Fahmi H. Quradaa, Sajid Anwar
Semantic code clone detection plays an essential role in software maintenance and quality assurance, as it helps uncover fragments of code that express the same logic even when their syntax has been altered or deliberately obfuscated. In this study, we propose a framework that combines hybrid representation learning…
Authors not listed
Metal hydrides play a pivotal role in a wide range of applications, including hydrogen storage, compression, heat management, and catalysis, making them a central focus of interdisciplinary research spanning chemistry, materials science, and engineering. The performance of the metal hydride based systems is strongly…
Melih Peker, Ozcan Ozturk
Selecting a good set of optimization flags requires extensive effort and expert input. While most of the prior research considers using static, spatial, or dynamic features, some of the latest research directly applied deep neural networks to source code. We combined the static features, spatial features, and deep…
Yanshuo Chen, Yuming Zhang, Joshua Li, Boxue Tian + 1 more
Codon optimization involves selecting synonymous codons to match host-specific preferences. It is critical for heterologous expression but remains challenging due to the combinatorial design space. Under long-term evolutionary selection, natural coding sequences are near-optimal compromises between translational…
Vlastimil Martinek, Andrea Gariboldi, Dimosthenis Tzimotoudis, Mark Galea + 7 more
The past decades have witnessed the transformation of molecular biology into a truly data-driven science, in large part due to the growth in the quantity and variety of molecular biology data generated by technologies such as mass spectrometry and high-throughput sequencing (, ). An important step in the analysis…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Fanny Bodart, Adrien De Voeght, Frédéric Baron, Gilles Louppe
Flow cytometry produces high-dimensional single-cell protein measurements central to immunophenotyp ing and clinical monitoring. Yet analysis still relies largely on manual gating, which is labour-intensive, poorly reproducible, and ill-suited to large marker panels. Existing computational approaches address…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Authors not listed
Realizing the promise of artificial intelligence (AI) to accelerate scientific progress and deliver technological impact depends on how effectively AI can be integrated into real-world decision- making processes. As Peter Norvig states, “Somewhat remarkably, almost all AI research until very recently has assumed that…