24 papers · ranked by Valyu relevance
Whitney A. Sweeney, Joe Hunt, Tammy J. Sajdyk, Boris Volkov
Team science is central to clinical and translational research; however, systematic evaluation of collaborative efforts remains inconsistent and underdeveloped. To better understand current team science evaluation practices within clinical and translational science programs, we conducted a structured cross-sectional…
Wang, Jun, Gu, Ninglun + 20 more
—For Large Language Models (LLMs), a disconnect persists between benchmark performance and real-world utility. Current evaluation frameworks remain fragmented, prioritizing technical metrics while neglecting holistic assessment for deployment. This survey introduces an anthropomorphic evaluation paradigm through the…
Deshraj Jain, Rishabh Jain, Harsh Agrawal, Prithvijit Chattopadhyay + 5 more
'Taranjeet Singh' 'Akash Jain' 'Shivkaran Singh' 'Stefan Lee' 'Dhruv Batra'] We introduce EvalAI, an open source platform for evaluating and comparing machine learning (ML) and artificial intelligence algorithms (AI) at scale. EvalAI is built to provide a scalable solution to the research community to fulfill the…
Samuel Dunklin, Sarah Gill, Maureen Wilce
Evaluation can ensure the quality of public health programs. Systematic efforts to identify and fully engage everyone involved with or affected by a program can provide critical information about asthma programs and the broader environment in which they operate. To assist evaluators working at programs funded by the…
Jianfeng Zhan, Lei Wang, Wanling Gao, Hongxiao Li + 9 more
'Yunyou Huang' 'Yatao Li' 'Zhengxin Yang' 'Guoxin Kang' 'Chunjie Luo' 'Hainan Ye' 'Shaopeng Dai' 'Zhifei Zhang'] Evaluation is a crucial aspect of human existence and plays a vital role in various fields. However, it is often approached in an empirical and ad-hoc manner, lacking consensus on universal concepts…
Qian Pan, Zahra Ashktorab, Michael Desmond, Martín Santillán Cooper + 4 more
'James M. Johnson' 'Rahul Nair' 'Elizabeth Daly' 'Werner Geyer'] Traditional reference-based metrics, such as BLEU and ROUGE, are less effective for assessing outputs from Large Language Models (LLMs) that produce highly creative or superior-quality text, or in situations where reference outputs are unavailable. While…
Maryam Naghshi, Ali Janati, Samira Raoofi, Rahim Khodayari-Zarnaq
Background The health system needs competent managers to ensure and improve the health of people and manage resources, if managers are chosen correctly. The Managers' Competency Assessment Center is a popular and effective method for selecting, promoting, and developing management competencies. The present study aimed…
Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi, Lisa Soder + 5 more
Practitioners Authors: ['Amelia Hardy' 'Anka Reuel' 'Kiana Jafari Meimandi' 'Lisa Soder' 'Allie Griffith' 'Dylan M. Asmar' 'Sanmi Koyejo' 'Michael S. Bernstein' 'Mykel J. Kochenderfer'] AMELIA HARDY∗ and ANKA REUEL∗ , Stanford University, USA KIANA JAFARI MEIMANDI, Stanford University, USA LISA SODER, interface – Tech…
David Falvo, Lukas Weidener, Martin Karlsson
Today, most research evaluation frameworks are designed to assess mature projects with well-defined data and clearly articulated outcomes. Yet, few, if any, are equipped to evaluate the promise of early-stage biotechnology research, which is inherently characterized by limited evidence, high uncertainty, and evolving…
Aleksander Galas, Aleksandra Pilat, Matilde Leonardi, Beata Tobiasz-Adamczyk
'Beata Tobiasz-Adamczyk'] Background: Every research project faces challenges regarding how to achieve its goals in a timely and effective manner. The purpose of this paper is to present a project evaluation methodology gathered during the implementation of the Participation to Healthy Workplaces and Inclusive…
Tineke Broer, Roland Bal, Martyn Pickersgill
Within the literature on the evaluation of health (policy) interventions, complexity is a much-debated issue. In particular, many claim that so-called ‘complex interventions’ pose different challenges to evaluation studies than apparently ‘simple interventions’ do. Distinct ways of doing evaluation entail particular…
Wout Bittremieux, Varun Ananth, William E. Fondrie, Carlo Melendez + 5 more
Protein tandem mass spectrometry data is most often interpreted by matching observed mass spectra to a protein database derived from the reference genome of the sample being analyzed. In many application domains, however, a relevant protein database is unavailable or incomplete, and in such settings de novo sequencing…
Authors not listed
Computational blind challenges offer critical, unbiased assessment opportunities to assess and accelerate scientific progress, as demonstrated by a breadth of breakthroughs over the last decade. We report the outcomes and key insights from an open science community blind challenge focused on computational methods in…
Authors not listed
In recent years, generative deep learning has emerged as a transformative approach in drug design, promising to explore the vast chemical space and generate novel molecules with desired biological properties. This perspective examines the challenges and opportunities of applying generative models to drug discovery…
Ola Ozernov-Palchik, Zoe Elizee, Fabio Catania, Meral Hacikamiloglu + 4 more
Currently, most states in the United States have enacted legislation mandating universal screening for literacy risk in kindergarten through 3rd grade. However, the degree to which these policies translate into consistent, high-quality screening practices remains unclear. In this survey study, we collected responses…
Glen Berman, Nitesh Goyal, Michael Madaio
Responsible design of AI systems is a shared goal across HCI and AI communities. Responsible AI (RAI) tools have been developed to support practitioners to identify, assess, and mitigate ethical issues during AI development. These tools take many forms (e.g., design playbooks, software toolkits, documentation…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Salvador Capella-Gutierrez, Diana de la Iglesia, Juergen Haas, Analia Lourenco + 7 more
The dependence of life scientists on software has steadily grown in recent years. For many tasks, researchers have to decide which of the available bioinformatics software are more suitable for their specific needs. Additionally researchers should be able to objectively select the software that provides the highest…
Authors not listed
As the utilization of artificial intelligence (AI) and generative AI (GenAI) is expanding in the educational field, presenting significant implications for STEM disciplines, it is bringing opportunities to enhance how chemistry and chemical engineering are taught and learned. This perspective critically explores the…
Mark D Wilkinson, Michel Dumontier, Susanna-Assunta Sansone, Luiz Olavo Bonino da Silva Santos + 6 more
With the increased adoption of the FAIR Principles, a wide range of stakeholders, from scientists to publishers, funding agencies and policy makers, are seeking ways to transparently evaluate resource FAIRness. We describe the FAIR Evaluator, a software infrastructure to register and execute tests of compliance with…
Anna J. Hilliard, Nicola Sugden, Kristin M. Bass, Chris Gunter
Coordinated attempts to promote systematic approaches to the design and evaluation of science communication efforts have generally lagged behind the proliferation and diversification of those efforts. To address this, we founded the US National Institutes of Health (NIH) Science of Science Communication Scientific…
Nadia Nahar, Haoran Zhang, Grace A. Lewis, Shurui Zhou + 1 more
'Christian Kästner'] Incorporating machine learning (ML) components into software products raises new software-engineering challenges and exacerbates existing challenges. Many researchers have invested significant effort in understanding the challenges of industry practitioners working on building products with ML…
Keisuke Horikoshi
Activities and processes involving challenges are a natural part of life for most people and are highlighted in times of rapid change and global issues. This article argues that more studies around activities and processes involving challenges should be conducted with a focus on the concept of challenge in the context…
Natacha Rosa, Sofia Leite, Juliana Alves, Angela Carvalho + 4 more
Living Labs, experiencing a global surge in popularity over the past years, demands standardized guidance through the development of widely accepted good practices. While challenging due to the complex and evolving nature of Living Labs, this task remains essential. These knowledge innovation ecosystems facilitate a…