Cureus. 2026 Jul 30;18(7):e113658. doi: 10.7759/cureus.113658. eCollection 2026 Jul.
ABSTRACT
Background Recruiting families for neonatal clinical trials is challenging due to narrow enrollment windows, high clinical acuity, and parental stress, which contributes to concerns regarding representativeness in neonatology research. The Better Research Interactions for Every Family (BRIEF) intervention was developed to improve recruitment experiences and equity. We aimed to gather validity evidence using Messick’s framework for the BRIEF assessment tool, a component of the BRIEF educational intervention, designed to evaluate research team members’ consent and recruitment interactions. Methods In this pilot prospective cohort study, physicians participating in the BRIEF educational intervention during a neonatal clinical trial completed recorded consent discussions (n = 22). Two blinded expert raters then evaluated sessions using the 10-item BRIEF assessment tool, a component of the BRIEF educational intervention. Participants completed self-assessments and surveys. Validity evidence for the BRIEF assessment tool was evaluated using weighted kappa for inter-rater agreement, Wilcoxon rank-sum tests for associations with self-reported skill, and Cronbach’s alpha for internal consistency. Results Five physicians contributed 22 recorded encounters using the BRIEF assessment tool. Internal consistency was strong across items (α > 0.70). Inter-rater agreement ranged from poor to almost perfect, with half of the items demonstrating moderate or higher agreement. Agreement between self-assessment and expert ratings was poor. Only one performance objective (PO) correlated with self-reported skill. Lower expert-rated scores were associated with families declining trial participation. Most participants rated themselves higher than expert raters. Conclusions In this pilot study, the BRIEF assessment tool demonstrates preliminary validity as an observer-based measure of neonatal research consent interactions, with strong internal consistency and variable inter-rater reliability. Self-assessment utility appears limited. Further validation in larger, more diverse samples is needed before high-stakes use.
PMID:42669018 | PMC:PMC13525908 | DOI:10.7759/cureus.113658