Predicting Mental Health Adverse Events from Clinical Notes – Can AI Make a Difference?
Abstract
Predicting mental health adverse events from clinical notes – can AI make a difference? Ioana Danciu
Dissertation under the direction of Dr. Colin Walsh
Suicide is the third leading cause of death among young people aged 15-24 years old. In 2021 this age group experienced the highest statistically significant percentage increase among all other age groups. A plethora of information is contained in notes that is not well captured in structured data. Natural language processing (NLP) using deep learning techniques intrinsically captures relationships between features, leading to models that more accurately represent the data. These models, belonging to a class of generically termed artificial intelligence (AI) methods, are very complex, with many parameters and are not immediately transparent to humans. To address these gaps, we developed predictive models of suicide attempts and adverse psychiatric outcomes using deep learning NLP methods and investigated the interpretability of these algorithms in three steps: First we created NLP models to predict mental health adverse events within 90 days of an emergency department (ED) stay using clinical notes. In step 2 we developed and evaluated a qualitative framework for model interpretability using expert-informed literature review and a survey instrument. Lastly, we developed and evaluated two commonly used interpretability visualizations. Our study found that models that predict suicide attempts and psychiatric adverse outcomes in a general population can perform above a threshold that justifies intervention. An AI model performance above a preset threshold is no guarantee this tool will be adopted and implemented in clinical care. As evidenced by the results of our survey, informaticians want to understand how AI models make recommendations. Previous experiences with AI tools, the types of risk involved, the degree of trust in the model’s prediction, the alignment with clinician’s assessment of the case all affect the utility of interpretable AI tools. The type of practice, the timing in providers’ workflows, and type of display also play a big role in adoption. Our results indicate that more work needs to be done to visualize AI model explanations in ways that fit into clinical workflows.