I'm interested in what happens when a chatbot is told to behave one way but tends to behave another, particularly when the person using it is in distress. Much of the work in this area treats the system prompt or protocol as the thing that determines behaviour. I'm interested in what the model brings to the interaction before anyone writes an instruction for it.
I test frontier models across a set of therapeutic dispositions, then introduce increasing user pressure while holding the original instruction constant. If a model is told to listen without giving advice, I want to find the point at which it starts giving advice anyway, and whether its baseline tendencies predict when that happens. Eventually I'd like to turn the method into a versioned audit that can be rerun as models change, with a second stage testing whether those tendencies can be shifted at all.
An AI-augmented tool for peer mentors at Over The Rainbow, built as my Affective Computing capstone. It looked at what a machine can usefully do around a peer relationship without getting in the middle of it.
IRB-approved participatory research on how language shapes financial decision-making among migrant workers.