Yu Ying Chiu (Kelly Chiu)

PhD student at UC Berkeley working on AI safety and alignment.

icon.jpg

kellycyy [at] berkeley [dot] edu

Hi, I’m Kelly, a first-year PhD student at UC Berkeley, where I work with Prof. Serina Chang and Prof. Stuart Russell. I’m affiliated with the Center for Human-Compatible Artificial Intelligence (CHAI) and Berkeley Artificial Intelligence Research (BAIR).

I work on AI safety and alignment. My research asks how AI systems should navigate conflicting human values and how pluralistic input can inform concrete alignment objectives and evaluations.

Before Berkeley, I worked on evaluating AI values and moral reasoning, cultural understanding, and AI behavior in mental-health settings at the University of Washington, through the ML Alignment & Theory Scholars (MATS) program with Anthropic, and at New York University.

selected publications

moral values

  1. ICLR 2026
    MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
    Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul Font-Reaulx, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, and Sydney Levine
    2025
  2. ICLR 2026
    Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
    Yu Ying Chiu, Zhilin Wang, Sharan Maiya, Yejin Choi, Kyle Fish, Sydney Levine, and Evan Hubinger
    2025
  3. ICLR 2025 (Spotlight)
    DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
    Yu Ying Chiu, Liwei Jiang, and Yejin Choi
    In International Conference on Learning Representations 2025 (Spotlight), 2024

cultural values

  1. ACL 2025
    CulturalBench: a Robust, Diverse and Challenging Benchmark on Measuring the (Lack of) Cultural Knowledge of LLMs Through Human-AI Red-Teaming
    Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, and Yejin Choi
    2024

empathetic values

  1. Nature Communications (Accepted in Principle)
    A Computational Framework for Behavioral Assessment of LLM Therapists
    Yu Ying Chiu*, Ashish Sharma*, Inna Wanyin Lin, and Tim Althoff
    2024

AI agent simulations

  1. EMNLP 2023 System Demo
    humanoidagent_gif.gif
    Humanoid Agents: Platform for Simulating Human-like Generative Agents
    Zhilin Wang*, Yu Ying Chiu*, and Yu Cheung Chiu
    In Empirical Methods in Natural Language Processing 2023 (System Demonstrations), Dec 2023