Aaron

Aaron J. Li

PhD Student in Computer Science

University of California, Berkeley

I am a second-year CS PhD student at UC Berkeley advised by Prof. Ion Stoica and Prof. Bin Yu, affiliated with Sky Computing Lab and BAIR. I completed my Master’s degree in Computational Science and Engineering at Harvard University, where I was fortunate to be advised by Prof. Hima Lakkaraju. Prior to that, I earned my Bachelor’s degree from UC Berkeley, double majoring in Computer Science and Psychology.

I’ve been collaborating with LM Arena since Spring 2026. In summer 2026, I joined Microsoft Research (Redmond) as a Research Intern.

My research centers on LLM-based agentic systems, with the overarching goal of making them more useful, affordable, and reliable in practice. Here are several concrete directions I’ve been actively working on:

· Improving the efficiency of agentic systems through developing open-source frameworks compatible with mainstream proprietary AI products, such as frontier coding agents.

· Developing evaluation paradigms for agentic systems that are realistic (closing the gap between benchmarks and real-world utility), sustainable (addressing saturation and contamination), and informative (surfacing concrete failure modes that guide model development and alignment).

· Building agentic systems with genuine real-world value, focusing on domain-specific self-improvement and reliable customizability, leveraging techniques such as test-time scaling and post-training.

· Operationalizing AI safety risks as concrete, measurable behaviors, and studying how post-training and alignment techniques can systematically address them in practice.

I’m always open to collaborations and happy to discuss all kinds of research ideas, and the best way to contact me is through email. For undergraduate students interested in working with me, I’m happy to have you either leading your own project or joining an existing one as a contributor, if there’s a good fit.

Selected Publications * equal contribution  · † equal advising

  1. dualeval.png
    DualEval: Joint Model-Item Calibration for Unified LLM Evaluation
    Aaron J. Li, Hao Huang, Youngmin Park, Yitong Ma, Wei-Lin Chiang,
    Li Chen, Cho-Jui Hsieh, Bin Yu, and Ion Stoica
    EMNLP 2026 Findings
  2. benchevolver.png
    BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
    Yangzhen Wu*, Aaron J. Li*, Wenjie Ma, Li Cao, Ziheng Zhou, Mert Cemri,
    Shu Liu, Yuran Xiu, Chenxiao Yan, Haikun Zhao, Bin Yu,
    Ion Stoica†, and Dawn Song†
    NeurIPS 2026
  3. greenshielding.png
    Green Shielding: A User-Centric Approach Towards Trustworthy AI
    Aaron J. Li*, Nicolas Sanchez*, Hao Huang, Ruijiang Dong, Jaskaran Bains,
    Katrin Jaradeh, Zhen Xiang, Bo Li, Feng Liu, Aaron Kornblith, and Bin Yu
    arXiv Preprint, 2026
  4. sae_robustness.png
    Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
    Aaron J. Li, Suraj Srinivas, Usha Bhalla, and Himabindu Lakkaraju
    EACL 2026
  5. rlhf_trust.png
    More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
    Aaron J. Li, Satyapriya Krishna, and Himabindu Lakkaraju
    ICLR 2025  ·  Oral, Top 1.8%
  6. r3_ppnet.png
    Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining
    Aaron J. Li, Robin Netzorg, Zhihan Cheng, Zhuoqin Zhang, and Bin Yu
    ICML 2024