I am a research scientist at the Princeton Center for Information Technology Policy, where I work with Arvind Narayanan and Zeynep Tufekci on multiple projects relating to AI agent evaluation and the societal impacts of AI. This past spring, I also worked with Tim Hua as part of the SPAR program on a project to study mechanisms for value reflection and systematization in frontier AI models.
Current Projects
- Open-world evaluations: I lead harness engineering, manage a team of six analysts, and direct reporting for CRUX (Collaborative Research for Updating AI eXpectations), a series of long-horizon, open-world evaluations of frontier AI agents on real-world tasks such as autonomous software development and autonomous AI research. Read the position paper introducing the project.
- AI agent reliability: Alongside Stephan Rabanser and the rest of the HAL team at Princeton, I am working to develop an index of AI agent reliability. Learn more at the project website.
- Anthropomorphization and attachment: With the AI and Society Lab at Princeton, I am working to document the prevalence of anthropomorphization, attachment, and engagement with AI chatbots using a combination of synthetic benchmarks, big data analytics, and population surveys.
- LLM value reflection: With Tim Hua, I am exploring how different LLMs iteratively refine their own constitutions, model specs, and system prompts.
News
- July 2026: Released Can AI Agents Conduct Open-Ended AI Research?, a first-author preprint introducing “shadow evaluations” of frontier agents on open-ended research tasks.
- May 2026: Log Analysis is Necessary for Credible Evaluation of AI Agents was accepted at the ICML 2026 FAGEN Workshop.
- April 2026: Launched CRUX (Collaborative Research for Updating AI eXpectations) with our first open-world evaluation, tasking an AI agent with autonomously developing and publishing an iOS app to the Apple App Store.
- March 2026: LLM Spirals of Delusion, an audit study of sycophancy and delusion reinforcement in ChatGPT, was accepted at IASEAI 2026.
- February 2026: Towards a Science of AI Agent Reliability was accepted at ICML 2026.
- January 2026: Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation was accepted at ICLR 2026.
- December 2025: Gave an invited talk at the NeurIPS AI Evaluator Forum on insights from the analysis of over 2,000 AI agent logs.
- November 2025: Presented “Is Consciousness Prerequisite for Moral Patienthood?” at the Eleos AI Consciousness Conference.
- August 2025: Joined the Princeton Center for Information Technology Policy as a research scientist working with Arvind Narayanan and Zeynep Tufekci.
- June 2025: Graduated from the Princeton School of Public and International Affairs with an MPA and a certificate in statistics and machine learning.
Background
I am a recent MPA graduate from Princeton, where I also completed a graduate certificate in statistics and machine learning. My broad research interests are in artificial intelligence, moral philosophy, intergenerational mobility, and social safety net implementation. I have also worked in civic technology, most recently at the US Census Bureau as a Coding it Forward Data Science Fellow, and prior to that, as a data analyst with the Massachusetts Digital Service. I received my BA from Williams College in 2020, where I studied political economy, philosophy, and cognitive science. Outside of work, I am an avid runner and cyclist and love spending long days in the mountains.