CV
Summary
Research scientist focused on rigorous empirical evaluation of the reliability, safety, and real-world capability of frontier AI systems and agents. Author of research at leading venues (ICML, ICLR, IASEAI), with hands-on experience designing and running large-scale evaluation and data pipelines. Combines technical expertise in machine learning and data science with public-policy training to translate AI evaluation into actionable insights for policymakers and the public.
Experience
Center for Information Technology Policy, Princeton University
August 2025 - Present
Research Scientist
- Lead harness engineering, manage a team of six analysts, and direct all reporting for a series of long-horizon assessments of frontier AI agents on real-world tasks such as autonomous software development and autonomous AI research.
- Led a multi-organization initiative defining and measuring threats to credible AI agent evaluation; demonstrated on τ-Bench Airline that reported agent performance was under-elicited by nearly 50%.
- Co-develop an AI agent reliability index that decomposes agent reliability into twelve metrics and benchmarks 14 frontier models.
- Led an audit study of sycophancy and delusion reinforcement across ChatGPT's chat and API interfaces, surfacing large safety-relevant behavioral gaps.
- Collaborate directly with Arvind Narayanan and Zeynep Tufekci on frontier AI research, and support their widely read public writing on AI, including columns in The New York Times and essays on Substack.
SPAR (Supervised Program for Alignment Research)
January 2026 - May 2026
Research Intern
- Designed and ran an experiment measuring value drift in frontier LLMs by having them iteratively revise their own constitutions, model specs, and system prompts across 20 rounds.
United States Census Bureau (Coding it Forward)
June 2024 - August 2024
Data Scientist Fellow
- Built sorting algorithms and k-nearest neighbors regression models in Python to improve existing imputation methods within the Economic Statistical Methods Division.
- Designed Python data visualizations comparing imputation strategies across several loss functions.
Commonwealth of Massachusetts
November 2021 - July 2023
Data Analyst
- Led inter-agency data sharing and program evaluation efforts on behalf of the Chief Data Officer.
- Designed and maintained PostgreSQL ETL pipelines and automated data workflows with Python, AWS Lambda, and Tableau.
- Implemented the Commonwealth's first differential privacy algorithm on longitudinal datasets related to early childhood education and workforce development.
Education
Princeton School of Public and International Affairs
May 2025
Master of Public Affairs (MPA), Certificate in Statistics and Machine Learning
- Cumulative GPA: 3.94
Williams College
June 2020
Bachelor of Arts with Honors in Political Economy and Philosophy
Selected Publications
Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies
2026
Kirgis, P., Kapoor, S., Schwartz, A., Rabanser, S., Africa, D., ... & Narayanan, A.
arXiv preprint arXiv:2607.27191
- Introduced "shadow evaluations," which task frontier agents with the open-ended research question of an unpublished paper and have its original authors grade the result. Initial results found that agents handled the engineering but could not make substantial research progress.
Log Analysis is Necessary for Credible Evaluation of AI Agents
2026
Kirgis, P., Kapoor, S., Rabanser, S., Nadgir, N., Ududec, C., Dubois, M., ... & Narayanan, A.
Accepted at ICML 2026 FAGEN Workshop
- Argued that systematic log analysis is necessary to overcome validity threats in agent evaluation, presenting a taxonomy of threats and guiding principles. Illustrated on τ-Bench Airline, revealing pass5 performance was under-elicited by nearly 50%.
Towards a Science of AI Agent Reliability
2026
Rabanser, S., Kapoor, S., Kirgis, P., Liu, K., Utpala, S., & Narayanan, A.
Accepted at ICML 2026
- Proposed twelve metrics decomposing AI agent reliability along four dimensions (consistency, robustness, predictability, safety) and evaluated 14 models, finding that recent capability gains have yielded only small improvements in reliability.
LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
2026
Kirgis*, P., Hawriluk*, B., Feng, S., Bilimer, A., Paech, S., & Tufekci, Z. (* Equal contribution)
Accepted at IASEAI 2026
- Conducted an audit study comparing LLM behavior across API and chat interfaces, documenting large differences in sycophancy, escalation, and delusion reinforcement between environments and across models.
Open-World Evaluations for Measuring Frontier AI Capabilities
2026
Kapoor, S., Kirgis, P., Schwartz, A., Rabanser, S., Allaire, JJ., Bommasani, R., ... & Narayanan, A.
Accepted at ICML 2026 AIWILD Workshop
- Advocated for open-world evaluations: long-horizon, real-world tasks assessed through small-sample qualitative analysis.
Life After Benchmark Saturation: A Case Study of CORE-Bench
2026
Nadgir, N., Kapoor, S., Liu, K., Kirgis, P., Orona, M., Rabanser, S., ... & Narayanan, A.
Accepted at ICML 2026 AIWILD Workshop
- Showed that when a benchmark's accuracy saturates, six other dimensions of agent performance remain informative. Measured a ~2× speedup from human-agent collaboration on reproducibility tasks.
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
2025
Kapoor, S., Stroebl, B., Kirgis, P., Nadgir, N., Siegel, Z. S., Wei, B., ... & Narayanan, A.
Accepted at ICLR 2026
- Executed large-scale data analysis of 21,000+ AI agent rollouts to evaluate the capabilities and limitations of frontier AI agents.
Differences in the Moral Foundations of Large Language Models
2025
Kirgis, Peter
arXiv preprint arXiv:2511.11790
- Administered synthetic experiments to sixteen frontier LLMs to elicit moral judgments, using PCA and NLP methods to visualize bias and clustering relative to a human baseline.
Talks & Presentations
"A Few Insights from the Analysis of Over 2,000 AI Agent Logs"
Dec 2025
NeurIPS AI Evaluator Forum
"Is Consciousness Prerequisite for Moral Patienthood?"
Nov 2025
ELEOS AI Consciousness Conference
"Differences in the Moral Foundations of Large Language Models"
Apr 2025
PICSciE/CSML Colloquium