Zehan Li

MA Student, Computational Social Science · Knowledge Lab, UChicago

Zehan Li

About Me

I'm a second-year MA student in Computational Social Science at the University of Chicago, advised by James Evans at the Knowledge Lab. I also work closely with Bernard Koch (Sociology and Data Science), and collaborate with Jiaxin Pei (UT Austin) and Xinyu Liang (INSEAD) on AI in science and society.

My research centers on three questions.

1. How do language models shape narrative, culture, and creativity?

Models now sit between people: they write stories, pass information along, and read work on our behalf. I want to understand how that changes what gets created, how it travels, and how it is judged.

Narrative Flattening (EMNLP 2026 Oral) · Event Causal Graph (WNU @ NAACL 2025)

2. What can't models do that people assume they can?

People increasingly hand tasks to models on the assumption that they can handle them. Often they can't, or not in the way people expect. I want to find these gaps, understand why they happen, and improve models where they fall short.

Simulated Ignorance Fails (IJCAI 2026) · Temporal Leakage (ACL 2026)

3. What happens when AI enters science and society?

AI is quickly becoming part of how science and society work. I want to understand what changes when it arrives, whether it is fit to be used there, and whether these communities are ready for it.

Ongoing: LLMs and Scientific Attention (with Bernard Koch) · AI Resume Review (with Jiaxin Pei and Xinyu Liang)

Before UChicago, I earned B.S. degrees in Cognitive Science (Machine Learning and Neural Computation) and Applied Mathematics from UC San Diego, graduating with Highest Distinction and completing the Cognitive Science Honors Program under Zhuowen Tu and Alex Warstadt. I also minored in History and Linguistics.

I'm applying to PhD programs (CS / Information Science) for Fall 2027.

News

  • Sep 2026 Narrative Flattening selected for an oral presentation at EMNLP 2026!
  • Aug 2026 Narrative Flattening accepted to EMNLP 2026 (Main Conference).
  • Apr 2026 Simulated Ignorance Fails accepted to IJCAI 2026 (Main Track)!
  • Apr 2026 Temporal Leakage in Search-Engine Date-Filtered Web Retrieval accepted to ACL 2026 (Main Conference).
  • Dec 2025 Our autonomous forecasting agent placed 15th globally in the Metaculus AI Forecasting Bot Tournament (Fall 2025).
  • Sep 2025 Started the MA in Computational Social Science at UChicago and joined the Knowledge Lab.

Publications

Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction

Zehan Li, Yutong Zhu, Siyang Wu, Honglin Bao, James A. Evans

EMNLP 2026, Main Conference Oral

Comparing matched story continuations from four OLMo 32B checkpoints (Base, SFT, DPO, RLVR) against human text, we show that post-training progressively compresses thematic motion, affective intensity, and stylistic diversity in LLM fiction, with professional literary fiction compressed the most.

Simulated Ignorance Fails: A Systematic Study of LLM Behaviors on Forecasting Problems Before Model Knowledge Cutoff

Zehan Li, Yuxuan Wang, Ali El Lahib, Ying-Jieh Xia, Xinyu Pi

IJCAI 2026, Main Track

Can prompting a model to "forget" what happened after a cutoff approximate true ignorance? Across 477 competition-level questions and 9 models, we find it cannot: a 52% performance gap remains, chain-of-thought does not suppress prior knowledge, and reasoning-optimized models are worse at it.

Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study

Ali El Lahib, Ying-Jieh Xia, Zehan Li, Yuxuan Wang, Xinyu Pi

ACL 2026, Main Conference

Search-engine date filters are widely used to keep retrieval "before the cutoff," but they leak: auditing Google and DuckDuckGo, we find major post-cutoff leakage for 71% and 81% of questions, which substantially inflates forecasting accuracy.

Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts

Zehan Li, Ruhua Pan, Xinyu Pi

7th Workshop on Narrative Understanding (WNU @ NAACL 2025)

A hybrid framework that combines LLM-based summarization with linguistically grounded features to generate causal graphs from narratives, outperforming GPT-4o and Claude 3.5 baselines on causal-link precision.

Ongoing Projects

LLMs as Intermediaries: How They Shape Real-World Discovery and Evaluation

LLMs and the Structure of Scientific Attention (working title)

How do large language models "remember" the scientific literature, and how might they reshape it as researchers increasingly rely on them to find work? We find that LLM citation memory is organized by textual co-occurrence (semantic similarity and co-citation) rather than by citation-graph topology or the social processes that generate citations. We are now comparing LLM-generated bibliographies against real human citation behavior.

with Bernard Koch and Sherry Ding In progress · Preprint forthcoming

AI Resume Review Instability (working title)

People increasingly use AI to revise their resumes and personal narratives, while institutions increasingly use AI to screen, audit, and evaluate them. What problems arise when AI sits on both sides of the same decision, and is this a system people can broadly accept? Starting from a single resume as the unit of analysis and scaling up to realistic end-to-end workflows, we examine what part AI actually plays at each stage and where its judgments become unstable in practice.

with Jiaxin Pei and Xinyu Liang In progress · Preprint forthcoming

Language, Meaning, and Model Behavior

Semantic Delta: What Translation Loses, Adds, and Forces (working title)

Translation quality metrics tell us how good a translation is, not what meaning changed. We build a per-dimension semantic atlas across 13 lexical, grammatical, and pragmatic axes (e.g., kinship, clusivity, evidentiality, honorifics) and six LLMs from six model families. We show that translation compresses meaning in one direction and forces new commitments in the other, and that surface-form loss is not the same as semantic loss. We are extending this to track how meaning drifts across multi-step translation chains.

with Junsol Kim and James Evans In progress · Preprint forthcoming

Curriculum Vitae

Education

2025 – present

M.A. in Computational Social Science

University of Chicago

2021 – 2025

B.S. in Cognitive Science (Machine Learning), with Highest Distinction

B.S. in Applied Mathematics

University of California, San Diego · Minors in History and Linguistics

Research Experience

06/2026 – present

Project Lead · AI-Mediated Resume Review

With Jiaxin Pei (UT Austin) and Xinyu Liang (INSEAD)

06/2026 – present

Project Lead · Semantic Delta: Meaning Change in LLM Translation

University of Chicago · With Junsol Kim and James A. Evans

01/2026 – 06/2026

Project Lead · Narrative Flattening in LLM Fiction

Knowledge Lab, University of Chicago · Supervised by James A. Evans

10/2025 – present

Researcher · LLMs and Inequality in Scientific Attention

University of Chicago · Supervised by Bernard Koch

09/2025 – 06/2026

Research Assistant · Automating Mechanistic Interpretability Reports

University of Chicago · Supervised by James A. Evans

07/2025 – 01/2026

Project Lead · Simulated Ignorance in LLM Forecasting Evaluation

UC San Diego · Supervised by Zhiting Hu · Resulted in IJCAI 2026 and ACL 2026 papers

08/2024 – 06/2025

Undergraduate Honors Thesis · Quantifying Conceptual Density and Vacuity in Text

UC San Diego · Supervised by Zhuowen Tu and Alex Warstadt

02/2024 – 09/2024

Project Lead · Linguistic Causal Graph Generation from Narratives

UC San Diego · Supervised by Zhiting Hu · Resulted in WNU @ NAACL 2025 paper

Teaching

01/2025 – 04/2025

Instructional Assistant

Deep Learning | UCSD Cognitive Science Department

Under Prof. Zhuowen Tu

09/2024 – 12/2024

Instructional Assistant

Supervised Machine Learning | UCSD Cognitive Science Department

Under Prof. Zhuowen Tu

01/2023 – 04/2023

Supplemental Instructor (SI)

Calculus and Analytic Geometry | UCSD Math Department

Awards & Honors

Fall 2025

15th Place Globally, Metaculus AI Forecasting Bot Tournament

Built an autonomous forecasting agent · $1,500 prize

2021 – 2025

Provost Honors (6 terms)

UC San Diego

Spring 2024

Triton Research & Experiential Learning Scholars Award

UCSD Council of Provosts

Summer 2023

Triton Research & Experiential Learning Scholars Award

UCSD Council of Provosts

Blog & Thoughts

January 2025

Thinking of Communication Alignment

A framework for understanding information transmission and bias in human communication, exploring how to minimize loss between speaker intent and receiver understanding...

Read More →
January 2025

Land Distribution and Military Organization: Lessons from the An Lushan Rebellion

An interdisciplinary analysis exploring the intricate relationship between land distribution and military organization in classical civilizations, using the Tang Dynasty's An Lushan Rebellion as a case study...

Read More →

Contact

Academic Email

zehan@uchicago.edu

Personal Email

lizehan@gmail.com