Thinking of Communication Alignment
A framework for understanding information transmission and bias in human communication, exploring how to minimize loss between speaker intent and receiver understanding...
Read More →I'm a second-year MA student in Computational Social Science at the University of Chicago, advised by James Evans at the Knowledge Lab. I also work closely with Bernard Koch (Sociology and Data Science), and collaborate with Jiaxin Pei (UT Austin) and Xinyu Liang (INSEAD) on AI in science and society.
My research centers on three questions.
Models now sit between people: they write stories, pass information along, and read work on our behalf. I want to understand how that changes what gets created, how it travels, and how it is judged.
Narrative Flattening (EMNLP 2026 Oral) · Event Causal Graph (WNU @ NAACL 2025)
People increasingly hand tasks to models on the assumption that they can handle them. Often they can't, or not in the way people expect. I want to find these gaps, understand why they happen, and improve models where they fall short.
Simulated Ignorance Fails (IJCAI 2026) · Temporal Leakage (ACL 2026)
AI is quickly becoming part of how science and society work. I want to understand what changes when it arrives, whether it is fit to be used there, and whether these communities are ready for it.
Ongoing: LLMs and Scientific Attention (with Bernard Koch) · AI Resume Review (with Jiaxin Pei and Xinyu Liang)
Before UChicago, I earned B.S. degrees in Cognitive Science (Machine Learning and Neural Computation) and Applied Mathematics from UC San Diego, graduating with Highest Distinction and completing the Cognitive Science Honors Program under Zhuowen Tu and Alex Warstadt. I also minored in History and Linguistics.
I'm applying to PhD programs (CS / Information Science) for Fall 2027.
EMNLP 2026, Main Conference Oral
Comparing matched story continuations from four OLMo 32B checkpoints (Base, SFT, DPO, RLVR) against human text, we show that post-training progressively compresses thematic motion, affective intensity, and stylistic diversity in LLM fiction, with professional literary fiction compressed the most.
IJCAI 2026, Main Track
Can prompting a model to "forget" what happened after a cutoff approximate true ignorance? Across 477 competition-level questions and 9 models, we find it cannot: a 52% performance gap remains, chain-of-thought does not suppress prior knowledge, and reasoning-optimized models are worse at it.
ACL 2026, Main Conference
Search-engine date filters are widely used to keep retrieval "before the cutoff," but they leak: auditing Google and DuckDuckGo, we find major post-cutoff leakage for 71% and 81% of questions, which substantially inflates forecasting accuracy.
7th Workshop on Narrative Understanding (WNU @ NAACL 2025)
A hybrid framework that combines LLM-based summarization with linguistically grounded features to generate causal graphs from narratives, outperforming GPT-4o and Claude 3.5 baselines on causal-link precision.
How do large language models "remember" the scientific literature, and how might they reshape it as researchers increasingly rely on them to find work? We find that LLM citation memory is organized by textual co-occurrence (semantic similarity and co-citation) rather than by citation-graph topology or the social processes that generate citations. We are now comparing LLM-generated bibliographies against real human citation behavior.
People increasingly use AI to revise their resumes and personal narratives, while institutions increasingly use AI to screen, audit, and evaluate them. What problems arise when AI sits on both sides of the same decision, and is this a system people can broadly accept? Starting from a single resume as the unit of analysis and scaling up to realistic end-to-end workflows, we examine what part AI actually plays at each stage and where its judgments become unstable in practice.
Translation quality metrics tell us how good a translation is, not what meaning changed. We build a per-dimension semantic atlas across 13 lexical, grammatical, and pragmatic axes (e.g., kinship, clusivity, evidentiality, honorifics) and six LLMs from six model families. We show that translation compresses meaning in one direction and forces new commitments in the other, and that surface-form loss is not the same as semantic loss. We are extending this to track how meaning drifts across multi-step translation chains.
University of Chicago
University of California, San Diego · Minors in History and Linguistics
With Jiaxin Pei (UT Austin) and Xinyu Liang (INSEAD)
University of Chicago · With Junsol Kim and James A. Evans
Knowledge Lab, University of Chicago · Supervised by James A. Evans
University of Chicago · Supervised by Bernard Koch
University of Chicago · Supervised by James A. Evans
UC San Diego · Supervised by Zhiting Hu · Resulted in IJCAI 2026 and ACL 2026 papers
UC San Diego · Supervised by Zhuowen Tu and Alex Warstadt
UC San Diego · Supervised by Zhiting Hu · Resulted in WNU @ NAACL 2025 paper
Deep Learning | UCSD Cognitive Science Department
Under Prof. Zhuowen Tu
Supervised Machine Learning | UCSD Cognitive Science Department
Under Prof. Zhuowen Tu
Calculus and Analytic Geometry | UCSD Math Department
Built an autonomous forecasting agent · $1,500 prize
UC San Diego
UCSD Council of Provosts
UCSD Council of Provosts
A framework for understanding information transmission and bias in human communication, exploring how to minimize loss between speaker intent and receiver understanding...
Read More →An interdisciplinary analysis exploring the intricate relationship between land distribution and military organization in classical civilizations, using the Tang Dynasty's An Lushan Rebellion as a case study...
Read More →zehan@uchicago.edu
lizehan@gmail.com