Peer-Reviewed Publications
-
Soonhong Cho, Doeun Kim, and Chad Hazlett. "Inference at the Data's Edge: Gaussian Processes for Estimation and Inference in the Face of Extrapolation Uncertainty." Political Analysis (2026) [Journal Page] [arXiv] [R package (Available on CRAN)] [replication]
Abstract
Many inferential tasks involve fitting models to observed data and predicting outcomes at new covariate values, requiring interpolation or extrapolation. Conventional methods select a single best-fitting model, discarding fits that were similarly plausible in-sample but would yield sharply different predictions out-of-sample. Gaussian Processes (GPs) offer a principled alternative. Rather than committing to one conditional expectation function, GPs deliver a posterior distribution over outcomes at any covariate value. This posterior effectively retains the range of models consistent with the data, widening uncertainty intervals where extrapolation magnifies divergence. In this way, the GP's uncertainty estimates reflect the implications of extrapolation on our predictions, helping to tame the ``dangers of extreme counterfactuals'' (King and Zeng 2006). The approach requires (i) specifying a covariance function linking outcome similarity to covariate similarity, and (ii) assuming Gaussian noise around the conditional expectation. We provide an accessible introduction to GPs with emphasis on this property, along with a simple, automated procedure for hyperparameter selection implemented in the R package gpss. We illustrate the value of GPs for capturing counterfactual uncertainty in three settings: (i) treatment effect estimation with poor overlap, (ii) interrupted time series requiring extrapolation beyond pre-intervention data, and (iii) regression discontinuity designs where estimates hinge on boundary behavior.
Under Review
-
"Non-existent Outcomes in Research on Inequality: A Causal Approach." (with Ian Lundberg) [Minor revision, Sociological Methods & Research] [arXiv] [R package]
Abstract
Scholars of social stratification often study exposures that shape life outcomes. But some outcomes (such as wage) only exist for some people (such as those who are employed). We show how a common practice---dropping cases with non-existent outcomes---can obscure causal effects when a treatment affects both outcome existence and outcome values. The effects of both beneficial and harmful treatments can be underestimated. Drawing on existing approaches for principal stratification, we show how to study (1) the average effect on whether an outcome exists and (2) the average effect on the outcome among the latent subgroup whose outcome would exist in either treatment condition. We develop a framework involving parametric regression and simulation to enable principal stratification estimates that adjust for measured confounders, using standard regression-based tools. We illustrate through an empirical example about the effects of parenthood on labor market outcomes. -
"Beyond Measurement-of-Mediation: Aligning Theory, Estimands, and Designs for More Credible Mediation-based Claims in Communication Research" (with Je Hoon Chae and Hyunjin (Jin) Song) [2nd R&R, Human Communication Research]
Abstract
Communication research uses mediation analysis to substantiate claims about causal mechanisms, yet a statistical indirect effect is not, by itself, evidence of process. We argue that credible mechanism claims require alignment among the causal structure implied by theory, the target causal estimand, and the design-based assumptions under which that estimand can be identified. This article advances a metatheoretical framework for establishing and evaluating such alignment. We distinguish causal mechanisms from statistical mediators and explain why the field’s default “measurement-of-mediation” designs often leave mechanism claims underidentified despite treatment randomization. Rather than advocating universal methodological or statistical solutions, we illustrate theory–estimand–design fit through parallel and crossover experimental architectures, showing how added design leverage can narrow the range of causal quantities consistent with observed data under theoretical conditions. Using Coleman et al. (2025) as a proof-of-principle, we discuss scope conditions, trade-offs, and diagnostics for credible causal mechanism evidence in communication research. -
"When Help Seems Optional: Institutional Projection Bias and Climate Refugee Disadvantage." (with Jieun S. Park and Margaret Peters; co-first author with Park)
Abstract
Why do climate refugees receive less public support than war refugees, especially in the developed countries best positioned to assist them? We argue this reflects a cognitive mechanism we term the \textit{institutional projection bias}. Those whose climate risk is buffered by robust disaster management infrastructure use their own experience as a benchmark, reading climate displacement as economically motivated rather than genuinely forced. Three studies test this argument. Cross-national survey evidence shows that the climate refugee disadvantage is concentrated among respondents suspicious of refugee motives, and specifically in countries with strong disaster management capacity. Two original U.S. experiments locate the mechanism in cognitive judgments of displacement legitimacy rather than affective prejudice, with displacement cause, not refugees’ origin, driving evaluations. The institutional projection bias identifies a general pattern in which effective governance inadvertently narrows public support for those who lack equivalent institutions. -
"Penalizing the Perpetual Foreigner: Partisanship, Gender, and Attitudes toward Asian Americans." (with Jieun S. Park and Jessica HyunJeong Lee; co-first author with Park)
Abstract
Americans' favorability toward a racial outgroup considers both affect toward the category itself and inferences about what the label implies, and group-level measures cannot separate the two. Three decades of American National Election Studies data show the aggregate gap in White Americans' ratings of Asian Americans closing while White Republicans and White Democrats diverge sharply. We separate them experimentally in the case of Asian Americans, racialized through two long-coexisting frames, the model minority and the perpetual foreigner. Respondents evaluated short profiles whose race (signaled by name), immigration status, and stereotyped traits are randomized independently, which distinguishes reactions to race itself from reactions to the attributes inferred from it. We find that the penalty against Asian profiles is (i) concentrated among White Republicans rather than being a general feature of the public; (ii) driven by inferences about foreignness rather than resentment of model minority success; and (iii) sharply gendered, falling on Asian male profiles and most heavily among Republican women. Partisan polarization of racial affect thus operates through group-specific stereotype content, here the presumption of national belonging, rather than through diffuse hostility. -
"Let Time Tell: Identification and Gaussian Process Estimation for Interrupted Time Series." [Job Market Paper] [arXiv]
Abstract
We study causal inference in interrupted time series designs where a treatment affects every unit simultaneously, so that the contemporaneous controls used by difference-in-differences and synthetic control are unavailable and the counterfactual must be extrapolated from a unit's own pre-treatment history. We establish identification within the potential outcomes framework and estimate the counterfactual by Gaussian process regression. Rather than committing to a single best-fitting trend, the estimator retains the functions consistent with the pre-treatment series and widens its intervals where extrapolation magnifies their divergence. Connecting it to reproducing kernel Hilbert space theory, we derive a bias decomposition that isolates the component extrapolation inflates and a worst-case bound on that component, justifying the Gaussian process estimator's posterior variance as extrapolation-aware uncertainty quantification. In closed form, the band equals the worst-case divergence the model class permits among functions consistent with the pre-treatment data. The method is illustrated with calibrated simulations and an analysis of handgun purchases after the Supreme Court's Heller decision, a universal treatment whose practical effect concentrates in a single jurisdiction. An R package, gpss, implements the approach.
Working Papers/Work in Progress
- "Covariate Imbalance and Non-Overlap in Interrupted Survey Designs: A Gaussian Process Approach." (with Chad Hazlett)
- "Estimating Time-Varying Treatment Effects under Misspecification." (with Ian Lundberg)
- "A GP Synthetic Control Method." (with Chad Hazlett and Wenlu Xu)