Research

My research focuses on causal inference and missing data methods, with applications to text -data and online discourse. Below are brief overviews of projects I’m currently working on or have recently completed.

The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text

Under Review

with Graham Tierney, Alex Volfovsky

Estimating causal effects of textual properties is challenging because adjustment representations learned from the full text can directly encode the treatment, creating overlap violations even when the underlying causal problem is well-posed. We propose masking-based adjustment representations that remove treatment-defining lexical signals before representation learning, and we formalize when and why this preserves overlap. Across simulations, masking improves overlap, stabilizes effect estimates, and reduces bias compared to methods that learn from unmasked text.

Can Platform Design Encourage Curiosity? Evidence from an Independent Social Media Experiment

Under Review

with Markus Reiter-Haas, Ben Rochford, Max Allamong, Christopher Bail, Sunshine Hillygus, Alexander Volfovsky (Duke Polarization Lab)

Testing interventions to promote prosocial behavior on social media has been difficult because researchers lack control over commercial platform features. We address this with a randomized controlled trial on a custom research platform that uses AI bots to simulate social media dynamics, exposing 2,282 U.S. adults to curiosity-priming interventions through modified norms, interface affordances, or both. Curiosity priming increased question-asking and reduced toxicity without harming user experience, suggesting that platform designs prioritizing curiosity can foster prosocial behavior.

Missing Data with Auxilliary Margins: Categorical Data with Item and Unit Missingness

Working Paper

with Jerry Reiter

Standard multiple imputation methods can produce biased estimates when data are missing not at random, since the missingness mechanism depends on the unobserved values themselves. We develop a Bayesian multiple imputation approach that incorporates known population margins for categorical variables as auxiliary information, helping to anchor imputations even under such non-ignorable missingness. Implemented in JAGS, the method demonstrates reduced bias and improved coverage probability compared to standard MICE procedures across a range of missing data mechanisms.

Prior Research

Below are older projects from my Master’s and undergraduate years at the University of Alabama.


Stochastic Automata Networks and Tensors with Application to Chemical Kinetics

Published; Master's Thesis

with Roger B. Sidje

This work uses tensor representations to address the curse of dimensionality in solving the chemical master equation for biochemical reaction systems, and establishes the differences and similarities between two prominent modeling methods through computational examples and a mathematical proof/

Detecting Small Multi-Set Differences Efficiently

Conference Presentation

with Anh Doan, Daniel Meskill, Yiyao Zhang (at RIPS REU @ UCLA IPAM)

This work develops efficient algorithms—a frequency-based maximum-ID method and a linear algebra-based RREF method—for detecting multi-set overlaps and differences in privacy-centric advertising environments, with theoretical guarantees on catching privacy violations and experimental results highlighting additional use cases.