Research

Working Papers

[1] Machine-Learning-Assisted Empirical Bayes Estimation for Regression on Grouped Network Data [Job Market Paper]
[Paper]

We study the latent group structure in a two-way effect model for matched data represented by a bipartite network. A directed acyclic graph (DAG) encodes the assumed relationships among the observed outcomes, network links, latent effects, group memberships, and prior parameters, thereby clarifying the data-generating process and facilitating estimation. Unobserved heterogeneity is modeled using group-specific prior distributions, allowing individual effects to vary both across and within groups. Although the two-way effects can be estimated by ordinary least squares (OLS), this estimator may be imprecise when the underlying bipartite network is sparse or connected across latent groups. We therefore develop a machine-learning-assisted EB estimator that first learns the latent group structure from the network and then exploits this structure to shrink the estimated effects toward their corresponding group-specific distributions. We establish identification conditions, consistency of prior-parameter estimates, oracle equivalence, and the conditional MSE/bias–variance characterization. The Monte Carlo simulations show that it substantially reduces estimation error relative to OLS under the data-generating processes considered. Finally, in an empirical application to Amazon user–product ratings, we use the proposed method to estimate and characterize heterogeneity in the user- and product-specific components of ratings.

[2] Mixed Membership Estimation in Bayesian Network Autoregression
with Endong Wang , revise & resubmit@The Econometrics Journal
[Paper]

We propose a generative network autoregressive model in which the network coefficient is a random matrix following a joint distribution specified by a Bayesian belief network. Rather than treating the coefficient as fixed and observed, we model it as the output of a latent generative process that simultaneously governs the network structure and the resulting panel data. A central feature is that units may exhibit mixed membership across latent groups rather than belonging to a single group. We develop an estimator that recovers the latent membership structure from panel observations, establish identification and consistency as both panel dimensions grow, and assess finite-sample performance through simulations. An application to the U.S. financial system documents substantial mixed exposure across latent groups, showing that mixed membership materially shapes the dependence structure among financial firms.

[3] Mixed Membership Estimation in Partial Correlation Network
with Endong Wang, under review
[Paper]

Latent group structure is a standard device for modeling cross-sectional dependence in large panels, yet existing partial-correlation community methods typically assign each unit to a single group and do not distinguish the formation of conditional-dependence links from their strength. We propose a two-path mixed-membership partial-correlation network (2PMM-PCN). Mixed membership governs link formation along an extensive margin, while node-specific sociability governs link strength along an intensive margin. Although the two paths have distinct economic interpretations, their node-specific components enter the population signal only through a common scale, allowing the mixed memberships to be recovered by a single spectral procedure. We develop PartialCorr-Mixed-SCORE and establish a convergence rate for the average membership estimation error under weak dependence. In an application to U.S. state employment growth, the estimated memberships reproduce the broad group structure found in earlier studies while revealing economically meaningful overlaps that a single-group assignment cannot represent.

[4] The Horizon Structure of Contagion: High-Dimensional Measurement and Inference
with Abderrahim Taamouti and Endong Wang
[Available upon request]

This paper develops a framework for measuring and conducting inference on horizon-specific contagion networks in high-dimensional restricted predictive systems. We propose a contagion measure that combines statistical evidence of directional predictive transmission with a scale-free measure of its magnitude. For a given source, receiver, and horizon, contagion strength is defined as the log ratio of forecast-error variances with and without the source’s predictive contribution, conditional on the remaining system; edge weights retain this strength only when the corresponding horizon-specific Granger-noncausality hypothesis is rejected. We develop estimation and inference for a broad class of restricted predictive regressions, including heterogeneous autoregressive and grouped-index specifications. Since aggregation and multi-horizon propagation may fail to preserve sparsity, we impose regularity conditions on the grouped predictive state and construct regularized estimators for the high-dimensional regression and covariance objects underlying the relevant forecast-error variances. Applications to large U.S. equity systems document substantial heterogeneity in the frequency and magnitude of predictive transmission across horizons, revealing features of financial connectedness obscured by horizon-agnostic measures.

[5] A New Empirical Bayes Estimation for Network Peer Effect Model
[Paper]

We propose a novel network-based peer effect model with unobserved unit-specific heterogeneity, in which the regression structure is augmented by a probabilistic graph that characterizes the latent dependence mechanism among all components of the model. This probabilistic graph, in turn, provides the foundation for the empirical Bayes estimator developed in the paper. The group structure embedded in the unit heterogeneity arises from a group prior, which we introduce to capture the latent heterogeneity across units. Although a least squares estimator exists for the proposed model, it performs poorly in the presence of sparsity in network. To address these challenges, we develop a machine learning–assisted empirical Bayes estimator. Building on this estimation framework and the proposed group prior, we further introduce a methodology for learning latent group structures in regression models. Simulation results demonstrate that the proposed empirical Bayes estimator substantially outperforms least squares based alternatives. Finally, we apply our framework to an empirical analysis of democratic spillovers across countries.

[6] Overlapping Community Detection in Mixed Membership Vector Autoregression
[Paper]

In this work, we introduce a novel cluster affiliation vector autoregressive model, termed the mixed membership stochastic block vector autoregression (MMSB-VAR(p)). This model assumes that the VAR coefficients at all lags are random matrices following a joint distribution defined by a probabilistic generative framework. The panel data generated by this model can exhibit either single or multiple membership structures with the estimated VAR coefficients recovering the corresponding non-overlapping or overlapping block structures. To uncover the latent group structure, we propose a two-step algorithm leveraging the key insight that spillover effects within groups are significantly stronger than those between groups. We establish the consistency of the proposed algorithm and the simulation study demonstrates its effectiveness. Finally, we apply our methodology to analyze public opinions about the Ukraine-Russia war on Reddit.

Work in Progress

[1] Bi-group Detection in Large Matrix-Variate Factor Model
with Endong Wang

We propose a bi-group matrix factor model that accommodates distinct latent group structures along both the row and column dimensions. The numbers and compositions of the row and column groups are allowed to differ. We develop a suite of singular-value-decomposition-based algorithms to estimate the group structure and recover the corresponding memberships on both sides of the matrix.

[2] Mixture Prior on Random Coefficient Model

Traditional random-coefficient models often assume that random slopes or other sources of parameter heterogeneity are drawn from a single common population distribution. In contrast, we assume that the parameter vector follows a finite mixture distribution, allowing both its location and dispersion to vary across latent groups. We propose a new structural model together with an iterative estimation procedure that jointly estimates the parameters of the mixture prior and the individual parameters of interest.