[1] Machine-Learning-Assisted Empirical Bayes Estimation for Regression on Grouped Network Data
[Job Market Paper]
[Paper]
We study the latent group structure in a two-way effect model for matched data
represented by a bipartite network. A directed acyclic graph (DAG) encodes the assumed relationships among the
observed outcomes, network links, latent effects, group memberships, and prior parameters, thereby clarifying
the data-generating process and facilitating estimation. Unobserved heterogeneity is modeled using group-specific
prior distributions, allowing individual effects to vary both across and within groups. Although the two-way
effects can be estimated by ordinary least squares (OLS), this estimator may be imprecise when the underlying
bipartite network is sparse or connected across latent groups. We therefore develop a machine-learning-assisted
EB estimator that first learns the latent group structure from the network and then exploits this structure to
shrink the estimated effects toward their corresponding group-specific distributions. We establish identification
conditions, consistency of prior-parameter estimates, oracle equivalence, and the conditional MSE/bias–variance
characterization. The Monte Carlo simulations show that it substantially reduces estimation error relative to
OLS under the data-generating processes considered. Finally, in an empirical application to Amazon user–product
ratings, we use the proposed method to estimate and characterize heterogeneity in the user- and product-specific
components of ratings.
[2] Mixed Membership Estimation in Bayesian Network Autoregression
with Endong Wang , revise & resubmit@The Econometrics Journal
[Paper]
We propose a generative network autoregressive model in which the
network coefficient is a random matrix following a joint distribution specified by a Bayesian belief network.
Rather than treating the coefficient as fixed and observed, we model it as the output of a latent generative
process that simultaneously governs the network structure and the resulting panel data. A central feature is that
units may exhibit mixed membership across latent groups rather than belonging to a single group. We develop an estimator
that recovers the latent membership structure from panel observations, establish identification and consistency as both
panel dimensions grow, and assess finite-sample performance through simulations. An application to the U.S. financial
system documents substantial mixed exposure across latent groups, showing that mixed membership materially
shapes the dependence structure among financial firms.
[3] Mixed Membership Estimation in Partial Correlation Network
with Endong Wang, under review
[Paper]
Latent group structure is a standard device for modeling cross-sectional
dependence in large panels, yet existing partial-correlation community methods
typically assign each unit to a single group and do not distinguish the
formation of conditional-dependence links from their strength. We propose a
two-path mixed-membership partial-correlation network (2PMM-PCN). Mixed
membership governs link formation along an extensive margin, while node-specific sociability governs link
strength along an intensive margin.
Although the two paths have distinct economic interpretations, their
node-specific components enter the population signal only through a common
scale, allowing the mixed memberships to be recovered by a single spectral
procedure. We develop PartialCorr-Mixed-SCORE and establish a convergence rate
for the average membership estimation error under weak dependence. In an
application to U.S. state employment growth, the estimated memberships reproduce
the broad group structure found in earlier studies while revealing economically
meaningful overlaps that a single-group assignment cannot represent.
This paper develops a framework for measuring and conducting inference on
horizon-specific contagion networks in high-dimensional restricted predictive systems. We propose a contagion
measure that combines statistical evidence of directional predictive transmission with a scale-free measure of
its magnitude. For a given source, receiver, and horizon, contagion strength is defined as the log ratio of
forecast-error variances with and without the source’s predictive contribution, conditional on the remaining system;
edge weights retain this strength only when the corresponding horizon-specific Granger-noncausality hypothesis is rejected.
We develop estimation and inference for a broad class of restricted predictive regressions, including heterogeneous
autoregressive and grouped-index specifications. Since aggregation and multi-horizon propagation may fail to preserve
sparsity, we impose regularity conditions on the grouped predictive state and construct regularized estimators for the
high-dimensional regression and covariance objects underlying the relevant forecast-error variances. Applications
to large U.S. equity systems document substantial heterogeneity in the frequency and magnitude of predictive
transmission across horizons,
revealing features of financial connectedness obscured by horizon-agnostic measures.
[5] A New Empirical Bayes Estimation for Network Peer Effect Model
[Paper]
We propose a novel network-based peer effect model with unobserved
unit-specific heterogeneity, in which the regression structure is augmented by a probabilistic graph that
characterizes the latent dependence mechanism among all components of the model. This probabilistic graph,
in turn, provides the foundation for the empirical Bayes estimator developed in the paper. The group structure
embedded in the unit heterogeneity arises from a group prior, which we introduce to capture the latent heterogeneity
across units. Although a least squares estimator exists for the proposed model, it performs poorly in the presence
of sparsity in network. To address these challenges, we develop a machine learning–assisted empirical Bayes estimator.
Building on this estimation framework and the proposed group prior, we further introduce a methodology for learning
latent group structures in regression models. Simulation results demonstrate that the proposed empirical Bayes estimator
substantially outperforms least squares based alternatives.
Finally, we apply our framework to an empirical analysis of democratic spillovers across countries.
[6] Overlapping Community Detection in Mixed Membership Vector Autoregression
[Paper]
In this work, we introduce a novel cluster affiliation vector autoregressive model,
termed the mixed membership stochastic block vector autoregression (MMSB-VAR(p)). This model assumes that the VAR
coefficients at all lags are random matrices following a joint distribution defined by a probabilistic generative
framework. The panel data generated by this model can exhibit either single or multiple membership structures with
the estimated VAR coefficients recovering the corresponding non-overlapping or overlapping block structures.
To uncover the latent group structure, we propose a two-step algorithm leveraging the key insight that spillover
effects within groups are significantly stronger than those between groups. We establish the consistency of the
proposed algorithm and the simulation study demonstrates its effectiveness.
Finally, we apply our methodology to analyze public opinions about the Ukraine-Russia war on Reddit.
Work in Progress
[1] Bi-group Detection in Large Matrix-Variate Factor Model
with Endong Wang
We propose a bi-group matrix factor model that accommodates distinct
latent group structures along both the row and column dimensions. The numbers and compositions of the
row and column groups are allowed to differ. We develop a suite of singular-value-decomposition-based
algorithms to estimate the group structure and recover the corresponding memberships on both sides of the matrix.
[2] Mixture Prior on Random Coefficient Model
Traditional random-coefficient models often assume that random slopes or
other sources of parameter heterogeneity are drawn from a single common population distribution.
In contrast, we assume that the parameter vector follows a finite mixture distribution, allowing both
its location and dispersion to vary across latent groups. We propose a new structural model together with
an iterative estimation procedure that jointly estimates the parameters of the mixture prior and the
individual parameters of interest.