Essays on Big Data and High-Dimensional Models: Machine Learning Approaches in Econometrics

dc.contributor.advisorSasaki, Yuya
dc.contributor.committeeChairSasaki, Yuya
dc.creatorLi, Jiatong
dc.creator.orcid0009-0005-1013-8307
dc.date.accessioned2025-06-05T12:46:19Z
dc.date.available2025-06-05T12:46:19Z
dc.date.created2025-05
dc.date.issued2025-03-22
dc.date.submittedMay 2025
dc.date.updated2025-06-05T12:46:19Z
dc.description.abstractThe increasing availability of high-dimensional and complex economic data presents both opportunities and challenges for econometric analysis. This dissertation, Essays on Big Data and High-Dimensional Models: Machine Learning Approaches in Econometrics, explores novel statistical and computational techniques for high-dimensional inference in threshold regression models and efficient data processing under multiway clustering. The three chapters collectively contribute to the advancement of econometric theory and practice by addressing critical methodological challenges in big data econometrics. The first chapter develops a uniform inference theory for high-dimensional slope parameters in threshold regression models, allowing for either cross-sectional or time series data. It establishes oracle inequalities for Lasso estimators and introduces a debiased Lasso estimator for threshold models. The results allow researchers to perform uniform inference without specifying whether the model is a linear or threshold regression. The empirical applications to economic growth and the effect of military news shocks on U.S. government spending demonstrate the practical importance of this method in economic analysis. The second chapter focuses on inference for the threshold parameter in high-dimensional threshold regression models. It derives the asymptotic distribution of the threshold parameter under different specifications, including kink, diminishing jump, and fixed jump models. The chapter rigorously proves the continuity of these asymptotic distributions and validates the use of subsampling for inference. The methods are applied to empirical settings where capturing nonlinearities and regime shifts is essential for understanding economic dynamics and informing policy decisions. The third chapter proposes a novel method of algorithmic subsampling (data sketching) for multiway cluster-dependent data. It develops a new uniform weak law of large numbers and a central limit theorem for multiway algorithmic subsample means, demonstrating that algorithmic subsampling ensures robustness against potential degeneracy, and even non-Gaussian degeneracy, of the asymptotic distribution under multiway clustering at the cost of efficiency and power loss due to algorithmic subsampling. An empirical application using scanner data from Dominick’s Finer Foods illustrates the method’s effectiveness in demand estimation for differentiated product markets. Collectively, this dissertation advances the intersection of big data econometrics and machine learning by developing theoretically rigorous methods for inference in high-dimensional threshold models and efficient data processing under multiway clustering. The findings have broad applications in economic policy and market analysis where large datasets and complex structural relationships are prevalent.
dc.format.mimetypeapplication/pdf
dc.identifier.urihttps://hdl.handle.net/1803/19640
dc.language.isoen
dc.subjectBig Data Econometrics, High-Dimensional Models, Machine Learning Methods
dc.titleEssays on Big Data and High-Dimensional Models: Machine Learning Approaches in Econometrics
dc.typeThesis
dc.type.materialtext
thesis.degree.disciplineEconomics
thesis.degree.grantorVanderbilt University Graduate School
thesis.degree.levelDoctoral
thesis.degree.namePhD

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
LI-DISSERTATION-2025.pdf
Size:
4.07 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
LICENSE.txt
Size:
1.92 KB
Format:
Plain Text
Description: