Essays on Big Data and High-Dimensional Models: Machine Learning Approaches in Econometrics

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

The increasing availability of high-dimensional and complex economic data presents both opportunities and challenges for econometric analysis. This dissertation, Essays on Big Data and High-Dimensional Models: Machine Learning Approaches in Econometrics, explores novel statistical and computational techniques for high-dimensional inference in threshold regression models and efficient data processing under multiway clustering. The three chapters collectively contribute to the advancement of econometric theory and practice by addressing critical methodological challenges in big data econometrics. The first chapter develops a uniform inference theory for high-dimensional slope parameters in threshold regression models, allowing for either cross-sectional or time series data. It establishes oracle inequalities for Lasso estimators and introduces a debiased Lasso estimator for threshold models. The results allow researchers to perform uniform inference without specifying whether the model is a linear or threshold regression. The empirical applications to economic growth and the effect of military news shocks on U.S. government spending demonstrate the practical importance of this method in economic analysis. The second chapter focuses on inference for the threshold parameter in high-dimensional threshold regression models. It derives the asymptotic distribution of the threshold parameter under different specifications, including kink, diminishing jump, and fixed jump models. The chapter rigorously proves the continuity of these asymptotic distributions and validates the use of subsampling for inference. The methods are applied to empirical settings where capturing nonlinearities and regime shifts is essential for understanding economic dynamics and informing policy decisions. The third chapter proposes a novel method of algorithmic subsampling (data sketching) for multiway cluster-dependent data. It develops a new uniform weak law of large numbers and a central limit theorem for multiway algorithmic subsample means, demonstrating that algorithmic subsampling ensures robustness against potential degeneracy, and even non-Gaussian degeneracy, of the asymptotic distribution under multiway clustering at the cost of efficiency and power loss due to algorithmic subsampling. An empirical application using scanner data from Dominick’s Finer Foods illustrates the method’s effectiveness in demand estimation for differentiated product markets. Collectively, this dissertation advances the intersection of big data econometrics and machine learning by developing theoretically rigorous methods for inference in high-dimensional threshold models and efficient data processing under multiway clustering. The findings have broad applications in economic policy and market analysis where large datasets and complex structural relationships are prevalent.

Description

Keywords

Big Data Econometrics, High-Dimensional Models, Machine Learning Methods

Citation

Endorsement

Review

Supplemented By

Referenced By