Rank-Based Analyses and Designs with Clustered Data

dc.contributor.advisorShepherd, Bryan E.
dc.contributor.committeeChairHarrell, Frank E.
dc.creatorTu, Shengxin
dc.creator.orcid0000-0002-3937-8903
dc.date.accessioned2024-08-15T18:20:12Z
dc.date.available2024-08-15T18:20:12Z
dc.date.created2024-08
dc.date.issued2024-05-21
dc.date.submittedAugust 2024
dc.date.updated2024-08-15T18:20:12Z
dc.description.abstractClustered data are common in biomedical research. It is often of interest to evaluate the correlations within clusters and between variables with clustered data. Conventional approaches, including intraclass correlation coefficients (ICCs) and Pearson correlations, are commonly used in analyses with clustered data. However, these conventional approaches are sensitive to extreme values and skewness. They also depend on the scale of the data and are not applicable to ordered categorical data. In this dissertation, we define population parameters for the rank ICC and between- and within-cluster Spearman rank correlations. These definitions are natural extensions of the conventional correlations to the rank scale. We show that the total Spearman rank correlation approximates a weighted sum of between- and within-cluster Spearman rank correlations, with weights determined by the rank ICCs of the two random variables. We also describe estimation and inference for these four rank-based correlations, conduct simulations to evaluate the performance of our estimators, and illustrate their use with real data examples. Furthermore, we apply the rank ICC in the design of clustered randomized controlled trials (RCTs), proposing unified and simple sample size calculations for cluster RCTs with skewed or ordinal outcomes. Our calculation involves inflating the sample size for an adequately powered individual RCT for an ordinal outcome with a design effect that incorporates the rank ICC. For continuous outcomes, our calculation sets the number of distinct ordinal levels to the sample size. We show that with continuous data, our calculations closely approximate more complicated sample size calculations based on clustered Wilcoxon rank-sum tests. We conduct simulations to evaluate our calculations' performance and illustrate their use in the design of two cluster RCTs, one with a skewed continuous outcome and a non-inferiority trial with an irregularly distributed count outcome.
dc.format.mimetypeapplication/pdf
dc.identifier.urihttp://hdl.handle.net/1803/19166
dc.language.isoen
dc.subjectClustered data
dc.subjectRank intraclass correlation
dc.subjectSpearman rank correlation
dc.subjectCluster randomized controlled trial
dc.subjectSample size
dc.titleRank-Based Analyses and Designs with Clustered Data
dc.typeThesis
dc.type.materialtext
thesis.degree.disciplineBiostatistics
thesis.degree.grantorVanderbilt University Graduate School
thesis.degree.levelDoctoral
thesis.degree.namePhD

Files

Original bundle

Now showing 1 - 2 of 2
Loading...
Thumbnail Image
Name:
TU-DISSERTATION-2024.pdf
Size:
7.26 MB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
Vanderbilt University Dissertation Template.zip
Size:
13.4 MB
Format:
Unknown data format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
LICENSE.txt
Size:
1.93 KB
Format:
Plain Text
Description: