Comparison of Three Clustering Methods for Dissecting Trait Heterogeneity in Genotypic Data

dc.contributor.committeeChairMike McDonald
dc.contributor.committeeMemberThomas J. Palmeri
dc.contributor.committeeMemberJason H. Moore
dc.contributor.committeeMemberConstantin F. Aliferis
dc.contributor.committeeMemberJonathan L. Haines
dc.creatorThornton-Wells, Tricia Ann
dc.date.accessioned2020-08-22T17:33:39Z
dc.date.available2006-07-23
dc.date.issued2005-07-23
dc.description.abstractTrait heterogeneity, which exists when a trait has been defined with insufficient specificity such that it is actually two or more distinct traits, has been implicated as a confounding factor in traditional statistical genetics of complex human disease. In the absence of detailed phenotypic data collected consistently in combination with genetic data, unsupervised computational methodologies offer the potential for discovering underlying trait heterogeneity. The performance of three such methods—Bayesian Classification, Hypergraph-Based Clustering, and Fuzzy k-Modes Clustering—that are appropriate for categorical data were compared. Also tested was the ability of these methods to additionally detect trait heterogeneity in the presence of locus heterogeneity and gene-gene interaction, which are two other complicating factors in discovering genetic models of complex human disease. Bayesian Classification performed well under the simplest of genetic models simulated, and it outperformed the other two methods, with the exception that the Fuzzy k-Modes Clustering performed best on the most complex genetic model. Permutation testing showed that Bayesian Classification controlled Type I error very well but produced less desirable Type II error rates. Methodological limitations and future directions are discussed.
dc.format.mimetypeapplication/pdf
dc.identifier.urihttps://etd.library.vanderbilt.edu/etd-07182005-122343
dc.identifier.urihttp://hdl.handle.net/1803/13157
dc.subjectunsupervised learning
dc.subjectsimulation study
dc.subjectmethod comparison
dc.subjectclustering
dc.subjectcomplex disease
dc.subjectgenetics
dc.subjecttrait heterogenetiy
dc.subjectlocus heterogeneity
dc.titleComparison of Three Clustering Methods for Dissecting Trait Heterogeneity in Genotypic Data
dc.typethesis
dc.type.materialtext
local.embargo.lift2006-07-23
local.embargo.terms2006-07-23
thesis.degree.disciplineBiomedical Informatics
thesis.degree.grantorVanderbilt University
thesis.degree.levelthesis
thesis.degree.nameMS

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
MS_Thesis_17June2005_eversion_edited18July2005.pdf
Size:
599.67 KB
Format:
Adobe Portable Document Format