Using Statistical Learning Methods for Better Spectrum Classification

dc.contributor.committeeChairMing Li
dc.contributor.committeeMemberChris Fonnesbeck
dc.creatorChen, Yaoyi
dc.date.accessioned2020-08-22T17:24:17Z
dc.date.available2014-07-16
dc.date.issued2014-07-16
dc.description.abstractShotgun proteomics has become a widely used technology for identifying a large number of peptides and proteins in complex biological samples. However, any single score function from most search algorithms to evaluate the quality of peptide-spectrum matches (PSMs) is not adequate to discriminate between correct and incorrect spectrum identification. Here, we used and compared multiple logistic regression models with different flexibilities and support vector machines with various kernel functions and random forests to incorporate multiple scores from search engines. New features, such as retention time differences and a number of other modifications, were also incorporated to build a better binary classifier. We validated these methods through bootstrapping and compared their performance to each other. My study has shown that these methods, with their unique strengths, have improved performance - specifically with higher area under ROC curve and better discrimination indices - to classify correct from incorrect peptide spectrum matches.
dc.format.mimetypeapplication/pdf
dc.identifier.urihttps://etd.library.vanderbilt.edu/etd-07142014-072802
dc.identifier.urihttp://hdl.handle.net/1803/12984
dc.subjectstatistical learning
dc.subjectdata mining
dc.subjectmachine learning
dc.titleUsing Statistical Learning Methods for Better Spectrum Classification
dc.typethesis
dc.type.materialtext
local.embargo.lift2014-07-16
local.embargo.terms2014-07-16
thesis.degree.disciplineBiostatistics
thesis.degree.grantorVanderbilt University
thesis.degree.levelthesis
thesis.degree.nameMS

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Chen.pdf
Size:
3.01 MB
Format:
Adobe Portable Document Format