Prediction whether a human cDNA sequence contains initiation codon by combining statistical information and similarity with protein sequences
Citation
T. Nishikawa et al., Prediction whether a human cDNA sequence contains initiation codon by combining statistical information and similarity with protein sequences, BIOINFORMAT, 16(11), 2000, pp. 960-967
Categorie Soggetti
Multidisciplinary
Journal title
BIOINFORMATICS
SICI code
1367-4803(200011)16:11<960:PWAHCS>2.0.ZU;2-T
Abstract
Motivation: In the previous works, we developed ATGpr, a computer program f
or predicting the fullness of a cDNA, i.e, whether it contains an initiatio
n codon or not. Statistical information of short nucleotide fragments was f
ully exploited in the prediction algorithm. However sequence similarities t
o known proteins, which are becoming increasingly available due to recent r
apid growth of protein database, were not used in the prediction. In this w
ork, we present a new prediction algorithm based on both statistical and si
milarity information, which provides better performance in sensitivity and
specificity.
Results: We evaluated the accuracy of ATGpr for predicting fullness of cDNA
sequences from human clustered ESTs of UniGene, and we obtained specificit
y sensitivity, and correlation coefficient of this prediction. Specificity
and sensitivity crossed at 46% over the ATGpr score threshold of 0.33 and t
he maximum correlation coefficient of 0.34 was obtained at this threshold.
Without ATGpr we found it effective to use alignments with known proteins f
or predicting the fullness of cDNA sequences. That is, specificity increase
d monotonously as similarity (identity of the alignments) increased. Specif
icity was achieved greater than 80% if identity was greater than 40%. For m
ore effective prediction of fullness of cDNA sequences we combined the simi
larity (identity of query sequence) with known proteins and ATGpr score. As
a result, specificity became greater than 80% if identity was greater than
20%.