Prediction whether a human cDNA sequence contains initiation codon by combining statistical information and similarity with protein sequences

Citation
T. Nishikawa et al., Prediction whether a human cDNA sequence contains initiation codon by combining statistical information and similarity with protein sequences, BIOINFORMAT, 16(11), 2000, pp. 960-967
Citations number
6
Categorie Soggetti
Multidisciplinary
Journal title
BIOINFORMATICS
ISSN journal
13674803 → ACNP
Volume
16
Issue
11
Year of publication
2000
Pages
960 - 967
Database
ISI
SICI code
1367-4803(200011)16:11<960:PWAHCS>2.0.ZU;2-T
Abstract
Motivation: In the previous works, we developed ATGpr, a computer program f or predicting the fullness of a cDNA, i.e, whether it contains an initiatio n codon or not. Statistical information of short nucleotide fragments was f ully exploited in the prediction algorithm. However sequence similarities t o known proteins, which are becoming increasingly available due to recent r apid growth of protein database, were not used in the prediction. In this w ork, we present a new prediction algorithm based on both statistical and si milarity information, which provides better performance in sensitivity and specificity. Results: We evaluated the accuracy of ATGpr for predicting fullness of cDNA sequences from human clustered ESTs of UniGene, and we obtained specificit y sensitivity, and correlation coefficient of this prediction. Specificity and sensitivity crossed at 46% over the ATGpr score threshold of 0.33 and t he maximum correlation coefficient of 0.34 was obtained at this threshold. Without ATGpr we found it effective to use alignments with known proteins f or predicting the fullness of cDNA sequences. That is, specificity increase d monotonously as similarity (identity of the alignments) increased. Specif icity was achieved greater than 80% if identity was greater than 40%. For m ore effective prediction of fullness of cDNA sequences we combined the simi larity (identity of query sequence) with known proteins and ATGpr score. As a result, specificity became greater than 80% if identity was greater than 20%.