6
views
0
recommends
+1 Recommend
0 collections
    0
    shares
      • Record: found
      • Abstract: found
      • Article: not found

      Estimating the number of protein folds and families from complete genome data.

      Journal of Molecular Biology
      Conserved Sequence, Databases, Factual, Genome, Genome, Archaeal, Genome, Bacterial, Genome, Fungal, Protein Folding, Protein Structure, Tertiary, Proteins, chemistry, classification, metabolism, Sampling Studies, Solubility, Statistical Distributions, Water

      Read this article at

      ScienceOpenPublisherPubMed
      Bookmark
          There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

          Abstract

          Using the data on proteins encoded in complete genomes, combined with a rigorous theory of the sampling process, we estimate the total number of protein folds and families, as well as the number of folds and families in each genome. The total number of folds in globular, water- soluble proteins is estimated at about 1000, with structural information currently available for about one-third of the number. The sequenced genomes of unicellular organisms encode from approximately 25%, for the minimal genomes of the Mycoplasmas, to 70-80% for larger genomes, such as Escherichia coli and yeast, of the total number of folds. The number of protein families with significant sequence conservation was estimated to be between 4000 and 7000, with structures available for about 20% of these. Copyright 2000 Academic Press.

          Related collections

          Author and article information

          Comments

          Comment on this article