Welcome! For Biotech stuff click on "4 Biotech" in the Top Menu. Similarly for stuff related to Biomedial, B.Pharmacy and Other branches click on "4 Biomedical", "4 B.Pharmacy", "4 Other Branches" respectively.

Earn Money from Mobile

Earn Money from Mobile
Click on the Image to Register in mGinger
Showing posts with label COMPUTATIONAL MOLECULAR BIOLOGY (CMB). Show all posts
Showing posts with label COMPUTATIONAL MOLECULAR BIOLOGY (CMB). Show all posts

COMPUTATIONAL MOLECULAR BIOLOGY Question Papers (2008,Reg)

Posted by m.s.chowdary at 7:42 AM

Saturday, November 22, 2008

SET : 1

  1. Ouline the steps in the BLAST algorithm.
  2. a) Eplain why neither FASTA nor BLAST are guaranteed to return a sequence that has maximum similarity score with respect to the query sequence. b) Explain why hueristic programs such as FASTA or BLASTA are used for database searching instead of comparing the query sequence with the database sequences using dynamic programming.
  3. a) What are the differences and similarities between cladograms, rooted and, unrooted phylograms. what information each tree conveys. b) Describe the 3 methods for constructing phylogenies: Maximum likelyhood, Neighbour Joining, and Maximum Parsimony.
  4. Explain in detail about the symmetric model of DNA evolution.
  5. What are the advantages of phylogenetic analysis .
  6. Explain why every genome is different.
  7. How many chain conformations are described by Ramachandran Plot?
  8. What are the difficulties faced from ab initio protein structure prediction. How can we solve it.
SET : 2
  1. a) What are the major extensions of BLAST? b) Discuss the areas of application of these programs.
  2. Name some situations in which a molecular biologist wishes to perform a pair wise sequence alignment, multiple sequence alignment and sequence database searching. b) Briefly describe how the PAM and BLOSUM scoring matrices are derived and how they are different.
  3. Write short jnotes on : a) Searching tree space b) Distance correction.
  4. a) What are DNA substitution models? Expalin. b) What are DNA substitution models good for? Discuss.
  5. What are the methods used to studying gene expression?
  6. Describe the comparitive genomics of bacteria.
  7. Protein structures are more highly conserved than sequences than sequences. Explain.
  8. Transmembrane proteins are important in drug discovery. What are their properties and how can we generate their 3-D structures.
SET : 3
  1. Explain the steps used by the BLAST algorithm and mention blast related programs.
  2. a) Write about the significance of substitution scores and gap penalties in sequence alignment. b) Explain FASTA database similarity searching program.
  3. Write short notes on : a) Rooted and Unrooted trees. b) Genes vs. species trees.
  4. Explain about the strategy employed by Kimuras Two Parameter model in estimating substitution numbers.
  5. Give a note on variety of polymorphic DNA markers and how they can be used for linkage studies.
  6. Describe the organisation of nuclear DNA in eukaryotes.
  7. Comment on the evolution of PDB. What are its classifications.
  8. Explain the importance of protein designing in the field of drug designing.
SET : 4
  1. Describe the following : a) General idea behind the PSI-BLAST approach. b) What problem does it solve (Input/Output)? HOw does it go about solving it?
  2. What type of scoring matrix is used by FASTA? Explain about different substitution matrices?
  3. Describe the following : a) Various tree building methods. b) Concept of Evolutionary trees.
  4. Explain a) What is a change point, and how could you fetect it? b) What is the importance of the KA/KS ratio?
  5. Write a short notes on : a) RFLP b) VNTRs c) STS d) EST
  6. Discuss the application of microarray technology.
  7. Write short notes : a) Protein sequence data b) Sequence data base search.
  8. Describe the diffferent databases available for storage of protein resources.

PROTEIN THREADING

Posted by m.s.chowdary at 1:48 AM

Tuesday, November 4, 2008

Protein threading also referred to as Fold Recognition is used in the prediction of protein structures from aminoacid sequence.
Target sequence (the protein structure for which the stucture is being predicted) is threaded into the backbone structures of a collection of template protiens. The collection of template proteins is referred to as Fold Library.
Then a 'Goodness of Fit' score is calculated for each structure sequence alignment. Goodness of Fit is often derived in terms of an empirical energy function, based on statistics derived from known protein structures . Many other scoring function have been proposed and the most useful among them are Pairwise terms (interaction between pairs of aminoacids) and solvation terms.
Threading methods share some of the characteristics of both homology modelling and ab initio prediction methods.

Fold recognition methods can be broadly divided into 2 types:

  • Methods that derive a one-dimensional (I-D) profile for each of the protein structure in the fold library and align target sequence to these profiles.
  • Methods that consider the full three dimensional structure of the protein template.
A simple way of profile representation would be to take each aminoacid in the structure and simply label it according to whether it is buried in the core structure of the protein or exposed on
the surface. More eloborate profiles might take into account the local secondary structure (ex: whether the aminoacid is part of an alpha helix) and/or evolutionary information (ex: how conserved the amino acid is).
In the three dimensional representation, the structure is modelled as a set of interatomic distances i.e. distances are calculated betweem some or all the atom pairs in the structure. This is a much better and flexible description of the structure, but is is much harder to use in calculating an alignment.
The profile based fold recognition aproach was first described by Bowie, Luthy and Eissenberg in 1991.
The term Threading was first coined by Jones, Tailor and Thornton in 1992, and originally referred specifically to the use of a full three dimensional structure atomic representation of the protein template in fold recognition. Today the two terms are frequently used interchangeably.
Fold recognition methods are more frequently used effective because it is believed that there are a strictly limited number of different protein folds in nature, mostly as a result of evolution and also due to constraints imposed by the physics and chemistry of polypeptide chains. there is, therefore a good chance that a protein which has a similar fold to the target protein has already been studied by X-ray cristallography/NMR spectroscopy and can be found in the ProteinDataBank (PDB). Currently there are just about 1100 different protein folds known and , but new follds are still being discovered every year.
Many different algorithms have been proposed for finding the correct threading of a sequence onto a structure though many make use of Dynamic Programming in some form. For full threedimensional threading, the problem of identifiying the best alignment is very difficult and researchers have made use of many combinatorial optimization methods such as simulated annealing or branch and bound searching to arrive at heuristic solutions.

RAPTOR is an integer programming based protein threading software

STRUCTURAL CLASSIFICATION OF PROTEINS (SCOP)

Posted by m.s.chowdary at 7:15 AM

Monday, November 3, 2008

The complexities of tertiary structure of proteins is decreased by considering substructures. Taking this idea forward researchers have organised complete content of the databases according to hierarchical levels of structure. SCOP, CATH & FSSP are the databases that were developed keeping this in view.

SCOP: Structural Classification Of Proteins

SCOP is a hierarchical classificatiov of proteins. I was first published in 1995 and is usually yearly updated by Alexei G. Murzin and his colleagues upon whose expertise the classification rests.

SCOP is mostly manually curated unlike CATH and FSSP, which make use of automated methods.

Hierarchical Structure:

SCOP has the following Hierarchical levels:

  • Class
  • Fold
  • Superfamily
  • Family
The top two levels of organization, Class and Fold are purely structural. Below the fold level, categorization is based on eolutionary relationships.

Class
Class is the top most level in the hierarchical classification. All the protein structures are divided into 4 classes. They are:
  1. All α
  2. All β
  3. α/β (α and β segments are either interspersed or present alternatively)
  4. α + β (α and β segments are sggregated)
Fold
With in every class there are ten to hundreds of folds. Similar arrangement of regular secondary structures is observed.

Family
Proteins with significant primary sequence similarity and demonstrable structural similarity are grouped into one family. Protens within a family have significant structural realtionships.
Proteins within a single family may found in all the 3 domains of life, suggesting a very ancient origin; and some times found only within a group of organisms, suggesting a recent origin.

Superfamily
Protein families that share a little similarity in primary sequence and make use of same major structural motif and have functional similarities are grouped as superfamilies.

SCOP is a power tool to sequence analysis in tracing many evolutionary relationships.


PROTEIN HOMOLOGY MODELING

Posted by m.s.chowdary at 12:56 AM

Homology modeling is also known as comparative modeling. It is a class of methods in protein structure prediction for constructing an atomic-resolution model of a protein from its amino acid sequence (the "query sequence" or "target").

Most of the homology modeling techniques rely on :

  • Identification of one or more known protein structures likely to resemble the structure of the query sequence &
  • Production of an alignment that maps residues in the query sequence to residues in the template sequence.

The sequence alignment and template structure are then used to produce a structural model of the target.

Because protein structures are more conserved than DNA sequences, detectable levels of sequence similarity usually imply significant structural similarity.

The quality of the homology model is dependent on the quality of the sequence alignment and template structure. The approach can be complicated by the presence of alignment gaps (commonly called indels) that indicate a structural region present in the target but not in the template, and by structure gaps in the template that arise from poor resolution in the experimental procedure (like X-ray crystallography) used to solve the structure. Model quality declines with decreasing sequence identity. Regions of the model that were constructed without a template, usually by loop modeling, are generally much less accurate than the rest of the model, particularly if the loop is long. Errors in side chain packing and position also increase with decreasing identity, and variations in these packing configurations have been suggested as a major reason for poor model quality at low identity. Taken together, these various atomic-position errors are significant and impede the use of homology models for purposes that require atomic-resolution data, such as drug design and protein-protein interaction predictions; even the quaternary structure of a protein may be difficult to predict from homology models of its subunit(s). However, homology models can be useful in reaching qualitative conclusions about the biochemistry of the query sequence, especially in formulating hypotheses about why certain residues are conserved, which may in turn lead to experiments to test those hypotheses. For example, the spatial arrangement of conserved residues may suggest whether a particular residue is conserved to stabilize the folding, to participate in binding some small molecule, or to foster association with another protein or nucleic acid.

Homology modeling can produce high-quality structural models when the target and template are closely related. The chief inaccuracies in homology modeling, which worsen with lower sequence identity, derive from errors in the initial sequence alignment and from improper template selection. Current practice in homology modeling is assessed in a biannual large-scale experiment known as the Critical Assessment of Techniques for Protein Structure Prediction( CASP).

The method of homology modeling is based on the observation that protein tertiary structure is better conserved than amino acid sequence. So, even proteins that have diverged appreciably in sequence but still share detectable similarity will also share common structural properties, particularly the overall fold. Because it is difficult and time-consuming to obtain experimental structures from methods such as X-ray crystallography and protein NMR for every protein of interest, homology modeling can provide useful structural models for generating hypotheses about a protein's function and directing further experimental work.

There are exceptions to the general rule that proteins sharing significant sequence identity will share a fold.

Two proteins may share a similar fold even if their evolutionary relationship is so distant that it cannot be discerned reliably.

The function of a protein is conserved much less than the protein sequence, since relatively few changes in amino-acid sequence are required to take on a related function.

Steps in Model Production

The homology modeling procedure can be broken down into four sequential steps:

  • Template selection
  • Target-template alignment
  • Model construction &
  • Model assessment

Template selection and sequence alignment

Template selection and sequence alignment are often essentially performed together . The critical first step in homology modeling is the identification of the best template structure. The simplest method of template identification relies on serial pairwise sequence alignments aided by database search techniques such as FASTA and BLAST. More sensitive methods based on multiple sequence alignment (of which PSI-BLAST is the most common example). This family of methods has been shown to produce a larger number of potential templates and to identify better templates for sequences that have only distant relationships to any solved structure. Protein threading is also used as a search technique for identifying templates to be used in homology modeling methods. When performing a BLAST search, a reliable first approach is to identify hits with a sufficiently low E-value, which are considered sufficiently close in evolution to make a reliable homology model.Often several candidate template structures are identified by these approaches. Choosing the best template from among the candidates is a key step, and can affect the final accuracy of the structure significantly. This choice is guided by several factors, such as the similarity of the query and template sequences, of their functions, and of the predicted query and observed template secondary structures. Perhaps most importantly, the coverage of the aligned regions: the fraction of the query sequence structure that can be predicted from the template, and the plausibility of the resulting model. Thus, sometimes several homology models are produced for a single query sequence, with the most likely candidate chosen only in the final step.It is possible to use the sequence alignment generated by the database search technique as the basis for the subsequent model production; however, more sophisticated approaches have also been explored. One proposal generates an ensemble of stochastically defined pairwise alignments between the target sequence and a single identified template as a means of exploring "alignment space" in regions of sequence with low local similarity. "Profile-profile" alignments that first generate a sequence profile of the target and systematically compare it to the sequence profiles of solved structures; the coarse-graining inherent in the profile construction is thought to reduce noise introduced by sequence drift in nonessential regions of the sequence.

Model generation

Given a template and an alignment, the information contained therein must be used to generate a three-dimensional structural model of the target. Three major classes of model generation methods have been proposed.

Fragment assembly

The original method of homology modeling relied on the assembly of a complete model from conserved structural fragments identified in closely related solved structures.

Segment matching

The segment-matching method divides the target into a series of short segments, each of which is matched to its own template fitted from the PDB. Thus, sequence alignment is done over segments rather than over the entire protein. Selection of the template for each segment is based on sequence similarity, comparisons of alpha carbon coordinates, and predicted steric conflicts arising from the van der Waals radii of the divergent atoms between target and template.

Satisfaction of spatial restraints

The most common current homology modeling method takes its inspiration from calculations required to construct a three-dimensional structure from data generated by NMR spectroscopy. One or more target-template alignments are used to construct a set of geometrical criteria that are then converted to probability density functions for each restraint. Restraints applied to the main protein internal coordinates - protein backbone distances and dihedral angles - serve as the basis for a global optimization procedure that originally used conjugate gradient energy minimization to iteratively refine the positions of all heavy atoms in the protein.

This method had been dramatically expanded to apply specifically to loop modeling, which can be extremely difficult due to the high flexibility of loops in proteins in aqueous solution. A more recent expansion applies the spatial-restraint model to electron density maps derived from cryoelectron microscopy studies, which provide low-resolution information that is not usually itself sufficient to generate atomic-resolution structural models. To address the problem of inaccuracies in initial target-template sequence alignment, an iterative procedure has also been introduced to refine the alignment on the basis of the initial structural fit. The most commonly user software in spatial restraint-based modeling is MODELLER and a database called ModBase has been established for reliable models generated with it.

Loop modeling

Regions of the target sequence that are not aligned to a template are modeled by loop modeling; they are the most susceptible to major modeling errors and occur with higher frequency when the target and template have low sequence identity. The coordinates of unmatched sections determined by loop modeling programs are generally much less accurate than those obtained from simply copying the coordinates of a known structure, particularly if the loop is longer than 10 residues. The first two sidechain dihedral angles (χ1 and χ2) can usually be estimated within 30° for an accurate backbone structure; however, the later dihedral angles found in longer side chains such as lysine and arginine are notoriously difficult to predict. Moreover, small errors in χ1 (and, to a lesser extent, in χ2) can cause relatively large errors in the positions of the atoms at the terminus of side chain; such atoms often have a functional importance, particularly when located near the active site.

Model assessment

Assessment of homology models without reference to the true target structure is usually performed with two methods: statistical potentials or physics-based energy calculations. Both methods produce an estimate of the energy (or an energy-like analog) for the model or models being assessed; independent criteria are needed to determine acceptable cutoffs. Neither of the two methods correlates exceptionally well with true structural accuracy, especially on protein types underrepresented in the PDB, such as membrane proteins.

Statistical potentials are empirical methods based on observed residue-residue contact frequencies among proteins of known structure in the PDB.

Physics-based energy calculations aim to capture the interatomic interactions that are physically responsible for protein stability in solution, especially van der Waals and electrostatic interactions.

One newer method for model assessment relies on machine learning techniques such as neural nets, which may be trained to assess the structure directly or to form a consensus among multiple statistical and energy-based methods. Very recent results using support vector machine regression on a jury of more traditional assessment methods outperformed common statistical, energy-based, and machine learning methods.

Structural comparison methods:

The assessment of homology models' accuracy is straightforward when the experimental structure is known. The most common method of comparing two protein structures uses the root-mean-square deviation (RMSD) metric to measure the mean distance between the corresponding atoms in the two structures after they have been superimposed. A method introduced for the modeling assessment experiment CASP is known as the global distance test (GDT) and measures the total number of atoms whose distance from the model to the experimental structure lies under a certain distance cutoff. Both methods can be used for any subset of atoms in the structure, but are often applied to only the alpha carbon or protein backbone atoms to minimize the noise created by poorly modeled side chain rotameric states, which most modeling methods are not optimized to predict.

Benchmarking

Several large-scale benchmarking efforts have been made to assess the relative quality of various current homology modeling methods. CASP is a community-wide prediction experiment that runs every two years during the summer months and challenges prediction teams to submit structural models for a number of sequences whose structures have recently been solved experimentally but have not yet been published. Its partner CAFASP has run in parallel with CASP but evaluates only models produced via fully automated servers. Continuously running experiments that do not have prediction 'seasons' focus mainly on benchmarking publicly available webservers. LiveBench and EVA run continuously to assess participating servers' performance in prediction of imminently released structures from the PDB. CASP and CAFASP serve mainly as evaluations of the state of the art in modeling, while the continuous assessments seek to evaluate the model quality that would be obtained by a non-expert user employing publicly available tools.

Accuracy

The accuracy of the structures generated by homology modeling is highly dependent on the sequence identity between target and template. Above 50% sequence identity, models tend to be reliable, with only minor errors in side chain packing and rotameric state, and an overall RMSD between the modeled and the experimental structure falling around 1 Â. This error is comparable to the typical resolution of a structure solved by NMR. In the 30-50% identity range, errors can be more severe and are often located in loops. Below 30% identity, serious errors occur, sometimes resulting in the basic fold being mis-predicted. This low-identity region is often referred to as the "twilight zone" within which homology modeling is extremely difficult, and to which it is possibly less suited than fold recognition methods.

At high sequence identities, the primary source of error in homology modeling derives from the choice of the template or templates on which the model is based, while lower identities exhibit serious errors in sequence alignment that inhibit the production of high-quality models. It has been suggested that the major impediment to quality model production is inadequacies in sequence alignment, since "optimal" structural alignments between two proteins of known structure can be used as input to current modeling methods to produce quite accurate reproductions of the original experimental structure.

Attempts have been made to improve the accuracy of homology models built with existing methods by subjecting them to molecular dynamics simulation in an effort to improve their RMSD to the experimental structure. However, current force field parameterizations may not be sufficiently accurate for this task, since homology models used as starting structures for molecular dynamics tend to produce slightly worse structures. Slight improvements have been observed in cases where significant restraints were used during the simulation.

Sources of error

The two most common and large-scale sources of error in homology modeling are poor template selection and inaccuracies in target-template sequence alignment. Controlling for these two factors by using a structural alignment, or a sequence alignment produced on the basis of comparing two solved structures, dramatically reduces the errors in final models; these "gold standard" alignments can be used as input to current modeling methods to produce quite accurate reproductions of the original experimental structure. Results from the most recent CASP experiment suggest that "consensus" methods collecting the results of multiple fold recognition and multiple alignment searches increase the likelihood of identifying the correct template; similarly, the use of multiple templates in the model-building step may be less optimal than the use of the single correct template but more optimal than the use of a single suboptimal one. Alignment errors may be minimized by the use of a multiple alignment even if only one template is used, and by the iterative refinement of local regions of low similarity. A lesser source of model errors are errors in the template structure. The [PDBREPORT] database lists several million, mostly very small but occasionally dramatic, errors in experimental (template) structures that have been deposited in the PDB.

Serious local errors can arise in homology models where an insertion or deletion mutation or a gap in a solved structure result in a region of target sequence for which there is no corresponding template. This problem can be minimized by the use of multiple templates, but the method is complicated by the templates' differing local structures around the gap and by the likelihood that a missing region in one experimental structure is also missing in other structures of the same protein family. Missing regions are most common in loops where high local flexibility increases the difficulty of resolving the region by structure-determination methods. Although some guidance is provided even with a single template by the positioning of the ends of the missing region, the longer the gap, the more difficult it is to model. Loops of up to about 9 residues can be modeled with moderate accuracy in some cases if the local alignment is correct. Larger regions are often modeled individually using ab initio structure prediction techniques, although this approach has met with only isolated success.

The rotameric states of side chains and their internal packing arrangement also present difficulties in homology modeling, even in targets for which the backbone structure is relatively easy to predict. This is partly due to the fact that many side chains in crystal structures are not in their "optimal" rotameric state as a result of energetic factors in the hydrophobic core and in the packing of the individual molecules in a protein crystal. One method of addressing this problem requires searching a rotameric library to identify locally low-energy combinations of packing states. It has been suggested that a major reason that homology modeling so difficult when target-template sequence identity lies below 30% is that such proteins have broadly similar folds but widely divergent side chain packing arrangements.

Applications/Uses

Uses of the structural models include protein-protein interaction prediction, protein-protein docking, molecular docking, and functional annotation of genes identified in an organism's genome. Even low-accuracy homology models can be useful for these purposes, because their inaccuracies tend to be located in the loops on the protein surface, which are normally more variable even between closely related proteins. The functional regions of the protein, especially its active site, tend to be more highly conserved and thus more accurately modeled.

Homology models can also be used to identify subtle differences between related proteins that have not all been solved structurally. For example, the method was used to identify cation binding sites on the Na+/K+ ATPase and to propose hypotheses about different ATPases' binding affinity. Used in conjunction with molecular dynamics simulations, homology models can also generate hypotheses about the kinetics and dynamics of a protein, as in studies of the ion selectivity of a potassium channel. Large-scale automated modeling of all identified protein-coding regions in a genome has been attempted for the yeast Saccharomyces cerevisiae, resulting in nearly 1000 quality models for proteins whose structures had not yet been determined at the time of the study, and identifying novel relationships between 236 yeast proteins and other previously solved structures.