Showing posts with label Sequencing. Show all posts
Showing posts with label Sequencing. Show all posts

Generation of detailed maps using sequence tagged sites


A map designed for a sequencing project to accurately locate the genes and assign the gene function adds direct value to mapping studies. The smallest functional units or genes are present on the chromosome. Hence it is important to study their position and locate them. The recombination frequencies or the number of recombinants found out of the total number of progenies guide in assessing the relative distances between the loci. The mathematical relationship between the map distance and recombination frequencies help to determine map function. Linkage of genes mainly tests whether the genes are present on the same chromosome. Therefore a map is constructed based on the relative positions of the genes. For large genomes, constructing a physical map would add value to the sequencing projects. However, if the purpose is to clone the individual genes, a different approach may be used. The physical mapping techniques do not rely on the presence of alleles to map the genome. The genetic mapping requires the alleles for a given marker. Unavailability of a map leads to an error-prone assembly of the genome sequence. The genetic maps have a poor resolution and inaccuracy. These properties are refined using a physical map.


Image 1: STS Mapping (Cloned DNA fragments showing markers)

Molecular biology techniques are used to construct a genome map. The plethora of physical mapping involves restriction mapping, FISH, and STS mapping. Restriction sites in small DNA molecules are possible in restriction mapping with a limitation in eukaryotic chromosomes. FISH is a good choice for mapping genomes. However, it takes a lot of time for mapping large genomes. More specific approach for mapping large genomes is required. Sequence tagged sites (STS) guide very accurately in physical mapping. This type of mapping is known as STS mapping. A sequence tagged site though seems complicated by its name, is not so complicated. In fact, it is a very specific approach. It is any site in the chromosome that is identified by a known unique DNA sequence. Mapping large genomes require a high resolution and rapid technique with less demand in technicality.
The limitations of FISH being difficult to conduct and accumulate the data in a single experiment, STS mapping may meet the requirement of providing map positions more than three markers in one go. A short DNA sequence with 100-500 base pairs is easy to recognize because it occurs only once in the chromosome or a genome. Such a site or a sequence is useful for physical mapping. STS mapping requires such a collection of overlapping fragments from the genome. The entire chromosome is used to obtain the mapping reagent or a collection of overlapping fragments. It helps to identify the marker position.

Properties of Sequence Tagged Sites:
As discussed earlier, sequence tagged sites are small DNA sequences that are unique. Two main properties determine sequence tagged sites. The DNA sequence must be known. Hence, a PCR assay can be set up. A PCR assay of known DNA sequence determines the STS on different fragments. The DNA fragment must have a unique location on a chromosome. If the STS are positioned more than once in the genome, it becomes difficult to map. Mainly repetitive DNA has high chances of having more than once positioned sequences. Thus STS mapping does not include repetitive DNA sequences. The reason for using PCR instead of hybridization lies in its automatic mechanisms and efficiency. It is difficult to include those fragments whose sequence is not known to us. Therefore the foremost criteria for a genome mapping are to know a sequence. The probability that the two closely linked markers determined by the fragments found on the same chromosome.
Distantly linked markers may be present in different fragments. A collection of fragments involves many fragments with closely linked or distantly linked markers. The frequency at which breaks occur between the two markers decides the map distance.




Image 2: Fragment collection (STS mapping)

Sources of sequence tagged sites (STS):
Three main sources of sequence tagged sites include expressed sequence tags (EST), SSLPs, and random genomic sequences.
·        Expressed sequence tags (ESTs): The cDNA clones are analyzed to obtain short sequences known as expressed sequence tags or ESTs. The sequence derived from the cDNA library is unique. A sequence transcribed in some tissue or at some stage of the developmental process is used to derive a unique sequence. Thus, an EST mapped through a specific mapping procedure identifies a unique gene locus. EST markers are produced using PCR. It involves oligonucleotide primers based on cDNA sequence. They correspond to protein-coding genes. So, the unique ESTs are capable of becoming STS.
·        SSLPs: Simple sequence length polymorphisms are arrays of repeat sequences displaying length variations. SSLPs contain alleles with a different number of repeat units that are multi-allelic. Two main types of SSLPs include minisatellites and microsatellites. Polymorphic SSLPs are usually preferred.
·        Random genomic sequences: They are randomly cloned sequences. Randomly spanning the available online databases help to obtain random genomic sequences. However, they are known sequences.

Fragment collection for STS mapping:
The collection of DNA fragments are known as a mapping reagent. The fragments are present in the entire chromosome. Each point has an average of five fragments. The markers may be near or far on the fragments. Mapping reagents assemble in two ways such as clone library and radiation hybrids. Rodent cell lines support the radiation mapping techniques. Different fragments of the second genome consist of the rodent cell line. Irradiation techniques are used to construct these cell lines. Hence they are mapping reagents in studying large genomes. A clone library is another mapping reagent. It is a collection of clones representing an entire genome. These clone collections supply individual clones of interest. Let us know the two mapping reagents in detail.

Radiation hybrids:
The human chromosomes paved way in the development of radiation hybrids. During the early 1970’s, experimenters exposed human cells to 3000-8000 rad doses of X-rays, leading to chromosome breakage. However, this treatment was lethal for human cells. The irradiated cells propagated on fusion with non-irradiated hamster cells. Polyethylene glycol or Sendai viral exposure led to the fusion of both the cells. However, all hamster cells are not capable of accepting human chromosomes. The hamster cells, therefore, need to undergo a selection process. Some hamster cells are unable to make thymidine kinase or hypoxanthine phosphoribosyl transferase. So the cells are grown in a medium containing hypoxanthine, aminopterin and thymidine medium (HAT). The fused cells are cultured in the HAT medium.

Hybrid hamster cells grow on this medium. It indicates that these cells have accepted human chromosomes. The hybrid cells consist of human DNA inserted into hamster chromosomes. These fragments are 5-10 Mb in size. The collection of hybrids is known as radiation hybrid panels. It is a mapping reagent. The rodent cells may also be used to obtain radiation hybrid panels. The rodent cell lines consist of human DNA fragments in the rodent nucleus. Sometimes the hybrid rodent cells are fused with hamster cells. Such hybrids may contain both human and mouse chromosomes or a mixture of both. Specific probes help to identify hybrid cells containing human DNA. The probes are specific sequences identifying human DNA. Examples include SINEs or short interspersed nuclear element called as Alu. The Alu elements have a copy number of over a million. Two types of radiation hybrid panels are known so far. They include single chromosome panels and whole genome panels. Let us know the difference between the two.

Single chromosome panel
Whole genome panel
Only a few hybrids are required
A few hundred hybrids are required.
PCR screening involved convenience in handling
Involved less convenience in handling
Involved irradiation of mouse cell containing more mice DNA and less human DNA.       
Irradiation of human DNA
Human DNA hybrid is less in content
Higher human DNA hybrid content
The human genome project avoided the approach due to less human DNA hybrid content.
The human genome project utilized the approach due to high human DNA hybrid content.
 Table: Difference between single chromosome panel and the whole genome panel
Clone library:
A collection of clones represents the human genome. They are used to obtain individual clones of interest. A clone library is obtained by breaking the genome into fragments and thereby cloning them into the vector. Hence, a clone library consisting of large genome fragments is a mapping reagent. A clone library is a chromosome specific library. It is possible to separate chromosomes using a clone library. They have sufficient information for STS mapping. The STS analysis determines the clones consisting of overlapping DNA fragments enabling clone contigs to build-up.

References:
[1] Molecular Biology, David P. Clark, Nanette J. Pazdernik
[2] Human Molecular Genetics 3, Volume 3, T. Strachan, Andrew P. Read




© Copyright, 2018 All Rights Reserved.

Human Genome Project

The human genome is the total genetic material in the cells of a human being. It contains billions of nucleotide base pairs. The genetic material is present in the chromosomes.  Human Genome Project involves sequencing of all the DNA base pairs and mapping of several genes. The project started in 1991 in the USA. It was the largest project in human genetics. It was completed by April 14, 2003. The National Institutes of Health-funded primarily for this project. The project covered 99% euchromatic genome with 99.99% accuracy. Human genome sequencing benefitted a lot. It helped in understanding the diseases, genotyping of specific microbes, identification of mutations, cancer genetics, drug designing, and biotechnology. Specific databases were designed to store the sequences of the DNA. The National Center for Biotechnology Information (NCBI) consists of all the database information in GenBank. It is a hub for gene sequence information, protein sequences, and related information.

Image: Human Genome Project

Objectives of the Human Genome Project:
1.     Human Genome Sequencing:
It is used to figure out the order of nucleotides or bases such as adenine, guanine, cytosine or thymine.
2.     Human Gene Mapping:
It is used to identify a locus of a gene and the distance between the genes. Human gene mapping places a collection of molecular markers on their respective genome positions.
3.     Mapping of human inherited Diseases:
It helps to identify genes and biological processes. It is used to understand the molecular basis of inheritance.
4.     Development of new DNA technologies:
Human genome project developed new DNA technologies. They consisted of recombinant techniques for studying diseases and drug designing.
5.     Development of bioinformatics:
It helped in assembling DNA sequences, finding genetic landscape features, genome mining, the study of genetic variation and disease.
6.     Comparative genomics:
A computer-based analysis is used to compare the entire genome sequence. Comparative genomics is used to study similarity and difference regions.

Techniques used in the Human Genome Project:
Genome annotation technique used in the human genome project identified boundaries between genes and other features in a DNA. The bioinformatics domain stores the genome annotated sequences. RNA-Seq is a new technology introduced to sequence a messenger RNA in the cells. RNA-Seq was more accurate than annotation.

Highlights of the findings:
·        The Human genome project identified approximately 22,300 protein-coding genes.
·        Human genome sequencing identified 3.2 billion base pairs.
·        A human body consists of 26,000 to 35,000 genes.
·     Genes constitute only 5% of the human genome. Over 95% of the human genome is known as junk DNA. The non-coding DNA is known as junk DNA.
·        The junk DNA constitutes repeating DNA segments.
·        Only 7% of protein families were vertebrate specific.
·        Genes function as complex networks.

The Mapping Phase of Human Genome Project:
The gene mapping involved restriction fragment length polymorphisms (RFLPs). These RFLPs are highly polymorphic DNA markers. It comprised of 393 RFLPs. It also consisted of ten polymorphic markers. The marker density was 10 Mb. The RFLP map consisted of single strand length polymorphisms (SSLPs). Clone contigs were primarily used to develop physical mapping. Methods used were STS screening and clone fingerprinting.
A clone contig map consisted of 33,000 Yeast Artificial Chromosomes (YACs). However, there was a limitation of YAC. It contained few pieces of non-contiguous DNA. STS markers were mapped using radiation hybrid mapping. STS maps included 7000 polymorphic SSLPs.

Human Genome Sequencing:
Since YACs consisted of non-contiguous DNA, the entire focus switched over to BACs. The scientists cloned and mapped the Bacterial artificial chromosomes (BACs). A library of three lakh BAC clones was generated and mapped into the genome. The ready map of BAC was a primary foundation for the sequencing project. The shotgun sequencing method was used to replace the clone contig method.
There are many advantages to the human genome project. It helps to study the genes and mutations associated with them. It is easier to diagnose, predict and prevent the disease using HGP principles. Medicines can be developed based on the individual’s response to treatment. There is a wide scope for personalized medicine and drug designing. Using the databases generated during HGP, it is easy to carry out research activities. The gene databases help the molecular and cytogeneticists to study specific gene mutations, chromosomal abnormalities, and sequencing strategies. DNA fingerprinting techniques are useful in forensic medicine and crime investigation. With the help of gene mapping, it becomes easier to study inherited diseases and conduct efficient genetic counseling and prenatal diagnosis. However, there are very few disadvantages to the human genome project. Knowing one’s genome and possible risks that could result in a disease in the future may create an environment for genetic discrimination. However certain laws have been implemented such as GINA act to prevent genetic discrimination.

Cancer genomics:
With approximately millions of cancer cases worldwide, the survival rates of cancer patients are decreasing every year. The response of patients to the current treatment methodologies is not satisfactory. Moreover, the drugs and radiations given to the patients have been fraught with severe toxicity and side effects thereby limiting their applications in the cancer therapeutics. The human genome project had an objective of genome sequencing and mapping. Thus the research can be directed toward the development of new anti-cancer products. These products might target the signaling pathways, apoptosis, metastasis or migration of cancer cells. An increasingly rapid DNA analysis with the help of the human genome project may establish new therapeutic targets and facilitate effectiveness.

Enable Technologies:
An evolutionary improvement in the existing genome sequencing technologies through Human genome project may have a revolutionary impact on the genetic research. The human genome project would be helpful in designing treatment strategies for various diseases. The past, present and the future are genome based. Overcoming those few disadvantages of HGP might equip the human population to adopt the changes in the future. Though the human genome project is a well-established one, we know that studying the entire genome would never be complete. As the environment changes, the conditions may change and so the genes. Thus human genome would always be a topic for study. Chances of new mutations would be high. Thus, studying a human genome is a coordinated effort of the researchers to carry out the mapping and sequencing. High-throughput revolutionary technologies developed with the synergy of advanced computerization, automated machines, and robotics work excellent with microarrays, modern screening, and imaging techniques.

References:
[1] Genomes, T.A. Brown
[2] Human Genome Project- Wikipedia
[3] Genetics Home Reference
© Copyright, 2018 All Rights Reserved.

Genomics and Proteomics for Cancer Research

The uncontrolled division of cells creates an abnormal environment in the body, leading to a condition known as cancer. It is the b...