Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Tetrahymena thermophila, a widely studied model for cellular and molecular biology, is a binucleated single-celled organism with a germline micronucleus (MIC) and somatic macronucleus (MAC). The recent draft MAC genome assembly revealed low sequence repetitiveness, a result of the epigenetic removal of invasive DNA elements found only in the MIC genome. Such low repetitiveness makes complete closure of the MAC genome a feasible goal, which to achieve would require standard closure methods as well as removal of minor MIC contamination of the MAC genome assembly. Highly accurate preliminary annotation of Tetrahymena's coding potential was hindered by the lack of both comparative genomic sequence information from close relatives and significant amounts of cDNA evidence, thus limiting the value of the genomic information and also leaving unanswered certain questions, such as the frequency of alternative splicing.
Pubmed ID: 19036158
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
A curated database that provides comprehensive integrated biological information for Saccharomyces cerevisiae along with search and analysis tools to explore these data. SGD allows researchers to discover functional relationships between sequence and gene products in fungi and higher organisms. The SGD also maintains the S. cerevisiae Gene Name Registry, a complete list of all gene names used in S. cerevisiae which includes a set of general guidelines to gene naming. Protein Page provides basic protein information calculated from the predicted sequence and contains links to a variety of secondary structure and tertiary structure resources. Yeast Biochemical Pathways allows users to view and search for biochemical reactions and pathways that occur in S. cerevisiae as well as map expression data onto the biochemical pathways. Literature citations are provided where available.
View all literature mentionsReconfigurable eukaryotic gene finder based on the Generalized Hidden Markov Model framework. The run time and memory requirements are linear in the sequence length. Genezilla utilizes Interpolated Markov Models (IMMs), Maximal Dependence Decomposition (MDD), and includes states for signal peptides, branch points, TATA boxes, and CAP sites.
View all literature mentionsSoftware tool for designing PCR primers on aligned groups of DNA sequences. The most important application is the design of "group-specific" PCR primer sets that amplify a DNA region from a given taxonomic group but do not amplify orthologous regions from other taxonomic groups. It is written in Python 2.3 and Tkinter 8.4. The current script was created for Windows and an executable is available. Future versions of the script should be able to run on Linux and Mac
View all literature mentionsRoche NimbleGen, Inc. is a leading innovator, manufacturer and supplier of a proprietary suite of DNA microarrays, consumables, instruments and services. Roche NimbleGen uniquely produces high-density arrays of long oligo probes that provide greater information content and higher data quality necessary for studying the full diversity of genomic and epigenomic variation. Roche NimbleGen is enabling a new era of High-Definition Genomics by providing scientists with cost-effective, high-throughput tools for extracting and integrating complex data on important forms of genomic and epigenomic variation not previously accessible on a genome-wide scale. Scientists can thus obtain a clearer understanding of genomic and epigenomic structure and function and how they impact biology and medicine. This improved performance is made possible by Roche NimbleGen''s proprietary Maskless Array Synthesis (MAS) technology, which uses digital light processing and rapid, high-yield photochemistry to synthesize long oligo, high-density DNA microarrays with extreme flexibility. NimbleGen Systems was established in 1999. The MAS technology is the result of research collaborations between the departments of biotechnology, genetics, physics, and semiconductor engineering at the University of Wisconsin - Madison. Roche NimbleGen has the exclusive worldwide license to the MAS technology from the Wisconsin Alumni Research Foundation (WARF).
View all literature mentionsSoftware designed to quickly find sequences of 95% and greater similarity of length 25 bases or more.
View all literature mentionsSoftware tool for automated eukaryotic gene structure annotation that reports eukaryotic gene structures as weighted consensus of all available evidence. Used to combine ab intio gene predictions and protein and transcript alignments into weighted consensus gene structures. Inputs include genome sequence, gene predictions, and alignment data (in GFF3 format).
View all literature mentionsThe TIGR database is a collection of plant transcript sequences. Transcript assemblies are searchable using BLAST and accession number. The construction of plant transcript assemblies (TAs) is similar to the TIGR gene indices. The sequences that are used to build the plant TAs are expressed transcripts collected from dbEST (ESTs) and the NCBI GenBank nucleotide database (full length and partial cDNAs). "Virtual" transcript sequences derived from whole genome annotation projects are not included. All plant species for which more than 1,000 ESTs or cDNA sequences are available are included in this project. TAs are clustered and assembled using the TGICL tool (Pertea et al., 2003), Megablast (Zhang et al., 2000) and the CAP3 assembler (Huang and Madan, 1999). TGICL is a wrapper script which invokes Megablast and CAP3. Sequences are initially clustered based on an all-against-all comparisons using Megablast. The initial clusters are assembled to generate consensus sequences using CAP3. Assembly criteria include a 50 bp minimum match, 95% minimum identity in the overlap region and 20 bp maximum unmatched overhangs. Any EST/cDNA sequences that are not assembled into TAs are included as singletons. All singletons retain their GenBank accession numbers as identifiers. Plant TA identifiers are of the form TAnumber_taxonID, where number is a unique numerical identifier of the transcript assembly and taxonID represents the NCBI taxon id. In order to provide annotation for the TAs, each TA/singleton was aligned to the UniProt Uniref database. For release 1 TAs, a masked version of the Uniref90 database was used. For release 2 and onwards, a masked version of the UniRef100 database is used. Alignments were required to have at least 20% identity and 20% coverage. The annotation for the protein with the best alignment to each TA or singleton was used as the annotation for that sequence. Additionally, the relative orientation of each TA/singleton to the best matching protein sequence was used to determine the orientation of each TA/singleton. Some sequences did not have alignments to the protein database that met our quality criteria, and those sequences have neither annotation nor orientation assignments. The release number for the plant TAs refers to the release version for a particular species. For the initial build, all TA sets are of version 1. Subsequent TA updates for new releases will be carried out when the percentage increase of the EST and cDNA counts exceeds 10% of the previous release and when the increase contains more than 1,000 new sequences. New releases will also include additional plant species with more than 1,000 EST or cDNA sequences that have become publicly available.
View all literature mentionsTetrahymena thermophila with name TTMN CU428 from TSC.
View all literature mentions