Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://www.neuroepigenomics.org/methylomedb/
A database containing genome-wide brain DNA methylation profiles for human and mouse brains. The DNA methylation profiles were generated by Methylation Mapping Analysis by Paired-end Sequencing (Methyl-MAPS) method and analyzed by Methyl-Analyzer software package. The methylation profiles cover over 80% CpG dinucleotides in human and mouse brains in single-CpG resolution. The integrated genome browser (modified from UCSC Genome Browser allows users to browse DNA methylation profiles in specific genomic loci, to search specific methylation patterns, and to compare methylation patterns between individual samples. Two species were included in the Brain Methylome Database: human and mouse. Human postmortem brain samples were obtained from three distinct cortical regions, i.e., dorsal lateral prefrontal cortex (dlPFC), ventral prefrontal cortex (vPFC), and auditory cortex (AC). Human samples were selected from our postmortem brain collection with extensive neuropathological and psychopathological data, as well as brain toxicology reports. The Department of Psychiatry of Columbia University and the New York State Psychiatric Institute have assembled this brain collection, where a validated psychological autopsy method is used to generate Axis I and II DSM IV diagnoses and data are obtained on developmental history, history of psychiatric illness and treatment, and family history for each subject. The mouse sample (strain 129S6/SvEv) DNA was collected from the entire left cerebral hemisphere. The three human brain regions were selected because they have been implicated in the neuropathology of depression and schizophrenia. Within each cortical region, both disease and non-psychiatric samples have been profiled (matching subjects by age and sex in each group). Such careful matching of subjects allows one to perform a wide range of queries with the ability to characterize methylation features in non-psychiatric controls, as well as detect differentially methylated domains or features between disease and non-psychiatric samples. A total of 14 non-psychiatric, 9 schizophrenic, and 6 depression methylation profiles are included in the database.
Proper citation: MethylomeDB (RRID:SCR_005583) Copy
http://llama.mshri.on.ca/funcassociate/
A web-based tool that accepts as input a list of genes, and returns a list of GO attributes that are over- (or under-) represented among the genes in the input list. Only those over- (or under-) representations that are statistically significant, after correcting for multiple hypotheses testing, are reported. Currently 37 organisms are supported. In addition to the input list of genes, users may specify a) whether this list should be regarded as ordered or unordered; b) the universe of genes to be considered by FuncAssociate; c) whether to report over-, or under-represented attributes, or both; and d) the p-value cutoff. A new version of FuncAssociate supports a wider range of naming schemes for input genes, and uses more frequently updated GO associations. However, some features of the original version, such as sorting by LOD or the option to see the gene-attribute table, are not yet implemented. Platform: Online tool
Proper citation: FuncAssociate: The Gene Set Functionator (RRID:SCR_005768) Copy
http://bioinfo.iitk.ac.in/MIPModDB/
This is a database of comparative protein structure models of MIP (Major Intrinsic Protein) family of proteins. The nearly completed sets of MIPs have been identified from the completed genome sequence of organisms available at NCBI. The structural models of MIP proteins were created by defined protocol. The database aims to provide key information of MIPs in particular based on sequence as well as structures. This will further help to decipher the function of uncharacterized MIPs. For each MIP entry, this database contains information about the source, gene structure, sequence features, substitutions in the conserved NPA motifs, structural model, the residues forming the selectivity filter and channel radius profile. For selected set of MIPs, it is possible to derive structure-based sequence alignment and evolutionary relationship. Sequences and structures of selected MIPs can be downloaded from MIPModDB database.
Proper citation: MIPModDB (RRID:SCR_006058) Copy
http://wego.genomics.org.cn/cgi-bin/wego/index.pl
Web Gene Ontology Annotation Plot (WEGO) is a simple but useful tool for plotting Gene Ontology (GO) annotation results. Different from other commercial software for chart creating, WEGO is designed to deal with the directed acyclic graph (DAG) structure of GO to facilitate histogram creation of GO annotation results. WEGO has been widely used in many important biological research projects, such as the rice genome project and the silkworm genome project. It has become one of the useful tools for downstream gene annotation analysis, especially when performing comparative genomics tasks. Platform: Online tool
Proper citation: WEGO - Web Gene Ontology Annotation Plot (RRID:SCR_005827) Copy
http://snps-and-go.biocomp.unibo.it/snps-and-go/
A server for the prediction of single point protein mutations likely to be involved in the insurgence of diseases in humans.
Proper citation: SNPsandGO (RRID:SCR_005788) Copy
A curated repository of more than 206000 regulatory associations between transcription factors (TF) and target genes in Saccharomyces cerevisiae, based on more than 1300 bibliographic references. It also includes the description of 326 specific DNA binding sites shared among 113 characterized TFs. Further information about each Yeast gene has been extracted from the Saccharomyces Genome Database (SGD). For each gene the associated Gene Ontology (GO) terms and their hierarchy in GO was obtained from the GO consortium. Currently, YEASTRACT maintains a total of 7130 terms from GO. The nucleotide sequences of the promoter and coding regions for Yeast genes were obtained from Regulatory Sequence Analysis Tools (RSAT). All the information in YEASTRACT is updated regularly to match the latest data from SGD, GO consortium, RSA Tools and recent literature on yeast regulatory networks. YEASTRACT includes DISCOVERER, a set of tools that can be used to identify complex motifs found to be over-represented in the promoter regions of co-regulated genes. DISCOVERER is based on the MUSA algorithm. These algorithms take as input a list of genes and identify over-represented motifs, which can then be compared with transcription factor binding sites described in the YEASTRACT database.
Proper citation: Yeast Search for Transcriptional Regulators And Consensus Tracking (RRID:SCR_006076) Copy
http://prism.ccbb.ku.edu.tr/hotregion/index.php
Hot spots are energetically important residues at protein interfaces and they are not randomly distributed across the interface but rather clustered. These clustered hot spots form hot regions. Hot regions are important for the stability of protein complexes, as well as providing specificity to binding sites. HotRegion provides the hot region information of the interfaces by using predicted hot spot residues, and structural properties of these interface residues such as pair potentials of interface residues, accessible surface area (ASA) and relative ASA values of interface residues of both monomer and complex forms of proteins. Also, the 3D visualization of the interface and interactions among hot spot residues are provided. The number of interfaces in the database is 147909 and still growing.
Proper citation: HotRegion - A Database of Cooperative Hotspots (RRID:SCR_006022) Copy
GlycomeDB is a database of all known carbohydrate structures. This was achieved by crosslinking several other databases of carbohydrate structures by using the GlycoCT XML language specification. We have analyzed all of the existing public databases and defined a sequence format based on XML (GlycoCT) capable of storing all structural information of carbohydrate sequences. We have implemented a library of parsers for the interpretation of the different encoding schemes for carbohydrates. With this library we have translated the carbohydrate sequences of all freely available databases (CFG , KEGG, GLYCOSCIENCES.de, BCSDB and Carbbank) to GlycoCT, and created a new database (GlycomeDB) containing all structures and annotations. During the process of data integration we found multiple inconsistencies in the existing databases which were corrected in collaboration with the responsible curators. With the new database, GlycomeDB, it is possible to get an overview of all carbohydrate structures in the different databases and to crosslink common structures in the different databases. Scientists are now able to search for a particular structure in the meta database and get information about the occurrence of this structure in the five carbohydrate structure databases.
Proper citation: glycomedb (RRID:SCR_005717) Copy
http://tropgenedb.cirad.fr/tropgene/JSP/index.jsp
A database that manages genetic and genomic information about tropical crops studied by Cirad. The database is organised into crop specific modules. Each module includes data on genetic ressources (agro-morphological data, parentages, allelic diversity), information on molecular markers, genetics maps, result of QTL analyses, data from physical mapping, sequences, genes, as well as corresponding references. GENE DB interface has been designed to allow quick consultations as well as complex queries. Nine modules are presently on line.
Proper citation: TropGENE DB (RRID:SCR_005716) Copy
DOMMINO is a comprehensive structural database on macromolecular interactions. As of June, 2011, it contains more than 407,000 binary interactions. The distinctive features of DOMMINO are: # Automated updates: DOMMINO is fully automated and is designed to update itself on a weekly basis, one day after a PDB weekly update. Thus, the community will be able to study macromolecular interactions almost immediately after they are released by PDB. # Coverage of non-domain mediated interactions: In addition to domain-domain and domain-peptide interactions the database characterizes the interaction between domains and unstructured protein regions that are not parts of a domain, such as inter-domain linkers and N- and C-termini. The interactions that involve the latter unstructured parts of proteins have been included to the database for the first time providing additional ~186,000 interactions (~45% of the total number of interactions, as of June, 2011). # Coverage of new structural domains: DOMMINO employs one of the most accurate structural classifications of proteins, SCOP. In addition to the existing SCOP-annotated domains, we employ a state-of-the-art machine learning approach to classify newer protein structures into existing SCOP families. With the progress of structural genomics, we do not expect a significant growth of the number of structurally novel folds or protein families and therefore our method allows covering almost all new protein structures. In total, using this predictive approach has allowed us to add more than 261,000 new interactions, almost twice as many as existing SCOP-annotated interactions. # The web-interface is designed to give the user a possibility of a flexible search as well as the capability to study macromolecular interactions in a PDB structure at the interaction network level and at the individual interface level. The web interface of the DOMMINO database includes a comprehensive list of help topics linked to the specific actions. In addition, we have designed a step-by-step tutorial that covers all aspects of working with the data from DOMMINO using the web interface.
Proper citation: DOMMINO - Database Of MacroMolecular INteractiOns (RRID:SCR_005958) Copy
http://www.jcvi.org/charprotdb/index.cgi/home
The Characterized Protein Database, CharProtDB, is designed and being developed as a resource of expertly curated, experimentally characterized proteins described in published literature. For each protein record in CharProtDB, storage of several data types is supported. It includes functional annotation (several instances of protein names and gene symbols) taxonomic classification, literature links, specific Gene Ontology (GO) terms and GO evidence codes, EC (Enzyme Commisssion) and TC (Transport Classification) numbers and protein sequence. Additionally, each protein record is associated with cross links to all public accessions in major protein databases as ��synonymous accessions��. Each of the above data types can be linked to as many literature references as possible. Every CharProtDB entry requires minimum data types to be furnished. They are protein name, GO terms and supporting reference(s) associated to GO evidence codes. Annotating using the GO system is of importance for several reasons; the GO system captures defined concepts (the GO terms) with unique ids, which can be attached to specific genes and the three controlled vocabularies of the GO allow for the capture of much more annotation information than is traditionally captured in protein common names, including, for example, not just the function of the protein, but its location as well. GO evidence codes implemented in CharProtDB directly correlate with the GO consortium definitions of experimental codes. CharProtDB tools link characterization data from multiple input streams through synonymous accessions or direct sequence identity. CharProtDB can represent multiple characterizations of the same protein, with proper attribution and links to database sources. Users can use a variety of search terms including protein name, gene symbol, EC number, organism name, accessions or any text to search the database. Following the search, a display page lists all the proteins that match the search term. Click on the protein name to view more detailed annotated information for each protein. Additionally, each protein record can be annotated.
Proper citation: CharProtDB: Characterized Protein Database (RRID:SCR_005872) Copy
http://pbildb1.univ-lyon1.fr/virhostnet/
Public knowledge base specialized in the management and analysis of integrated virus-virus, virus-host and host-host interaction networks coupled to their functional annotations. It contains high quality and up-to-date information gathered and curated from public databases (VirusMint, Intact, HIV-1 database). It allows users to search by host gene, host/viral protein, gene ontology function, KEGG pathway, Interpro domain, and publication information. It also allows users to browse viral taxonomy.
Proper citation: VirHostNet: Virus-Host Network (RRID:SCR_005978) Copy
A web-based tool that provides composite interpretations for microarray data comparing two sample groups as well as lists of genes from diverse sources of biological information. It provides multiple gene set analysis methods for microarray inputs as well as enrichment analyses for lists of genes. It screens redundant composite annotations when generating and prioritizing them. It also incorporates union and subtracted sets as well as intersection sets. Users can upload their gene sets (e.g. predicted miRNA targets) to generate and analyze new composite sets.
Proper citation: ADGO (RRID:SCR_006343) Copy
http://bioinformatics.biol.uoa.gr/HMM-TM/
A web tool using the Hidden Markov Model method for the topology prediction of alpha-helical membrane proteins that incorporates experimentally derived topological information. Hidden Markov Models (HMMs) have been extensively used in computational molecular biology, for modelling protein and nucleic acid sequences. In many applications, such as transmembrane protein topology prediction, the incorporation of limited amount of information regarding the topology, arising from biochemical experiments, has been proved a very useful strategy that increased remarkably the performance of even the top-scoring methods. However, no clear and formal explanation of the algorithms that retains the probabilistic interpretation of the models has been presented so far in the literature. We present here, a simple method that allows incorporation of prior topological information concerning the sequences at hand, while at the same time the HMMs retain their full probabilistic interpretation in terms of conditional probabilities. We present modifications to the standard Forward and Backward algorithms of HMMs and we also show explicitly, how reliable predictions may arise by these modifications, using all the algorithms currently available for decoding HMMs. A similar procedure may be used in the training procedure, aiming at optimizing the labels of the HMM''s classes, especially in cases such as transmembrane proteins where the labels of the membrane-spanning segments are inherently misplaced. We present an application of this approach developing a method to predict the transmembrane regions of alpha-helical membrane proteins, trained on crystallographically solved data. We show that this method compares well against already established algorithms presented in the literature, and it is extremely useful in practical applications.
Proper citation: HMM-TM (RRID:SCR_006186) Copy
http://bioinformatics.biol.uoa.gr/PRED-LIPO/
A web tool using the Hidden Markov Model method for the prediction of lipoprotein signal peptides of Gram-positive bacteria, trained on a set of 67 experimentally verified lipoproteins. The method outperforms LipoP and the methods based on regular expression patterns, in various data sets containing experimentally characterized lipoproteins, secretory proteins, proteins with an N-terminal TM segment and cytoplasmic proteins. The method is also very sensitive and specific in the detection of secretory signal peptides and in terms of overall accuracy outperforms even SignalP, which is the top-scoring method for the prediction of signal peptides.
Proper citation: PRED-LIPO (RRID:SCR_006187) Copy
http://bioinformatics.biol.uoa.gr/PRED-SIGNAL/
A web tool for prediction of signal peptides in archaea. Computational prediction of signal peptides (SPs) and their cleavage sites is of great importance in computational biology; however, currently there is no available method capable of predicting reliably the SPs of archaea, due to the limited amount of experimentally verified proteins with SPs. We performed an extensive literature search in order to identify archaeal proteins having experimentally verified SP and managed to find 69 such proteins, the largest number ever reported. A detailed analysis of these sequences revealed some unique features of the SPs of archaea, such as the unique amino acid composition of the hydrophobic region with a higher than expected occurrence of isoleucine, and a cleavage site resembling more the sequences of gram-positives with almost equal amounts of alanine and valine at the position-3 before the cleavage site and a dominant alanine at position-1, followed in abundance by serine and glycine. Using these proteins as a training set, we trained a hidden Markov model method that predicts the presence of the SPs and their cleavage sites and also discriminates such proteins from cytoplasmic and transmembrane ones.
Proper citation: PRED-SIGNAL (RRID:SCR_006181) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. Database for corrected read counts and genome mapping on NCBI's Short Read Archive. The corrected count was done using RECOUNT and the mapping with LAST. We also provide information of reference genome to which we aligned the short reads. We focus on transcriptomic data, specifically TSS-Seq and RNA-Seq. Because this is the type of data for which sequence count correction is most important. Hence we do not include the genomic reads. The current version contains 2,265 entries from 45 organisms, with read lengths from 17 to 100bp. Via a searchable and browseable interface users can obtain corrected data in formats useful for transcriptomic analysis. We provide the data grouped according to the genome, type of studies and submitter in TAB , PSL and BAM format. They contain the mapping position and annotation of reads observed and corrected counts.
Proper citation: RecountDB (RRID:SCR_006117) Copy
http://prorepeat.bioinformatics.nl/
ProRepeat is an integrated curated repository and analysis platform for in-depth research on the biological characteristics of amino acid tandem repeats. ProRepeat collects repeats from all proteins included in the UniProt knowledgebase, together with 85 completely sequenced eukaryotic proteomes contained within the RefSeq collection. It contains non-redundant perfect tandem repeats, approximate tandem repeats and simple, low-complexity sequences, covering the majority of the amino acid tandem repeat patterns found in proteins. The ProRepeat web interface allows querying the repeat database using repeat characteristics like repeat unit and length, number of repetitions of the repeat unit and position of the repeat in the protein. Users can also search for repeats by the characteristics of repeat containing proteins, such as entry ID, protein description, sequence length, gene name and taxon. ProRepeat offers powerful analysis tools for finding biological interesting properties of repeats, such as the strong position bias of leucine repeats in the N-terminus of eukaryotic protein sequences, the differences of repeat abundance among proteomes, the functional classification of repeat containing proteins and GC content constrains of repeats' corresponding codons.
Proper citation: ProRepeat (RRID:SCR_006113) Copy
The database of protein-chemical structural interactions includes all existing 3D structures of complexes of proteins with low molecular weight ligands. When one considers the proteins and chemical vertices of a graph, all these interactions form a network. Biological networks are powerful tools for predicting undocumented relationships between molecules. The underlying principle is that existing interactions between molecules can be used to predict new interactions. For pairs of proteins sharing a common ligand, we use protein and chemical superimpositions combined with fast structural compatibility screens to predict whether additional compounds bound by one protein would bind the other. The current version includes data from the Protein Data Bank as of August 2011. The database is updated monthly.
Proper citation: ProtChemSI (RRID:SCR_006115) Copy
High quality ribosomal RNA databases providing comprehensive, quality checked and regularly updated datasets of aligned small (16S/18S, SSU) and large subunit (23S/28S, LSU) ribosomal RNA (rRNA) sequences for all three domains of life (Bacteria, Archaea and Eukarya). Supplementary services include a rRNA gene aligner, online tools for probe and primer evaluation and optimized browsing, searching and downloading on the website. The extensively curated SILVA taxonomy and the new non-redundant SILVA datasets provide an ideal reference for high-throughput classification of data from next-generation sequencing approaches. Alignment tool, SINA, is available for download as well as available for use online.
Proper citation: SILVA (RRID:SCR_006423) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.