Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://purl.obolibrary.org/obo/flu/
An application ontology established by a collaborative group of influenza researchers that includes consolidated influenza sequence and surveillance terms from resources such as the BioHealthBase (BHB), a Bioinformatics Resource Center (BRC) for Biodefense and Emerging and Re-emerging Infectious Diseases, the Centers for Excellence in Influenza Research and Surveillance (CEIRS)
Proper citation: Influenza Ontology (RRID:SCR_003346) Copy
http://www.hgsc.bcm.tmc.edu/content/honey-bee-genome-project
The HGSC has sequenced the honey bee, Apis mellifera. The version 4.0 assembly was released in March 2006 and published in October 2006. The genome sequence is being upgraded with additional sequence coverage. The honey bee is important in the agricultural community as a producer of honey and as a facilitator of pollination. It is a model organism for studying the following human health issues: immunity, allergic reaction, antibiotic resistance, development, mental health, longevity and diseases of the X chromosome. In addition, biologists are interested in the honey bee's social organization and behavioral traits. This project was proposed to the HGSC by a group of dedicated insect biologists, headed by Gene Robinson. Following a workshop at the HGSC and a honey bee white paper, the HGSC began the project in 2002. A 6-fold coverage WGS, BAC sequence from pooled arrays, and an initial genome assembly (Amel_v1.0) were released beginning in 2003. This has been a challenging project with difficulty in recovering AT-rich regions. The WGS data had lower coverage in AT-rich regions and BAC data from clones showed evidence of internal deletions. Additional reads from AT enriched DNA addressed these underrepresented regions. The current assembly Amel_4.0 was produced with Atlas and includes 2.7 million reads (1.8 Gb) or 7.5x coverage of the (clonable) genome. About 97% of STSs, 98% of ESTs, and 96% of cDNAs are represented in the 231 Mb assembly. About 2,500 reads were also produced from a strain of Africanized honey bee and SNPs were extracted. These were released in dbSNP and the NCBI Trace Archive. Analysis of the genome by a consortium of 20 labs has been completed. This produced a gene list derived from five different methods melded through the GLEAN software. Publications include a main paper in Nature and up to forty companion papers in Genome Research and Insect Molecular Biology. Sponsors: Sequencing of the honey bee is jointly funded by National Human Genome Research Institute (NHGRI) and the Department of Agriculture (USDA). Multiple drones from the same queen (strain DH4) were obtained from Danny Weaver of B. Weaver Apiaries. All libraries were made from DNA isolated from these drones. The honey bee BAC library (CHORI-224) was prepared by Pieter de Jong and Katzutoyo Osoegawa at the Children's Hospital Oakland Research Institute.
Proper citation: Honey Bee Genome Project (RRID:SCR_002890) Copy
http://www.brenda-enzymes.org/
Database for functional enzyme and ligand-related information maintained as part of the German ELIXIR Node. Provides advanced query systems, evaluation tools, and various visualization options for the detailed assessment of enzyme properties. Enzyme data in BRENDA are classified according to the Enzyme Commission (EC) nomenclature of IUBMB.
Proper citation: BRENDA (RRID:SCR_002997) Copy
A database of hierarchical classification of enzymes that relates specific sequence-structure features to specific chemical capabilities. The SFLD classifies evolutionarily related enzymes according to shared chemical functions and maps these shared functions to conserved active site features. The classification is hierarchical, where broader levels encompass more distantly related proteins with fewer shared features. It thus serves as the analysis and archive site for superfamilies targeted by the Enzyme Function Initiative, and is developed by the Babbitt Laboratory in collaboration with the UCSF Resource for Biocomputing, Visualization, and Informatics. The resource also provides a collection of tools and data for investigating sequence-structure-function relationships and hypothesizing function.
Proper citation: Structure-function linkage database (RRID:SCR_001375) Copy
https://docs.python.org/2/library/random.html
This module implements pseudo-random number generators for various distributions. For integers, uniform selection from a range. For sequences, uniform selection of a random element, a function to generate a random permutation of a list in-place, and a function for random sampling without replacement. On the real line, there are functions to compute uniform, normal (Gaussian), lognormal, negative exponential, gamma, and beta distributions. For generating distributions of angles, the von Mises distribution is available. Sponsors: This resource is supported by ASTi logo Advanced Simulation Technology Inc. (ASTi); Array BioPharma Inc.; BizRate.com; Canonical Ltd.; CCP Games; cPacket Networks; EarnMyDegree.com; Enthought Inc.; Exoweb Ltd.; Google; HitMeister Inc.; IronPort Systems; KNMP; Lucasfilm; Madison Tyler LLC.; Merfin, LLC.; Microsoft; OpenEye Scientific Software; Opsware, Inc.; O''Reilly & Associates, Inc.; PropertySold.ca; Rogue Wave; SEO Moves; Strakt Holdings, Inc.; Sun Microsystems; Tabblo; ZeOmega, LLC., and Zope Corporation.
Proper citation: Generate Pseudo-Random Numbers (RRID:SCR_006535) Copy
This service offers a gateway to well-benchmarked protein structure and function prediction methods. Structural models collected from the prediction servers are assessed using the powerful 3D-jury consensus approach. The Structure Prediction Meta Server provides access to various fold recognition, function prediction and local structure prediction methods. The Server takes the amino acid sequence of the query protein, the reference name for the prediction job, and the E-mail address as input. The E-mail address is used only for notification about errors during the execution of the job. The query sequence and the reference name are placed in the process queue. The Meta Server accepts only sequences, which have not been submitted before. In case of duplicate sequences the second user will be notified with a link to the previous submission. Sequences longer than 800 amino acids are not accepted by some services. The internal SQL database offers the possibility to find any previous jobs processed by the Meta Server using regular expressions addressing field like E-mail, Job Name and the host name, from which the job was initiated. Each server has its own process queuing system managed by the Meta Server. All results of fold recognition servers are translated into uniform formats. The information extracted from the raw output of the servers includes the PDB codes of the hits, the alignments and the similarity (reliability) scores specific for every server. Mapping of the hits to the SCOP and FSSP classifications are made either using known PDB representatives or alignment of the template sequence with the databases of proteins in both classifications. The secondary structure assignments for all hits are taken from the mapped FSSP (red for helices and blue for strands). Underscored amino acids indicate the first residue after an insertion in the template sequence. The Meta server provides translation of the alignments in standard formats like FASTA, PDB or CASP. The Meta Server is coupled to consensus servers. They provide jury predictions based on the results collected from other services. Not all fold recognition servers are used by the jury system. The data stored on the meta server is available through http://meta.bioinfo.pl/data/JOBID/. Jobs older than 2 months are not shown. The Meta Server is only a set of programs aimed to process and manage biological data, while the predictive power of the service comes from (mostly) remote prediction providers. Sponsors: This resource is supported by The BioInfoBank Institute.
Proper citation: BioInfoBank Meta Server (RRID:SCR_007181) Copy
http://net.icgeb.org/benchmark/
It was created in order to create standard datasets on which the performance of machine learning methods can be compared. The collection contains datasets of sequences and structures, each subdivided into positive/negative training/test sets. Such a subdivision is called a classification task. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences, as fell as various functional and taxonomic classification tasks. Running a performance evaluation test on an entire database can include many different classification tasks. These ensembles of classification tasks are encoded in a simple matrix format - called the cast matrix or membership table - that specifies the role of each sequence (or structure) in the different calculations. Each column of this matrix is a subdivision of the objects (rows) into positive/negative training/test sets. Typically, a database record contains such an ensemble of classification tasks, encoded in a single cast matrix. In addition, there is a collection of distance matrices that contain an all vs. all comparison of the datasets using methods as BLAST, Smith-Waterman, 3D-comparisons etc. Evaluation of a method on a given database consists of calculating a performance measure such as a receiver operating curve (ROC) AUC value. Results of evaluation are deposited along with the data, each dataset is evaluated at least by one classification method, such as 1NN (nearest neighbour) or SVM (support vector machines), ANN (artificial neural networks), RF (random forests) etc.. There are small datasets meant for program developers, as well as downloadable programs for various classification algorithms.
Proper citation: Protein Classification Benchmark Collection (RRID:SCR_007561) Copy
http://cmckb.cellmigration.org
It is a database of keys facts about proteins, families, and complexes involved in cell migration. This ongoing project provides a large amount of automated and curated data, collected from numerous online resources that are updated monthly. These data include names, synonyms, sequence information, summaries, CMC research data, reagents, structures, as well as protein family and complex details. CMKB''s ultimate goal is to create a database that will enable the cell migration community to conveniently access significant information about molecules of interest. This will also serve as a stepping stone to pathway analysis and demonstrate how these molecules coordinate with one another during cell adhesion and movement. Sponsors: This resource is supported by the Cell Migration Consortium.
Proper citation: CMKB (RRID:SCR_007229) Copy
Alternative splicing essentially increases the diversity of the transcriptome and has important implications for physiology, development and the genesis of diseases. This resource uses a different approach to investigate alternative splicing (instead of the conventional case-by case fashion) and integrates all transcripts derived from a gene into a single splicing graph. ASG is a database of splicing graphs for human genes, using transcript information from various major sources (Ensembl, RefSeq, STACK, TIGR and UniGene). Each transcript corresponds to a path in the graph, and alternative splicing is displayed by bifurcations. This representation preserves the relationships between different splicing variants and allows us to investigate systematically all possible putative transcripts. Web interface allows users to display the splicing graphs, to interactively assemble transcripts and to access their sequences as well as neighboring genomic regions. ASG also provide for each gene, an exhaustive pre-computed catalog of putative transcriptsin total more than 1.2 million sequences. It has found that ~65 of the investigated genes show evidence for alternative splicing, and in 5 of the cases, a single gene might produce over 100 transcripts.
Proper citation: Alternate splicing gallery (RRID:SCR_008129) Copy
http://kinasedb.ontology.ims.u-tokyo.ac.jp
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 23, 2016. KinasePathwayDatabase is an integrated database concerning completed sequenced major eukaryotes, which contains the classification of protein kinases and their functional conservation and orthologous tables among species, protein-protein interaction data, domain information, structural information, and automatic pathway graph image interface. The protein-protein interactions are extracted by natural language processing (NLP) from abstracts using basic word pattern and protein name dictionary GENA: developed by our group. In this system, pathways are easily compared among species using protein interactions data more than 47,000 and orthologous tables.
Proper citation: Kinase Pathway Database (RRID:SCR_008199) Copy
http://www.roselab.jhu.edu/coil/
The Protein Coil Library is a library of protein structure fragments derived from the Protein Data Bank (PDB). The fragments in this library are those fragments in the PDB that cannot be classified as either alpha-helix or beta-strand. Three-dimensional structures as well as side-chain and backbone torsion angles are stored in the database. The Protein Coil Library allows rapid and comprehensive access to non-alpha-helix and non-beta-strand fragments contained in the Protein Data Bank (PDB). The library contains both sequence and structure information together with calculated torsion angles for both the backbone and side chains. Several search options are implemented, including a query function that uses output from popular PDB-culling servers directly. Additionally, several popular searches are stored and updated for immediate access. The library is a useful tool for exploring conformational propensities, turn motifs, and a recent model of the unfolded state. The library stores the complete torsion angle descriptions for the fragments as well as the three dimensional structures of the fragments themselves. The goal of extracting and pre-calculating this data is to allow for more straightforward investigation of peptide structure without the background of secondary structure elements. In addition to searching by PDB ID, it is possible to download a particular size class, perform a batch search of PDB/chain ID''s, or download precompiled lists of PDB ID''s of interest (PDB Select, etc.). For users interested in browsing the entire database at once or maintaining their own locally-updated copy of the library, FTP access instructions are also provided. The files stored in the coil library FTP site or returned after a batch search are organized heirarchically by PDB ID. This is done to reduce filesystem access times and fascilitate searches using the UNIX find utility. At the lowest directory level in the heirarchy, files are further sorted by fragment length. As a result, the number of files in a particular directory is generally less then 50, yielding relatively fast access on UNIX/Linux filesystems. The heirarchical organization is based on the middle two letters of the PDB ID. For example, hen egg lysozyme, which has a PDB ID of 1HEL, will be located in the directory h/he/. At the final level, fragments of varying sizes are stored in directories that correspond to their fragment length. Again, using lysozyme as an example, any seven-residue fragments, if they exist, will reside in the directory h/he/7/. Similarly, seven-residue fragments from 2HEX and 1HE0 will also be in this location. Sponsors: The Protein Coil Library is funded by Johns Hopkins University.
Proper citation: The Protein Coil Library (RRID:SCR_008233) Copy
http://pbil.univ-lyon1.fr/acuts/ACUTS.html
THIS RESOURCE IS NO LONGER IN SERVICE, Documented on August 12, 2014. Database that identifies new regulatory elements in untranslated regions of protein-coding genes (5 prime flanks, 5 prime UTRs, introns, 3 prime UTRs and 3 prime flanks). The analyses is focused on genes from metazoan species (essentially vertebrates, insects and nematodes). Information on highly conserved regions (sequences, alignments, annotations, bibliographic references) are compiled. Currently 176 out of 326 detected highly conserved regions (HCRs) have been analyzed and incorporated in the database. You can also access the list of annotated conserved elements and the list of conserved elements that remain to be processed. Their approach is based on comparative sequence analysis, for the identification of phylogenetic footprints.
Proper citation: Ancient conserved untranslated sequences (RRID:SCR_008130) Copy
http://animal.dna.affrc.go.jp/agp/index.html
Database of comparative gene mapping between species to assist the mapping of the genes related to phenotypic traits in livestock. The linkage maps, cytogenetic maps, polymerase chain reaction primers of pig, cattle, mouse and human, and their references have been included in the database, and the correspondence among species have been stipulated in the database. AGP is an animal genome database developed on a Unix workstation and maintained by a relational database management system. It is a joint project of National Institute of Agrobiological Sciences (NIAS) and Institute of the Society for Techno-innovation of Agriculture, Forestry and Fisheries (STAFF-Institute), under cooperation with other related research institutes. AGP also contains the Pig Expression Data Explorer (PEDE), a database of porcine EST collections derived from full-length cDNA libraries and full-length sequences of the cDNA clones picked from the EST collection. The EST sequences have been clustered and assembled, and their similarity to sequences in RefSeq, and UniGene determined. The PEDE database system was constructed to store sequences and similarity data of swine full-length cDNA libraries and to make them available to users. It provides interfaces for keyword and ID searches of BLAST results and enables users to obtain sequence data and names of clones of interest. Putative SNPs in EST assemblies have been classified according to breed specificity and their effect on coding amino acids, and the assemblies are equipped with an SNP search interface. The database contains porcine nucleotide sequences and cDNA clones that are ready for analyses such as expression in mammalian cells, because of their high likelihood of containing full-length CDS. PEDE will be useful for researchers who want to explore genes that may be responsible for traits such as disease susceptibility. The database also offers information regarding major and minor porcine-specific antigens, which might be investigated in regard to the use of pigs as models in various medical research applications.
Proper citation: Animal Genome Database (RRID:SCR_008165) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented on August 20,2019.The COG-database has become a powerful tool in the field of comparative genomics. The construction of this data-base is based on sequence homologies of proteins from different completely sequenced genomes. Highly homologous proteins are assigned to clusters of orthologous groups. The updated collection of orthologous protein sets for prokaryotes and eukaryotes is expected to be a useful platform for functional annotation of newly sequenced genomes, including those of complex eukaryotes, and genome-wide evolutionary studies. The availability of multiple, essentially complete genome sequences of prokaryotes and eukaryotes spurred both the demand and the opportunity for the construction of an evolutionary classification of genes from these genomes. Such a classification system based on orthologous relationships between genes appears to be a natural framework for comparative genomics and should facilitate both functional annotation of genomes and large-scale evolutionary studies. Here is a major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes and the construction of clusters of predicted orthologs for 7 eukaryotic genomes, which we named KOGs after eukaryotic orthologous groups. The COG collection currently consists of 138,458 proteins, which form 4873 COGs and comprise 75% of the 185,505 (predicted) proteins encoded in 66 genomes of unicellular organisms. The eukaryotic orthologous groups (KOGs) include proteins from 7 eukaryotic genomes: three animals (the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster and Homo sapiens), one plant, Arabidopsis thaliana, two fungi (Saccharomyces cerevisiae and Schizosaccharomyces pombe), and the intracellular microsporidian parasite Encephalitozoon cuniculi. The current KOG set consists of 4852 clusters of orthologs, which include 59,838 proteins, or approximately 54% of the analyzed eukaryotic 110,655 gene products. Compared to the coverage of the prokaryotic genomes with COGs, a considerably smaller fraction of eukaryotic genes could be included into the KOGs; addition of new eukaryotic genomes is expected to result in substantial increase in the coverage of eukaryotic genomes with KOGs. Examination of the phyletic patterns of KOGs reveals a conserved core represented in all analyzed species and consisting of approximately 20% of the KOG set. This conserved portion of the KOG set is much greater than the ubiquitous portion of the COG set (approximately 1% of the COGs). In part, this difference is probably due to the small number of included eukaryotic genomes, but it could also reflect the relative compactness of eukaryotes as a clade and the greater evolutionary stability of eukaryotic genomes.
Proper citation: Phylogenetic Clusters of Orthologous Groups Ranking (RRID:SCR_008223) Copy
http://www.nisc.nih.gov/projects/comp_seq.html
Generates data for use in developing and refining computational tools for comparing genomic sequence from multiple species. The NISC Comparative Sequencing Program's goal is to establish a data resource consisting of sequences for the same set of targeted genomic regions derived from multiple animal species. The broader program includes plans for a diverse set of analytical studies using the generated sequence and the publication of a series of papers describing the results of those analysis in peer-reviewed journals in a timely fashion. Experimentally, this project involves the shotgun sequencing of mapped BAC clones. For each BAC, an assembly is first performed when a sufficient number of sequence reads have been generated to provide full shotgun coverage of the clone. At that time, the assembled sequence is submitted to the HTGS division of GenBank. Subsequent refinements of the sequence, including the generation of higher-accuracy finished sequence, results in the updating of the sequence record in GenBank. By immediately submitting our BAC-derived sequences to GenBank, it makes their data available as a public service to allow colleagues to speed up their research, consistent with the now well-established routine of sequencing centers participating in the Human Genome Project. However, at the same time, it has made considerable investment in acquiring these mapping and sequence data, including sizable efforts of graduate students, postdoctoral fellows, and other trainees. Furthermore, in most cases, large data sets involving multiple BAC sequences from multiple species must first be generated, often taking many months to accumulate, before the planned analysis can be performed and the resulting papers written and submitted for publication.
Proper citation: Comparative Vertebrate Sequencing (RRID:SCR_008213) Copy
http://lemur.amu.edu.pl/share/php/mirnest/home.php
A database of animal, plant and virus microRNA data maintained at the University of Poznan. The database provides: * 9980 miRNA candiates from 420 animal and plant species predicted in Expressed Sequence Tags * predicted targets for plant candidates * RNA-seq reads mapped to candidates from 29 species * external data from 12 databases that includes sequences, polymorphism, expression and regulation. miRNEST 1.0, it contains miRNA from 563 animals, plants and viruses plant species.
Proper citation: miRNEST (RRID:SCR_008907) Copy
http://genome.jgi.doe.gov/programs/metagenomes/index.jsf
Portal providing access to metagenomics projects, data and tools supported by the DOE Joint Genome Institute (JGI). A primary motivation for metagenomics is that most microbes found in nature exist in complex, interdependent communities and cannot readily be grown in isolation in the laboratory. One can, however, isolate DNA or RNA from the community as a whole, and studies of such communities have revealed a diversity of microbes far beyond those found in culture collections. It is suspected that these uncultivated organisms must harbor considerable as-yet undiscovered genomic, functional, and metabolic features and capabilities. Thus to fully explore microbial genomics, it is imperative that we access the genomes of these elusive players.
Proper citation: Metagenomics Program at JGI (RRID:SCR_008804) Copy
http://ecoliwiki.net/colipedia/index.php/T4-like_genome_database
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. A database of information on bacterial phages. It contains multiple phage genomes, which users can BLAST and MegaBLAST, and also hosts a Phage Forum in which users can discuss phage data. Interactive browsing of completed phage genomes is available using the program. The browser allows users to scan the genome for particular features and to download sequence information plus analyses of those features. Views of the genome are generated showing named genes BLAST similarities to other phages predicted tRNAs and other sequence features.
Proper citation: T4-like genome database (RRID:SCR_005367) Copy
http://genome.jgi.doe.gov/programs/plants/index.jsf
The goal of the DOE JGI Plant Genome Program is to shed light on the fundamental biology of photosynthesis and transduction of solar to chemical energy. Other areas of interest include characterizing: * Ecosystems and the role of terrestrial plants and oceanic phytoplankton-in carbon sequestration. * The role of plants in coping with toxic pollutants in soils by hyper-accumulation and detoxification. * Feedstocks for biofuels, e.g., biodiesel from soybean; cellulosic ethanol from perennial grasses. * The ability to respond to environmental change (e.g., loss of diversity from monoculture produces vulnerabilities; nitrogen fixing nodules in legumes reduce fertilizer need). * The generation of useful secondary metabolites (produced largely for disease resistance)- for positive/negative control in agriculture, with attendant influence on global carbon cycle. The Plant Genome Program accomplishes the above through the following activities: # Sequence. Produce genome sequences of key plant (and algal) species to accelerate biofuel development and understand response to climate change. # Function. Develop datasets (and synthetic biology tools) to elucidate functional elements in plant genomes, with special focus on handful of flagship genomes. # Variation. Characterize natural genomic variation in plants (and their associated microbiomes), and relate to biofuel sustainability and adaptation to climate change. # Integration. Provide a centralized hub for the retrieval and deep integrated analysis of plant genome datasets.
Proper citation: Plant Genome Resource at JGI (RRID:SCR_005315) Copy
Database of known and predicted protein interactions. The interactions include direct (physical) and indirect (functional) associations and are derived from four sources: Genomic Context, High-throughput experiments, (Conserved) Coexpression, and previous knowledge. STRING quantitatively integrates interaction data from these sources for a large number of organisms, and transfers information between these organisms where applicable. The database currently covers 5''214''234 proteins from 1133 organisms. (2013)
Proper citation: STRING (RRID:SCR_005223) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.