Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
Collection of genome databases for vertebrates and other eukaryotic species with DNA and protein sequence search capabilities. Used to automatically annotate genome, integrate this annotation with other available biological data and make data publicly available via web. Ensembl tools include BLAST, BLAT, BioMart and the Variant Effect Predictor (VEP) for all supported species.
Proper citation: Ensembl (RRID:SCR_002344) Copy
http://fged.org/projects/miame/
Standard specification for the Minimum Information About a Microarray Experiment that is needed to enable the interpretation of the results of the experiment unambiguously and potentially to reproduce the experiment.
Proper citation: MIAME (RRID:SCR_002349) Copy
Web service that tags gene, protein, and small molecule names in any web page. Clicking on a tagged term opens a small popup showing summary information, and allows the user to quickly link to more detailed information. For each protein or gene, Reflect provides domain structure, sub-cellular localization, 3D structure, and interaction partners. For small molecules, it provides the chemical structure and interaction partners. Reflect can be installed as a plugin to Firefox or Internet Explorer, or can be used by entering a URL in the field provided. It can also be accessed programmatically via a REST or SOAP API, and a Reflect button can easily be added to any web page using Javascript or using a CGI proxy. Reflect was first-prize winner out of over 70 submissions in the Elsevier Grand Challenge, an international competition for systems that improve the way scientific information is communicated and used. Reflect can be edited and improved by the community.
Proper citation: Reflect (RRID:SCR_002714) Copy
http://www.ebi.ac.uk/Tools/msa/clustalw2/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on January 19, 2022. Command line version of multiple sequence alignment program Clustal for DNA or proteins. Alignment is progressive and considers sequence redundancy. No longer being maintained. Please consider using Clustal Omega instead which accepts nucleic acid or protein sequences in multiple sequence formats NBRF/PIR, EMBL/UniProt, Pearson (FASTA), GDE, ALN/ClustalW, GCG/MSF, RSF.
Proper citation: Clustal W2 (RRID:SCR_002909) Copy
http://www.bioinformatics.org/peakanalyzer/wiki/
A set of standalone software programs for the automated processing of any genomic loci, with an emphasis on datasets consisting of ChIP-derived signal peaks. The software is able to identify individual binding / modification sites from enrichment loci, retrieve peak region sequences for motif discovery, and integrate experimental data with different classes of annotated elements throughout the genome. PeakAnalyzer requires a peak file and a feature annotation file in BED or GTF format. Complete annotation files for the current builds of the human (HG19) and mouse (MM9) genomes are provided with the software distribution.
Proper citation: PeakAnalyzer (RRID:SCR_001194) Copy
https://www.ebi.ac.uk/jdispatcher/msa/clustalo?stype=protein
Software package as multiple sequence alignment tool that uses seeded guide trees and HMM profile-profile techniques to generate alignments between three or more sequences. Accepts nucleic acid or protein sequences in multiple sequence formats NBRF/PIR, EMBL/UniProt, Pearson (FASTA), GDE, ALN/Clustal, GCG/MSF, RSF.
Proper citation: Clustal Omega (RRID:SCR_001591) Copy
http://www.bioconductor.org/packages/release/bioc/html/ArrayExpress.html
Software to access the ArrayExpress Repository at EBI and build Bioconductor data structures: ExpressionSet, AffyBatch, NChannelSet
Proper citation: ArrayExpress (R) (RRID:SCR_000120) Copy
Open source database system and analysis tools for molecular interaction data. All interactions are derived from literature curation or direct user submissions. Direct user submissions of molecular interaction data are encouraged, which may be deposited prior to publication in a peer-reviewed journal. The IntAct Database contains (Jun. 2014): * 447368 Interactions * 33021 experiments * 12698 publications * 82745 Interactors IntAct provides a two-tiered view of the interaction data. The search interface allows the user to iteratively develop complex queries, exploiting the detailed annotation with hierarchical controlled vocabularies. Results are provided at any stage in a simplified, tabular view. Specialized views then allows "zooming in" on the full annotation of interactions, interactors and their properties. IntAct source code and data are freely available.
Proper citation: IntAct (RRID:SCR_006944) Copy
Public archive providing a comprehensive record of the world''''s nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation. All submitted data, once public, will be exchanged with the NCBI and DDBJ as part of the INSDC data exchange agreement. The European Nucleotide Archive (ENA) captures and presents information relating to experimental workflows that are based around nucleotide sequencing. A typical workflow includes the isolation and preparation of material for sequencing, a run of a sequencing machine in which sequencing data are produced and a subsequent bioinformatic analysis pipeline. ENA records this information in a data model that covers input information (sample, experimental setup, machine configuration), output machine data (sequence traces, reads and quality scores) and interpreted information (assembly, mapping, functional annotation). Data arrive at ENA from a variety of sources including submissions of raw data, assembled sequences and annotation from small-scale sequencing efforts, data provision from the major European sequencing centers and routine and comprehensive exchange with their partners in the International Nucleotide Sequence Database Collaboration (INSDC). Provision of nucleotide sequence data to ENA or its INSDC partners has become a central and mandatory step in the dissemination of research findings to the scientific community. ENA works with publishers of scientific literature and funding bodies to ensure compliance with these principles and to provide optimal submission systems and data access tools that work seamlessly with the published literature. ENA is made up of a number of distinct databases that includes the EMBL Nucleotide Sequence Database (Embl-Bank), the newly established Sequence Read Archive (SRA) and the Trace Archive. The main tool for downloading ENA data is the ENA Browser, which is available through REST URLs for easy programmatic use. All ENA data are available through the ENA Browser. Note: EMBL Nucleotide Sequence Database (EMBL-Bank) is entirely included within this resource.
Proper citation: European Nucleotide Archive (ENA) (RRID:SCR_006515) Copy
Service providing functional analysis of proteins by classifying them into families and predicting domains and important sites. They combine protein signatures from a number of member databases into a single searchable resource, capitalizing on their individual strengths to produce a powerful integrated database and diagnostic tool. This integrated database of predictive protein signatures is used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures. You can access the data programmatically, via Web Services. The member databases use a number of approaches: # ProDom: provider of sequence-clusters built from UniProtKB using PSI-BLAST. # PROSITE patterns: provider of simple regular expressions. # PROSITE and HAMAP profiles: provide sequence matrices. # PRINTS provider of fingerprints, which are groups of aligned, un-weighted Position Specific Sequence Matrices (PSSMs). # PANTHER, PIRSF, Pfam, SMART, TIGRFAMs, Gene3D and SUPERFAMILY: are providers of hidden Markov models (HMMs). Your contributions are welcome. You are encouraged to use the ''''Add your annotation'''' button on InterPro entry pages to suggest updated or improved annotation for individual InterPro entries.
Proper citation: InterPro (RRID:SCR_006695) Copy
International collaboration producing an extensive public catalog of human genetic variation, including SNPs and structural variants, and their haplotype contexts, in an effort to provide a foundation for investigating the relationship between genotype and phenotype. The genomes of about 2500 unidentified people from about 25 populations around the world were sequenced using next-generation sequencing technologies. Redundant sequencing on various platforms and by different groups of scientists of the same samples can be compared. The results of the study are freely and publicly accessible to researchers worldwide. The consortium identified the following populations whose DNA will be sequenced: Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States. The goal Project is to find most genetic variants that have frequencies of at least 1% in the populations studied. Sequencing is still too expensive to deeply sequence the many samples being studied for this project. However, any particular region of the genome generally contains a limited number of haplotypes. Data can be combined across many samples to allow efficient detection of most of the variants in a region. The Project currently plans to sequence each sample to about 4X coverage; at this depth sequencing cannot provide the complete genotype of each sample, but should allow the detection of most variants with frequencies as low as 1%. Combining the data from 2500 samples should allow highly accurate estimation (imputation) of the variants and genotypes for each sample that were not seen directly by the light sequencing. All samples from the 1000 genomes are available as lymphoblastoid cell lines (LCLs) and LCL derived DNA from the Coriell Cell Repository as part of the NHGRI Catalog. The sequence and alignment data generated by the 1000genomes project is made available as quickly as possible via their mirrored ftp sites. ftp://ftp.1000genomes.ebi.ac.uk ftp://ftp-trace.ncbi.nlm.nih.gov/1000genomes
Proper citation: 1000 Genomes: A Deep Catalog of Human Genetic Variation (RRID:SCR_006828) Copy
BioPerl is a community effort to produce Perl code which is useful in biology. This toolkit of perl modules is useful in building bioinformatics solutions in Perl. It is built in an object-oriented manner so that many modules depend on each other to achieve a task. The collection of modules in the bioperl-live repository consist of the core of the functionality of bioperl. Additionally auxiliary modules for creating graphical interfaces (bioperl-gui), persistent storage in RDMBS (bioperl-db), running and parsing the results from hundreds of bioinformatics applications (Run package), software to automate bioinformatic analyses (bioperl-pipeline) are all available as Git modules in our repository. The BioPerl toolkit provides a library of hundreds of routines for processing sequence, annotation, alignment, and sequence analysis reports. It often serves as a bridge between different computational biology applications assisting the user to construct analysis pipelines. This chapter illustrates how BioPerl facilitates tasks such as writing scripts summarizing information from BLAST reports or extracting key annotation details from a GenBank sequence record. BioPerl includes modules written by Sohel Merchant of the GO Consortium for parsing and manipulating OBO ontologies. Platform: Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible
Proper citation: BioPerl (RRID:SCR_002989) Copy
http://www.ebi.ac.uk/research/enright/software/kraken
A set of software tools ( Reaper, Tally and Sequence Imp) designed to streamline the analysis of next-generation sequencing data. Although designed with small RNA sequence analysis in mind the tools can be used to address issues facing next-generation sequencing in general.
Proper citation: Kraken (RRID:SCR_005484) Copy
http://www.bioconductor.org/packages/2.14/bioc/html/h5vc.html
Software package that contains functions to interact with tally data from Next-Generation Sequencing (NGS) experiments that is stored in HDF5 files.
Proper citation: h5vc (RRID:SCR_006039) Copy
http://www.ebi.ac.uk/ena/about/cram_toolkit
A framework technology comprising file format and toolkit in which we combine highly efficient and tunable reference-based compression of sequence data with a data format that is directly available for computational use.
Proper citation: CRAM (RRID:SCR_012975) Copy
http://www.ebi.ac.uk/Rebholz-srv/ebimed/
A web application that combines Information Retrieval and Extraction from Medline. EBIMed finds Medline abstracts in the same way PubMed does. Then it goes a step beyond and analyses them to offer a complete overview on associations between UniProt protein/gene names, GO annotations, Drugs and Species. The results are shown in a table that displays all the associations and links to the sentences that support them and to the original abstracts. By selecting relevant sentences and highlighting the biomedical terminology EBIMed enhances your ability to acquire knowledge, relate facts, discover implications and, overall, have a good overview economizing the effort in reading.
Proper citation: EBIMed (RRID:SCR_005314) Copy
http://www.ebi.ac.uk/Tools/msa/kalign/
A fast and accurate multiple sequence alignment algorithm.
Proper citation: Kalign (RRID:SCR_011810) Copy
http://www.ebi.ac.uk/enright-srv/microcosm/htdocs/targets/v5/
Database of computationally predicted targets for microRNAs across many species.
Proper citation: MicroCosm Targets (RRID:SCR_010846) Copy
http://www.mged.org/Workgroups/MAGE/mage.html
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on July 27,2023. Group providing a standard for the representation of microarray expression data that would facilitate the exchange of microarray information between different data systems.
Proper citation: MAGE (RRID:SCR_002313) Copy
http://www.ebi.ac.uk/Tools/sss/fasta/
Software package for DNA and protein sequence alignment to find regions of local or global similarity between Protein or DNA sequences, either by searching Protein or DNA databases, or by identifying local duplications within a sequence.
Proper citation: FASTA (RRID:SCR_011819) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.