Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
Open source database system and analysis tools for molecular interaction data. All interactions are derived from literature curation or direct user submissions. Direct user submissions of molecular interaction data are encouraged, which may be deposited prior to publication in a peer-reviewed journal. The IntAct Database contains (Jun. 2014): * 447368 Interactions * 33021 experiments * 12698 publications * 82745 Interactors IntAct provides a two-tiered view of the interaction data. The search interface allows the user to iteratively develop complex queries, exploiting the detailed annotation with hierarchical controlled vocabularies. Results are provided at any stage in a simplified, tabular view. Specialized views then allows "zooming in" on the full annotation of interactions, interactors and their properties. IntAct source code and data are freely available.
Proper citation: IntAct (RRID:SCR_006944) Copy
http://www.ebi.ac.uk/thornton-srv/databases/enzymes/
Database of known enzyme structures that have been deposited in the Protein Data Bank (PDB). The enzyme structures are classified by their E.C. number of the ENZYME Data Bank. Browse the classification hierarchy or enter an EC number or search-string. There are currently 45,638 PDB-enzyme entries in the PDB (as at 23 February, 2013) involving 38,109 separate PDB files - some files having more than one E.C. number associated with them.
Proper citation: Enzyme Structures Database (RRID:SCR_007125) Copy
Gene Expression Atlas is a semantically enriched database of meta-analysis based summary statistics over a curated subset of ArrayExpress Archive, servicing queries for condition-specific gene expression patterns as well as broader exploratory searches for biologically interesting genes/samples. The EBI Gene Expression Atlas Blog discusses ideas, features and problems of creating a large scale meta-analytical atlas of gene expression from publicly available microarray data. Atlas REST API provides all the results available in the main web application in a pragmatic, easy to use form - simple HTTP GET queries as input and either JSON or XML formats as output. Gene Expression Atlas goals: 1. Provision of a statistically robust framework for integration of gene expression experiment results across different platforms at a meta-analytical level 2. A simple interface for identifying strong differential expression candidate genes in conditions of interest 3. Integration of ontologies for high quality annotation of gene and sample attributes 4. Construction of new gene expression summarized views, with a view to analysis of putative signaling pathway targets, discovery of correlated gene expression patterns and the identification of condition/tissue-specific patterns of gene expression.
Proper citation: Gene Expression Atlas (RRID:SCR_007989) Copy
http://www.ebi.ac.uk/parasites/parasite-genome.html
This website contains information about the genomic sequence of parasites. It also contains multiple search engines to search six frame translations of parasite nucleotide databases for motifs, parasite protein databases for motifs, and parasite protein databases for keywords and text terms. * Guide to Internet Access to Parasite Genome Information * Guide to web-based analysis tools * Parasite Genome BLAST Server: Search a range of parasite specific nucleotide sequence databases with your own sequence. * Parasite Proteome Keyword Search Facility: Search parasite protein databases for keywords and text terms * Parasite Proteome Motif Search Facility: Search parasite protein databases for motifs * Parasite Six Frame Translation Motif Search Facility: Search six frame translations of parasite nucleotide databases for motifs * Genome computing resources: A list of ftp and gopher sites where genome computing applications and other resources can be found.
Proper citation: Parasite genome databases and genome research resources (RRID:SCR_008150) Copy
http://www.ebi.ac.uk/Tools/webservices/psicquic/registry/registry?action=STATUS
Web service with well defined methods to enable programmatic access to molecular interactions. Standard for computational access to molecular interaction data resources.
Proper citation: PSICQUIC Registry (RRID:SCR_006392) Copy
Public archive providing a comprehensive record of the world''''s nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation. All submitted data, once public, will be exchanged with the NCBI and DDBJ as part of the INSDC data exchange agreement. The European Nucleotide Archive (ENA) captures and presents information relating to experimental workflows that are based around nucleotide sequencing. A typical workflow includes the isolation and preparation of material for sequencing, a run of a sequencing machine in which sequencing data are produced and a subsequent bioinformatic analysis pipeline. ENA records this information in a data model that covers input information (sample, experimental setup, machine configuration), output machine data (sequence traces, reads and quality scores) and interpreted information (assembly, mapping, functional annotation). Data arrive at ENA from a variety of sources including submissions of raw data, assembled sequences and annotation from small-scale sequencing efforts, data provision from the major European sequencing centers and routine and comprehensive exchange with their partners in the International Nucleotide Sequence Database Collaboration (INSDC). Provision of nucleotide sequence data to ENA or its INSDC partners has become a central and mandatory step in the dissemination of research findings to the scientific community. ENA works with publishers of scientific literature and funding bodies to ensure compliance with these principles and to provide optimal submission systems and data access tools that work seamlessly with the published literature. ENA is made up of a number of distinct databases that includes the EMBL Nucleotide Sequence Database (Embl-Bank), the newly established Sequence Read Archive (SRA) and the Trace Archive. The main tool for downloading ENA data is the ENA Browser, which is available through REST URLs for easy programmatic use. All ENA data are available through the ENA Browser. Note: EMBL Nucleotide Sequence Database (EMBL-Bank) is entirely included within this resource.
Proper citation: European Nucleotide Archive (ENA) (RRID:SCR_006515) Copy
Pictorial database of an at-a-glance overview of the contents of each 3D structure deposited in the Protein Data Bank (PDB). It shows the molecule(s) that make up the structure (ie protein chains, DNA, ligands and metal ions) and schematic diagrams of their interactions. Extensive use is made of the freely available RasMol molecular graphics program to view the molecules and their interactions in 3D. Entries are accessed either by their 4-character PDB code, or by one of the two search boxes provided on the PDBsum home page: text search or sequence search. The information given on each PDBsum entry is spread across several pages, as listed below and accessible from the tabs at the top of the page. Only the relevant tabs will be present on any given page. * Top page - summary information including thumbnail image of structure, molecules in structure, enzyme reaction diagram (where relevant), GO functional assignments, and selected figures from key reference * Protein - wiring diagram, topology diagram(s) by CATH domain, and residue conservation (where available) * DNA/RNA - DNA/RNA sequence and NUCPLOT showing interactions made with protein * Ligands - description of bound molecule and LIGPLOT showing interactions made with protein * Prot-prot - schematic diagrams of any protein-protein interfaces and the residue-residue interactions made across them * Clefts - listing of top ten clefts in the surface of the protein, listed by volume with any bound ligands shown * Links - links to external databases Additionally, it accepts users'''' own PDB format files and generates a private set of analyses for each uploaded structure.
Proper citation: PDBsum (RRID:SCR_006511) Copy
Service providing functional analysis of proteins by classifying them into families and predicting domains and important sites. They combine protein signatures from a number of member databases into a single searchable resource, capitalizing on their individual strengths to produce a powerful integrated database and diagnostic tool. This integrated database of predictive protein signatures is used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures. You can access the data programmatically, via Web Services. The member databases use a number of approaches: # ProDom: provider of sequence-clusters built from UniProtKB using PSI-BLAST. # PROSITE patterns: provider of simple regular expressions. # PROSITE and HAMAP profiles: provide sequence matrices. # PRINTS provider of fingerprints, which are groups of aligned, un-weighted Position Specific Sequence Matrices (PSSMs). # PANTHER, PIRSF, Pfam, SMART, TIGRFAMs, Gene3D and SUPERFAMILY: are providers of hidden Markov models (HMMs). Your contributions are welcome. You are encouraged to use the ''''Add your annotation'''' button on InterPro entry pages to suggest updated or improved annotation for individual InterPro entries.
Proper citation: InterPro (RRID:SCR_006695) Copy
International collaboration producing an extensive public catalog of human genetic variation, including SNPs and structural variants, and their haplotype contexts, in an effort to provide a foundation for investigating the relationship between genotype and phenotype. The genomes of about 2500 unidentified people from about 25 populations around the world were sequenced using next-generation sequencing technologies. Redundant sequencing on various platforms and by different groups of scientists of the same samples can be compared. The results of the study are freely and publicly accessible to researchers worldwide. The consortium identified the following populations whose DNA will be sequenced: Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States. The goal Project is to find most genetic variants that have frequencies of at least 1% in the populations studied. Sequencing is still too expensive to deeply sequence the many samples being studied for this project. However, any particular region of the genome generally contains a limited number of haplotypes. Data can be combined across many samples to allow efficient detection of most of the variants in a region. The Project currently plans to sequence each sample to about 4X coverage; at this depth sequencing cannot provide the complete genotype of each sample, but should allow the detection of most variants with frequencies as low as 1%. Combining the data from 2500 samples should allow highly accurate estimation (imputation) of the variants and genotypes for each sample that were not seen directly by the light sequencing. All samples from the 1000 genomes are available as lymphoblastoid cell lines (LCLs) and LCL derived DNA from the Coriell Cell Repository as part of the NHGRI Catalog. The sequence and alignment data generated by the 1000genomes project is made available as quickly as possible via their mirrored ftp sites. ftp://ftp.1000genomes.ebi.ac.uk ftp://ftp-trace.ncbi.nlm.nih.gov/1000genomes
Proper citation: 1000 Genomes: A Deep Catalog of Human Genetic Variation (RRID:SCR_006828) Copy
An open source JavaScript library of components for visualisation of biological data on the web.
Proper citation: BioJS (RRID:SCR_003119) Copy
http://www.ebi.ac.uk/imgt/hla/
Database for sequences of the human major histocompatibility complex (HLA) and includes the official sequences for the WHO Nomenclature Committee For Factors of the HLA System. It currently contains 9,310 allele sequences (2013) along with detailed information concerning the material from which the sequence was derived and data on the validation of the sequences. It is established procedure for authors to submit the sequences directly to the IMGT/HLA Database for checking and assignment of an official name prior to publication, this avoids the problems associated with renaming published sequences and the confusion of multiple names for the same sequence. The need for reasonably rapid publication of new HLA allele sequences has necessitated an annual meeting of the WHO Nomenclature Committee for Factors of the HLA System. Additionally they now publish monthly HLA nomenclature updates both in journals and online to provide quick and easy access to new sequence information. The IMGT/HLA Database is part of the international ImMunoGeneTics project. In collaboration with the Imperial Cancer Research Fund (ICRF) and European Bioinformatics Institute (EBI) they have developed an Oracle database to house the HLA sequences in such a way as to allow users to present complex queries about the sequence, sequence features, references, contacts and allele designations to the database via a graphical user interface over the web. The IMGT/HLA Database Submission Tool allows direct submission of sequences to the WHO HLA Nomenclature Committee for Factors of the HLA System. The IMGT/HLA Database provides an FTP site for the retrieval of sequences in a number of pre-formatted files.
Proper citation: IMGT/HLA (RRID:SCR_002971) Copy
http://www.ebi.ac.uk/Tools/dalilite/indexhtml
Tool that computes optimal and suboptimal structural alignments between two protein structures. It will compare all chains in the first structure against all chains in the second (unless specific chain IDs are given). The resulting superimposed coordinate files can be downloaded or viewed interactively in Jmol. The Dali method optimizes a weighted sum of similarities of intramolecular distances. Suboptimal alignments do not overlap the optimal alignment or each other. Suboptimal alignments detected by the program are reported if the Z-score is above 2; they may be of interest if there are internal repeats in either structure. SOAP Web services are also available.
Proper citation: DaliLite Pairwise comparison of protein structures (RRID:SCR_003047) Copy
BioPerl is a community effort to produce Perl code which is useful in biology. This toolkit of perl modules is useful in building bioinformatics solutions in Perl. It is built in an object-oriented manner so that many modules depend on each other to achieve a task. The collection of modules in the bioperl-live repository consist of the core of the functionality of bioperl. Additionally auxiliary modules for creating graphical interfaces (bioperl-gui), persistent storage in RDMBS (bioperl-db), running and parsing the results from hundreds of bioinformatics applications (Run package), software to automate bioinformatic analyses (bioperl-pipeline) are all available as Git modules in our repository. The BioPerl toolkit provides a library of hundreds of routines for processing sequence, annotation, alignment, and sequence analysis reports. It often serves as a bridge between different computational biology applications assisting the user to construct analysis pipelines. This chapter illustrates how BioPerl facilitates tasks such as writing scripts summarizing information from BLAST reports or extracting key annotation details from a GenBank sequence record. BioPerl includes modules written by Sohel Merchant of the GO Consortium for parsing and manipulating OBO ontologies. Platform: Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible
Proper citation: BioPerl (RRID:SCR_002989) Copy
https://github.com/egonw/semanticchemistry
An ontology that aims to establish a standard in representing chemical information including chemical structure and the ability to richly describe chemical properties, whether intrinsic or computed. It includes terms for the descriptors commonly used in cheminformatics software applications and the algorithms which generate them.
Proper citation: Chemical Information Ontology (RRID:SCR_003290) Copy
An ontology for describing software tools, their types, tasks, versions, provenance and data associated (the input and output data types and the uses the software can be put to).
Proper citation: Software Ontology (RRID:SCR_003493) Copy
http://www.ebi.ac.uk/research/enright/software/kraken
A set of software tools ( Reaper, Tally and Sequence Imp) designed to streamline the analysis of next-generation sequencing data. Although designed with small RNA sequence analysis in mind the tools can be used to address issues facing next-generation sequencing in general.
Proper citation: Kraken (RRID:SCR_005484) Copy
Free access to biomedical literature resources including all of PubMed and PubMed Central, agricultural abstracts (from AGRICOLA), over 4 million international life science patents abstracts, National Health Service (NHS) clinical guidelines, and is supplemented with Chinese Biological Abstracts and the Citeseer database. As well as powerful search of abstracts and full text articles, it also includes: * article citations and sort order based on citation count * data citations mined from full text articles * links to and from related databases and institutional repositories * a tool to create bibliographies linked to your ORCID * named entity recognition of keywords and text-mining-based applications showcased in Europe PMC Labs * Tools for recipients of grants from one of the Europe PMC funders to deposit full-text manuscripts and link them to those specific grants. * Web services for programmatic access to all the above bibliographic information and 50,000 grants. * Search by publication date, relevance, or the number of times an article has been cited. * Links to public databases such as UniProt, Protein Data Bank (PDBe), and the European Nucleotide Archive (ENA) are provided. * Through textmining technologies, you can highlight and browse keywords such as gene names, organisms and diseases. * Search 40,000 biomedical research grants awarded to the 18,000 PIs supported by the Europe PMC funders. * Roadtest new tools based on Europe PMC content in Europe PMC labs. * In Europe PMC plus, PIs supported by the Europe PMC funders can link grants to publication information, view article citation and download statistics, and submit manuscripts.
Proper citation: Europe PubMed Central (RRID:SCR_005901) Copy
http://www.bioconductor.org/packages/2.14/bioc/html/h5vc.html
Software package that contains functions to interact with tally data from Next-Generation Sequencing (NGS) experiments that is stored in HDF5 files.
Proper citation: h5vc (RRID:SCR_006039) Copy
The CREATE consortium represents a core of major European and international mouse database holders and research groups involved in conditional mutagenesis, primarily to develop a strategy for the integration and dissemination of Cre driver strains for modelling aspects of complex human diseases in the mouse. Collectively the participants have amassed a significant number of these strains in their respective databases. Therefore one of the goals of CREATE is to provide a unified portal for worldwide access to these critical resources. The portal can either be searched through an advanced BioMart interface, by driver name, or by anatomical site of expression using Embryonic Mouse Anatomy Project (EMAP) and Mouse Anatomy (MA) ontology terms. Search results link back to the original source of the data for more detailed information and to IMSR to order mice if available. The ontology browser is particularly useful as it enables the CREATE consortium to identify cell and tissues that are not currently covered by existing lines. CREATE also aims to coordinate the production of suitable lines by the Cre generation projects described above. Through the CREATE portal, the CREATE consortium aims to develop a strategy for the production, integration and dissemination of new Cre driver strains for modelling aspects of complex human diseases in the mouse. CREATE is also developing a roadmap for harnessing emerging technologies and methods for improving Cre-mediated recombination in vivo through targeted, intensive workshops and discussion forums on the portal. This will entail review of construct design options for classical transgenic constructs (promoter/enhancer used, small size <2025 Kb) vs large transgenic constructs (BAC, P1, YAC etc.); methods used for Cre transgenic lines including random vs targeted integration, position independent expression loci, or replacement of endogenous coding sequences with Cre recombinase under the control of the endogenous locus. CREATE provides a platform for discussion of additional issues specific to inducible Cre strategies including background activity before induction, inducibility (kinetics), efficiency, and protocols used for induction of Cre recombinase activity. Additional components of the technology roadmap will be the cataloguing of other existing methodologies (rtTA, FLP, Dre) of mouse genome modification, sharing information on validated Cre mutant lines as well as identification and assessment of new methods of mutagenesis such as RNAi and other emerging technologies. Other discussion topics addressed through surveys on the CREATE portal include the characterization of Cre lines (specificity of expression/deletion; efficiency of expression/ deletion; reproducibility of deletion from animal to animal for the same floxed allele; reproducibility with different floxed alleles; timing of expression/deletion, etc.), the extent to which Cre expression changes upon backcrossing to specific genetic backgrounds through variegation and silencing; potential phenotypes caused by either integration- mediated mutagenesis or Cre ''toxicity''; and other factors affecting the specificity of Cre-mediated expression/deletion. CREATE regularly integrates common fields from the Cre-X, CreZOO and the MGI recombinase portal resources described below. The data in common consists of: * Transgene or Knock-in name. * MGI ID of allele. * Driver. * Anatomical site of expression. * Pubmed ID. * IMSR strain name and link. * Inducibility (YES/NO).
Proper citation: CREATE (RRID:SCR_006133) Copy
http://www.ebi.ac.uk/ena/about/cram_toolkit
A framework technology comprising file format and toolkit in which we combine highly efficient and tunable reference-based compression of sequence data with a data format that is directly available for computational use.
Proper citation: CRAM (RRID:SCR_012975) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.