Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://bioafrica.mrc.ac.za/index.html
The BioAfrica HIV-1 Proteomics Resource is a website that contains detailed information about the HIV-1 proteome and protease cleavage sites, as well as data-mining tools that can be used to manipulate and query protein sequence data, a BLAST tool for initiating structural analyses of HIV-1 proteins, and a proteomics tools directory. HIV Proteomics Resource contains information about each HIV-1 gene product in regard to expression, post-transcriptional / post-translational modifications, localization, functional activities, and potential interactions with viral and host macromolecules. The Proteome section contains extensive data on each of 19 HIV-1 proteins, including their functional properties, a sample analysis of HIV-1HXB2, structural models and links to other online resources. The HIV-1 Protease Cleavage Sites section provides information on the position, subtype variation and genetic evolution of Gag, Gag-Pol and Nef cleavage sites.
Proper citation: BioAfrica HIV Informatics in Africa (RRID:SCR_002295) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 14,2026. Integrated database of genomic, expression and protein data for Drosophila, Anopheles, C. elegans and other organisms. You can run flexible queries, export results and analyze lists of data. FlyMine presents data in categories, with each providing information on a particular type of data (for example Gene Expression or Protein Interactions). Template queries, as well as the QueryBuilder itself, allow you to perform searches that span data from more than one category. Advanced users can use a flexible query interface to construct their own data mining queries across the multiple integrated data sources, to modify existing template queries or to create your own template queries. Access our FlyMine data via our Application Programming Interface (API). We provide client libraries in the following languages: Perl, Python, Ruby and & Java API
Proper citation: FlyMine (RRID:SCR_002694) Copy
A package of over twenty mass spectrometry-based tools primarily geared toward proteomic data analysis and database mining. It can be run from the command line, but is primarily used through a web browser, and there is a public website that allows anyone to use the software without local installation. Tandem mass spectrometry analysis tools are used for database searching and identification of peptides, including post-translationally modified peptides and cross-linked peptides. Support for isotope and label-free quantification from this type of data is provided. MS-Viewer software allows sharing and displaying of annotated spectra from many different tandem mass spectrometry data analysis packages. Other tools include software for analyzing peptide mass fingerprinting data (MS-Fit); prediction of theoretical fragmentation of peptides (MS-Product); theoretical chemical or enzymatic digestion of proteins (MS-Digest); and theoretical modeling of the isotope distribution of any chemical, including peptides (MS-Isotope). Searches using amino acid sequence can be used to identify homologous peptides in a database (MS-Pattern); the use of the combination of amino acid sequence and masses can be used for homologous peptide and protein identification using MS-Homology. Tandem mass spectrometry peak list files can be filtered for the presence of certain peaks or neutral losses using MS-Filter. Given a list of proteins, MS-Bridge can report all potential cross-linked peptide combinations of a specified mass. Given a precursor peptide mass and information about known amino acid presence, absence, or modifications, MS-Comp can report all amino acid combinations that could lead to the observed mass.
Proper citation: Protein Prospector (RRID:SCR_014558) Copy
Curated protein-protein and genetic interaction repository of raw protein and genetic interactions from major model organism species, with data compiled through comprehensive curation efforts.
Proper citation: Biological General Repository for Interaction Datasets (BioGRID) (RRID:SCR_007393) Copy
http://www.gene-regulation.com/pub/databases.html
In an effort to strongly support the collaborative nature of scientific research, BIOBASE offers academic and non-profit organizations free access to reduced functionality versions of their products. TRANSFAC Professional provides gene regulation analysis solutions, offering the most comprehensive collection of eukaryotic gene regulation data. The professional paid subscription gives customers access to up-to-date data and tools not available in the free version. The public databases currently available for academic and non-profit organizations are: * TRANSFAC: contains data on transcription factors, their experimentally-proven binding sites, and regulated genes. Its broad compilation of binding sites allows the derivation of positional weight matrices. * TRANSPATH: provides data about molecules participating in signal transduction pathways and the reactions they are involved in, resulting in a complex network of interconnected signaling components.TRANSPATH focuses on signaling cascades that change the activities of transcription factors and thus alter the gene expression profile of a given cell. * PathoDB: is a database on pathologically relevant mutated forms of transcription factors and their binding sites. It comprises numerous cases of defective transcription factors or mutated transcription factor binding sites, which are known to cause pathological defects. * S/MARt DB: presents data on scaffold or matrix attached regions (S/MARs) of eukaryotic genomes, as well as about the proteins that bind to them. S/MARs organize the chromatin in the form of functionally independent loop domains gained increasing support. Scaffold or Matrix Attached Regions (S/MARs) are genomic DNA sequences through which the chromatin is tightly attached to the proteinaceous scaffold of the nucleus. * TRANSCompel: is a database on composite regulatory elements affecting gene transcription in eukaryotes. Composite regulatory elements consist of two closely situated binding sites for distinct transcription factors, and provide cross-coupling of different signaling pathways. * PathoSign Public: is a database which collects information about defective cell signaling molecules causing human diseases. While constituting a useful data repository in itself, PathoSign is also aimed at being a foundational part of a platform for modeling human disease processes.
Proper citation: Gene Regulation Databases (RRID:SCR_008033) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented May 26, 2016. Search engine that integrates over 100 curated and publicly contributed data sources and provides integrated views on the genomic, proteomic, transcriptomic, genetic and functional information currently available. Information featured in the database includes gene function, orthologies, gene expression, pathways and protein-protein interactions, mutations and SNPs, disease relationships, related drugs and compounds.
Proper citation: IntegromeDB (RRID:SCR_004620) Copy
http://www.ebi.ac.uk/biosamples/
Database that aggregates sample information for reference samples (e.g. Coriell Cell lines) and samples for which data exist in one of the EBI''''s assay databases such as ArrayExpress, the European Nucleotide Archive or PRoteomics Identificates DatabasE. It provides links to assays for specific samples, and accepts direct submissions of sample information. The goals of the BioSample Database include: # recording and linking of sample information consistently within EBI databases such as ENA, ArrayExpress and PRIDE; # minimizing data entry efforts for EBI database submitters by enabling submitting sample descriptions once and referencing them later in data submissions to assay databases and # supporting cross database queries by sample characteristics. The database includes a growing set of reference samples, such as cell lines, which are repeatedly used in experiments and can be easily referenced from any database by their accession numbers. Accession numbers for the reference samples will be exchanged with a similar database at NCBI. The samples in the database can be queried by their attributes, such as sample types, disease names or sample providers. A simple tab-delimited format facilitates submissions of sample information to the database, initially via email to biosamples (at) ebi.ac.uk. Current data sources: * European Nucleotide Archive (424,811 samples) * PRIDE (17,001 samples) * ArrayExpress (1,187,884 samples) * ENCODE cell lines (119 samples) * CORIELL cell lines (27,002 samples) * Thousand Genome (2,628 samples) * HapMap (1,417 samples) * IMSR (248,660 samples)
Proper citation: BioSample Database at EBI (RRID:SCR_004856) Copy
http://www.biosino.org/bodyfluid/
A database of bodily fluid proteome data. It contains information on proteins from humanplasma/serum, urine, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, synovial fluid, nipple aspirate fluid, tear fluid, seminal fluid, human milk, and amniotic fluid. Our body fluid protein database, Sys-BodyFluid, contains 11 body fluid proteomes, including plasma/serum, urine, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, synovial fluid, nipple aspirate fluid, tear fluid, seminal fluid, human milk, and amniotic fluid. Over 10,000 proteins are included in the Sys-BodyFluid. These body fluid proteome data come from 50 peer-review publications of different laboratories all over the world. Protein annotation are provided including protein description, Gene ontology, Domain information, Protein sequence and involved pathway. User can access the proteome data by protein name, protein accession number, sequence similarity. In addition, user could perform query cross different body fluids to get more comprehensive understanding. The difference and similarity between these 11 body fluids are also analyzed. Thus , the Sys-BodyFluid database could serve as a reference database for body fluid research and disease proteomics. plasm, serum, urine, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, synovial fluid, nipple aspirate fluid, tear fluid, seminal fluid, human milk, and amniotic fluid, protein, proteomics
Proper citation: Sys-BodyFluid (RRID:SCR_005335) Copy
The SSD has been developed to address the need for resources and tools for understanding large sets of superpositions in order to understand evolutionary relationships and to make predictions of function. We have therefore created the Structure Superposition Database (SSD) for accessing, viewing and understanding large sets of structure superposition data. It contains the results of pairwise, all-by-all superpositions of a representative set of 115 (beta/alpha) barrel structures (TIM barrels). The initial implementation of the SSD contains the results of pairwise, all-by-all superpositions of a representative set of 115 (/alpha)8 barrel structures (TIM barrels). Future plans call for extending the database to include representative structure superpositions for many additional folds. The SSD can be browsed with a user interface module developed as an extension to Chimera, an extensible molecular modeling program. Features of the user interface module facilitate viewing multiple superpositions together.
Proper citation: Structure Superposition Database (RRID:SCR_005236) Copy
A publicly available database of Transposed elements (TEs) which are located within protein-coding genes of 7 organisms: human, mouse, chicken, zebrafish, fruilt fly, nematode and sea squirt. Using TranspoGene the user can learn about the many aspects of the effect these TEs have on their hosting genes, such as: exonization events (including alternative splicing-related data), insertion of TEs into introns, exons, and promoters, specific location of the TE over the gene, evolutionary divergence of the TE from its consensus sequence and involvement in diseases. TranspoGene database is quickly searchable through its website, enables many kinds of searches and is available for download. TranspoGene contains information regarding specific type and family of the TEs, genomic and mRNA location, sequence, supporting transcript accession and alignment to the TE consensus sequence. The database also contains host gene specific data: gene name, genomic location, Swiss-Prot and RefSeq accessions, diseases associated with the gene and splicing pattern. The TranspoGene and microTranspoGene databases can be used by researchers interested in the effect of TE insertion on the eukaryotic transcriptome.
Proper citation: TranspoGene (RRID:SCR_005634) Copy
A knowledgebase of Biochemically, Genetically and Genomically structured genome-scale metabolic network reconstructions. BiGG integrates several published genome-scale metabolic networks into one resource with standard nomenclature which allows components to be compared across different organisms. BiGG can be used to browse model content, visualize metabolic pathway maps, and export SBML files of the models for further analysis by external software packages. Users may follow links from BiGG to several external databases to obtain additional information on genes, proteins, reactions, metabolites and citations of interest.
Proper citation: BiGG Database (RRID:SCR_005809) Copy
http://the_brain.bwh.harvard.edu/uniprobe/
Database that hosts experimental data from universal protein binding microarray (PBM) experiments (Berger et al., 2006) and their accompanying statistical analyses from prokaryotic and eukaryotic organisms, malarial parasites, yeast, worms, mouse, and human. It provides a centralized resource for accessing comprehensive data on the preferences of proteins for all possible sequence variants ("words") of length k ("k-mers"), as well as position weight matrix (PWM) and graphical sequence logo representations of the k-mer data. The database's web tools include a text-based search, a function for assessing motif similarity between user-entered data and database PWMs, and a function for locating putative binding sites along user-entered nucleotide sequences.
Proper citation: UniPROBE (RRID:SCR_005803) Copy
http://indel.bioinfo.sdu.edu.cn/gridsphere/gridsphere
THIS RESOURCE IS NO LONGER IN SERVCE, documented September 2, 2016. Indel Flanking Region Database is an online resource for indels and the flanking regions of proteins in SCOP superfamilies, including amino acid sequences, lengths, locations, secondary structure constitutions, hydrophilicity / hydrophobicity, domain information, 3D structures and so on. It aims at providing a comprehensive dataset for analyzing the qualities of amino acid insertion/deletions(indels), substitutions and the relationship between them. The indels were obtained through the pairwise alignment of homologous structures in SCOP superfamilies. The IndelFR database contains 2,925,017 indels with flanking regions extracted from 373,402 structural alignment pairs of 12,573 non-redundant domains from 1053 superfamilies. IndelFR has already been used for molecular evolution studies and may help to promote future functional studies of indels and their flanking regions.
Proper citation: IndelFR - Indel Flanking Region Database (RRID:SCR_006050) Copy
ProPortal is a database containing genomic, metagenomic, transcriptomic and field data for the marine cyanobacterium Prochlorococcus. Our goal is to provide a source of cross-referenced data across multiple scales of biological organization--from the genome to the ecosystem--embracing the full diversity of ecotypic variation within this microbial taxon, its sister group, Synechococcus and phage that infect them. The site currently contains the genomes of 13 Prochlorococcus strains, 11 Synechococcus strains and 28 cyanophage strains that infect one or both groups. Cyanobacterial and cyanophage genes are clustered into orthologous groups that can be accessed by keyword search or through a genome browser. Users can also identify orthologous gene clusters shared by cyanobacterial and cyanophage genomes. Gene expression data for Prochlorococcus ecotypes MED4 and MIT9313 allow users to identify genes that are up or downregulated in response to environmental stressors. In addition, the transcriptome in synchronized cells grown on a 24-h light-dark cycle reveals the choreography of gene expression in cells in a ''natural'' state. Metagenomic sequences from the Global Ocean Survey from Prochlorococcus, Synechococcus and phage genomes are archived so users can examine the differences between populations from diverse habitats. Finally, an example of cyanobacterial population data from the field is included.
Proper citation: ProPortal (RRID:SCR_006112) Copy
Relational database of all the discovered similar pairs in a huge number of protein-ligand binding sites with annotations of various types (e.g., CATH, SCOP, EC number, Gene ontology). They used a tremendously fast algorithm called SketchSort that enables the enumeration of similar pairs in a huge number of protein-ligand binding sites. They conducted all-pair similarity searches for 3.4 million known and potential binding sites using the proposed method and discovered over 24 million similar pairs of binding sites. PoSSuM enables rapid exploration of similar binding sites among structures with different global folds as well as similar ones. Moreover, PoSSuM is useful for predicting the binding ligand for unbound structures. Basically, the users can search similar binding pockets using two search modes: # Search K is useful for finding similar binding sites for a known ligand-binding site. Post a known ligand-binding site (a pair of PDB ID and HET code) in the PDB, and PoSSuM will search similar sites for the query site. # Search P is useful for predicting ligands that potentially bind to a structure of interest. Post a known protein structure (PDB ID) in the PDB, and PoSSuM will search similar known-ligand binding sites for the query structure.
Proper citation: PoSSuM (RRID:SCR_006109) Copy
http://mint.bio.uniroma2.it/virusmint/
A virus protein interactions database that collects and annotates all the interactions between human and viral proteins and integrates this information in the human protein interaction network. It uses the PSI-MI standard and is fully integrated with the MINT database. You can search for any viral or human protein by entering either common names or database identifiers or display a complete viral interactome.
Proper citation: VirusMINT (RRID:SCR_005987) Copy
The Kabat Database determines the combining site of antibodies based on the available amino acid sequences. The precise delineation of complementarity determining regions (CDR) of both light and heavy chains provides the first example of how properly aligned sequences can be used to derive structural and functional information of biological macromolecules. The Kabat database now includes nucleotide sequences, sequences of T cell receptors for antigens (TCR), major histocompatibility complex (MHC) class I and II molecules, and other proteins of immunological interest. The Kabat Database searching and analysis tools package is an ASP.NET web-based portal containing lookup tools, sequence matching tools, alignment tools, length distribution tools, positional correlation tools and much more. The searching and analysis tools are custom made for the aligned data sets contained in both the SQL Server and ASCII text flat file formats. The searching and analysis tools may be run on a single PC workstation or in a distributed environment. The analysis tools are written in ASP.NET and C# and are available in Visual Studio .NET 2003/2005/2008 formats. The Kabat Database was initially started in 1970 to determine the combining site of antibodies based on the available amino acid sequences at that time. Bence Jones proteins, mostly from human, were aligned, using the now-known Kabat numbering system, and a quantitative measure, variability, was calculated for every position. Three peaks, at positions 24-34, 50-56 and 89-97, were identified and proposed to form the complementarity determining regions (CDR) of light chains. Subsequently, antibody heavy chain amino acid sequences were also aligned using a different numbering system, since the locations of their CDRs (31-35B, 50-65 and 95-102) are different from those of the light chains. CDRL1 starts right after the first invariant Cys 23 of light chains, while CDRH1 is eight amino acid residues away from the first invariant Cys 22 of heavy chains. During the past 30 years, the Kabat database has grown to include nucleotide sequences, sequences of T cell receptors for antigens (TCR), major histocompatibility complex (MHC) class I and II molecules and other proteins of immunological interest. It has been used extensively by immunologists to derive useful structural and functional information from the primary sequences of these proteins.
Proper citation: Kabat Database of Sequences of Proteins of Immunological Interest (RRID:SCR_006465) Copy
http://tardis.nibio.go.jp/homstrad/
A curated database of structure-based alignments for homologous protein families. All known protein structure are clustered into homologous families (i.e., common ancestry), and the sequences of representative members of each family are aligned on the basis of their 3D structures using the programs MNYFIT, STAMP and COMPARER. These structure-based alignments are annotated with JOY and examined individually.
Proper citation: HOMSTRAD - Homologous Structure Alignment Database (RRID:SCR_006544) Copy
Public global Protein Data Bank archive of macromolecular structural data overseen by organizations that act as deposition, data processing and distribution centers for PDB data. Members are: RCSB PDB (USA), PDBe (Europe) and PDBj (Japan), and BMRB (USA). This site provides information about services provided by individual member organizations and about projects undertaken by wwPDB. Data available via websites of its member organizations.
Proper citation: Worldwide Protein Data Bank (wwPDB) (RRID:SCR_006555) Copy
http://chemistry.st-andrews.ac.uk/staff/jbom/group/databases.html
It is a publicly available web-based database that aims to provide further understanding of protein-ligand interactions. It''s a resource containing biomolecular data, including binding energies, Tanimoto ligand similarity scores and protein sequence similarities of protein-ligand complexes. The PLD contains biomolecular data including calculated binding energies, Tanimoto ligand similarity scores and protein percentage sequence similarities. The database has potential for application as a tool in molecular design.
Proper citation: Protein Ligand Database (RRID:SCR_006980) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.