Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
This project encompasses development of novel biological network analysis methods and infrastructure for querying biological data in a semantically-enabled format, and aims to create a semantic interactome model. Research within the BioMANTA project will focus on computational modelling and analysis, primarily using Semantic Web technologies and Machine Learning methods, of large-scale protein-protein interaction and compound activity networks across a wide variety of species. A range of information such as kinetic activity, tissue expression, and subcellular localization and disease state attributes will be included in the resulting data model. Protein interactions are a fundamental component of biological processes. Many proteins are functional only in multimeric complexes, or require interaction partners to achieve their correct localisation or function. For this reason, the study of protein-protein interaction (PPI) networks has become an area of growing interest in computational biology. Through the use of Semantic Web technologies such as Resource Description Framework (RDF) and Web Ontology Language (OWL), interaction data is modelled to create a knowledge representation in which meaning is vested in the ontology rather than instances of data. Stochastic and computational intelligence methods are applied to this data to infer high coverage networks. Semantic inferencing is used to infer previously unknown and meaningful pathways. Major project components: - The BioMANTA Ontology:- An OWL DL ontology incorporating the PSI-MI Ontology, the NCBI Taxonomy, and elements of BioPax ontology and Gene Ontology (describing subcellular localisation). This allows us to re-use existing ontologies, thereby reducing overheads associated with knowledge acquisition in the ontology development process. We are able to integrate existing public data that contain annotation in these formats. - Data conversion & semantic protein integration:- A set of software components that convert protein-protein databases (DIP, MPact, IntAct, etc.) from PSI-MI XML to RDF compliant with the BioMANTA ontology. These software allow us to make these protein-protein interaction datasets (and more generally, any PSI-MI XML data) semantically available for querying and inference within BioMANTA. - A RDF triple store based on RDF Molecules and the MapReduce architecture:- A proof-of-concept RDF triple store using RDF molecules and Hadoop scale-out architectures. Regular RDF graphs are deconstructed into RDF molecules, which are distributed over distributed compute nodes in the MapReduce architecture, and are subsequently combined to form equivalent RDF graphs. Such an approach makes the distributed SPARQL querying and reasoning on RDF triple stores possible. - A quantitative framework to integrate networks extracted from independent data sources (gene expression, subcellular localization, and ortholog mapping):- The model is multi-layer, with a first layer based on Decision Trees where each Decision tree is built on each dataset independently. The tree nodes are cut using Shannon''s entropy (mutual information); the decision of these independent trees is integrated using logistic regression, and the parameters are optimised using maximum likelihood. Sponsors: This resource is supported by the Pfizer Global Research and Development, the Institute for Molecular Bioscience (IMB), and the University of Queensland, Australia.
Proper citation: BioMANTA (RRID:SCR_007177) Copy
Center that acquires, maintains, and distributes genetic stocks and information about stocks of the small free-living nematode Caenorhabditis elegans for use by investigators initiating or continuing research on this genetic model organism. A searchable strain database, general information about C. elegans, and links to key Web sites of use to scientists, including WormBase, WormAtlas, and WormBook are available.
Proper citation: Caenorhabditis Genetics Center (RRID:SCR_007341) Copy
This service offers a gateway to well-benchmarked protein structure and function prediction methods. Structural models collected from the prediction servers are assessed using the powerful 3D-jury consensus approach. The Structure Prediction Meta Server provides access to various fold recognition, function prediction and local structure prediction methods. The Server takes the amino acid sequence of the query protein, the reference name for the prediction job, and the E-mail address as input. The E-mail address is used only for notification about errors during the execution of the job. The query sequence and the reference name are placed in the process queue. The Meta Server accepts only sequences, which have not been submitted before. In case of duplicate sequences the second user will be notified with a link to the previous submission. Sequences longer than 800 amino acids are not accepted by some services. The internal SQL database offers the possibility to find any previous jobs processed by the Meta Server using regular expressions addressing field like E-mail, Job Name and the host name, from which the job was initiated. Each server has its own process queuing system managed by the Meta Server. All results of fold recognition servers are translated into uniform formats. The information extracted from the raw output of the servers includes the PDB codes of the hits, the alignments and the similarity (reliability) scores specific for every server. Mapping of the hits to the SCOP and FSSP classifications are made either using known PDB representatives or alignment of the template sequence with the databases of proteins in both classifications. The secondary structure assignments for all hits are taken from the mapped FSSP (red for helices and blue for strands). Underscored amino acids indicate the first residue after an insertion in the template sequence. The Meta server provides translation of the alignments in standard formats like FASTA, PDB or CASP. The Meta Server is coupled to consensus servers. They provide jury predictions based on the results collected from other services. Not all fold recognition servers are used by the jury system. The data stored on the meta server is available through http://meta.bioinfo.pl/data/JOBID/. Jobs older than 2 months are not shown. The Meta Server is only a set of programs aimed to process and manage biological data, while the predictive power of the service comes from (mostly) remote prediction providers. Sponsors: This resource is supported by The BioInfoBank Institute.
Proper citation: BioInfoBank Meta Server (RRID:SCR_007181) Copy
Portal for Macromolecular X-Ray Crystallography to produce and support an integrated suite of programs that allows researchers to determine macromolecular structures by X-ray crystallography, and other biophysical techniques. Used in the education and training of scientists in experimental structural biology for determination and analysis of protein structure.
Proper citation: CCP4 (RRID:SCR_007255) Copy
https://services.healthtech.dtu.dk/services/NetNGlyc-1.0/
Server that predicts N-Glycosylation sites in human proteins using artificial neural networks that examine the sequence context of Asn-Xaa-Ser/Thr sequons. NetNGlyc 1.0 is also available as a stand-alone software package, with the same functionality as the service above. Ready-to-ship packages exist for the most common UNIX platforms.
Proper citation: NetNGlyc (RRID:SCR_001570) Copy
https://services.healthtech.dtu.dk/services/YinOYang-1.2/
Server that produces neural network predictions for O-beta-GlcNAc attachment sites in eukaryotic protein sequences. This server can also use NetPhos, to mark possible phosphorylated sites and hence identify Yin-Yang sites. YinOYang 1.2 is available as a stand-alone software package, with the same functionality. Ready-to-ship packages exist for the most common UNIX platforms.
Proper citation: YinOYang (RRID:SCR_001605) Copy
https://www.ebi.ac.uk/jdispatcher/msa/clustalo?stype=protein
Software package as multiple sequence alignment tool that uses seeded guide trees and HMM profile-profile techniques to generate alignments between three or more sequences. Accepts nucleic acid or protein sequences in multiple sequence formats NBRF/PIR, EMBL/UniProt, Pearson (FASTA), GDE, ALN/Clustal, GCG/MSF, RSF.
Proper citation: Clustal Omega (RRID:SCR_001591) Copy
Web application to search protein databases using a translated nucleotide query. Translated BLAST services are useful when trying to find homologous proteins to a nucleotide coding region. Blastx compares translational products of the nucleotide query sequence to a protein database. Because blastx translates the query sequence in all six reading frames and provides combined significance statistics for hits to different frames, it is particularly useful when the reading frame of the query sequence is unknown or it contains errors that may lead to frame shifts or other coding errors. Thus blastx is often the first analysis performed with a newly determined nucleotide sequence and is used extensively in analyzing EST sequences. This search is more sensitive than nucleotide blast since the comparison is performed at the protein level.
Proper citation: BLASTX (RRID:SCR_001653) Copy
http://www.ncbi.nlm.nih.gov/projects/homology/maps/
This page provides quick access to the Comparative mapping functions available in the Map Viewer. Currently, comparative maps are calculated using HomoloGene orthology predictions. Once the gene pairs have been established, blocks of conserved syteny can be established using the positions of each gene object in their respective builds. Sponsors: This resource is supported by NCBI.
Proper citation: Homology Maps Page (RRID:SCR_001666) Copy
http://matrixdb.univ-lyon1.fr/
Freely available database focused on interactions established by extracellular proteins and polysaccharides, taking into account the multimeric nature of the extracellular proteins (e.g. collagens, laminins and thrombospondins are multimers). MatrixDB is an active member of the International Molecular Exchange (IMEx) consortium and has adopted the PSI-MI standards for annotating and exchanging interaction data. It includes interaction data extracted from the literature by manual curation, and offers access to relevant data involving extracellular proteins provided by the IMEx partner databases through the PSICQUIC webservice, as well as data from the Human Protein Reference Database. The database reports mammalian protein-protein and protein-carbohydrate interactions involving extracellular molecules. Interactions with lipids and cations are also reported. MatrixDB is focused on mammalian interactions, but aims to integrate interaction datasets of model organisms when available. MatrixDB provides direct links to databases recapitulating mutations in genes encoding extracellular proteins, to UniGene and to the Human Protein Atlas that shows expression and localization of proteins in a large variety of normal human tissues and cells. MatrixDB allows researchers to perform customized queries and to build tissue- and disease-specific interaction networks that can be visualized and analyzed with Cytoscape or Medusa. Statistics (2013): 2283 extracellular matrix interactions including 2095 protein-protein and 169 protein-glycosaminoglycan interactions.
Proper citation: MatrixDB (RRID:SCR_001727) Copy
http://datahub.io/dataset/kupkb
A collection of omics datasets (mRNA, proteins and miRNA) that have been extracted from PubMed and other related renal databases, all related to kidney physiology and pathology giving KUP biologists the means to ask queries across many resources in order to aggregate knowledge that is necessary for answering biological questions. Some microarray raw datasets have also been downloaded from the Gene Expression Omnibus and analyzed by the open-source software GeneArmada. The Semantic Web technologies, together with the background knowledge from the domain's ontologies, allows both rapid conversion and integration of this knowledge base. SPARQL endpoint http://sparql.kupkb.org/sparql The KUPKB Network Explorer will help you visualize the relationships among molecules stored in the KUPKB. A simple spreadsheet template is available for users to submit data to the KUPKB. It aims to capture a minimal amount of information about the experiment and the observations made.
Proper citation: Kidney and Urinary Pathway Knowledge Base (RRID:SCR_001746) Copy
The Physiome Project is a worldwide public domain effort to provide a computational framework for understanding human and other eukaryotic physiology. It aims to develop integrative models at all levels of biological organization, from genes to the whole organism via gene regulatory networks, protein pathways, integrative cell function, and tissue and whole organ structure/function relations. Additionally, an important goal of the project is to develop applications for teaching physiology. Current projects include the development of: - ontologies to organize biological knowledge and access to databases - markup languages to encode models of biological structure and function in a standard format for sharing between different application programs and for re-use as components of more comprehensive models - databases of structure at the cell, tissue and organ levels - software to render computational models of cell function such as ion channel electrophysiology, cell signaling and metabolic pathways, transport, motility, the cell cycle, etc. in 2 & 3D graphical form - software for displaying and interacting with the organ models which will allow the user to move across all spatial scales Sponsors: This project is supported by the International Union of Physiological Sciences (IUPS), the IEEE Engineering. in Medicine and Biology (EMBS), and the International Federation for Medical and Biological Engineering (IFMBE)
Proper citation: International Union of Physiological Sciences: Physiome Project (RRID:SCR_001760) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025. Bioinformatics resource system including web server and web service for functional annotation and enrichment analyses of gene lists. Consists of comprehensive knowledgebase and set of functional analysis tools. Includes gene centered database integrating heterogeneous gene annotation resources to facilitate high throughput gene functional analysis.
Proper citation: DAVID (RRID:SCR_001881) Copy
http://learn.genetics.utah.edu/
Educational resources that provide accurate and unbiased information about topics in genetics, bioscience and health for global and local audiences. They are jargon-free, target multiple learning styles, and often convey concepts through animation and interactivity. The Genetic Science Learning Center is a science and health education program located in the midst of the bioscience research being carried out at the University of Utah. Our mission is making science easy for everyone to understand. * Two websites, available free of charge to Internet users worldwide: ** Learn.Genetics delivers educational materials on genetics, bioscience and health topics. They are designed to be used by students, teachers and members of the public. The materials meet selected US education standards for science and health. ** Teach.Genetics provides resources for K-12 teachers, higher education faculty, and public educators. These include PDF-based Print-and-Go™ activities, unit plans and other supporting resources. The materials are designed to support and extend the materials on Learn.Genetics. *Professional development programs that update K-16 teachers' expertise in bioscience and health topics as well as prepare them to implement the materials on our websites. * Community programs that engage with diverse communities in discussions about genetics and health, and in developing culturally and linguistically-appropriate educational materials. Some topics in genetics and bioscience research are controversial. The Center does not take sides in political or ethical controversies. Rather, our goal is to provide comprehensive information that promotes a lively discussion of these topics, so that individuals can arrive at their own informed decisions.
Proper citation: University of Utah Genetic Science Learning Center - Learn Genetics (RRID:SCR_001910) Copy
http://dynamicbrain.neuroinf.jp/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on January 19. 2022. Platform to promote studies on dynamic principles of brain functions through unifying experimental and computational approaches in cellular, local circuit, global network and behavioral levels. Provides services such as data sets, popular research findings and articles and current developments in field. This site has been archived since FY2019 and is no longer updated.
Proper citation: Dynamic Brain Platform (RRID:SCR_001754) Copy
A manually curated database of both known and predicted metabolic pathways for the laboratory mouse. It has been integrated with genetic and genomic data for the laboratory mouse available from the Mouse Genome Informatics database and with pathway data from other organisms, including human. The database records for 1,060 genes in Mouse Genome Informatics (MGI) are linked directly to 294 pathways with 1,790 compounds and 1,122 enzymatic reactions in MouseCyc. (Aug. 2013) BLAST and other tools are available. The initial focus for the development of MouseCyc is on metabolism and includes such cell level processes as biosynthesis, degradation, energy production, and detoxification. MouseCyc differs from existing pathway databases and software tools because of the extent to which the pathway information in MouseCyc is integrated with the wealth of biological knowledge for the laboratory mouse that is available from the Mouse Genome Informatics (MGI) database.
Proper citation: MouseCyc (RRID:SCR_001791) Copy
Data analysis service that searches PubMed literature database (abstracts) about specific relationships between proteins, genes, or keywords using a NLP-based text-mining approach. The results are returned as a graph. The synonym database used in Chilibot is available, without fee, for academic use only. Several different search methods are supported including: * searching for relationship between two genes, proteins or keywords * searching for relationships between many genes, proteins, or keywords * searching for relationships between two lists of genes, proteins, or keywords Advanced options include: * Automated hypothesis generation (graph) * Restricting context using keywords * Providing your own synonyms * Modifying synonyms provided by Chilibot * Color coding nodes with gene expression values * Special search: modulation
Proper citation: Chilibot: Gene and Protein relationships from MEDLINE (RRID:SCR_001705) Copy
http://www.megabionet.org/atpid/webfile/
Centralized platform to depict and integrate the information pertaining to protein-protein interaction networks, domain architecture, ortholog information and GO annotation in the Arabidopsis thaliana proteome. The Protein-protein interaction pairs are predicted by integrating several methods with the Naive Baysian Classifier. All other related information curated is manually extracted from published literature and other resources from some expert biologists. You are welcomed to upload your PPI or subcellular localization information or report data errors. Arabidopsis proteins is annotated with information (e.g. functional annotation, subcellular localization, tissue-specific expression, phosphorylation information, SNP phenotype and mutant phenotype, etc.) and interaction qualifications (e.g. transcriptional regulation, complex assembly, functional collaboration, etc.) via further literature text mining and integration of other resources. Meanwhile, the related information is vividly displayed to users through a comprehensive and newly developed display and analytical tools. The system allows the construction of tissue-specific interaction networks with display of canonical pathways.
Proper citation: Arabidopsis thaliana Protein Interactome Database (RRID:SCR_001896) Copy
http://www.structuralgenomics.org/
The Structural Genomics Project aims at determination of the 3D structure of all proteins. It also aims to reduce the cost and time required to determine three-dimensional protein structures. It supports selection, registration, and tracking of protein families and representative targets. This aim can be achieved in four steps : -Organize known protein sequences into families. -Select family representatives as targets. -Solve the 3D structure of targets by X-ray crystallography or NMR spectroscopy. -Build models for other proteins by homology to solved 3D structures. PSI has established a high-throughput structure determination pipeline focused on eukaryotic proteins. NMR spectroscopy is an integral part of this pipeline, both as a method for structure determinations and as a means for screening proteins for stable structure. Because computational approaches have estimated that many eukaryotic proteins are highly disordered, about 1 year into the project, CESG began to use an algorithm. The project has been organized into two separate phases. The first phase was dedicated to demonstrating the feasibility of high-throughput structure determination, solving unique protein structures, and preparing for a subsequent production phase. The second phase, PSI-2, has focused on implementing the high-throughput structure determination methods developed in PSI-1, as well as homology modeling and addressing bottlenecks like modeling membrane proteins. The first phase of the Protein Structure Initiative (PSI-1) saw the establishment of nine pilot centers focusing on structural genomics studies of a range of organisms, including Arabidopsis thaliana, Caenorhabditis elegans and Mycobacterium tuberculosis. During this five-year period over 1,100 protein structures were determined, over 700 of which were classified as unique due to their < 30% sequence similarity with other known protein structures. The primary goal of PSI-1 was to develop methods to streamline the structure determination process, resulted in an array of technical advances. Several methods developed during PSI-1 enhanced expression of recombinant proteins in systems like Escherichia coli, Pichia pastoris and insect cell lines. New streamlined approaches to cell cloning, expression and protein purification were also introduced, in which robotics and software platforms were integrated into the protein production pipeline to minimize required manpower, increase speed, and lower costs. The goal of the second phase of the Protein Structure Initiative (PSI-2) is to use methods introduced in PSI-1 to determine a large number of proteins and continue development in streamlining the structural genomics pipeline. Currently, the third phase of the PSI is being developed and will be called PSI: Biology. The consortia will propose work on substantial biological problems that can benefit from the determination of many protein structures Sponsors: PSI is funded by the U.S. National Institute of General Medical Sciences (NIGMS),
Proper citation: Protein Structure Initiative (RRID:SCR_002161) Copy
Database of genetic and molecular biological information about Candida albicans. Contains information about genes and proteins, descriptions and classifications of their biological roles, molecular functions, and subcellular localizations, gene, protein, and chromosome sequence information, tools for analysis and comparison of sequences and links to literature information. Each CGD gene or open reading frame has an individual Locus Page. Genetic loci that are not tied to DNA sequence also have Locus Pages. Provides Gene Ontology, GO, to all its users. Three ontologies that comprise GO (Molecular Function, Cellular Component, and Biological Process) are used by multiple databases to annotate gene products, so that this common vocabulary can be used to compare gene products across species. Development of ontologies is ongoing in order to incorporate new information. Data submissions are welcome.
Proper citation: Candida Genome Database (RRID:SCR_002036) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.