Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
https://github.com/vgteam/vg#vg
Software toolkit to improve read mapping by representing genetic variation in reference.Provides succinct encoding of sequences of many genomes.
Proper citation: variation graph (RRID:SCR_024369) Copy
https://cab.spbu.ru/software/spades/
Software package for assembling single cell genomes and mini metagenomes. Uses short read sets as input. Used for genomes of uncultivatable bacteria that vastly exceeds what may be obtained via traditional metagenomics studies. Works with Illumina or IonTorrent reads and can provide hybrid assemblies using PacBio, Oxford Nanopore and Sanger reads. Intended for small genomes like bacterial or fungal., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: SPAdes (RRID:SCR_000131) Copy
https://cell-innovation.nig.ac.jp/maser/Tools/visualization_top_en.html
One stop platform for NGS big data from analysis to visualization. There are about 400 analysis pipelines integrated on Maser. List of all analysis pipelines, including descriptions and approximate execution times, can be found on page for ‘All pipelines’ in the User Guide.. Regist custom genome software registers custom genomes to Genome Explorer (IN: FASTA).
Proper citation: regist custom genome (RRID:SCR_015999) Copy
http://noble.gs.washington.edu/proj/genomedata/
A format for efficient storage of multiple tracks of numeric data anchored to a genome. The format allows fast random access to hundreds of gigabytes of data, while retaining a small disk space footprint. They have also developed utilities to load data into this format. Retrieving data from this format is more than 2900 times faster than a naive approach using wiggle files. A reference implementation in Python and C components is available here under the GNU General Public License. The software has only been tested on Linux and Mac systems.
Proper citation: Genomedata (RRID:SCR_004544) Copy
https://github.com/gatech-genemark/ProtHint
Software pipeline for predicting and scoring hints (in form of introns, start and stop codons) in genome of interest by mapping and spliced aligning predicted genes to database of reference protein sequences.
Proper citation: ProtHint (RRID:SCR_021167) Copy
https://github.com/MicrosoftGenomics/FaST-LMM
FaST-LMM (Factored Spectrally Transformed Linear Mixed Models) is a set of tools for efficiently performing genome-wide association studies (GWAS), prediction, and heritability estimation on large data sets.
Proper citation: FaST LMM (RRID:SCR_015506) Copy
https://github.com/HMPNK/CSA2.6
Software pipeline for high-throughput chromosome level vertebrate genome assembly. Pipeline, which after contig assembly performs post assembly improvements by ordering assembly and closing gaps, as well as splitting of low supported regions.
Proper citation: Chromosome Scale Assembler (RRID:SCR_017960) Copy
https://github.com/BackofenLab/HVSeeker/tree/main
Software tool for distinguishing between bacterial and phage sequences. Consists of two separate models: one analyzing DNA sequences and the other focusing on proteins.
Proper citation: HVSeeker (RRID:SCR_026120) Copy
A C++ application designed for compression of genome collections from the same species.
Proper citation: GDC (RRID:SCR_001007) Copy
https://www.mc.vanderbilt.edu/victr/dcc/projects/acc/index.php/Main_Page
A national consortium formed to develop, disseminate, and apply approaches to research that combine DNA biorepositories with electronic medical record (EMR) systems for large-scale, high-throughput genetic research. The consortium is composed of seven member sites exploring the ability and feasibility of using EMR systems to investigate gene-disease relationships. Themes of bioinformatics, genomic medicine, privacy and community engagement are of particular relevance to eMERGE. The consortium uses data from the EMR clinical systems that represent actual health care events and focuses on ethical issues such as privacy, confidentiality, and interactions with the broader community.
Proper citation: eMERGE Network: electronic Medical Records and Genomics (RRID:SCR_007428) Copy
http://genome.wustl.edu/projects/detail/human-gut-microbiome/
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on August 19,2022. Human Gut Microbiome Initiative (HGMI) seeks to provide simply annotated, deep draft genome sequences for 100 cultured representatives of the phylogenetic diversity documented by 16S rRNA surveys of the human gut microbiota. Humans are supra-organisms, composed of 10 times more microbial cells than human cells. Therefore, it seems appropriate to consider ourselves as a composite of many species - human, bacterial, and archaeal - and our genome as an amalgamation of human genes and the genes in ''our'' microbial genomes (''microbiome''). In the same sense, our metabolome can be considered to be a synthesis of co-evolved human and microbial traits. The total number of genes present in the human microbiome likely exceeds the number of our H. sapiens genes by orders of magnitude. Thus, without an understanding of our microbiota and microbiome, it not possible to obtain a complete picture of our genetic diversity and of our normal physiology. Our intestine is home to our largest collections of microbes: bacterial densities in the colon (up to 1 trillion cells/ml of luminal contents) are the highest recorded for any known ecosystem. The vast majority of phylogenetic types in the distal gut microbiota belong to just two divisions (phyla) of the domain Bacteria - the Bacteroidetes and the Firmicutes. Members of eight other divisions have also been identified using culture-independent 16S rRNA gene-based surveys. Metagenomic studies of complex microbial communities residing in our various body habitats are limited by the availability of suitable reference genomes for confident assignment of short sequence reads generated by highly parallel DNA sequencers, and by knowledge of the professions (niches) of community members. Therefore, HGMI, which represents a collaboration between Washington University''s Genome Center and its Center for Genome Sciences, seeks to provide simply annotated, deep draft genome sequences for 100 cultured representatives of the phylogenetic diversity documented by 16S rRNA surveys of the human gut microbiota.
Proper citation: Human Gut Microbiome Initiative (RRID:SCR_008137) Copy
http://www.gene-regulation.com/pub/databases.html
In an effort to strongly support the collaborative nature of scientific research, BIOBASE offers academic and non-profit organizations free access to reduced functionality versions of their products. TRANSFAC Professional provides gene regulation analysis solutions, offering the most comprehensive collection of eukaryotic gene regulation data. The professional paid subscription gives customers access to up-to-date data and tools not available in the free version. The public databases currently available for academic and non-profit organizations are: * TRANSFAC: contains data on transcription factors, their experimentally-proven binding sites, and regulated genes. Its broad compilation of binding sites allows the derivation of positional weight matrices. * TRANSPATH: provides data about molecules participating in signal transduction pathways and the reactions they are involved in, resulting in a complex network of interconnected signaling components.TRANSPATH focuses on signaling cascades that change the activities of transcription factors and thus alter the gene expression profile of a given cell. * PathoDB: is a database on pathologically relevant mutated forms of transcription factors and their binding sites. It comprises numerous cases of defective transcription factors or mutated transcription factor binding sites, which are known to cause pathological defects. * S/MARt DB: presents data on scaffold or matrix attached regions (S/MARs) of eukaryotic genomes, as well as about the proteins that bind to them. S/MARs organize the chromatin in the form of functionally independent loop domains gained increasing support. Scaffold or Matrix Attached Regions (S/MARs) are genomic DNA sequences through which the chromatin is tightly attached to the proteinaceous scaffold of the nucleus. * TRANSCompel: is a database on composite regulatory elements affecting gene transcription in eukaryotes. Composite regulatory elements consist of two closely situated binding sites for distinct transcription factors, and provide cross-coupling of different signaling pathways. * PathoSign Public: is a database which collects information about defective cell signaling molecules causing human diseases. While constituting a useful data repository in itself, PathoSign is also aimed at being a foundational part of a platform for modeling human disease processes.
Proper citation: Gene Regulation Databases (RRID:SCR_008033) Copy
http://ncv.unl.edu/Angelettilab/HPV/Database.html
THIS RESOURCE IS NO LONGER IN SERVICE, documented May 10, 2017. A pilot effort that has developed a centralized, web-based biospecimen locator that presents biospecimens collected and stored at participating Arizona hospitals and biospecimen banks, which are available for acquisition and use by researchers. Researchers may use this site to browse, search and request biospecimens to use in qualified studies. The development of the ABL was guided by the Arizona Biospecimen Consortium (ABC), a consortium of hospitals and medical centers in the Phoenix area, and is now being piloted by this Consortium under the direction of ABRC. You may browse by type (cells, fluid, molecular, tissue) or disease. Common data elements decided by the ABC Standards Committee, based on data elements on the National Cancer Institute''s (NCI''s) Common Biorepository Model (CBM), are displayed. These describe the minimum set of data elements that the NCI determined were most important for a researcher to see about a biospecimen. The ABL currently does not display information on whether or not clinical data is available to accompany the biospecimens. However, a requester has the ability to solicit clinical data in the request. Once a request is approved, the biospecimen provider will contact the requester to discuss the request (and the requester''s questions) before finalizing the invoice and shipment. The ABL is available to the public to browse. In order to request biospecimens from the ABL, the researcher will be required to submit the requested required information. Upon submission of the information, shipment of the requested biospecimen(s) will be dependent on the scientific and institutional review approval. Account required. Registration is open to everyone., documented August 23, 2016. The Human Papillomaviruses Database collects, curates, analyzes, and publishes genetic sequences of papillomaviruses and related cellular proteins. It includes molecular biologists, sequence analysts, computer technicians, post-docs and graduate research assistants. This Web site has two main branches. The first contains our four annual data books of papillomavirus information, called Human Papillomaviruses: A Compilation and Analysis of Nucleic Acid and Amino Acid Sequences. and the second contains papillomavirus genetic sequence data. There is also a New Items location where we store the latest changes to the database or any other current news of interest. Besides the compendium, we also provide genetic sequence information for papilloma viruses and related cellular proteins. Each year they publish a compendium of papillomavirus information called Human Papillomaviruses: A Compilation and Analysis of Nucleic Acid and Amino Acid Sequences. which can now be downloaded from this Web site.
Proper citation: HPV Sequence Database (RRID:SCR_008154) Copy
http://www.animalgenome.org/pigs/nagrp.html
Database and resources on the pig genome.
Proper citation: U.S. Pig Genome Project (RRID:SCR_008151) Copy
The project began as a pilot study to identify inherited genetic susceptibility to prostate and breast cancer. CGEMS has developed into a robust research program involving genome-wide association studies (GWASs) for a number of cancers to identify common genetic variants that affect a person''s risk of developing cancer. In collaboration with extramural scientists, NCI''s Division of Cancer Epidemiology and Genetics (DCEG) has carried out genome-wide scans for breast, prostate, pancreatic, and lung cancers, while a GWAS of bladder cancer is currently underway. By making the data available to both intramural and extramural research scientists, as well as those in the private sector through rapid posting, NIH can leverage its resources to ensure that the dramatic advances in genomics are incorporated into rigorous population-based studies. Ultimately, findings from these studies may yield new preventive, diagnostic, and therapeutic interventions for cancer. Sponsors: This resource is supported by the U.S. National Institues Of Health.
Proper citation: CGEMS (RRID:SCR_008445) Copy
Genome wide map of putative transcription factor binding sites in Arabidopsis thaliana genome.Data in AthaMap is based on published transcription factor (TF) binding specificities available as alignment matrices or experimentally determined single binding sites.Integrated transcriptional and post transcriptional data.Provides web tools for analysis and identification of co-regulated genes. Provides web tools for database assisted identification of combinatorial cis-regulatory elements and the display of highly conserved transcription factor binding sites in Arabidopsis thaliana.
Proper citation: AthaMap (RRID:SCR_006717) Copy
A database and interactive web site for manipulating and displaying annotations on genomes. Features include: detailed views of the genome; use of a variety of premade or personally made glyphs ; customizable order and appearance of tracks by administrators and end-users; search by annotation ID, name, or comment; support of third party annotation using GFF formats; DNA and GFF dumps; connectivity to different databases, including BioSQL and Chado; and a customizable plug-in architecture (e.g. run BLAST, find oligonucleotides, design primers, etc.). GBrowse is distributed as source code for Macintosh OS X, UNIX and Linux platforms, and as pre-packaged binaries for Windows machines. It can be installed using the standard Perl module build procedure, or automated using a network-based install script. In order to use the net installer, you will need to have Perl 5.8.6 or higher and the Apache web server installed. The wiki portion accepts data submissions.
Proper citation: GBrowse (RRID:SCR_006829) Copy
http://bond.unleashedinformatics.com/
THIS RESOURCE IS NO LONGER IN SERVICE, documented May 10, 2017. A pilot effort that has developed a centralized, web-based biospecimen locator that presents biospecimens collected and stored at participating Arizona hospitals and biospecimen banks, which are available for acquisition and use by researchers. Researchers may use this site to browse, search and request biospecimens to use in qualified studies. The development of the ABL was guided by the Arizona Biospecimen Consortium (ABC), a consortium of hospitals and medical centers in the Phoenix area, and is now being piloted by this Consortium under the direction of ABRC. You may browse by type (cells, fluid, molecular, tissue) or disease. Common data elements decided by the ABC Standards Committee, based on data elements on the National Cancer Institute''s (NCI''s) Common Biorepository Model (CBM), are displayed. These describe the minimum set of data elements that the NCI determined were most important for a researcher to see about a biospecimen. The ABL currently does not display information on whether or not clinical data is available to accompany the biospecimens. However, a requester has the ability to solicit clinical data in the request. Once a request is approved, the biospecimen provider will contact the requester to discuss the request (and the requester''s questions) before finalizing the invoice and shipment. The ABL is available to the public to browse. In order to request biospecimens from the ABL, the researcher will be required to submit the requested required information. Upon submission of the information, shipment of the requested biospecimen(s) will be dependent on the scientific and institutional review approval. Account required. Registration is open to everyone.. Documented on August 19,2019.BOND, which requires registration of a free account, is a resource used to perform cross-database searches of available sequence, interaction, complex and pathway information. BOND integrates a range of component databases including GenBank and BIND, the Biomolecular Interaction Network Database. BOND contains 70+ million biological sequences, 33,000 structures, 38,000 GO terms, and over 200,000 human curated interactions contained in BIND, and is open access. BOND serves the interests of the developing global interactome effort encompassing the genomic, proteomic and metabolomic research communities. BOND is the first open access search resource to integrate sequence and interaction information. BOND integrates BLAST functionality, and contains a well-documented API. BOND also stores annotation links for sequences, including links to Genome Ontology descriptions, MedLine abstracts, taxon identifiers, associated structures, redundant sequences, sequence neighbors, conserved domains, data base cross-references, Online Mendalian Inheritance in Man identifiers, LocusLink identifiers and complete genomes. BIND on BOND The Biomolecular Interaction Network Database (BIND), a component database of BOND, is a collection of records documenting molecular interactions. The contents of BIND include high-throughput data submissions and hand-curated information gathered from the scientific literature. BIND is an interaction database with three classifications for molecular associations: molecules that associate with each other to form interactions, molecular complexes that are formed from one or more interaction(s) and pathways that are defined by a specific sequence of two or more interactions.Interactions A BIND record represents an interaction between two or more objects that is believed to occur in a living organism. A biological object can be a protein, DNA, RNA, ligand, molecular complex, gene, photon or an unclassified biological entity. BIND records are created for interactions which have been shown experimentally and published in at least one peer-reviewed journal. A record also references any papers with experimental evidence that support or dispute the associated interaction. Interactions are the basic units of BIND and can be linked together to form molecular complexes or pathways. The BIND interaction viewer is a tool to visualize and analyze molecular interactions, complexes and pathways. The BIND interaction viewer uses Ontoglyphs to display information about a protein via attributes such as molecular function, biological process and sub-cellular localization. Ontoglyphs allow to graphically and interactively explore interaction networks, by visualizing interactions in the context of 34 functional, 25 binding specificity and 24 sub-cellular localization Ontoglyphs categories. We will continue to provide an open access version of BOND, providing its subscribers with free, unlimited access to a core content set. But we are confident you will soon want to upgrade to BONDplus.
Proper citation: Biomolecular Object Network Databank (RRID:SCR_007433) Copy
http://genolist.pasteur.fr/Colibri/
Database dedicated to the analysis of the genome of Escherichia coli. Its purpose is to collate and integrate various aspects of the genomic information from E. coli, the paradigm of Gram-negative bacteria. Colibri provides a complete dataset of DNA and protein sequences derived from the paradigm strain E. coli K-12, linked to the relevant annotations and functional assignments. It allows one to easily browse through these data and retrieve information, using various criteria (gene names, location, keywords, etc.). The data contained in Colibri originates from two major sources of information, the reference genomic DNA sequence from the E. coli Genome Project and the feature annotations from the EcoGene data collection., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: Colibri (RRID:SCR_007606) Copy
http://mips.gsf.de/genre/proj/ustilago/
The MIPS Ustilago maydis Genome Database aims to present information on the molecular structure and functional network of the entirely sequenced, filamentous fungus Ustilago maydis. The underlying sequence is the initial release of the high quality draft sequence of the Broad Institute. The goal of the MIPS database is to provide a comprehensive genome database in the Genome Research Environment in parallel with other fungal genomes to enable in depth fungal comparative analysis. The specific aims are to: 1. Generate and assemble Whole Genome Shotgun sequence reads yielding 10X coverage of the U. maydis genome 2. Integrate the genomic sequence assembly with physical maps generated by Bayer CropScience 3. Perform automated annotation of the sequence assembly 4. Align the strain 521 assembly with the FB1 assembly provided by Exelixis 5. Release the sequence assembly and results of our annotation and analysis to public Ustilago maydis is a basidiomycete fungal pathogen of maize and teosinte. The genome size is approximately 20 Mb. The fungus induces tumors on host plants and forms masses of diploid teliospores. These spores germinate and form haploid meiotic products that can be propagated in culture as yeast-like cells. Haploid strains of opposite mating type fuse and form a filamentous, dikaryotic cell type that invades plant tissue to reinitiate infection. Ustilago maydis is an important model system for studying pathogen-host interactions and has been studied for more than 100 years by plant pathologists. Molecular genetic research with U. maydis focuses on recombination, the role of mating in pathogenesis, and signaling pathways that influence virulence. Recently, the fungus has emerged as an excellent experimental model for the molecular genetic analysis of phytopathogenesis, particularly in the characterization of infection-specific morphogenesis in response to signals from host plants. Ustilago maydis also serves as an important model for other basidiomycete plant pathogens that are more difficult to work with in the laboratory, such as the rust and bunt fungi. Genomic sequence of U. maydis will also be valuable for comparative analysis of other fungal genomes, especially with respect to understanding the host range of fungal phytopathogens. The analysis of U. maydis would provide a framework for studying the hundreds of other Ustilago species that attack important crops, such as barley, wheat, sorghum, and sugarcane. Comparisons would also be possible with other basidiomycete fungi, such as the important human pathogen C. neoformans. Commercially, U. maydis is an excellent model for the discovery of antifungal drugs. In addition, maize tumors caused by U. maydis are prized in Hispanic cuisine and there is interest in improving commercial production. The complete putative gene set of the Broad Institute''s second release is loaded into the database and in addition all deviating putative genes from a putative gene set produced by MIPS with different gene prediction parameters are also loaded. The complete dataset will then be analysed, gene predictions will be manually corrected due to combined information derived from different gene prediction algorithms and, more important, protein and EST comparisons. Gene prediction will be restricted to ORFs larger than 50 codons; smaller ORFs will be included only if similarities to other proteins or EST matches confirm their existence or if a coding region was postulated by all prediction programs used. The resulting proteins will be annotated. They will be classified according to the MIPS classification catalogue receiving appropriate descriptions. All proteins with a known, characterized homolog will be automatically assigned to functional categories using the MIPS functional catalog. All extracted proteins are in addition automatically analysed and annotated by the PEDANT suite.
Proper citation: MIPS Ustilago maydis Database (RRID:SCR_007563) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.