Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
Non profit research organization for genome sequences to advance understanding of biology of humans and pathogens in order to improve human health globally. Provides data which can be translated for diagnostics, treatments or therapies including over 100 finished genomes, which can be downloaded. Data are publicly available on limited basis, and provided more extensively upon request.
Proper citation: Wellcome Trust Sanger Institute; Hinxton; United Kingdom (RRID:SCR_011784) Copy
https://github.com/xavierdidelot/clonalorigin
Software package for comparative analysis of the sequences of a sample of bacterial genomes in order to reconstruct the recombination events that have taken place in their ancestry.
Proper citation: ClonalOrigin (RRID:SCR_016061) Copy
http://www.xavierdidelot.xtreemhost.com/clonalframe.htm
Software package for the inference of bacterial microevolution using multilocus sequence data. It is used to identify the clonal relationships between the members of a sample, while also estimating the chromosomal position of homologous recombination events that have disrupted the clonal inheritance.
Proper citation: Clonalframe (RRID:SCR_016060) Copy
https://github.com/Teichlab/tracer
Software application for recovery of T cell receptor (TCR) data from single cell data. Used to reconstruct full-length, paired T cell receptor (TCR) sequences from T lymphocyte single-cell RNA sequence data. Links T cell specificity with functional response by revealing clonal relationships between cells alongside their transcriptional profiles.
Proper citation: TraCeR (RRID:SCR_016338) Copy
http://www.sanger.ac.uk/resources/software/lookseq/
A web-based application for alignment visualization, browsing and analysis of genome sequence data.
Proper citation: LookSeq (RRID:SCR_005625) Copy
http://www.sanger.ac.uk/resources/software/vagrent/
Software tool set for calculating the biological consequences of genomic variations. The suite of perl modules compares genomic variations with reference genome annotations and generates the possible effects each variant may have on the transcripts it overlaps. It evaluates each variation/transcript combination and describes the effects in the mRNA, CDS and protein sequence contexts. It provides details of the sequence and position of the change within the transcript / protein as well as Sequence Ontology terms to classify its consequences.
Proper citation: VAGrENT (RRID:SCR_005180) Copy
Collection of genome databases for vertebrates and other eukaryotic species with DNA and protein sequence search capabilities. Used to automatically annotate genome, integrate this annotation with other available biological data and make data publicly available via web. Ensembl tools include BLAST, BLAT, BioMart and the Variant Effect Predictor (VEP) for all supported species.
Proper citation: Ensembl (RRID:SCR_002344) Copy
http://www.sanger.ac.uk/science/tools/seqtools
Software for sequence alignments that displays multiple match sequences aligned against a single genomic reference sequence. It can be used for manipulation, display and annotation of genomic data, to check the quality of an alignment, to find missing/misaligned sequence, and to identify splice sites and polyA sites.
Proper citation: Blixem (RRID:SCR_015994) Copy
International collaboration producing an extensive public catalog of human genetic variation, including SNPs and structural variants, and their haplotype contexts, in an effort to provide a foundation for investigating the relationship between genotype and phenotype. The genomes of about 2500 unidentified people from about 25 populations around the world were sequenced using next-generation sequencing technologies. Redundant sequencing on various platforms and by different groups of scientists of the same samples can be compared. The results of the study are freely and publicly accessible to researchers worldwide. The consortium identified the following populations whose DNA will be sequenced: Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States. The goal Project is to find most genetic variants that have frequencies of at least 1% in the populations studied. Sequencing is still too expensive to deeply sequence the many samples being studied for this project. However, any particular region of the genome generally contains a limited number of haplotypes. Data can be combined across many samples to allow efficient detection of most of the variants in a region. The Project currently plans to sequence each sample to about 4X coverage; at this depth sequencing cannot provide the complete genotype of each sample, but should allow the detection of most variants with frequencies as low as 1%. Combining the data from 2500 samples should allow highly accurate estimation (imputation) of the variants and genotypes for each sample that were not seen directly by the light sequencing. All samples from the 1000 genomes are available as lymphoblastoid cell lines (LCLs) and LCL derived DNA from the Coriell Cell Repository as part of the NHGRI Catalog. The sequence and alignment data generated by the 1000genomes project is made available as quickly as possible via their mirrored ftp sites. ftp://ftp.1000genomes.ebi.ac.uk ftp://ftp-trace.ncbi.nlm.nih.gov/1000genomes
Proper citation: 1000 Genomes: A Deep Catalog of Human Genetic Variation (RRID:SCR_006828) Copy
https://www.sanger.ac.uk/collaboration/sequencing-idd-regions-nod-mouse-genome/
Genetic variations associated with type 1 diabetes identified by sequencing regions of the non-obese diabetic (NOD) mouse genome and comparing them with the same areas of a diabetes-resistant C57BL/6J reference mouse allowing identification of single nucleotide polymorphisms (SNPs) or other genomic variations putatively associated with diabetes in mice. Finished clones from the targeted insulin-dependent diabetes (Idd) candidate regions are displayed in the NOD clone sequence section of the website, where they can be downloaded either as individual clone sequences or larger contigs that make up the accession golden path (AGP). All sequences are publicly available via the International Nucleotide Sequence Database Collaboration. Two NOD mouse BAC libraries were constructed and the BAC ends sequenced. Clones from the DIL NOD BAC library constructed by RIKEN Genomic Sciences Centre (Japan) in conjunction with the Diabetes and Inflammation Laboratory (DIL) (University of Cambridge) from the NOD/MrkTac mouse strain are designated DIL. Clones from the CHORI-29 NOD BAC library constructed by Pieter de Jong (Children's Hospital, Oakland, California, USA) from the NOD/ShiLtJ mouse strain are designated CHORI-29. All NOD mouse BAC end-sequences have been submitted to the International Nucleotide Sequence Database Consortium (INSDC), deposited in the NCBI trace archive. They have generated a clone map from these two libraries by mapping the BAC end-sequences to the latest assembly of the C57BL/6J mouse reference genome sequence. These BAC end-sequence alignments can then be visualized in the Ensembl mouse genome browser where the alignments of both NOD BAC libraries can be accessed through the Distributed Annotation System (DAS). The Mouse Genomes Project has used the Illumina platform to sequence the entire NOD/ShiLtJ genome and this should help to position unaligned BAC end-sequences to novel non-reference regions of the NOD genome. Further information about the BAC end-sequences, such as their alignment, variation data and Ensembl gene coverage, can be obtained from the NOD mouse ftp site.
Proper citation: Sequencing of Idd regions in the NOD mouse genome (RRID:SCR_001483) Copy
https://github.com/tk2/RetroSeq
A tool for discovery and genotyping of transposable element variants (TEVs) (also known as mobile element insertions) from next-gen sequencing reads aligned to a reference genome in BAM format. The goal is to call TEVs that are not present in the reference genome but present in the sample that has been sequenced. It should be noted that RetroSeq can be used to locate any class of viral insertion in any species where whole-genome sequencing data with a suitable reference genome is available. RetroSeq is a two phase process, the first being the read pair discovery phase where discorandant mate pairs are detected and assigned to a TE class (Alu, SINE, LINE, etc.) by using either the annotated TE elements in the reference and/or aligned with Exonerate to the supplied library of viral sequences.
Proper citation: RetroSeq (RRID:SCR_005133) Copy
https://github.com/sanger-pathogens/Fastaq
Software application for diverse collection of scripts that perform useful and common FASTA/FASTQ manipulation tasks, such as filtering, merging, splitting, sorting, trimming, search/replace, etc. Input and output files can be gzipped (format is automatically detected) and individual Fastaq commands can be piped together.
Proper citation: Fastaq (RRID:SCR_016091) Copy
http://www.sanger.ac.uk/science/tools/ssaha2-0
A program designed for the efficient mapping of sequence reads onto genomic references. The software is capable of reading most sequencing platforms and giving a range of outputs are supported.
Proper citation: Sequence Search and Alignment by Hashing Algorithm (RRID:SCR_000544) Copy
https://www.sanger.ac.uk/science/tools/reapr
Software tool to identify errors in genome assemblies without need for reference sequence. Can be used in any stage of assembly pipeline to automatically break incorrect scaffolds and flag other errors in assembly for manual inspection. Reports mis-assemblies and other warnings, and produces new broken assembly based on error calls.
Proper citation: Recognition of Errors in Assemblies using Paired Reads (RRID:SCR_017625) Copy
http://www.sanger.ac.uk/Projects/Microbes/
This website includes a list of projects that the Sanger Institute is currently working on or completed. All projects consist of the genomic sequencing of different bacteria. Each description of the bacteria includes its classification, a description, and the types of diseases that the bacteria is likely to cause. The Sanger Institute bacterial sequencing effort is concentrated on pathogens and model organisms. Data is accessible in a number of ways; for each organism there is a BLAST server, allowing users to search the sequences with their own query and retrieve the matching contigs. Sequences can also be downloaded directly by FTP. Data is accessible in a number of ways; for each organism there is a BLAST server, allowing you to search the sequences with your own query and retrieve the matching contigs. Sequences can also be downloaded directly by FTP. The primary sequence viewer and annotation tool, Artemis is available for download. This is a portable Java program which is used extensively within the Microbial Genomes group for the analysis and annotation of sequence data from cosmids to whole genomes. The Artemis Comparison Tool (ACT) is also useful for interactive viewing of the comparisons between large and small sequences.
Proper citation: Bacterial Genomes (RRID:SCR_008141) Copy
http://www.sanger.ac.uk/science/tools/seqtools
Software for multiple sequence alignment viewing, editing and phylogeny. It includes a set of user-configurable modes to color residues used to create high-quality reference alignments.
Proper citation: Belvu (RRID:SCR_015989) Copy
https://sanger-pathogens.github.io/gubbins/
Software application as an algorithm that iteratively identifies loci containing elevated densities of base substitutions while concurrently constructing a phylogeny based on the putative point mutations outside of these regions. It is used for phylogenetic analysis of genome sequences and generating highly accurate reconstructions under realistic models of short-term bacterial evolution., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: Gubbins (RRID:SCR_016131) Copy
http://www.sanger.ac.uk/science/tools/seqtools
Software for sequence alignment that is a graphical dot-matrix program for detailed comparison of two sequences.
Proper citation: Dotter (RRID:SCR_016080) Copy
https://www.ebi.ac.uk/about/vertebrate-genomics/software/exonerate
Software package for sequence alignment of pairwise sequence comparison. Exonerate can be used to align sequences using many alignment models, exhaustive dynamic programming, or a variety of heuristics., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: Exonerate (RRID:SCR_016088) Copy
Original SAMTOOLS package has been split into three separate repositories including Samtools, BCFtools and HTSlib. Samtools for manipulating next generation sequencing data used for reading, writing, editing, indexing,viewing nucleotide alignments in SAM,BAM,CRAM format. BCFtools used for reading, writing BCF2,VCF, gVCF files and calling, filtering, summarising SNP and short indel sequence variants. HTSlib used for reading, writing high throughput sequencing data.
Proper citation: SAMTOOLS (RRID:SCR_002105) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.