Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
International collaboration producing an extensive public catalog of human genetic variation, including SNPs and structural variants, and their haplotype contexts, in an effort to provide a foundation for investigating the relationship between genotype and phenotype. The genomes of about 2500 unidentified people from about 25 populations around the world were sequenced using next-generation sequencing technologies. Redundant sequencing on various platforms and by different groups of scientists of the same samples can be compared. The results of the study are freely and publicly accessible to researchers worldwide. The consortium identified the following populations whose DNA will be sequenced: Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States. The goal Project is to find most genetic variants that have frequencies of at least 1% in the populations studied. Sequencing is still too expensive to deeply sequence the many samples being studied for this project. However, any particular region of the genome generally contains a limited number of haplotypes. Data can be combined across many samples to allow efficient detection of most of the variants in a region. The Project currently plans to sequence each sample to about 4X coverage; at this depth sequencing cannot provide the complete genotype of each sample, but should allow the detection of most variants with frequencies as low as 1%. Combining the data from 2500 samples should allow highly accurate estimation (imputation) of the variants and genotypes for each sample that were not seen directly by the light sequencing. All samples from the 1000 genomes are available as lymphoblastoid cell lines (LCLs) and LCL derived DNA from the Coriell Cell Repository as part of the NHGRI Catalog. The sequence and alignment data generated by the 1000genomes project is made available as quickly as possible via their mirrored ftp sites. ftp://ftp.1000genomes.ebi.ac.uk ftp://ftp-trace.ncbi.nlm.nih.gov/1000genomes
Proper citation: 1000 Genomes: A Deep Catalog of Human Genetic Variation (RRID:SCR_006828) Copy
http://www.sanger.ac.uk/resources/databases/exomiser/query/exomiser2
A Java program that functionally annotates variants from whole-exome sequencing data starting from a VCF (Variant Call Format) file (version 4). The functional annotation code is based on Annovar and uses UCSCKnownGene transcript definitions and hg19 genomic coordinates. Variants are prioritized according to user-defined criteria on variant frequency, pathogenicity, quality, inheritance pattern, phenotype data from human and model organisms, and proximity in the interactome to phenotypically similar genes.
Proper citation: Exomiser (RRID:SCR_002192) Copy
https://www.sanger.ac.uk/collaboration/sequencing-idd-regions-nod-mouse-genome/
Genetic variations associated with type 1 diabetes identified by sequencing regions of the non-obese diabetic (NOD) mouse genome and comparing them with the same areas of a diabetes-resistant C57BL/6J reference mouse allowing identification of single nucleotide polymorphisms (SNPs) or other genomic variations putatively associated with diabetes in mice. Finished clones from the targeted insulin-dependent diabetes (Idd) candidate regions are displayed in the NOD clone sequence section of the website, where they can be downloaded either as individual clone sequences or larger contigs that make up the accession golden path (AGP). All sequences are publicly available via the International Nucleotide Sequence Database Collaboration. Two NOD mouse BAC libraries were constructed and the BAC ends sequenced. Clones from the DIL NOD BAC library constructed by RIKEN Genomic Sciences Centre (Japan) in conjunction with the Diabetes and Inflammation Laboratory (DIL) (University of Cambridge) from the NOD/MrkTac mouse strain are designated DIL. Clones from the CHORI-29 NOD BAC library constructed by Pieter de Jong (Children's Hospital, Oakland, California, USA) from the NOD/ShiLtJ mouse strain are designated CHORI-29. All NOD mouse BAC end-sequences have been submitted to the International Nucleotide Sequence Database Consortium (INSDC), deposited in the NCBI trace archive. They have generated a clone map from these two libraries by mapping the BAC end-sequences to the latest assembly of the C57BL/6J mouse reference genome sequence. These BAC end-sequence alignments can then be visualized in the Ensembl mouse genome browser where the alignments of both NOD BAC libraries can be accessed through the Distributed Annotation System (DAS). The Mouse Genomes Project has used the Illumina platform to sequence the entire NOD/ShiLtJ genome and this should help to position unaligned BAC end-sequences to novel non-reference regions of the NOD genome. Further information about the BAC end-sequences, such as their alignment, variation data and Ensembl gene coverage, can be obtained from the NOD mouse ftp site.
Proper citation: Sequencing of Idd regions in the NOD mouse genome (RRID:SCR_001483) Copy
Interactive database which incorporates a suite of tools designed to aid the interpretation of submicroscopic chromosomal imbalance. Used to enhance clinical diagnosis by retrieving information from bioinformatics resources relevant to the imbalance found in the patient. Contributing to the DECIPHER database is a Consortium, comprising an international community of academic departments of clinical genetics. Each center maintains control of its own patient data (which are password protected within the center''''s own DECIPHER project) until patient consent is given to allow anonymous genomic and phenotypic data to become freely viewable within Ensembl and other genome browsers. Once data are shared, consortium members are able to gain access to the patient report and contact each other to discuss patients of mutual interest, thus facilitating the delineation of new microdeletion and microduplication syndromes.
Proper citation: DECIPHER (RRID:SCR_006552) Copy
Software R package as search tool for single cell RNA-seq data by gene lists. Builds index from scRNA-seq datasets which organizes information in suitable and compact manner so that datasets can be very efficiently searched for either cells or cell types in which given list of genes is expressed.
Proper citation: Scfind (RRID:SCR_017339) Copy
https://github.com/wtsi-npg/Illuminus
A fast and accurate algorithm for assigning single nucleotide polymorphism (SNP) genotypes to microarray data from the Illumina BeadArray technology.
Proper citation: ILLUMINUS (RRID:SCR_000388) Copy
http://www.sanger.ac.uk/science/tools/carol
Software application that is a combined functional annotation score of non-synonymous coding variants. A major challenge in interpreting whole-exome data is predicting which of the discovered variants are deleterious or neutral. To address this question in silico, they have developed a score called Combined Annotation scoRing toOL (CAROL), which combines information from two bioinformatics tools: PolyPhen-2 and SIFT, in order to improve the prediction of the effect of non-synonymous coding variants. The combination of annotation tools can help improve automated prediction of whole-genome/exome non-synonymous variant functional consequences. (entry from Genetic Analysis Software) The software should run on any UNIX or GNU/Linux system.
Proper citation: CAROL (RRID:SCR_001800) Copy
http://www.sanger.ac.uk/science/tools/dindel
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on March 7,2024. Software program for calling small indels from short-read sequence data ("next generation sequence data"). It is currently designed to handle only Illumina data. Dindel takes BAM files with mapped Illumina read data and enables researchers to detect small indels and produce a VCF file of all the variant calls. It has been written in C++ and can be used on Linux-based and Mac computers (it has not been tested on Windows operating systems).
Proper citation: DINDEL (RRID:SCR_001827) Copy
Consortium of 50 research groups across the UK to harness the power of newly-available genotyping technologies to improve our understanding of the aetiological basis of several major causes of global disease. The consortium has gathered genotype data for up to 500,000 sites of genome sequence variation (single nucleotide polymorphisms or SNPs) in samples ascertained for the disease phenotypes. Analysis of the genome-wide association data generated has lead to the identification of many SNPs and genes showing evidence of association with disease susceptibility, some of which will be followed up in future studies. In addition, the Consortium has gained important insights into the technical, analytical, methodological and biological aspects of genome-wide association analysis. The core of the study comprised an analysis of 2,000 samples from each of seven diseases (type 1 diabetes, type 2 diabetes, coronary heart disease, hypertension, bipolar disorder, rheumatoid arthritis and Crohn's disease). For each disease, the case samples have been ascertained from sites widely distributed across Great Britain, allowing us to obtain considerable efficiencies by comparing each of these case populations to a common set of 3,000 nationally-ascertained controls also from England, Scotland and Wales. These controls come from two sources: 1,500 are representative samples from the 1958 British Birth Cohort and 1,500 are blood donors recruited by the three national UK Blood Services. One of the questions that the WTCCC study has addressed relates to the relative merits of these alternative strategies for the generation of representative population cohorts. Genotyping for this main Case Control study was conducted by Affymetrix using the (commercial) Affymetrix 500K chip. As part of this study a total of 17,000 samples were typed for 500,000 SNPs. There are two additional components to the study. First, the WTCCC award is part-funding a study of host resistance to infectious diseases in African populations. The same approach has been used to type 2,000 cases of tuberculosis (TB) and 2,000 cases of malaria, as well as 2,000 shared controls. As well as addressing diseases of major global significance, and extending WTCCC coverage into the area of infectious disease, the inclusion of samples of African origin has obvious benefits with respect to methodological aspects of genome-wide association analysis. Second, the WTCCC has, for four additional diseases (autoimmune thyroid disease, breast cancer, ankylosing spondylitis, multiple sclerosis), completed an analysis of 15,000 SNPs designed to represent a large proportion of the known non-synonymous coding SNPs across the genome. This analysis has been performed at the WTSI using a custom Infinium chip (Illumina). Data release The genotypic data of the control samples (1958 British Birth Cohort and UK Blood Service) and from seven diseases analyzed in the main study are now available to qualified researchers. Summary genotype statistics for these collections are available directly from the website. Access to the individual-level genotype data and summary genotype statistics is by application to the Consortium Data Access Committee (CDAC) and approval subject to a Data Access Agreement. WTCCC2: A further round of GWA studies were funded in April 2008. These include 15 WTCCC-collaborative studies and 12 independent studies be supported totaling approximately 120,000 samples. Many of the studies represent major international collaborative networks that have together assembled large sample collections. WTCCC2 will perform genome-wide association studies in 13 disease conditions: Ankylosing spondylitis, Barrett's oesophagus and oesophageal adenocarcinoma, glaucoma, ischaemic stroke, multiple sclerosis, pre-eclampsia, Parkinson's disease, psychosis endophenotypes, psoriasis, schizophrenia, ulcerative colitis and visceral leishmaniasis. WTCCC2 will also investigate the genetics of reading and mathematics abilities in children and the pharmacogenomics of statin response. Over 60,000 samples will be analyzed using either the Affymetrix v6.0 chip or the Illumina 660K chip. The WTCCC2 will also genotype 3,000 controls each from the 1958 British Birth cohort and the UK Blood Service control group, and the 6,000 controls will be genotyped on both the Affymetrix v6.0 and Illumina 1.2M chips. WTCCC3: The Wellcome Trust has provided support for a further round of GWA studies in January 2009. These include 5 WTCCC-collaborative studies to be carried out in WTCCC3 and 5 independent studies, across a range of diseases. Many of the studies represent major international collaborative networks that have together assembled large sample collections. WTCCC3 will perform genome-wide association studies in the following 4 disease conditions: primary biliary cirrhosis, anorexia nervosa, pre-eclampsia in UK subjects, and the interactions between donor and recipient DNA related to early and late renal transplant dysfunction. The WTCCC3 will also carry out a pilot in a study of the genetics of host control of HIV-1 infection. Over 40,000 samples will be analyzed using the Illumina 660K chip. The WTCCC3 will utilize the 6,000 control genotypes generated by the WTCCC2.
Proper citation: Wellcome Trust Case Control Consortium (RRID:SCR_001973) Copy
http://www.sanger.ac.uk/mouseportal/
Database of mouse research resources at Sanger: BACs, targeting vectors, targeted ES cells, mutant mouse lines, and phenotypic data generated from the Institute''''s primary screen. The Wellcome Trust Sanger Institute generates, characterizes, and uses a variety of reagents for mouse genetics research. It also aims to facilitate the distribution of these resources to the external scientific community. Here, you will find unified access to the different resources available from the Institute or its collaborators. The resources include: 129S7 and C57BL6/J bacterial artificial chromosomes (BACs), MICER gene targeting vectors, knock-out first conditional-ready gene targeting vectors, embryonic stem (ES) cells with gene targeted mutations or with retroviral gene trap insertions, mutant mouse lines, and phenotypic data generated from the Institute''''s primary screen.
Proper citation: Sanger Mouse Resources Portal (RRID:SCR_006239) Copy
http://www.sanger.ac.uk/Projects/D_rerio/zmp/
Create knockout alleles in protein coding genes in the zebrafish genome, using a combination of whole exome enrichment and Illumina next generation sequencing, with the aim to cover them all. Each allele created is analyzed for morphological differences and published on the ZMP site. Transcript counting is performed on alleles with a morphological phenotype. Alleles generated are archived and can be requested from this site through the Zebrafish International Resource Center (ZIRC). You may register to receive updates on genes of interest, or browse a complete list, or search by Ensembl ID, gene name or human and mouse orthologue.
Proper citation: ZMP (RRID:SCR_006161) Copy
http://www.sanger.ac.uk/cgi-bin/teams/team30/arnie
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 1,2023. Database that integrates the extracellular protein interaction network generated in our lab using AVEXIS technology with spatiotemporal expression patterns for all genes in the network. The tool allows users to browse the network by clicking on individual proteins, or by specifying the spatiotemporal parameters. Clicking on connector lines will allow users to compare stage-matched expression patterns for genes encoding interacting proteins. Additionally, users can rapidly search for their genes in the network using the BLAST server provided.
Proper citation: ARNIE (RRID:SCR_000514) Copy
http://www.ncbi.nlm.nih.gov/CCDS/
Database (anonymous FTP) resulting from a collaborative effort to identify a core set of human and mouse protein coding regions that are consistently annotated and of high quality. The long term goal is to support convergence towards a standard set of gene annotations. Collaborators are EBI, NCBI, UCSC, WTSI and the initial results are also available from the participants'''' genome browser Web sites. In addition, CCDS identifiers are indicated on the relevant NCBI RefSeq and Entrez Gene records and in Map Viewer displays of RNA (RefSeq) and Gene annotations on the reference assembly.
Proper citation: Consensus CDS (RRID:SCR_006729) Copy
http://cancer.sanger.ac.uk/cancergenome/projects/cosmic/
Database to store and display somatic mutation information and related details and contains information relating to human cancers. The mutation data and associated information is extracted from the primary literature. In order to provide a consistent view of the data a histology and tissue ontology has been created and all mutations are mapped to a single version of each gene. The data can be queried by tissue, histology or gene and displayed as a graph, as a table or exported in various formats.
Some key features of COSMIC are:
* Contains information on publications, samples and mutations. Includes samples which have been found to be negative for mutations during screening therefore enabling frequency data to be calculated for mutations in different genes in different cancer types.
* Samples entered include benign neoplasms and other benign proliferations, in situ and invasive tumours, recurrences, metastases and cancer cell lines.
Proper citation: COSMIC - Catalogue Of Somatic Mutations In Cancer (RRID:SCR_002260) Copy
http://www.genes2cognition.org/
A neuroscience research program that studies genes, the brain and behavior in an integrated manner, established to elucidate the molecular mechanisms of learning and memory, and shed light on the pathogenesis of disorders of cognition. Central to G2C investigations is the NMDA receptor complex (NRC/MASC), that is found at the synapses in the central nervous system which constitute the functional connections between neurons. Changes in the receptor and associated components are thought to be in a large part responsible for the phenomenon of synaptic plasticity, that may underlie learning and memory. G2C is addressing the function of synapse proteins using large scale approaches combining genomics, proteomics and genetic methods with electrophysiological and behavioral studies. This is incorporated with computational models of the organization of molecular networks at the synapse. These combined approaches provide a powerful and unique opportunity to understand the mechanisms of disease genes in behavior and brain pathology as well as provide fundamental insights into the complexity of the human brain. Additionally, Genes to Cognition makes available its biological resources, including gene-targeting vectors, ES cell lines, antibodies, and transgenic mice, generated for its phenotyping pipeline. The resources are freely-available to interested researchers.
Proper citation: Genes to Cognition: Neuroscience Research Programme (RRID:SCR_007121) Copy
The Deciphering Developmental Disorders (DDD) study aims to find out if using new genetic technologies can help doctors understand why patients get developmental disorders. To do this we have brought together doctors in the 23 NHS Regional Genetics Services throughout the UK and scientists at the Wellcome Trust Sanger Institute, a charitably funded research institute which played a world-leading role in sequencing (reading) the human genome. The DDD study involves experts in clinical, molecular and statistical genetics, as well as ethics and social science. It has a Scientific Advisory Board consisting of scientists, doctors, a lawyer and patient representative, and has received National ethical approval in the UK. Over the next few years, we are aiming to collect DNA and clinical information from 12,000 undiagnosed children in the UK with developmental disorders and their parents. The results of the DDD study will provide a unique, online catalogue of genetic changes linked to clinical features that will enable clinicians to diagnose developmental disorders. Furthermore, the study will enable the design of more efficient and cheaper diagnostic assays for relevant genetic testing to be offered to all such patients in the UK and so transform clinical practice for children with developmental disorders. Over time, the work will also improve understanding of how genetic changes cause developmental disorders and why the severity of the disease varies in individuals. The Sanger Institute will contribute to the DDD study by performing genetic analysis of DNA samples from patients with developmental disorders, and their parents, recruited into the study through the Regional Genetics Services. Using microarray technology and the latest DNA sequencing methods, research teams will probe genetic information to identify mutations (DNA errors or rearrangements) and establish if these mutations play a role in the developmental disorders observed in patients. The DDD initiative grew out of the groundbreaking DECIPHER database, a global partnership of clinical genetics centres set up in 2004, which allows researchers and clinicians to share clinical and genomic data from patients worldwide. The DDD study aims to transform the power of DECIPHER as a diagnostic tool for use by clinicians. As well as improving patient care, the DDD team will empower researchers in the field by making the data generated securely available to other research teams around the world. By assembling a solid resource of high-quality, high-resolution and consistent genomic data, the leaders of the DDD study hope to extend the reach of DECIPHER across a broader spectrum of disorders than is currently possible.
Proper citation: Deciphering Developmental Disorders (RRID:SCR_006171) Copy
http://www.sanger.ac.uk/resources/software/peer/
Software collection of Bayesian approaches to infer hidden determinants and their effects from gene expression profiles using factor analysis methods. Applications of PEER have * detected batch effects and experimental confounders * increased the number of expression QTL findings by threefold * allowed inference of intermediate cellular traits, such as transcription factor or pathway activations This project offers an efficient and versatile C++ implementation of the underlying algorithms with user-friendly interfaces to R and python.
Proper citation: PEER (RRID:SCR_009326) Copy
http://linux1.softberry.com/spldb/SpliceDB.html
Database of canonical and non-canonical mammalian splice sites. The information about verified splice site sequences for canonical and non-canonical sites is presented with the supporting evidence. Weight matrices were built for the major splice groups, which can be incorporated into gene prediction programs.
Proper citation: SpliceDB (RRID:SCR_006262) Copy
http://www.sanger.ac.uk/science/tools/olorin
An interactive filtering tool for next generation sequencing data coming from the study of large complex disease pedigrees. It integrates gene flow output from Merlin and next generation sequencing data. Users can interactively filter and prioritize variants based on haplotype sharing across different sets of selected individuals and allele frequency in reference datasets. (entry from Genetic Analysis Software)
Proper citation: OLORIN (RRID:SCR_002015) Copy
A database of phylogenetic trees of animal genes. It aims at developing a curated resource that gives reliable information about ortholog and paralog assignments, and evolutionary history of various gene families. TreeFam defines a gene family as a group of genes that evolved after the speciation of single-metazoan animals. It also tries to include outgroup genes like yeast (S. cerevisiae and S. pombe) and plant (A. thaliana) to reveal these distant members.TreeFam is also an ortholog database. Unlike other pairwise alignment based ones, TreeFam infers orthologs by means of gene trees. It fits a gene tree into the universal species tree and finds historical duplications, speciations and losses events. TreeFam uses this information to evaluate tree building, guide manual curation, and infer complex ortholog and paralog relations.The basic elements of TreeFam are gene families that can be divided into two parts: TreeFam-A and TreeFam-B families. TreeFam-B families are automatically created. They might contain errors given complex phylogenies. TreeFam-A families are manually curated from TreeFam-B ones. Family names and node names are assigned at the same time. The ultimate goal of TreeFam is to present a curated resource for all the families. phylogenetic tree, animal, vertebrate, invertebrate, gene, ortholog, paralog, evolutionary history, gene families, single-metazoan animals, outgroup genes like yeast (S. cerevisiae and S. pombe), plant (A. thaliana), historical duplications, speciations, losses, Human, Genome, comparative genomics
Proper citation: Tree families database (RRID:SCR_013401) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.