Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
Maintains and provides archival, retrieval and analytical resources for biological information. Central DDBJ resource consists of public, open-access nucleotide sequence databases including raw sequence reads, assembly information and functional annotation. Database content is exchanged with EBI and NCBI within the framework of the International Nucleotide Sequence Database Collaboration (INSDC). In 2011, DDBJ launched two new resources: DDBJ Omics Archive and BioProject. DOR is archival database of functional genomics data generated by microarray and highly parallel new generation sequencers. Data are exchanged between the ArrayExpress at EBI and DOR in the common MAGE-TAB format. BioProject provides organizational framework to access metadata about research projects and data from projects that are deposited into different databases.
Proper citation: DNA DataBank of Japan (DDBJ) (RRID:SCR_002359) Copy
http://deweylab.biostat.wisc.edu/rsem/
Software package for quantifying gene and isoform abundances from single end or paired end RNA Seq data. Accurate transcript quantification from RNA Seq data with or without reference genome. Used for accurate quantification of gene and isoform expression from RNA-Seq data.
Proper citation: RSEM (RRID:SCR_000262) Copy
International consortium of six centers assembled to participate in the development and implementation of studies to identify infectious agents, dietary factors, or other environmental agents, including psychosocial factors, that trigger type 1 diabetes in genetically susceptible people. The coordinating centers recruit and enroll subjects, obtaining informed consent from parents prior to or shortly after birth, genetic and other types of samples from neonates and parents, and prospectively following selected neonates throughout childhood or until development of islet autoimmunity or T1DM. The study tracks child diet, illnesses, allergies and other life experiences. A blood sample is taken from children every 3 months for 4 years. After 4 years, children will be seen every 6 months until the age of 15 years. Children are tested for 3 different autoantibodies. The study will compare the life experiences and blood and stool tests of the children who get autoantibodies and diabetes with some of those children who do not get autoantibodies or diabetes. In this way the study hopes to find the triggers of T1DM in children with higher risk genes.
Proper citation: TEDDY (RRID:SCR_000383) Copy
http://www.epilepsygenetics.eu/
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on August 16,2023. Group of clinical care and epilepsy research centers who are committed to improving the lives of people with epilepsy through an understanding of the genetics of epilepsy. The consoritum was in an effort to speed discovery to epilepsy genetics by pooling the resources of several research centres., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: EPIGEN (RRID:SCR_000093) Copy
http://harvard.eagle-i.net/i/0000012e-58c7-d44f-55da-381e80000000
Core to provide gene expression data analysis service. Activities range from the provision of services to fully collaborative grant funded investigations.
Proper citation: Harvard Partners HealthCare Center for Personalized Genetic Medicine Bioinformatics Core Facility (RRID:SCR_000882) Copy
http://fulxie.0fees.us/?type=reference&ckattempt=1
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 1,2023. Web-based tool for evaluating and screening reference genes from extensive experimental datasets. It integrates major computational programs (geNorm, Normfinder, BestKeeper, and the comparative delta-Ct method) to compare and rank the tested candidate reference genes. Based on the rankings from each program, it assigns an appropriate weight to an individual gene and calculated the geometric mean of their weights for the overall final ranking., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: RefFinder (RRID:SCR_000472) Copy
http://www.yandell-lab.org/software/index.html
Sequenced genomes contain a treasure trove of information about how genes function and evolve. Getting at this information, however, is challenging and requires novel approaches that combine computer science and experimental molecular biology. My lab works at the intersection of both domains, and research in our group can be summarized as follows: generate hypotheses concerning gene function and evolution by computational means, and then test these hypotheses at the bench. This is easier said than done, as serious barriers still exist to using sequenced genomes and their annotations as starting points for experimental work. Some of these barriers lie in the computational domain, others in the experimental. Though challenging, overcoming these barriers offers exciting training opportunities in both computer science and molecular genetics, especially for those seeking a future at the intersection of both fields. Ongoing projects in the lab are centered on genome annotation and comparative genomics; exploring the relationships between sequence variation and human disease; and high-throughput biological image analysis. Current software tools available: VAAST (the Variant Annotation, Analysis & Search Tool) is a probabilistic search tool for identifying damaged genes and their disease-causing variants in personal genome sequences. VAAST builds upon existing amino acid substitution (AAS) and aggregative approaches to variant prioritization, combining elements of both into a single unified likelihood-framework that allows users to identify damaged genes and deleterious variants with greater accuracy, and in an easy-to-use fashion. VAAST can score both coding and non-coding variants, evaluating the cumulative impact of both types of variants simultaneously. VAAST can identify rare variants causing rare genetic diseases, and it can also use both rare and common variants to identify genes responsible for common diseases. VAAST thus has a much greater scope of use than any existing methodology. MAKER 2 (updated 01-16-2012) MAKER is a portable and easily configurable genome annotation pipeline. It's purpose is to allow smaller eukaryotic and prokaryotic genomeprojects to independently annotate their genomes and to create genome databases. MAKER identifies repeats, aligns ESTs and proteins to a genome, produces ab-initio gene predictions and automatically synthesizes these data into gene annotations having evidence-based quality values. MAKER is also easily trainable: outputs of preliminary runs can be used to automatically retrain its gene prediction algorithm, producing higher quality gene-models on seusequent runs. MAKER's inputs are minimal and its ouputs can be directly loaded into a GMOD database. They can also be viewed in the Apollo genome browser; this feature of MAKER provides an easy means to annotate, view and edit individual contigs and BACs without the overhead of a database. MAKER should prove especially useful for emerging model organism projects with minimal bioinformatics expertise and computer resources. RepeatRunner RepeatRunner is a CGL-based program that integrates RepeatMasker with BLASTX to provide a comprehensive means of identifying repetitive elements. Because RepeatMasker identifies repeats by means of similarity to a nucleotide library of known repeats, it often fails to identify highly divergent repeats and divergent portions of repeats, especially near repeat edges. To remedy this problem, RepeatRunner uses BLASTX to search a database of repeat encoded proteins (reverse transcriptases, gag, env, etc...). Because protein homologies can be detected across larger phylogenetic distances than nucleotide similarities, this BLASTX search allows RepeatRunner to identify divergent protein coding portions of retro-elements and retro-viruses not detected by RepeatMasker. RepeatRunner merges its BLASTX and RepeatMasker results to produce a single, comprehensive XML-based output. It also masks the input sequence appropriately. In practice RepeatRunner has been shown to greatly improve the efficacy of repeat identifcation. RepeatRunner can also be used in conjunction with PILER-DF - a program designed to identify novel repeats - and RepeatMasker to produce a comprehensive system for repeat identification, characterization, and masking in the newly sequenced genomes. CGL CGL is a software library designed to facilitate the use of genome annotations as substrates for computation and experimentation; we call it CGL, an acronym for Comparitive Genomics Library, and pronounce it Seagull. The purpose of CGL is to provide an informatics infrastructure for a laboratory, department, or research institute engaged in the large-scale analysis of genomes and their annotations.
Proper citation: Yandell Lab Portal (RRID:SCR_000807) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 23,2022. Interactive database of Drosophila melanogaster nervous system. Used by drosophila neuroscience community and by other researchers studying arthropod brain structure.
Proper citation: FlyBrain (RRID:SCR_000706) Copy
International collaborative research project and database of annotated mammalian genome. Used to improve estimates of total number of genes and their alternative transcript isoforms in both human and mouse. Consortium to assign functional annotations to full length cDNAs that were collected during Mouse Encyclopedia Project at RIKEN.
Proper citation: Functional Annotation of the Mammalian Genome (RRID:SCR_000788) Copy
http://franklin.imgen.bcm.tmc.edu/
The mission of the Baylor College of Medicine - Shaw Laboratory is to apply methods of statistics and bioinformatics to the analysis of large scale genomic data. Our vision is data integration to reveal the underlying connections between genes and processes in order to cure disease and improve healthcare.
Proper citation: Baylor College of Medicine - Shaw Laboratory (RRID:SCR_000604) Copy
Laboratory portal of the University of Sao Paulo Molecular Genetics and Bioinformatic Laboratory.
Proper citation: USP Molecular Genetics and Bioinformatics Laboratory (RRID:SCR_000605) Copy
http://ccb.jhu.edu/software/sim4cc/
Software tool as cross species spliced alignment program.Heuristic sequence alignment tool for comparing cDNA sequence with genomic sequence containing homolog of gene in another species.
Proper citation: sim4cc (RRID:SCR_001204) Copy
http://www.omixon.com/data-analysis-and-pro/
Software application suite to help clinical labs adopt next generation sequencing for the analysis of diagnostic gene targets.
Proper citation: Omixon Target Data Analysis (RRID:SCR_001207) Copy
http://www.genome.jp/kegg/expression/
Database for mapping gene expression profiles to pathways and genomes. Repository of microarray gene expression profile data for Synechocystis PCC6803 (syn), Bacillus subtilis (bsu), Escherichia coli W3110 (ecj), Anabaena PCC7120 (ana), and other species contributed by the Japanese research community.
Proper citation: Kyoto Encyclopedia of Genes and Genomes Expression Database (RRID:SCR_001120) Copy
http://archive.ics.uci.edu/ml/datasets/EEG+Database
Data set from a large study to examine EEG correlates of genetic predisposition to alcoholism. It contains measurements from 64 electrodes placed on the scalp sampled at 256 Hz (3.9-msec epoch) for 1 second. There were two groups of subjects: alcoholic and control. Each subject was exposed to either a single stimulus (S1) or to two stimuli (S1 and S2) which were pictures of objects chosen from the 1980 Snodgrass and Vanderwart picture set. When two stimuli were shown, they were presented in either a matched condition where S1 was identical to S2 or in a non-matched condition where S1 differed from S2. There were 122 subjects and each subject completed 120 trials where different stimuli were shown. The electrode positions were located at standard sites (Standard Electrode Position Nomenclature, American Electroencephalographic Association 1990). Zhang et al. (1995) describes in detail the data collection process. There are three versions of the EEG data set. * The Small Data Set (smni97_eeg_data.tar.gz) contains data for the 2 subjects, alcoholic a_co2a0000364 and control c_co2c0000337. For each of the 3 matching paradigms, c_1 (one presentation only), c_m (match to previous presentation) and c_n (no-match to previous presentation), 10 runs are shown. * The Large Data Set (SMNI_CMI_TRAIN.tar.gz and SMNI_CMI_TEST.tar.gz) contains data for 10 alcoholic and 10 control subjects, with 10 runs per subject per paradigm. The test data used the same 10 alcoholic and 10 control subjects as with the training data, but with 10 out-of-sample runs per subject per paradigm. * The Full Data Set contains all 120 trials for 122 subjects. The entire set of data is about 700 MBytes.
Proper citation: EEG Database (RRID:SCR_001581) Copy
The EBI genomes pages give access to a large number of complete genomes including bacteria, archaea, viruses, phages, plasmids, viroids and eukaryotes. Methods using whole genome shotgun data are used to gain a large amount of genome coverage for an organism. WGS data for a growing number of organisms are being submitted to DDBJ/EMBL/GenBank. Genome entries have been listed in their appropriate category which may be browsed using the website navigation tool bar on the left. While organelles are all listed in a separate category, any from Eukaryota with chromosome entries are also listed in the Eukaryota page. Within each page, entries are grouped and sorted at the species level with links to the taxonomy page for that species separating each group. Within each species, entries whose source organism has been categorized further are grouped and numbered accordingly. Links are made to: * taxonomy * complete EMBL flatfile * CON files * lists of CON segments * Project * Proteomes pages * FASTA file of Proteins * list of Proteins
Proper citation: EBI Genomes (RRID:SCR_002426) Copy
Multicenter observational study designed to identify genetic determinants of diabetic nephropathy. It is conducted in eleven U.S. clinical centers and a coordinating center, and with four ethnic groups (European Americans, African Americans, Mexican Americans, and American Indians). Two strategies are used to localize susceptibility genes: a family-based linkage study and a case-control study using mapping by admixture linkage disequilibrium (MALD). In the family-based study, probands with diabetic nephropathy are recruited with their parents and selected siblings. Linkage analyses will be conducted to identify chromosomal regions containing genes that influence the development of diabetic nephropathy or related quantitative traits such as serum creatinine concentration, urinary albumin excretion, and plasma glucose concentrations. Regions showing evidence of linkage will be examined further with both genetic linkage and association studies to identify genes that influence diabetic nephropathy or related traits. Two types of MALD studies are being done. One is a case-control study of unrelated individuals of Mexican American heritage in which both cases and controls have diabetes, but only the case has nephropathy. The other is a case-control study of African American patients with nephropathy (cases) and their spouses (controls) unaffected by diabetes and nephropathy; offspring are genotyped when available to provide haplotype data. The specific goals of this program: * Delineate genomic regions associated with the development and progression of renal disease(s) * Evaluate whether there is a genetic link between diabetic nephropathy and diabetic retinopathy * Improve outcomes * Provide protection for people at risk and slow the progression of renal disease * Help establish a resource for genetic studies of kidney disease and diabetic complications by creating a repository of genetic samples and a database * Encourage studies of the genetics of progressive renal disease
Proper citation: Family Investigation of Nephropathy of Diabetes (RRID:SCR_001525) Copy
http://bpg.utoledo.edu/~afedorov/lab/eid.html
Data sets of protein-coding intron-containing genes that contain gene information from humans, mice, rats, and other eukaryotes, as well as genes from species whose genomes have not been completely sequenced. This is a comprehensive and convenient dataset of sequences for computational biologists who study exon-intron gene structures and pre-mRNA splicing. The database is derived from GenBank release 112, and it contains protein-coding genes that harbor introns, along with extensive descriptions of each gene and its DNA and protein sequences, as well as splice motif information. They have created subdatabases of genes whose intron positions have been experimentally determined. The collection also contains data on untranslated regions of gene sequences and intron-less genes. For species with entirely sequenced genomes, species-specific databases have been generated. A novel Mammalian Orthologous Intron Database (MOID) has been introduced which includes the full set of introns that come from orthologous genes that have the same positions relative to the reading frames.
Proper citation: EID: Exon-Intron Database (RRID:SCR_002469) Copy
http://www.sci.unisannio.it/docenti/rampone/
Data set of Homo Sapiens Exons, Introns and Splice regions extracted from GenBank Rel.123 with an aim of giving standardized material to train and to assess the prediction accuracy of computational approaches for gene identification and characterization. From the complete GenBank (Primate Sequences Division) Rel.123 (162,557 entries), entries of Human Nuclear DNA including Complete CDS and more than one Exon have been selected, and 4523 exons and 3802 introns have been extracted from these entries. Details about extracted exons and introns are reported (Locus, number, Start and End position in the entry, sequence, length, G+C content, presence of not AGCT data (nucleotide scan check)). Statistics are also reported (overall nucleotides, average G+C content, nucleotide scan check results, number of not GT starting / AG ending introns, minimum / maximum / average length, length standard deviation). 3799+3799 donor and acceptor sites, as windows of 140 nucleotides around each splice site have been extracted. After discarding sequences not including canonical GTAG junctions (65+74), including insufficient data (not enough material for a 140 nucleotide window) (686+589), including not AGCT bases (29+30), and redundant (218+226) there are 2796+ 2880 windows. Finally, there are 271,937 + 332,296 windows of false splice sites, selected by searching canonical GTAG pairs in not splicing positions. The false sites in a range of +/- 60 from a true splice site are marked as proximal.
Proper citation: HS3D - Homo Sapiens Splice Sites Dataset (RRID:SCR_002939) Copy
Curated lists of genes associated to speech / language phenotypes and structural or functional abnormalities observed in patient populations. Entrez ID gene information, as well as gene expression profiles from the Allen Brain Atlas are available. You can also download expression data for a given gene in JSON or XML format.
Proper citation: Speech Language Disorders Database (RRID:SCR_003655) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.