Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
https://github.com/MikkelSchubert/adapterremoval
Software program to remove residual adapter sequences from next generation sequencing reads. Used for cleaning of next-generation sequencing reads. AdapterRemoval v2 introduces improvements in throughput, through use of single instruction, multiple data (SIMD; SSE1 and SSE2) instructions and multi-threading support; handles datasets containing reads or read-pairs with different adapters or adapter pairs; provides simultaneous demultiplexing and adapter trimming; has ability to reconstruct adapter sequences from paired-end reads for poorly documented data sets; provides native gzip and bzip2 support.
Proper citation: AdapterRemoval (RRID:SCR_011834) Copy
Core facility provides researchers with access to high-throughput sequencing technologies. The staff provide consultation on experimental design, library preparation, and data analysis. The Sequencing Core Facility works closely with Bioinformatics staff in the Center for Quantitative Biology to provide researchers with computing power and consulting services to analyze sequencing data.
Proper citation: Princeton High Throughput Sequencing and Microarray Facility (RRID:SCR_012619) Copy
http://sift.bii.a-star.edu.sg/
Data analysis service to predict whether an amino acid substitution affects protein function based on sequence homology and the physical properties of amino acids. SIFT can be applied to naturally occurring nonsynonymous polymorphisms and laboratory-induced missense mutations. (entry from Genetic Analysis Software) Web service is also available.
Proper citation: SIFT (RRID:SCR_012813) Copy
http://www.ornl.gov/sci/techresources/Human_Genome/home.shtml
This resource gives information about the U.S. Human Genome Project, which was was a 13-year effort to to discover all the estimated 20,000-25,000 human genes and make them accessible for further biological study. The primary project goals were to: - identify all the approximately 20,000-25,000 genes in human DNA, - determine the sequences of the 3 billion chemical base pairs that make up human DNA, - store this information in databases, - improve tools for data analysis, - transfer related technologies to the private sector, and - address the ethical, legal, and social issues (ELSI) that may arise from the project. To help achieve these goals, researchers also studied the genetic makeup of several nonhuman organisms. These include the common human gut bacterium Escherichia coli, the fruit fly, and the laboratory mouse. These parallel studies helped to develop technology and interpret human gene function. Sponsors: The DOE Human Genome Program and the NIH National Human Genome Research Institute (NHGRI) together sponsored the U.S. Human Genome Project.
Proper citation: Human Genome Project Information (RRID:SCR_013028) Copy
http://genetics.bwh.harvard.edu/pph2/
Software tool which predicts possible impact of amino acid substitution on structure and function of human protein using straightforward physical and comparative considerations. PolyPhen-2 is new development of PolyPhen tool for annotating coding nonsynonymous SNPs.
Proper citation: PolyPhen: Polymorphism Phenotyping (RRID:SCR_013189) Copy
http://www.mrc-lmb.cam.ac.uk/genomes/dolop/
DOLOP is an exclusive knowledge base for bacterial lipoproteins by processing information from 510 entries to provide a list of 199 distinct lipoproteins with relevant links to molecular details. Features include functional classification, predictive algorithm for query sequences, primary sequence analysis and lists of predicted lipoproteins from 43 completed bacterial genomes along with interactive information exchange facility. This website along will have additional information on the biosynthetic pathway, supplementary material and other related figures. DOLOP also contains information and links to molecular details for about 278 distinct lipoproteins and predicted lipoproteins from 234 completely sequenced bacterial genomes. Additionally, the website features a tool that applies a predictive algorithm to identify the presence or absence of the lipoprotein signal sequence in a user-given sequence. The experimentally verified lipoproteins have been classified into different functional classes and more importantly functional domain assignments using hidden Markov models from the SUPERFAMILY database that have been provided for the predicted lipoproteins. Other features include: primary sequence analysis, signal sequence analysis, and search facility and information exchange facility to allow researchers to exchange results on newly characterized lipoproteins.
Proper citation: DOLOP: A Database of Bacterial Lipoproteins (RRID:SCR_013487) Copy
http://biologylabs.utah.edu/jorgensen/wayned/ape/
Software tool for plasmid and sequence editing, annotating and drawing plasmid sequences. Used to view circular or linear maps of DNA sequences. Users can perform virtual digests whereby they select predefined DNA ladder, or specify their own, and visualize theoretical DNA fragments. Used to highlight restriction sites in editing window, accurately reflect Dam/Dcm blocking of enzyme sites, highlighting and drawing graphic maps using feature annotations from genbank and embl files, highlighting text using pre-defined and custom feature libraries, and directly BLASTing selected sequence at NCBI or Wormbase. Runs across Windows, OS X, and Linux/Unix.
Proper citation: A plasmid Editor (RRID:SCR_014266) Copy
https://www.sanger.ac.uk/collaboration/sequencing-idd-regions-nod-mouse-genome/
Genetic variations associated with type 1 diabetes identified by sequencing regions of the non-obese diabetic (NOD) mouse genome and comparing them with the same areas of a diabetes-resistant C57BL/6J reference mouse allowing identification of single nucleotide polymorphisms (SNPs) or other genomic variations putatively associated with diabetes in mice. Finished clones from the targeted insulin-dependent diabetes (Idd) candidate regions are displayed in the NOD clone sequence section of the website, where they can be downloaded either as individual clone sequences or larger contigs that make up the accession golden path (AGP). All sequences are publicly available via the International Nucleotide Sequence Database Collaboration. Two NOD mouse BAC libraries were constructed and the BAC ends sequenced. Clones from the DIL NOD BAC library constructed by RIKEN Genomic Sciences Centre (Japan) in conjunction with the Diabetes and Inflammation Laboratory (DIL) (University of Cambridge) from the NOD/MrkTac mouse strain are designated DIL. Clones from the CHORI-29 NOD BAC library constructed by Pieter de Jong (Children's Hospital, Oakland, California, USA) from the NOD/ShiLtJ mouse strain are designated CHORI-29. All NOD mouse BAC end-sequences have been submitted to the International Nucleotide Sequence Database Consortium (INSDC), deposited in the NCBI trace archive. They have generated a clone map from these two libraries by mapping the BAC end-sequences to the latest assembly of the C57BL/6J mouse reference genome sequence. These BAC end-sequence alignments can then be visualized in the Ensembl mouse genome browser where the alignments of both NOD BAC libraries can be accessed through the Distributed Annotation System (DAS). The Mouse Genomes Project has used the Illumina platform to sequence the entire NOD/ShiLtJ genome and this should help to position unaligned BAC end-sequences to novel non-reference regions of the NOD genome. Further information about the BAC end-sequences, such as their alignment, variation data and Ensembl gene coverage, can be obtained from the NOD mouse ftp site.
Proper citation: Sequencing of Idd regions in the NOD mouse genome (RRID:SCR_001483) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on October 28,2025. A chicken EST Web site has been created to provide access to the data, and a set of unique sequences has been deposited with GenBank. This site contains over 40,000 EST sequences from the chicken cDNA libraries in the University of Delaware collection. Users can perform keyword searches, BLAST nucleotide sequences against our database, view clusters of similar or overlapping clones, and order clones. The cDNA and gene sequences of many mammalian cytokines and their receptors are known. However, corresponding information on avian cytokines is limited due to the lack of cross-species activity at the functional level or strong homology at the molecular level. To improve the efficiency of identifying cytokines and novel chicken genes, a directionally cloned cDNA library from T-cell-enriched activated chicken splenocytes was constructed, and the partial sequence of 5251 clones was obtained. Sequence clustering indicates that 2357 (42%) of the clones are present as a single copy, and 2961 are distinct clones, demonstrating the high level of complexity of this library. Comparisons of the sequence data with known DNA sequences in GenBank indicate that approximately 25% of the clones match known chicken genes, 39% have similarity to known genes in other species, and 11% had no match to any sequence in the database. Several previously uncharacterized chicken cytokines and their receptors were present in our library. This collection provides a useful database for cataloging genes expressed in T cells and a valuable resource for future investigations of gene expression in avian immunology. Therefore, the Chick EST database was created.
Proper citation: UD Chick EST Project (RRID:SCR_002236) Copy
NIH initiative project to provide full-length open reading frame (FL-ORF) clones for human, mouse, and rat genes, cow. MGC cDNA clones were obtained by screening of cDNA libraries, by transcript-specific RT-PCR cloning, and by DNA synthesis of cDNA inserts. All MGC sequences are deposited in GenBank and clones can be purchased from distributors of IMAGE consortium. With conclusion of MGC project in March 2009, GenBank records of MGC sequences will be frozen, without further updates. Since definition of what constitutes full-length coding region for some of genes and transcripts for which they have MGC clones will likely change in future, users planning to order MGC clones will need to monitor for these changes. Users can make use of genome browsers and gene-specific databases, such as the UCSC Genome browser, NCBI's Map Viewer, and Entrez Gene, to view relevant regions of genome (browsers) or gene-related information (Entrez Gene).
Proper citation: Mammalian Gene Collection (RRID:SCR_007024) Copy
Offer biorepository services to public and private research institutes, to the highest standards of quality and safety with the aim of contributing to the advancement of medical research and scientific discovery. The BioRep Cell Repository establishes, maintains and distributes cell line cultures as well as DNA derived from these cultures. The scientific and business affiliation between BioRep and Coriell allows access to more than a million types of cell vials, stored in liquid nitrogen. Cells that have been stored for nearly 50 years, are still viable and available for research purposes today. Thanks to an exclusive agreement with the Coriell Institute for Medical Research, the oldest and largest biorepository of the world, BioRep is specialized in cell lines preparation, in nucleic acid extraction and long term storage in liquid nitrose (-196 degrees C) and in refrigerators (-80 degrees C) of any kind of biosamples, using procedures and standards developed by the Coriell in over 50 years of activity. BioRep and Coriell together constitute one of the few Global Biorepository able to serve the pharmaceutical industries for world wide clinical trials. BioRep facility is specifically designed to give the utmost efficiency and security by implementing Coriell procedures and standards. The BioRep Tissue Repository provides safe and secure storage of tissue specimens as required for medical research and scientific investigation. All tissues are preserved with the most current preservation techniques and processes. In addition to the storage service, BioRep provides Cell Biology, Molecular Biology, Microbiology services developed in ISO 9001:2008 certified laboratories.
Proper citation: BioRep (RRID:SCR_004907) Copy
http://www.ebi.ac.uk/Tools/blast2/index.html
It is used to compare a novel sequence with those contained in nucleotide and protein databases by aligning the novel sequence with previously characterized genes.
Proper citation: Washington University Basic Local Alignment Search Tool (RRID:SCR_008285) Copy
http://www.uwstructuralgenomics.org/
It is a specialized research center supported by the Protein Structure Initiative (PSI) of the National Institute of General Medical Sciences (NIGMS), one of the National Institutes of Health (NIH). PSI is a federal, university, and industry effort aimed at dramatically reducing the costs and lessening the time it takes to determine a three-dimensional protein structure. The long-range goal of PSI is to solve 10,000 protein structures in 10 years and to make the three-dimensional atomic-level structures of most proteins easily obtainable from knowledge of their corresponding DNA sequences. CESG is located within the Department of Biochemistry at the University of Wisconsin-Madison (Madison, WI) and the Department of Biochemistry at the Medical College of Wisconsin (Milwaukee, WI). CESG develops new methods and technologies to address unique eukaryotic bottlenecks and disseminates its methodologies and experimental results to the scientific community worldwide through: :- Cell-Free Protein Production Workshops :- Plasmids at PSI Materials Repository :- Posters Presented at Scientific Meetings :- Publications in PubMed / PubMed Central :- Sesame (LIMS) Available for Researchers :- Solved Structures in the Protein Data Bank :- Technology Dissemination Reports They have welcomed requests by researchers to solve eukaryotic protein structures, particularly medically relevant proteins, through our Online Structure Request System for Researchers. They have solved many community-nominated targets and deposited information about these targets in public databases and published on our investigations and findings. Sponsors: CESG is supported by NIH / NIGMS Protein Structure Initiative grant numbers U54 GM074901 and P50 GM064598.
Proper citation: CESG (RRID:SCR_008451) Copy
http://hihg.med.miami.edu/software-download/seqem-version-1.0
Online tool for utilizing a genotype calling algorithm for next-generation sequence data.
Proper citation: SeqEM (RRID:SCR_002021) Copy
http://code.google.com/p/rnao/
An ontology to capture all aspects of RNA - from primary sequence to alignments, secondary and tertiary structure from base pairing and base stacking to sophisticated motifs.
Proper citation: RNA Ontology (RRID:SCR_003470) Copy
http://purl.obolibrary.org/obo/flu/
An application ontology established by a collaborative group of influenza researchers that includes consolidated influenza sequence and surveillance terms from resources such as the BioHealthBase (BHB), a Bioinformatics Resource Center (BRC) for Biodefense and Emerging and Re-emerging Infectious Diseases, the Centers for Excellence in Influenza Research and Surveillance (CEIRS)
Proper citation: Influenza Ontology (RRID:SCR_003346) Copy
http://blocks.fhcrc.org/blocks/codehop.html
This COnsensus-DEgenerate Hybrid Oligonucleotide Primer (CODEHOP) strategy has been implemented as a computer program that is accessible over the World-Wide Web and is directly linked from the BlockMaker multiple sequence alignment site for hybrid primer prediction beginning with a set of related protein sequences. This is a new primer design strategy for PCR amplification of unknown targets that are related to multiply-aligned protein sequences. Each primer consists of a short 3' degenerate core region and a longer 5' consensus clamp region. Only 3-4 highly conserved amino acid residues are necessary for design of the core, which is stabilized by the clamp during annealing to template molecules. During later rounds of amplification, the non-degenerate clamp permits stable annealing to product molecules. The researchers demonstrate the practical utility of this hybrid primer method by detection of diverse reverse transcriptase-like genes in a human genome, and by detection of C5 DNA methyltransferase homologs in various plant DNAs. In each case, amplified products were sufficiently pure to be cloned without gel fractionation. Sponsors: This work was supported in part by a grant from the M. J. Murdock Charitable Trust and by a grant from NIH. S. P. is a Howard Hughes Medical Institute Fellow of the Life Sciences Research Foundation., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 15,2026.
Proper citation: COnsensus-DEgenerate Hybride Oligonucleotide Primers (RRID:SCR_002875) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on February 28,2023. Software tool for aligning sequences, similar to BLAST 2 sequences that colour-codes the alignments by reliability. Another useful feature of LAST is that it can compare huge (vertebrate-genome-sized) datasets. Unfortunately, this only applies to the downloadable version of LAST, not the web service. The web service can just about handle bacterial genomes, but it will take a few minutes and the output will be large. LAST can: * Handle big sequence data, e.g: ** Compare two vertebrate genomes ** Align billions of DNA reads to a genome * Indicate the reliability of each aligned column. * Use sequence quality data properly. * Compare DNA to proteins, with frameshifts. * Compare PSSMs to sequences * Calculate the likelihood of chance similarities between random sequences. LAST cannot (yet): * Do spliced alignment., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: LAST (RRID:SCR_006119) Copy
http://www.ebi.ac.uk/Tools/sss/fasta/
Software package for DNA and protein sequence alignment to find regions of local or global similarity between Protein or DNA sequences, either by searching Protein or DNA databases, or by identifying local duplications within a sequence.
Proper citation: FASTA (RRID:SCR_011819) Copy
http://cgap.nci.nih.gov/Chromosomes/Mitelman
The web site includes genomic data for humans and mice, including transcript sequence, gene expression patterns, single-nucleotide polymorphisms, clone resources, and cytogenetic information. Descriptions of the methods and reagents used in deriving the CGAP datasets are also provided. An extensive suite of informatics tools facilitates queries and analysis of the CGAP data by the community. One of the newest features of the CGAP web site is an electronic version of the Mitelman Database of Chromosome Aberrations in Cancer. The data in the Mitelman Database is manually culled from the literature and subsequently organized into three distinct sub-databases, as follows: -The sub-database of cases contains the data that relates chromosomal aberrations to specific tumor characteristics in individual patient cases. It can be searched using either the Cases Quick Searcher or the Cases Full Searcher. -The sub-database of molecular biology and clinical associations contains no data from individual patient cases. Instead, the data is pulled from studies with distinct information about: -Molecular biology associations that relate chromosomal aberrations and tumor histologies to genomic sequence data, typically genes rearranged as a consequence of structural chromosome changes. -Clinical associations that relate chromosomal aberrations and/or gene rearrangements and tumor histologies to clinical variables, such as prognosis, tumor grade, and patient characteristics. It can be searched using the Molecular Biology and Clinical (MBC) Associations Searcher -The reference sub-database contains all the references culled from the literature i.e., the sum of the references from the cases and the molecular biology and clinical associations. It can be searched using the Reference Searcher. CGAP has developed six web search tools to help you analyze the information within the Mitelman Database: -The Cases Quick Searcher allows you to query the individual patient cases using the four major fields: aberration, breakpoint, morphology, and topography. -The Cases Full Searcher permits a more detailed search of the same individual patient cases as above, by including more cytogenetic field choices and adding search fields for patient characteristics and references. -The Molecular Biology Associations Searcher does not search any of the individual patient cases. It searches studies pertaining to gene rearrangements as a consequence of cytogenetic aberrations. -The Clinical Associations Searcher does not search any of the individual patient cases. It searches studies pertaining to clinical associations of cytogenetic aberrations and/or gene rearrangements. -The Recurrent Chromosome Aberrations Searcher provides a way to search for structural and numerical abnormalities that are recurrent, i.e., present in two or more cases with the same morphology and topography. -The Reference Searcher queries only the references themselves, i.e., the references from the individual cases and the molecular biology and clinical associations. Sponsors: This database is sponsored by the University of Lund, Sweden and have support from the Swedish Cancer Society and the Swedish Children''s Cancer Foundation
Proper citation: Mitelman Database of Chromosome Aberrations in Cancer (RRID:SCR_012877) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.