Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://weizhong-lab.ucsd.edu/cd-hit/
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on February 28,2023. Software program for clustering biological sequences with many applications in various fields such as making non-redundant databases, finding duplicates, identifying protein families, filtering sequence errors and improving sequence assembly etc. It is very fast and can handle extremely large databases. CD-HIT helps to significantly reduce the computational and manual efforts in many sequence analysis tasks and aids in understanding the data structure and correct the bias within a dataset. The CD-HIT package has CD-HIT, CD-HIT-2D, CD-HIT-EST, CD-HIT-EST-2D, CD-HIT-454, CD-HIT-PARA, PSI-CD-HIT, CD-HIT-OTU and over a dozen scripts. * CD-HIT (CD-HIT-EST) clusters similar proteins (DNAs) into clusters that meet a user-defined similarity threshold. * CD-HIT-2D (CD-HIT-EST-2D) compares 2 datasets and identifies the sequences in db2 that are similar to db1 above a threshold. * CD-HIT-454 identifies natural and artificial duplicates from pyrosequencing reads. * CD-HIT-OTU cluster rRNA tags into OTUs The usage of other programs and scripts can be found in CD-HIT user''s guide. CD-HIT was originally developed by Dr. Weizhong Li at Dr. Adam Godzik''s Lab at the Burnham Institute (now Sanford-Burnham Medical Research Institute)., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: CD-HIT (RRID:SCR_007105) Copy
http://math.mcb.berkeley.edu/~meromit/MetMap/
A computational pipeline for the analysis of MethylSeq experiments., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: MetMap (RRID:SCR_006954) Copy
http://bowtie-bio.sourceforge.net/myrna/index.shtml
A cloud computing tool for calculating differential gene expression in large RNA-seq datasets. It uses Bowtie for short read alignment and R/Bioconductor for interval calculations, normalization, and statistical testing. These tools are combined in an automatic, parallel pipeline that runs in the cloud (Elastic MapReduce in this case) on a local Hadoop cluster, or on a single computer, exploiting multiple computers and CPUs wherever possible.
Proper citation: Myrna (RRID:SCR_006951) Copy
https://github.com/jstjohn/SimSeq
An illumina paired-end and mate-pair short read simulator. This project attempts to model as many of the quirks that exist in Illumina data as possible. Some of these quirks include the potential for chimeric reads, and non-biotinylated fragment pull down in mate-pair libraries .
Proper citation: SimSeq (RRID:SCR_006947) Copy
http://www.bioconductor.org/packages/release/bioc/html/qvalue.html
R package that takes a list of p-values resulting from the simultaneous testing of hypotheses and estimates their q-values. It is designed to measure the proportion of false positives when a test is significant. The software is capable of generating plots for visualization. It can be applied to problems in genomics, brain imaging, astrophysics, and data mining.
Proper citation: Qvalue (RRID:SCR_001073) Copy
A tool for performing multi-cluster gene functional enrichment analyses on large scale data (microarray experiments with many time-points, cell-types, tissue-types, etc.). It facilitates co-analysis of multiple gene lists and yields as output a rich functional map showing the shared and list-specific functional features. The output can be visualized in tabular, heatmap or network formats using built-in options as well as third-party software. It uses the hypergeometric test to obtain functional enrichment achieved via the gene list enrichment analysis option available in ToppGene.
Proper citation: ToppCluster (RRID:SCR_001503) Copy
Software package for Bayesian analysis of protein, DNA and RNA sequences. It utilizes multiple alignments, phylogenetic trees and evolutionary parameters to quantify uncertainty in these analyses. It is written in Java.
Proper citation: StatAlign (RRID:SCR_001892) Copy
http://www.nactem.ac.uk/facta/
Text mining tool to discover associations between biomedical concepts from MEDLINE articles. Use the service from your browser or via a Web Service. The whole MEDLINE corpus containing more than 20 million articles is indexed with an efficient text search engine, and it allows you to navigate such associations and their textual evidence in a highly interactive manner - the system accepts arbitrary query terms and displays relevant concepts immediately. A broad range of important biomedical concepts are covered by the combination of a machine learning-based term recognizer and large-scale dictionaries for genes, proteins, diseases, and chemical compounds. There is also a FACTA+ visualization service that can be found here: http://www.nactem.ac.uk/facta-visualizer/
Proper citation: FACTA+. (RRID:SCR_001767) Copy
http://gmdd.shgmo.org/Computational-Biology/GRS/
A compression tool for efficient storage of Genome Re-Sequencing data. GRS processes genome sequence data without use of reference SNPs and other variants. It can also automatically rebuild the individual genome sequence data using the reference genome sequence.
Proper citation: GRS (RRID:SCR_001008) Copy
http://ntap.cbi.pku.edu.cn/usage.php
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 23,2022. Software for tiling array data analysis to survey the genome-wide binding sites of transcription factor HY5 in Arabidopsis and the genome-wide histone modifications/DNA methylation level in rice. It was developed in the process of generating NimbleGen analysis. Written in R and Perl.
Proper citation: NTAP (RRID:SCR_001488) Copy
https://github.com/ndaniel/fusioncatcher
Software that searches for novel/known fusion genes, translocations, and chimeras in RNA-seq data (paired-end reads from Illumina NGS platforms like Solexa and HiSeq) from diseased samples.
Proper citation: FusionCatcher (RRID:SCR_000060) Copy
http://dissect-trans.sourceforge.net/Home
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on July 31,2025. Software transcriptome-to-genome alignment tool, which can identify and characterize transcriptomic events such as duplications, inversions, rearrangements and fusions.
Proper citation: Dissect (RRID:SCR_000058) Copy
A curated collection of chaperonin sequence data collected from public databases or generated by a network of collaborators exploiting the cpn60 target in clinical, phylogenetic and microbial ecology studies. The database contains all available sequences for both group I and group II chaperonins. Users can search the database by Chaperonin type, group (I or II), BLAST, or other options, and can also enter and analyze FASTA sequences.
Proper citation: cpnDB: A Chaperonin Database (RRID:SCR_002263) Copy
http://bioconductor.org/packages/2.8/bioc/html/qrqc.html
Software R package to quickly scan reads and gather statistics on base and quality frequencies, read length, k-mers by position, and frequent sequences. Produces graphical output of statistics for use in quality control pipelines, and an optional HTML quality report. S4 SequenceSummary objects allow specific tests and functionality to be written around the data collected.
Proper citation: qrqc (RRID:SCR_006867) Copy
A database for phenotyping human single nucleotide polymorphisms (SNPs)that primarily focuses on the molecular characterization and annotation of disease and polymorphism variants in the human proteome. They provide a detailed variant analysis using their tools such as: * TANGO to predict aggregation prone regions * WALTZ to predict amylogenic regions * LIMBO to predict hsp70 chaperone binding sites * FoldX to analyse the effect on structure stability Further, SNPeffect holds per-variant annotations on functional sites, structural features and post-translational modification. The meta-analysis tool enables scientists to carry out a large scale mining of SNPeffect data and visualize the results in a graph. It is now possible to submit custom single protein variants for a detailed phenotypic analysis., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: SNPeffect (RRID:SCR_005091) Copy
http://jilab.biostat.jhsph.edu/database/cgi-bin/hmChIP.pl
A database of genome-wide chromatin immunoprecipitation (ChIP) data in human and mouse. Currently, the database contains >2000 samples from >500 ChIP-seq and ChIP-chip experiments, representing a total of >170 proteins and >10,000,000 protein-DNA interactions (March 2014). A web server provides an interface for database query. Protein-DNA binding intensities can be retrieved from individual samples for user-provided genomic regions. The retrieved intensities can be used to cluster samples and genomic regions to facilitate exploration of combinatorial patterns, cell type dependencies, and cross-sample variability of protein-DNA interactions.
Proper citation: hmChIP (RRID:SCR_005407) Copy
http://amp.pharm.mssm.edu/lib/chea.jsp
Data analysis service for gene-list enrichment analysis against a manual database. It allows users to input lists of mammalian gene symbols for which the program computes over-representation of transcription factor targets from the ChIP-X database. The database integrates interaction data from ChIP-chip, ChIP-seq, ChIP-PET and DamID studies and contains 189,933 interactions, manually extracted from 87 publications, describing the binding of 92 transcription factors to 31,932 target genes.
Proper citation: ChEA (RRID:SCR_005403) Copy
http://sourceforge.net/projects/hadoop-bam/
A Java library for the manipulation of files in common bioinformatics formats using the Hadoop MapReduce framework with the Picard SAM JDK, and command line tools similar to SAMtools. The file formats currently supported are BAM, SAM, FASTQ, FASTA, QSEQ, BCF, and VCF.
Proper citation: Hadoop-BAM (RRID:SCR_005516) Copy
Database on transcriptional regulation in Escherichia coli K-12 containing knowledge manually curated from original scientific publications, complemented with high throughput datasets and comprehensive computational predictions. Graphic and text-integrated environment with friendly navigation where regulatory information is always at hand. They provide integrated views to understand as well as organized knowledge in computable form. Users may submit data to make it publicly available.
Proper citation: RegulonDB (RRID:SCR_003499) Copy
Freely accessible phenotype-centered database with integrated analysis and visualization tools. It combines diverse data sets from multiple species and experiment types, and allows data sharing across collaborative groups or to public users. It was conceived of as a tool for the integration of biological functions based on the molecular processes that subserved them. From these data, an empirically derived ontology may one day be inferred. Users have found the system valuable for a wide range of applications in the arena of functional genomic data integration.
Proper citation: Gene Weaver (RRID:SCR_003009) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the dkNET Resources search. From here you can search through a compilation of resources used by dkNET and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that dkNET has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on dkNET then you can log in from here to get additional features in dkNET such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into dkNET you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within dkNET that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.