Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes
Norway

PMID:37626289  

Novel and improved Caenorhabditis briggsae gene models generated by community curation.

Nicolas D Moya | Lewis Stevens | Isabella R Miller | Chloe E Sokol | Joseph L Galindo | Alexandra D Bardas | Edward S H Koh | Justine Rozenich | Cassia Yeo | Maryanne Xu | Erik C Andersen
BMC genomics | 2023

The nematode Caenorhabditis briggsae has been used as a model in comparative genomics studies with Caenorhabditis elegans because of their striking morphological and behavioral similarities. However, the potential of C. briggsae for comparative studies is limited by the quality of its genome resources. The genome resources for the C. briggsae laboratory strain AF16 have not been developed to the same extent as C. elegans. The recent publication of a new chromosome-level reference genome for QX1410, a C. briggsae wild strain closely related to AF16, has provided the first step to bridge the gap between C. elegans and C. briggsae genome resources. Currently, the QX1410 gene models consist of software-derived gene predictions that contain numerous errors in their structure and coding sequences. In this study, a team of researchers manually inspected over 21,000 gene models and underlying transcriptomic data to repair software-derived errors.

Pubmed ID: 37626289

Associated grants

  • Agency: NIH HHS, United States
    Id: R21 OD030067
  • Agency: NIH HHS, United States
    Id: R21 OD30067
  • Agency: NIH HHS, United States
    Id: T32 GM008449

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


Apollo (tool)

RRID:SCR_001936

A standalone Java application with a GUI (graphical user interface) for editing genome annotations. Like GBrowse, it allows users to scroll and zoom in on areas of interest in a sequence; authorized users can edit annotations and write the changes back to the underlying database. Apollo can run off GFF3 or a Chado database, and it can also integrate with remote services, such as BLAST and Primer BLAST analyses.

View all literature mentions

STAR (tool)

RRID:SCR_004463

Software performing alignment of high-throughput RNA-seq data. Aligns RNA-seq reads to reference genome using uncompressed suffix arrays.

View all literature mentions

Pfam (tool)

RRID:SCR_004726

A database of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Users can analyze protein sequences for Pfam matches, view Pfam family annotation and alignments, see groups of related families, look at the domain organization of a protein sequence, find the domains on a PDB structure, and query Pfam by keywords. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families that may automatically generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans (collections of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM).

View all literature mentions

RepeatMasker (tool)

RRID:SCR_012954

Software tool that screens DNA sequences for interspersed repeats and low complexity DNA sequences. The output of the program is a detailed annotation of the repeats that are present in the query sequence as well as a modified version of the query sequence in which all the annotated repeats have been masked (default: replaced by Ns). Currently over 56% of human genomic sequence is identified and masked by the program. Sequence comparisons in RepeatMasker are performed by one of several popular search engines including nhmmer, cross_match, ABBlast/WUBlast, RMBlast and Decypher. RepeatMasker makes use of curated libraries of repeats and currently supports Dfam ( profile HMM library ) and RepBase ( consensus sequence library ).

View all literature mentions

New England Biolabs (tool)

RRID:SCR_013517

An Antibody supplier

View all literature mentions

BUSCO (tool)

RRID:SCR_015008

Software tool to quantitatively measure genome assembly and annotation completeness based on evolutionarily informed expectations of gene content.

View all literature mentions

RepeatModeler (tool)

RRID:SCR_015027

Sequence analysis software that performs repeat family identification and creates models for sequence data. RepeatModeler utilizes RepeatScout and RECON to identify repeat element boundaries and family relationships.

View all literature mentions

GenomeTools (tool)

RRID:SCR_016120

Software toolkit for biological sequence analysis and -presentation combined into a single binary. It is used for genome analysis, efficient processing of structured genome annotations and contains binaries for sequence and annotation handling, sequence compression, index structure generation and access, annotation visualization.

View all literature mentions

StringTie (tool)

RRID:SCR_016323

Software application for assembling of RNA-Seq alignments into potential transcripts. It enables improved reconstruction of a transcriptome from RNA-seq reads. This transcript assembling and quantification program is implemented in C++ .

View all literature mentions

OrthoFinder (tool)

RRID:SCR_017118

Software Python application for comparative genomics analysis. Finds orthogroups and orthologs, infers rooted gene trees for all orthogroups and identifies all of gene duplcation events in those gene trees, infers rooted species tree for species being analysed and maps gene duplication events from gene trees to branches in species tree, improves orthogroup inference accuracy. Runs set of protein sequence files, one per species, in FASTA format.

View all literature mentions

TransDecoder (tool)

RRID:SCR_017647

Software tool to identify candidate coding regions within transcript sequences, such as those generated by de novo RNA-Seq transcript assembly using Trinity, or constructed based on RNA-Seq alignments to genome using Tophat and Cufflinks.Starts from FASTA or GFF file. Can scan and retain open reading frames (ORFs) for homology to known proteins by using BlastP or Pfam search and incorporate results into obtained selection. Predictions can then be visualized by using genome browser such as IGV.

View all literature mentions

BRAKER (tool)

RRID:SCR_018964

Software tool as pipeline for accurate and automated gene prediction in novel eukaryotic genomes. Automated gene prediction training and gene prediction pipeline.BRAKER1 is eukaryotic genome annotation pipeline. BRAKER2 is extension of BRAKER1 which allows for fully automated training of gene prediction tools GeneMark EX R14, R15, R17, F1 and AUGUSTUS from RNA Seq and/or protein homology information, and that integrates extrinsic evidence from RNA-Seq and protein homology information into prediction.

View all literature mentions

VSEARCH (tool)

RRID:SCR_024494

Software versatile open source tool for metagenomics. Used for processing and preparing metagenomics, genomics and population genomics nucleotide sequence data.

View all literature mentions