NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Abstract

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

Keywords

RefSeqGenBankBiologySequence databaseSequence (biology)GenomeAnnotationEnsemblComputational biologyBioinformaticsGeneticsDatabaseGeneGenomicsComputer science

Affiliated Institutions

National Institutes of Health US

Related Publications

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Kim D. Pruitt

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) provides a non-redundant collection of sequen...

2004 Nucleic Acids Research 1622 citations

RefSeq: an update on mammalian reference sequences

Kim D. Pruitt , Garth Brown , Susan M. Hiatt +26 more

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database is a collection of annotated genomic, transcript and protein sequence records deriv...

2013 Nucleic Acids Research 994 citations

NCBI Reference Sequences (RefSeq): current status, new features and genome annotation policy

Kim D. Pruitt , Tatiana Tatusova , Garth Brown +1 more

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database is a collection of genomic, transcript and protein sequence records. These records ...

2011 Nucleic Acids Research 1166 citations

RefSeq and LocusLink: NCBI gene-centered resources

Kim D. Pruitt

Thousands of genes have been painstakingly identified and characterized a few genes at a time. Many thousands more are being predicted by large scale cDNA and genomic sequencing...

2001 Nucleic Acids Research 913 citations

Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation

Nuala A. O’Leary , Matt W. Wright , J. Rodney Brister +52 more

The RefSeq project at the National Center for Biotechnology Information (NCBI) maintains and curates a publicly available database of annotated genomic, transcript, and protein ...

2015 Nucleic Acids Research 6668 citations

Publication Info

Year: 2006
Type: article
Volume: 35
Issue: Database
Pages: D61-D65
Citations: 4555
Access: Closed

External Links

View on DOI.org

Social Impact

Altmetric

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

PlumX Metrics

Social media, news, blog, policy document mentions

Citation Metrics

4555

OpenAlex

Cite This

APA Style

                            
                                    Kim D. Pruitt, 
                                
                                    Tatiana Tatusova, 
                                
                                    D. R. Maglott
                                
                            (2006). 
                            NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins. 
                            Nucleic Acids Research
                            , 35
                            (Database)
                            , D61-D65.
                            https://doi.org/10.1093/nar/gkl842

Identifiers

DOI: 10.1093/nar/gkl842