<?xml version="1.0"?>
<!DOCTYPE book PUBLIC "-//NLM//DTD Book DTD v2.1 20050630//EN" "book.dtd">
<book>
<book-meta>
<book-id pub-id-type="other">handbook</book-id>
<book-title-group>
<book-title>The NCBI Handbook</book-title>
</book-title-group>
<edition>1st</edition>
<contrib-group>
<contrib contrib-type="editor" rid="bid.m.1">
<name><surname>McEntyre</surname>
<given-names>Jo</given-names></name>
</contrib>
<contrib contrib-type="editor" rid="bid.m.1">
<name><surname>Ostell</surname>
<given-names>Jim</given-names></name>
</contrib>
</contrib-group>
<aff id="bid.m.1">
<institution>National Center for Biotechnology Information 
(NCBI), National Library of Medicine, National Institutes 
of Health</institution>, 
<addr-line>Bethesda, MD 20892-6510</addr-line>
</aff>
<publisher>
<publisher-name>National Center for Biotechnology Information 
(NCBI), National Library of Medicine, National Institutes of 
Health</publisher-name>
<publisher-loc>Bethesda, MD</publisher-loc>
</publisher>
<pub-date><month>Nov.</month><year>2002</year></pub-date>
<counts>
<fig-count count="98"/>
<table-count count="40"/>
<equation-count count="0"/>
<ref-count count="115"/>
<page-count count="532"/>
<word-count count="149852"/>
</counts>
</book-meta>
<?mtl hellip?><book-front>
    <title>About this book</title>
    <sec sec-type="miscinfo">
    <title>The NCBI Handbook</title>
      <p>Bioinformatics consists of a computational approach to biomedical 
      information management and analysis. It is being used increasingly as 
      a component of research within both academic and industrial settings 
      and is becoming integrated into both undergraduate and postgraduate 
      curricula. The new generation of biology graduates is emerging with 
      experience in using bioinformatics resources and, in some cases, 
      programming skills.</p>
      <p>The National Center for Biotechnology Information (NCBI) is one 
      of the world's premier Web sites for biomedical and bioinformatics 
      research. Based within the National Library of Medicine at the National 
      Institutes of Health, USA, the NCBI hosts many databases used by 
      biomedical and research professionals. The services include PubMed, 
      the bibliographic database; GenBank, the nucleotide sequence database; 
      and the BLAST algorithm for sequence comparison, among many others. 
      The NCBI Web site is visited by about 250,000 people per day.</p>
      <p>Although each NCBI resource has online help documentation associated 
      with it, there is no cohesive approach to describing the databases and 
      search engines, nor any significant information on how the databases 
      work or how they can be leveraged, for bioinformatics research on a 
      larger scale. The NCBI Handbook is designed to address this information 
      gap.</p>
      <p>All of our users know how to execute a straightforward PubMed or 
      BLAST search. However, feedback from help desk personnel and booth 
      staff at scientific meetings suggests that people often want to know 
      how to use our resources in a more sophisticated manner and are 
      frequently unaware of less well-known databases that might be helpful 
      to them. The intended audience for The NCBI Handbook is, therefore, 
      the growing number of scientists and students who would like a more 
      in-depth guide to NCBI resources&mdash;powerusers and aspiring 
      powerusers.</p>
      <p>The NCBI Handbook is focused on the relatively stable information 
      about each resource; it is not a point-and-click user guide (this 
      type of information can be found in the online help documents, 
      referred to frequently but not repeated, in the Handbook). Each 
      chapter is devoted to one service; after a brief overview on using 
      the resource, there is an account of how the resource works, including 
      topics such as how data are included in a database, database design, 
      query processing, and how the different resources relate to each 
      other. For example, the BLAST chapter briefly describes what to use 
      BLAST for, the various varieties of the BLAST algorithm, and BLAST 
      statistics, before discussing output formats, query processing, and 
      tips for setting up a BLAST database. A certain amount of biological 
      knowledge is assumed.</p>
      <p>The online content will be updated when necessary, although major 
      changes are not expected to occur more than once every few years. 
      (For example, PubMed query processing does not change dramatically 
      year after year.) We hope that The NCBI Handbook will provide a 
      valuable reference for anyone who wants to use our resources more 
      effectively.</p>
    </sec>
  </book-front><?mtl /hellip?>
<body>
<book-part id="bid.1" book-part-type="part" book-part-number="Part 1">
<book-part-meta>
<title-group>
<title>The Databases</title>
</title-group>
</book-part-meta>
<body>
<?mtl hellip?><book-part id="bid.2" book-part-type="chapter" book-part-number="1">
          <book-part-meta>
            <title-group>
              <title>GenBank: The Nucleotide Sequence Database</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Mizrachi</surname>
                  <given-names>Ilene</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>27</day>
                <month>7</month>
                <year>2004</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch1d1"/>
            <abstract>
              <title>Summary</title>
              <p>The GenBank sequence database is an annotated collection of all publicly available nucleotide sequences and their protein translations. This database is produced at National Center for Biotechnology Information (<xref ref-type="other" rid="bid.1353">NCBI</xref>) as part of an international collaboration with the European Molecular Biology Laboratory (EMBL) Data Library from the European Bioinformatics Institute (EBI) and the DNA Data Bank of Japan (DDBJ). GenBank and its collaborators receive sequences produced in laboratories throughout the world from more than 100,000 distinct organisms. GenBank continues to grow at an exponential rate, doubling every 10 months. Release 134, produced in February 2003, contained over 29.3 billion nucleotide bases in more than 23.0 million sequences. GenBank is built by direct submissions from individual laboratories, as well as from bulk submissions from large-scale sequencing centers.</p>
              <p>Direct submissions are made to GenBank using <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BankIt/">BankIt</ext-link>, which is a Web-based form, or the stand-alone submission program, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/index.html">Sequin</ext-link>. Upon receipt of a sequence submission, the GenBank staff assigns an Accession number to the sequence and performs quality assurance checks. The submissions are then released to the public database, where the entries are retrievable by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Nucleotide">Entrez</ext-link> or downloadable by <xref ref-type="other" rid="bid.1295">FTP</xref>. Bulk submissions of Expressed Sequence Tag (<xref ref-type="other" rid="bid.1283">EST</xref>), Sequence Tagged Site (<xref ref-type="other" rid="bid.1410">STS</xref>), Genome Survey Sequence (<xref ref-type="other" rid="bid.1305">GSS</xref>), and High-Throughput Genome Sequence (<xref ref-type="other" rid="bid.1311">HTGS</xref>) data are most often submitted by large-scale sequencing centers. The GenBank direct submissions group also processes complete microbial genome sequences.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.3">
              <title>History</title>
              <p>Initially, GenBank was built and maintained at Los Alamos National Laboratory (<xref ref-type="other" rid="bid.1329">LANL</xref>). In the early 1990s, this responsibility was awarded to NCBI through congressional mandate. NCBI undertook the task of scanning the literature for sequences and manually typing the sequences into the database. Staff then added annotation to these records, based upon information in the published article. Scanning sequences from the literature and placing them into GenBank is now a rare occurrence. Nearly all of the sequences are now deposited directly by the labs that generate the sequences. This is attributable to, in part, a requirement by most journal publishers that nucleotide sequences are first deposited into publicly available databases (DDBJ/EMBL/GenBank) so that the Accession number can be cited and the sequence can be retrieved when the article is published. NCBI began accepting direct submissions to GenBank in 1993 and received data from LANL until 1996. Currently, NCBI receives and processes about 20,000 direct submission sequences per month, in addition to the approximately 200,000 bulk submissions that are processed automatically.</p>
            </sec>
            <sec id="bid.4">
              <title>International Collaboration</title>
              <p>In the mid-1990s, the GenBank database became part of the International Nucleotide Sequence Database Collaboration with the EMBL database (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ebi.ac.uk/">European Bioinformatics Institute</ext-link>, Hinxton, United Kingdom) and the Genome Sequence Database (GSDB; LANL, Los Alamos, NM). Subsequently, the GSDB was removed from the Collaboration (by the National Center for Genome Resources, Santa Fe, NM), and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ddbj.nig.ac.jp/">DDBJ</ext-link> (Mishima, Japan) joined the group. Each database has its own set of submission and retrieval tools, but the three databases exchange data daily so that all three databases should contain the same set of sequences. Members of the DDBJ, EMBL, and GenBank staff meet annually to discuss technical issues, and an international advisory board meets with the database staff to provide additional guidance. An entry can only be updated by the database that initially prepared it to avoid conflicting data at the three sites.</p>
              <p>The Collaboration created a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/collab/FT/index.html">Feature Table Definition</ext-link> that outlines legal features and syntax for the DDBJ, EMBL, and GenBank feature tables. The purpose of this document is to standardize annotation across the databases. The presentation and format of the data are different in the three databases, however, the underlying biological information is the same.</p>
            </sec>
            <sec id="bid.5">
              <title>Confidentiality of Data</title>
              <p>When scientists submit data to GenBank, they have the opportunity to keep their data confidential for a specified period of time. This helps to allay concerns that the availability of their data in GenBank before publication may compromise their work. When the article containing the citation of the sequence or its Accession number is published, the sequence record is released. The database staff request that submitters notify GenBank of the date of publication so that the sequence can be released without delay. The request to release should be sent to <email>gb-admin@ncbi.nlm.nih.gov</email>.</p>
            </sec>
            <sec id="bid.6">
              <title>Direct Submissions</title>
              <p>The typical GenBank submission consists of a single, contiguous stretch of DNA or RNA sequence with annotations. The annotations are meant to provide an adequate representation of the biological information in the record. The GenBank <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/collab/FT/index.html">Feature Table Definition</ext-link> describes the various features and subsequent qualifiers agreed upon by the International Nucleotide Sequence Database Collaboration.</p>
              <p>Currently, only nucleotide sequences are accepted for direct submission to GenBank. These include mRNA sequences with coding regions, fragments of genomic DNA with a single gene or multiple genes, and ribosomal RNA gene clusters. If part of the nucleotide sequence encodes a protein, a conceptual translation, called a <xref ref-type="other" rid="bid.1259">CDS</xref> (coding sequence), is annotated. The span of the CDS feature is mapped to the nucleotide sequence encoding the protein. A protein Accession number (/protein_id) is assigned to the translation product, which will subsequently be added to the protein databases.</p>
              <p>Multiple sequences can be submitted together. Such batch submissions of non-related sequences may be processed together but will be displayed in Entrez (<xref ref-type="other" rid="bid.588">Chapter 15</xref>) as single records. Alternatively, by using the Sequin submission tool (<xref ref-type="other" rid="bid.538">Chapter 12</xref>), a submitter can specify that several sequences are biologically related. Such sequences are classified as environmental sample sets, population sets, phylogenetic sets, mutation sets, or segmented sets. Each sequence within a set is assigned its own Accession number and can be viewed independently in Entrez. However, with the exception of segmented sets, each set is also indexed within the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Popset">PopSet</ext-link> division of Entrez, thus allowing scientists to view the relationship between the sequences.</p>
              <p>What defines a set? Environmental sample, population, phylogenetic, and mutation sets all contain a group of sequences that spans the same gene or region of the genome. Environmental samples are derived from a group of unclassified or unknown organisms. A population set contains sequences from different isolates of the same organism. A phylogenetic set contains sequences from different organisms that are used to determine the phylogenetic relationship between them. Sequencing multiple mutations within a single gene gives rise to a mutation set.</p>
              <p>All sets, except segmented sets, may contain an alignment of the sequences within them and might include external sequences already present in the database. In fact, the submitter can begin with an existing alignment to create a submission to the database using the Sequin submission tool. Currently, Sequin accepts <xref ref-type="other" rid="bid.1290">FASTA</xref>+GAP, <xref ref-type="other" rid="bid.1373">PHYLIP</xref>, <xref ref-type="other" rid="bid.1335">MACAW</xref>, <xref ref-type="other" rid="bid.1356">NEXUS</xref> Interleaved, and NEXUS Contiguous alignments. Submitted alignments will be displayed in the PopSet section of Entrez.</p>
              <p>Segmented sets are a collection of noncontiguous sequences that cover a specified genetic region. The most common example is a set of genomic sequences containing <xref ref-type="other" rid="bid.1287">exon</xref>s from a single gene where part or all of the intervening regions have not been sequenced. Each member record within the set contains the appropriate annotation, exon features in this case. However, the mRNA and CDS will be annotated as joined features across the individual records. Segmented sets themselves can be part of an environmental sample, population, phylogenetic, or mutation set.</p>
            </sec>
            <sec id="bid.7">
              <title>Bulk Submissions: High-Throughput Genomic Sequence (HTGS)</title>
              <p><xref ref-type="other" rid="bid.1311">HTGS</xref> entries are submitted in bulk by genome centers, processed by an automated system, and then released to GenBank. Currently, about 30 genome centers are submitting data for a number of organisms, including human, mouse, rat, rice, and <italic>Plasmodium falciparum</italic>, the malaria parasite.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/HTGS/">HTGS</ext-link> data are submitted in four phases of completion: 0, 1, 2, and 3. Phase 0 sequences are one-to-few reads of a single clone and are not usually assembled into <xref ref-type="other" rid="bid.1267">contig</xref>s. They are low-quality sequences that are often used to check whether another center is already sequencing a particular clone. Phase 1 entries are assembled into contigs that are separated by sequence gaps, the relative order and orientation of which are not known (<xref ref-type="fig" rid="bid.8">Figure 1</xref>). Phase 2 entries are also unfinished sequences that may or may not contain sequence gaps. If there are gaps, then the contigs are in the correct order and orientation. Phase 3 sequences are of finished quality and have no gaps. For each organism, the group overseeing the sequencing effort determines the definition of finished quality.<fig id="bid.8"><label>1</label><caption><p>Diagram showing the orientation and gaps that might be expected in high-throughput sequence from phases 1, 2, and 3.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch1f1" mime-subtype="gif"/></fig></p>
              <p>Phase 0, 1, and 2 records are in the HTG division of GenBank, whereas phase 3 entries go into the taxonomic division of the organism, for example, PRI (primate) for human. An entry keeps its Accession number as it progresses from one phase to another but receives a new Accession.Version number and a new gi number each time there is a sequence change.</p>
              <sec id="bid.9">
                <title>Submitting Data to the HTG Division</title>
                <p>To submit sequences in bulk to the HTG processing system, a center or group must set up an FTP account by writing to <email>htgs-admin@ncbi.nlm.nih.gov</email>. Submitters frequently use two tools to create HTG submissions, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/HTGS/sequininfo.html">Sequin</ext-link> or <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/HTGS/fa2htgsinfo.html">fa2htgs</ext-link>. Both of these tools require <xref ref-type="other" rid="bid.1290">FASTA</xref>-formatted sequence, i.e., a definition line beginning with a &ldquo;greater than&rdquo; sign (&ldquo;&gt;&rdquo;) followed by a unique identifier for the sequence. The raw sequence appears on the lines after the definition line. For sequences composed of contigs separated by gaps, a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/HTGS/sequininfo.html">modified FASTA format</ext-link> is used. In addition, Sequin users must modify the Sequin configuration file so that the HTG genome center features are enabled.</p>
                <p>fa2htgs is a command-line program that is downloaded to the user's computer. The submitter invokes a script with a series of parameters (arguments) to create a submission. It has an advantage over Sequin in that it can be set up by the user to create submissions in bulk from multiple files.</p>
                <p>Submissions to HTG must contain three identifiers that are used to track each HTG record: the genome center tag, the sequence name, and the Accession number. The genome center tag is assigned by NCBI and is generally the FTP account login name. The sequence name is a unique identifier that is assigned by the submitter to a particular clone or entry and must be unique within the group's submissions. When a sequence is first submitted, it has only a sequence name and genome center tag; the Accession number is assigned during processing. All updates to that entry must include the center tag, sequence name, and Accession number, or processing will fail.</p>
              </sec>
              <sec id="bid.10">
                <title>The HTG Processing Pathway</title>
                <p>Submitters deposit HTGS sequences in the form of Seq-submit files generated by Sequin, fa2htgs, or their own <xref ref-type="other" rid="bid.1242">ASN.1</xref> dumper tool into the SEQSUBMIT directory of their FTP account. Every morning, scripts automatically pick up the files from the FTP site and copy them to the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/HTGS/processing.html">processing</ext-link> pathway, as well as to an archive. Once processing is complete and if there are no errors in the submission, the files are automatically loaded into GenBank. The processing time is related to the number of submissions that day; therefore, processing can take from one to many hours.</p>
                <p>Entries can fail HTG processing because of three types of problems:
                  <list list-type="arabic">
                    <list-item id="bid.11">
                      <label>1</label>
                      <p>Formatting: submissions are not in the proper Seq-submit format.</p>
                    </list-item>
                    <list-item id="bid.12">
                      <label>2</label>
                      <p>Identification: submissions may be missing the genome center tag, 
                      sequence name, or Accession number, or this information is incorrect.</p>
                    </list-item>
                    <list-item id="bid.13">
                      <label>3</label>
                      <p>Data: submissions have problems with the data and therefore fail 
                      the validator checks.</p>
                    </list-item>
                  </list>
                </p>
                <p>When submissions fail HTG processing, a GenBank annotator sends email to the sequencing center, describing the problem and asking the center to submit a corrected entry. Annotators do not fix incorrect submissions; this ensures that the staff of the submitting genome center fixes the problems in their database as well.</p>
                <p>The processing pathway also generates reports. For successful submissions, two files are generated: one contains the submission in GenBank flat file format (without the sequence); and another is a status report file. The status report file, ac4htgs, contains the genome center, sequence name, Accession number, phase, create date, and update date for the submission. Submissions that fail processing receive an error file with a short description of the error(s) that prevented processing. The GenBank annotator also sends email to the submitter, explaining the errors in further detail.</p>
              </sec>
              <sec id="bid.14">
                <title>Additional Quality Assurance</title>
                <p>When successful submissions are loaded into GenBank, they undergo additional validation checks. If GenBank annotators find errors, they write to the submitters, asking them to fix these errors and submit an update.</p>
              </sec>
            </sec>
            <sec id="bid.1738">
              <title>Whole Genome Shotgun Sequences (WGS)</title>
              <p>Genome centers are taking multiple approaches to sequencing complete genomes from a number of organisms. In addition to the traditional clone-based sequencing whose data are being submitted to HTGS, these centers are also using a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/GenBank/wgs.html">WGS</ext-link> approach to sequence the genome. The shotgun sequencing reads are assembled into contigs, which are now being accepted for inclusion in GenBank. WGS contig assemblies may be updated as the sequencing project progresses and new assemblies are computed. WGS sequence records may also contain annotation, similar to other GenBank records.</p>
              <p>Each sequencing project is assigned a stable project ID, which is made up of four letters. The Accession number for a WGS sequence contains the project ID, a two-digit version number, and six digits for the contig ID. For instance, a project would be assigned an Accession number AAAX00000000. The first assembly version would be AAAX01000000. The last six digits of this ID identify individual contigs. A master record for each assembly is created. This master record contains information that is common among all records of the sequencing project, such as the biological source, submitter, and publication information. There is also a link to the range of Accession numbers for the individual contigs in this assembly.</p>
              <p>WGS submissions can be created using tbl12asn, a utility that is packaged with the Sequin submission software. Information on submitting these sequences can be found at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Genbank/wgs.html">Whole Genome Shotgun Submissions</ext-link>.</p>
            </sec>
            <sec id="bid.15">
              <title>Bulk Submissions: EST, STS, and GSS</title>
              <p>
			Expressed Sequence Tags (<xref ref-type="other" rid="bid.1283">EST</xref>), Sequence Tagged Sites (<xref ref-type="other" rid="bid.1410">STS</xref>s), and Genome Survey Sequences (<xref ref-type="other" rid="bid.1305">GSS</xref>s) sequences are generally submitted in a batch and are usually part of a large sequencing project devoted to a particular genome. These entries have a streamlined submission process and undergo minimal processing before being loaded to GenBank.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/dbEST/">EST</ext-link>s are generally short (&lt;1 kb), single-pass <xref ref-type="other" rid="bid.1258">cDNA</xref> sequences from a particular tissue and/or developmental stage. However, they can also be longer sequences that are obtained by differential display or Rapid Amplification of cDNA Ends (RACE) experiments. The common feature of all ESTs is that little is known about them; therefore, they lack feature annotation.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/dbSTS/">STS</ext-link>s are short genomic landmark sequences (<xref ref-type="bibr" rid="bid.41">1</xref>). They are operationally unique in that they are specifically amplified from the genome by PCR amplification. In addition, they define a specific location on the genome and are, therefore, useful for mapping.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/dbGSS/">GSS</ext-link>s are also short sequences but are derived from genomic DNA, about which little is known. They include, but are not limited to, single-pass GSSs, <xref ref-type="other" rid="bid.1243">BAC</xref> ends, <xref ref-type="other" rid="bid.1288">exon-trapped</xref> genomic sequences, and <xref ref-type="other" rid="bid.1239">Alu</xref><xref ref-type="other" rid="bid.1366">PCR</xref> sequences.</p>
              <p>EST, STS, and GSS sequences reside in their respective divisions within GenBank, rather than in the taxonomic division of the organism. The sequences are maintained within GenBank in the dbEST, dbSTS, and dbGSS databases.</p>
              <sec id="bid.16">
                <title>Submitting Data to dbEST, dbSTS, or dbGSS</title>
                <p>Because of the large numbers of sequences that are submitted at once, dbEST, dbSTS, and dbGSS entries are stored in relational databases where information that is common to all sequences can be shared. Submissions consist of several files containing the common information, plus a file of the sequences themselves. The three types of submissions have different requirements, but all include a Publication file and a Contact file. See the<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/dbEST/"> dbEST</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/dbSTS/"> dbSTS</ext-link>, and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/dbGSS/"> dbGSS</ext-link> pages for the specific requirements for each type of submission.</p>
                <p>In general, users generate the appropriate files for the submission type and then email the files to <email>batch-sub@ncbi.nlm.nih.gov</email>. If the files are too big for email, they can be deposited into a FTP account. Upon receipt, the files are examined by a GenBank annotator, who fixes any errors when possible or contacts the submitter to request corrected files. Once the files are satisfactory, they are loaded into the appropriate database and assigned Accession numbers. Additional formatting errors may be detected at this step by the data-loading software, such as double quotes anywhere in the file or invalid characters in the sequences. Again, if the annotator cannot fix the errors, a request for a corrected submission is sent to the user. After all problems are resolved, the entries are loaded into GenBank.</p>
              </sec>
            </sec>
            <sec id="bid.17">
              <title>Bulk Submissions: HTC and FLIC</title>
              <p>HTC records are High-Throughput cDNA/mRNA submissions that are similar to ESTs but often contain more information. For example, HTC entries often have a systematic gene name (not necessarily an official gene name) that is related to the lab or center that submitted them, and the longest open reading frame is often annotated as a coding region.</p>
              <p>FLIC records, Full-Length Insert cDNA, contain the entire sequence of a cloned cDNA/mRNA. Therefore, FLICs are generally longer, and sometimes even full-length, mRNAs. They are usually annotated with genes and coding regions, although these may be lab systematic names rather than functional names.</p>
              <sec id="bid.1739">
                <title>HTC Submissions</title>
                <p>HTC entries are usually generated with <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/index.html">Sequin</ext-link> or <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/table.html">tbl2asn</ext-link>, and the files are emailed to <email>gb-sub@ncbi.nlm.nih.gov</email>. If the files are too big for email, then by prior arrangement, the submitter can deposit the files by FTP and send a notification to <email>gb-admin@ncbi.nlm.nih.gov</email> that files are on the FTP site.</p>
                <p>HTC entries undergo the same validation and processing as non-bulk submissions. Once processing is complete, the records are loaded into GenBank and are available in Entrez and other retrieval systems.</p>
              </sec>
              <sec id="bid.1740">
                <title>FLIC Submissions</title>
                <p>FLICs are processed via an automated FLIC processing system that is based on the HTG automated processing system. Submitters use the program tbl2asn to generate their submissions. As with HTG submissions, submissions to the automated FLIC processing system must contain three identifiers: the genome center tag, the sequence name (SeqId), and the Accession number. The genome center tag is assigned by NCBI and is generally the FTP account login name. The sequence name is a unique identifier that is assigned by the submitter to a particular clone or entry and must be unique within the group's FLIC submissions. When a sequence is first submitted, it has only a sequence name and genome center tag; the Accession number is assigned during processing. All updates to that entry include the center tag, sequence name, and Accession number, or processing will fail.</p>
              </sec>
              <sec id="bid.1741">
                <title>The FLIC Processing Pathway</title>
                <p>The FLIC processing system is analogous to the HTG processing system. Submitters deposit their submissions in the FLICSEQSUBMIT directory of their FTP account and notify us that the submissions are there. We then run the scripts to pick up the files from the FTP site and copy them to the processing pathway, as well as to an archive. Once processing is complete and if there are no errors in the submission, the files are automatically loaded into GenBank.</p>
                <p>As with HTG submissions, FLIC entries can fail for three reasons: problems with the format, problems with the identification of the record (the genome center, the SeqId, or the Accession number), or problems with the data itself. When submissions fail FLIC processing, a GenBank annotator sends email to the sequencing center, describing the problem and asking the center to submit a corrected entry. Annotators do not fix incorrect submissions; this ensures that the staff of the submitting genome center fixes the problems in their database as well. At the completion of processing, reports are generated and deposited in the submitter's FTP account, as described for HTG submissions.</p>
              </sec>
            </sec>
            <sec id="bid.18">
              <title>Submission Tools</title>
              <p>Direct submissions to GenBank are prepared using one of two submission tools, BankIt or Sequin.</p>
              <sec id="bid.19">
                <title>BankIt</title>
                <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BankIt/">BankIt</ext-link> is a Web-based form that is a convenient and easy way to submit a small number of sequences with minimal annotation to GenBank. To complete the form, a user is prompted to enter submitter information, the nucleotide sequence, biological source information, and features and annotation pertinent to the submission. BankIt has extensive <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BankIt/help.html">Help</ext-link> documentation to guide the submitter. Included with the Help document is a set of annotation examples that detail the types of information that are required for each type of submission. After the information is entered into the form, BankIt transforms this information into a GenBank <xref ref-type="other" rid="bid.1294">flatfile</xref> for review. In addition, a number of quality assurance and validation checks ensure that the sequence submitted to GenBank is of the highest quality. The submitter is asked to include spans (sequence coordinates) for the coding regions and other features and to include amino acid sequence for the proteins that derive from these coding regions. The BankIt validator compares the amino acid sequence provided by the submitter with the conceptual translation of the coding region based on the provided spans. If there is a discrepancy, the submitter is requested to fix the problem, and the process is halted until the error is resolved. To prevent the deposit of sequences that contain cloning vector sequence, a <xref ref-type="other" rid="bid.1246">BLAST</xref> similarity search is performed on the sequence, comparing it to the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/VecScreen/VecScreen.html">VecScreen</ext-link> database. If there is a match to this database, the user is asked to remove the contaminating vector sequence from their submission or provide an explanation as to why the screen was positive. Completed forms are saved in ASN.1 format, and the entry is submitted to the GenBank processing queue. The submitter receives confirmation by email, indicating that the submission process was successful.</p>
              </sec>
              <sec id="bid.20">
                <title>Sequin</title>
                <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/index.html">Sequin</ext-link> is more appropriate for complicated submissions containing a significant amount of annotation or many sequences. It is a stand-alone application available on NCBI's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/sequin/">FTP</ext-link> site. Sequin creates submissions from nucleotide and amino acid sequences in FASTA format with tagged biological source information in the FASTA definition line. As in BankIt, Sequin has the ability to predict the spans of coding regions. Alternatively, a submitter can specify the spans of their coding regions in a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/table.html">five-column, tab-delimited table</ext-link> and import that table into Sequin. For submitting multiple, related sequences, e.g., those in a phylogenetic or population study, Sequin accepts the output of many popular multiple sequence-alignment packages, including FASTA+GAP, PHYLIP, MACAW, NEXUS Interleaved, and NEXUS Contiguous. It also allows users to annotate features in a single record or a set of records globally. For more information on Sequin, see <xref ref-type="other" rid="bid.538">Chapter 12</xref>.
			</p>
                <p>Completed Sequin submissions should be emailed to GenBank at <email>gb-sub@ncbi.nlm.nih.gov</email>. Larger files may be submitted by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="www.ncbi.nlm.nih.gov/LargeDirSubs/dir_submit.cgi">SequinMacrosend</ext-link>.</p>
              </sec>
            </sec>
            <sec id="bid.21">
              <title>Sequence Data Flow and Processing: From Laboratory to GenBank</title>
              <sec id="bid.22">
                <title>Triage</title>
                <p>All direct submissions to GenBank, created either by Sequin or BankIt, are processed by the GenBank annotation staff. The first step in processing submissions is called triage. Within 48 hours of receipt, the database staff reviews the submission to determine whether it meets the minimal criteria for incorporation into GenBank and then assigns an Accession number to each sequence. All sequences must be &gt;50 bp in length and be sequenced by, or on behalf of, the group submitting the sequence. GenBank will not accept sequences constructed <italic>in silico</italic>; noncontiguous sequences containing internal, unsequenced spacers; or sequences for which there is not a physical counterpart, such as those derived from a mix of genomic DNA and mRNA. Submissions are also checked to determine whether they are new sequences or updates to sequences submitted previously. After receiving Accession numbers, the sequences are put into a queue for more extensive processing and review by the annotation staff.</p>
              </sec>
              <sec id="bid.23">
                <title>Indexing</title>
                <p>Triaged submissions are subjected to a thorough examination, referred to as the indexing phase. Here, entries are checked for:
                  <list list-type="arabic">
                    <list-item id="bid.24">
                      <label>1</label>
                      <p>Biological validity. For example, does the conceptual translation of a coding region match the amino acid sequence provided by the submitter? Annotators also ensure that the source organism name and lineage are present, and that they are represented in NCBI's taxonomy database. If either of these is not true, the submitter is asked to correct the problem. Entries are also subjected to a series of BLAST similarity searches to compare the annotation with existing sequences in GenBank.</p>
                    </list-item>
                    <list-item id="bid.25">
                      <label>2</label>
                      <p>Vector contamination. Entries are screened against NCBI's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/VecScreen/UniVec.html">UniVec</ext-link> database to detect contaminating cloning vector.</p>
                    </list-item>
                    <list-item id="bid.26">
                      <label>3</label>
                      <p>Publication status. If there is a published citation, <xref ref-type="other" rid="bid.1387">PubMed</xref> and <xref ref-type="other" rid="bid.1339">MEDLINE</xref> identifiers are added to the entry so that the sequence and publication records can be linked in Entrez.</p>
                    </list-item>
                    <list-item id="bid.27">
                      <label>4</label>
                      <p>Formatting and spelling. If there are problems with the sequence or annotation, the annotator works with the submitter to correct them.</p>
                    </list-item>
                  </list>
                </p>
                <p>Completed entries are sent to the submitter for a final review before release into the public database. If the submitters requested that their sequences be released after processing, they have 5 days to make changes prior to release. The submitter may also request that GenBank hold their sequence until a future date. The sequence must become publicly available once the Accession number or the sequence has been published. The GenBank annotation staff currently processes about 1,900 submissions per month, corresponding to approximately 20,000 sequences.</p>
                <p>GenBank annotation staff must also respond to email inquiries that arrive at the rate of approximately 200 per day. These exchanges address a range of topics including:
                  <list list-type="bullet">
                    <list-item id="bid.28">
                      <p>updates to existing GenBank records, such as new annotation or sequence changes</p>
                    </list-item>
                    <list-item id="bid.29">
                      <p>problem resolution during the indexing phase</p>
                    </list-item>
                    <list-item id="bid.30">
                      <p>requests for release of the submitter's sequence data or an extension of the hold date</p>
                    </list-item>
                    <list-item id="bid.31">
                      <p>requests for release of sequences that have been published but are not yet available in GenBank</p>
                    </list-item>
                    <list-item id="bid.32">
                      <p>lists of Accession numbers that are due to appear in upcoming issues of a publisher's journals</p>
                    </list-item>
                    <list-item id="bid.33">
                      <p>reports of potential annotation problems with entries in the public database</p>
                    </list-item>
                    <list-item id="bid.34">
                      <p>requests for information on how to submit data to GenBank</p>
                    </list-item>
                  </list>
                </p>
                <p>One annotator is responsible for handling all email received in a 24-hour period, and all messages must be acted upon and replied to in a timely fashion. Replies to previous emails are forwarded to the appropriate annotator.</p>
              </sec>
              <sec id="bid.35">
                <title>Processing Tools</title>
                <p>The annotation staff uses a variety of tools to process and update sequence submissions. Sequence records are edited with Sequin, which allows staff to annotate large sets of records by global editing rather than changing each record individually. This is truly a time saver because more than 100 entries can be edited in a single step (see <xref ref-type="other" rid="bid.538">Chapter 12</xref> on Sequin for more details). Records are stored in a database that is accessed through a queue management tool that automates some of the processing steps, such as looking up taxonomy and PubMed data, starting BLAST jobs, and running automatic validation checks. Hence, when an annotator is ready to start working on an entry, all of this information is ready to view. In addition, all of the correspondence between GenBank staff and the submitter is stored with the entry. For updates to entries already present in the public database, the live version of the entry is retrieved from ID, and after making changes, the annotator loads the entry back into the public database. This entry is available to the public immediately after loading.</p>
              </sec>
            </sec>
            <sec id="bid.36">
              <title>Microbial Genomes</title>
              <p>The GenBank direct submissions group has processed more than 50 complete microbial genomes since 1996. These genomes are relatively small in size compared with their eukaryotic counterparts, ranging from five hundred thousand to five million bases. Nonetheless, these genomes can contain thousands of genes, coding regions, and structural RNAs; therefore, processing and presenting them correctly is a challenge. Currently, the DDBJ/EMBL/GenBank Nucleotide Sequence Database Collaboration has a 350-kilobase (kb) upper size limit for sequence entries. Because a complete bacterial genome is larger than this arbitrary limit, it must be split into pieces. GenBank routinely splits complete microbial genomes into 10-kb pieces with a 60-bp overlap between pieces. Each piece contains approximately 10 genes. A CON entry, containing instructions on how to put the pieces back together, is also made. The CON entry contains descriptor information, such as source organism and references, as well as a join statement providing explicit instructions on how to generate the complete genome from the pieces. The Accession number assigned to the CON record is also added as a secondary Accession number on each of the pieces that make up the complete genome (see <xref ref-type="fig" rid="bid.37">Figure 2</xref>).<fig id="bid.37"><label>2</label><caption><title>A GenBank CON entry for a complete bacterial genome.</title><p>The information toward the <italic>bottom</italic> of the record describes how to generate the complete genome from the pieces.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch1f2" mime-subtype="gif"/></fig></p>
              <sec id="bid.38">
                <title>Submitting and Processing Data</title>
                <p>Submitters of complete genomes are encouraged to contact us at <email>genomes@ncbi.nlm.nih.gov</email> before preparing their entries. A FTP account is required to submit large files, and the submission should be deposited at least 1 month before publication to allow for processing time and coordinated release before publication. In addition, submitters are required to follow certain guidelines, such as providing unique identifiers for proteins and systematic names for all genes. Entries should be prepared with the submission tool <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/table.html">tbl2asn</ext-link>, a utility that is part of the Sequin package (<xref ref-type="other" rid="bid.538">Chapter 12</xref>). This utility creates an ASN.1 submission file from a five-column, tab-delimited file containing feature annotation, a FASTA-formatted nucleotide sequence, and an optional FASTA-formatted protein sequence.</p>
                <p>Complete genome submissions are reviewed by a member of the GenBank annotation staff to ensure that the annotation and gene and protein identifiers are correct, and that the entry is in proper GenBank format. Any problems with the entry are resolved through communication with the submitter. Once the record is complete, the genome is carefully split into its component pieces. The genome is split so that none of the breaks occurs within a gene or coding region. A member of the annotation staff performs quality assurance checks on the set of genome pieces to ensure that they are correct and representative of the complete genome. The pieces are then loaded into GenBank, and the CON record is created.</p>
                <p>The microbial genome records in GenBank are the building blocks for the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/micr.html">Microbial Genome Resources</ext-link> in Entrez Genomes.</p>
              </sec>
            </sec>
            <sec id="bid.39">
              <title>Third Party Annotation (TPA) Sequence Database</title>
              <p>The vast amount of publicly available data from the human genome project and other genome sequencing efforts is a valuable resource for
scientists throughout the world. A laboratory studying a particular gene
or gene family may have sequenced numerous cDNAs but has neither the
resources nor inclination to sequence large genomic regions containing
the genes, especially when the sequence is available in public
databases. The researcher might choose then to download genomic
sequences from GenBank and perform experimental analyses on these
sequences. However, because this researcher did not perform the
sequencing, the sequence, with its new annotations, cannot be submitted
to DDBJ/EMBL/GenBank. This is unfortunate because important scientific
information is being excluded from the public databases. To address this
problem, the International Nucleotide Sequence Database Collaboration
established a separate section of the database for such TPA (see <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="www.ncbi.nlm.nih.gov/Genbank/tpa.html">Third Party Annotation Sequence Database</ext-link>).</p>
              <p>All sequences in the TPA database are derived from the publicly
available collection of sequences in DDBJ/EMBL/GenBank. Researchers can
submit both new and alternative annotations of genomic sequence to
GenBank. TPA entries can be also created by combining the exon sequences
from genomic sequences or by making contigs of EST sequences to make
mRNA sequences. TPA submissions must use sequence data that are already
represented in DDBJ/EMBL/GenBank, have annotation that is experimentally
supported, and appear in a peer-reviewed scientific journal. TPA
sequences will be released to the public database only when their
Accession numbers and/or sequence data appear in a peer-reviewed
publication in a biological journal. </p>
            </sec>
</body>
<back>
              <ref-list>
                <title>References</title>
                <ref id="bid.41">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Olson</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Hood</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Cantor</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Botstein</surname>
                        <given-names>D</given-names>
                      </name>
                    </person-group>
                    <article-title>A common language for physical mapping of the human genome</article-title>
                    <source>Science</source>
                    <year>1989</year>
                    <volume>245</volume>
                    <issue>4925</issue>
                    <fpage>1434</fpage>
                    <lpage>1435</lpage>
                    <pub-id pub-id-type="pmid">2781285</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.42" book-part-type="chapter" book-part-number="2">
          <book-part-meta>
            <title-group>
              <title>PubMed: The Bibliographic Database</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Canese</surname>
                  <given-names>Kathi</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Jentsch</surname>
                  <given-names>Jennifer</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Myers</surname>
                  <given-names>Carol</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch2d1"/>
            <abstract>
              <title>Summary</title>
              <p>PubMed is a database developed by the National Center for Biotechnology Information (NCBI) at the National Library of Medicine (NLM), one of the institutes of the National Institutes of Health (NIH). The database was designed to provide access to citations (with abstracts) from biomedical journals. Subsequently, a linking feature was added to provide access to full-text journal articles at Web sites of participating publishers, as well as to other related Web resources. PubMed is the bibliographic component of the NCBI's Entrez retrieval system.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.43">
              <title>Data Sources</title>
              <sec id="bid.44">
                <title>MEDLINE<sup>&reg;</sup></title>
                <p><xref ref-type="other" rid="bid.1387">PubMed</xref>'s primary data resource is <xref ref-type="other" rid="bid.1339">MEDLINE</xref>, the NLM's premier bibliographic database covering the fields of medicine, nursing, dentistry, veterinary medicine, the health care system, and the preclinical sciences, such as molecular biology. MEDLINE contains bibliographic citations and author abstracts from about 4,600 biomedical journals published in the United States and 70 other countries. The database contains about 12 million citations dating back to the mid-1960s. Coverage is worldwide, but most records are from English-language sources or have English abstracts.</p>
              </sec>
              <sec id="bid.45">
                <title>Non-MEDLINE</title>
                <p>In addition to MEDLINE citations, PubMed<sup>&reg;</sup> provides access to non-MEDLINE resources, such as out-of-scope citations, citations that precede MEDLINE selection, and PubMed Central (PMC; see <xref ref-type="other" rid="bid.457">Chapter 9</xref>) citations. Together, these are often referred to as &ldquo;PubMed-only citations.&rdquo; Out-of-scope citations are primarily from general science and chemistry journals that contain life sciences articles indexed for MEDLINE, e.g., the plate tectonics or astrophysics articles from <italic>Science</italic> magazine. Publishers can also submit citations with publication dates that precede the journal's selection for MEDLINE indexing, usually because they want to create links to older content. PMC citations are taken from life sciences journals (MEDLINE or non-MEDLINE) that submit full-text articles to PMC. In addition to the incorporation of PubMed-only citations, PubMed has been enhanced recently by the incorporation of citations from the following unique databases: HealthSTAR, AIDSLINE, HISTLINE, SPACELINE, BIOETHICSLINE, and POPLINE.</p>
                <p>In response to new approaches to electronic publishing, PubMed can now also accommodate articles published electronically in advance of being collected into an issue. We refer to these citations as &ldquo;ahead of print&rdquo; or &ldquo;epub&rdquo; citations.</p>
              </sec>
              <sec id="bid.46">
                <title>Journal Selection Criteria</title>
                <p>All content in PubMed ultimately comes from publishers of biomedical journals, and journals that are to be included in MEDLINE are subject to a selection process. The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/pubs/factsheets/jsel.html">Fact Sheet</ext-link> on <italic>Journal Selection for Index Medicus<sup>&reg;</sup>/MEDLINE<sup>&reg;</sup></italic> describes the journal selection policy, criteria, and procedures for data submission.</p>
              </sec>
            </sec>
            <sec id="bid.47">
              <title>Electronic Data Submission</title>
              <p>Electronic data submission benefits everyone: publishers, the NLM, and users. For the NLM, it eliminates the tremendous costs associated with entering data by hand. For publishers and users, it means that newly published data appear rapidly and accurately in PubMed. Some publishers are now making pre-publication material available before it is formally published (&ldquo;ahead of print&rdquo; or &ldquo;epub&rdquo; citations); others are publishing electronic-only journals. By close collaboration with the publisher, the citations for these publications can appear in PubMed on the same day as the article is published.</p>
              <p>Furthermore, electronic data submission allows publishers to create links from abstracts in PubMed to the full text of the appropriate articles available on their own Web site. This can be achieved using LinkOut (<xref ref-type="other" rid="bid.650">Chapter 17</xref>). Both subscribers to the journals and other PubMed users can access the full text according to criteria that are determined by the publishers, increasing traffic to their sites.</p>
              <p>Although the NLM works with many publishers directly, some publishers contract with commercial data aggregators, companies that prepare and submit the publisher's data to the NLM. Many aggregators also host publisher data on their Web sites.</p>
              <sec id="bid.48">
                <title>Electronic Data Submission Process</title>
                <p>All electronic data are supplied via <xref ref-type="other" rid="bid.1295">FTP</xref> to NCBI in XML format, in accordance with the NLM's specifications (document type definition, or <xref ref-type="other" rid="bid.1277">DTD</xref>). These specifications can be found in <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/spec.html">NLM Standard Publisher Data Format</ext-link> document. The document includes information on XML tag descriptions, how to handle special characters (e.g., &alpha; or &beta;), examples of tagged records, the PubMed DTD, and a FAQ section for participating or potential data providers. Publishers or other data providers who want to submit electronic data should write to: <email>publisher@ncbi.nlm.nih.gov</email>.</p>
                <p>NCBI staff will guide new data providers through the approval process for file submission. New providers are asked to submit test files, which are then checked for XML formatting and syntax and for bibliographic accuracy and completeness. The files are revised and resubmitted as many times as necessary until all criteria are met. Once approved, a private account is set up on our FTP site to receive new journal issues, or in the case of online publications, individual articles as they are added to the publisher's Web site. We run a file-loading script that automatically processes the files daily, Monday through Friday at approximately 9:00 a.m. (Eastern Time). The new citations are assigned a PubMed ID number (PMID), a confirmation report is sent to the provider, and the new citations usually become available in PubMed sometime after 11:00 a.m. the next day, Tuesday through Saturday.</p>
                <p>After posting in PubMed, the citations are forwarded to NLM's Indexing Section for bibliographic data verification and for the addition of subject indexing terms from Medical Subject Headings [<xref ref-type="other" rid="bid.1341">MeSH</xref>]. This process can take several weeks, after which time completed citations flow back into PubMed, replacing the originally submitted data.</p>
              </sec>
            </sec>
            <sec id="bid.49">
              <title>Database Management and Hardware</title>
              <p>PubMed is one of the NCBI databases within the relational database management system, <xref ref-type="other" rid="bid.1282">Entrez</xref> (see <xref ref-type="other" rid="bid.588">Chapter 15</xref>). Entrez is a text-based search and retrieval system based on in-house software that uses an indexing system for rapid retrieval of information.</p>
              <p>Requests for NCBI services, including PubMed, are first proxied through three load-balanced Dell PowerEdge 1650 servers, each with two central processing units. The proxy servers, in turn, load-balance requests forwarded on to the Web servers for PubMed and other NCBI services.</p>
              <p>The PubMed Web servers comprise eight Dell PowerEdge 8450 servers. The Dell servers have eight central processing units, 8 GB of memory, and about 300 GB of disk space and run the Linux operating system.</p>
              <p>The Web servers retrieve PubMed records from two Sybase SQL database servers, which run on Sun Enterprise 450s. To accommodate the data volume output by PubMed and other Web-based services, the NLM has a high-speed connection (OC-3, up to 155 Mbits/sec) to the Internet, as well as a 622 Mbits/sec connection (OC-12) to Internet2, the noncommercial network used by many leading research universities.</p>
            </sec>
            <sec id="bid.50">
              <title>Indexing</title>
              <sec id="bid.51">
                <title>PubMed Citation Status and Assignment of MeSH Terms</title>
                <p>Citations in PubMed are assigned one of three citation status tags that display next to the PubMed ID (PMID) numbers on all PubMed citations. The citation status tags indicate the citation's stage in the MEDLINE indexing process. The three tags are:</p>
                <p><bold>[PubMed - as supplied by publisher]</bold>: This tag is displayed on citations added recently to PubMed via electronic submission from a publisher (which may or may not move on for MEDLINE MeSH indexing).</p>
                <p><bold>[PubMed - in process]</bold>: This tag is displayed on citations that have had the first stage of quality review to verify that the journal, date, volume, and/or issue are correct. They will be reviewed for other accurate bibliographic data at the article level (e.g., pagination, authors, article title, and abstract) and indexed, i.e., the articles will be reviewed and MeSH vocabulary will be assigned (if the subject of the article is within the scope of MEDLINE).</p>
                <p><bold>[PubMed - indexed for MEDLINE]</bold>: This tag is displayed on citations that have been indexed with MeSH, Publication Types, Registry Numbers, etc., and have been completely reviewed for accurate bibliographic data. This is an intellectual process of assigning controlled vocabulary terms to describe the contents of the journal article and verifying other aspects of the citation data.</p>
                <p>Most citations that are received electronically from publishers progress through &ldquo;in process&rdquo; status to MEDLINE status. Those citations not indexed for MEDLINE remain tagged [PubMed - as supplied by publisher]. Citations with &ldquo;in process&rdquo; status proceed to MEDLINE status after MeSH terms, publication types, sequence Accession numbers, and other indexing data are added.</p>
                <p>All records are added to PubMed Monday through Friday and become available for viewing Tuesday through Saturday. For additional information, please see the NLM <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/pubs/factsheets/dif_med_pub.html">Fact Sheet</ext-link>: <italic>What's the Difference Between MEDLINE<sup>&reg;</sup> and PubMed<sup>&reg;</sup>?</italic></p>
              </sec>
              <sec id="bid.52">
                <title>The Automatic Computer Indexing Process</title>
                <p>The aim of the computer indexing process is to automatically create multiple machine-readable access points that refer to the different components of the journal citations for use when searching PubMed. The citations are loaded into PubMed from both the NLM Data Creation and Maintenance System (DCMS) and directly from journal publishers (<xref ref-type="fig" rid="bid.53">Figure 1</xref>). Both sources are in XML.<fig id="bid.53"><label>1</label><caption><title>A schematic representation of PubMed data flow.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch2f1" mime-subtype="gif"/></fig></p>
                <p>During the computer indexing process, the citation information is broken down into index fields such as Journal Name, Author Name, and Title/Abstract. The words in each of the fields are checked against the corresponding index (i.e., title words in a new citation are looked up in the Title/Abstract Index). If the word already exists, the PMID of the citation is listed with that index term. If the word is a new one for the Index, it is added as a new Index term, and the PMID is listed alongside it. (In the first instance that the term already exists, the new term will have only this one citation associated with it; this is how the PubMed indexes grow.)</p>
                <p>Each PubMed citation is, therefore, associated with several indexes, and in cases similar to the Title/Abstract Index, many different index terms can refer back to a single citation. Likewise, commonly used terms will refer to thousands of citations (the term &ldquo;cell&rdquo;, for example, is found in the Title/Abstract of 1,092,124 citations at the time of this writing). The Field Indexes can be browsed by using PubMed's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?CMD=IndexView&amp;DB=PubMed">Preview/Index</ext-link> function.</p>
              </sec>
            </sec>
            <sec id="bid.54">
              <title>How PubMed Queries Are Processed</title>
              <sec id="bid.55">
                <title>Automatic Term Mapping</title>
                <p>PubMed uses <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#AutomaticTermMapping">Automatic Term Mapping</ext-link> to process words entered in the query box by someone searching PubMed. Terms entered without a qualifier, i.e., a simple text phrase that does not specify a search field, are looked up against the following translation tables and indexes in a distinct order:
                <list list-type="arabic">
                <list-item><label>1</label><p>MeSH Translation Table</p>
                </list-item>
                <list-item><label>2</label><p>Journals Translation Table</p>
                </list-item>
                <list-item><label>3</label><p>3. Author Index</p>
                </list-item>
                </list>
                </p>
                <sec id="bid.56">
                  <title>MeSH Translation Table</title>
                  <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#MeSHTranslationTable">MeSH Translation Table</ext-link> contains:
                    <list list-type="bullet">
                      <list-item id="bid.57">
                        <p>MeSH Terms</p>
                      </list-item>
                      <list-item id="bid.58">
                        <p>
                          <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#subheadingslist">Subheadings</ext-link>
                        </p>
                      </list-item>
                      <list-item id="bid.59">
                        <p>See-Reference mappings (also known as entry terms) for MeSH terms</p>
                      </list-item>
                      <list-item id="bid.60">
                        <p>Mappings derived from the Unified Medical Language System (<xref ref-type="other" rid="bid.1422">UMLS</xref>) that have equivalent synonyms or lexical variants in English</p>
                      </list-item>
                      <list-item id="bid.61">
                        <p>Names of Substances and synonyms to the Names of Substances (now known as Supplementary Concept Substance Names)</p>
                      </list-item>
                    </list>
                  </p>
                  <p>If the search term is found in this translation table, the term will be mapped to the appropriate MeSH term, and the Indexes will be searched as both the text word entered by the user and the MeSH term:
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: gallstones</title>
                    <p>&ldquo;Gallstones&rdquo; is an entry term for the MeSH term &ldquo;cholelithiasis&rdquo; in the MeSH translation table.</p>
                    <p>Search translated to: &ldquo;cholelithiasis&rdquo; [MeSH Terms] OR gallstones [Text Word]</p>
                  </boxed-text></p>
                  <p>When a term is searched as a MeSH term, PubMed automatically searches that term plus the more specific terms underneath in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/meshbrowser.cgi?retrievestring=&amp;mbdetail=n&amp;term=breast+cancer">MeSH hierarchy</ext-link>:
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: breast cancer</title>
                    <p>&ldquo;Breast cancer&rdquo; is an entry term for the MeSH term &ldquo;breast neoplasms&rdquo; in the MeSH translation table.</p>
                    <p>&ldquo;Breast neoplasms&rdquo; has the specific headings &ldquo;breast neoplasms, male&rdquo;, &ldquo;mammary neoplasms&rdquo;, &ldquo;mammary neoplasms, experimental&rdquo;, and &ldquo;phyllodes tumor&rdquo;, all of which are also searched.</p>
                  </boxed-text>
                  </p>
                </sec>
                <sec id="bid.62">
                  <title>2. Journals Translation Table</title>
                  <p>If the search term(s) is not found in the MeSH Translation Table, the PubMed search algorithm then looks up the term in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#JournalsTranslationTable">Journals Translation Table</ext-link>, which contains the full journal title, MEDLINE abbreviation, and International Standard Serial Number (<xref ref-type="other" rid="bid.1327">ISSN</xref>):
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: New England Journal of Medicine</title>
                    <p>&ldquo;New England Journal of Medicine&rdquo; maps to N Engl J Med.</p>
                    <p>Search translated to: &ldquo;N Engl J Med&rdquo; [Journal Name]</p>
                  </boxed-text>
                  </p>
                  <p>If a journal name is also a MeSH term, PubMed will search the term as both a MeSH term and as a Text Word, but not as a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#JournalTitles">journal name</ext-link>, for a search that does not specify the &ldquo;journal&rdquo; field:
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: Cell</title>
                    <p>Search translated as: &ldquo;cells&rdquo; [MeSH Terms] OR cell [Text Word]</p>
                  </boxed-text>
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: Cell [Journal]</title>
                    <p>Search translated as: &ldquo;Cell&rdquo; [Journal]</p>
                  </boxed-text>
                </p>
                </sec>
                <sec id="bid.64">
                  <title>3. Author Index</title>
                  <p>If the phrase is not found in MeSH or the Journals Translation Table and is a word with one or two letters after it, PubMed then checks the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#AuthorIndex">Author Index</ext-link>. The author's name should be entered in the form: Last Name (space) Initials, e.g., o'malley f, smith jp, or gomez-sanchez m.</p>
                  <p>If only one initial is used, PubMed finds all names with that first initial, and if only an author's last name is entered, PubMed will search that name in All Fields. It will not default to the Author Index because an initial does not follow the last name:
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: o'malley f</title>
                    <p>Search translated as: o'malley fa, o'malley fb, o'malley fc, o'malley fd, o'malley f jr, etc.</p>
                  </boxed-text>
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: o'malley</title>
                    <p>Search translated as: &ldquo;o'malley&rdquo; [All Fields]</p>
                  </boxed-text>
                  </p>
                  <p>A history of the NLM's author indexing policy regarding the number of authors to include in a citation is outlined in <xref ref-type="table" rid="bid.65">Table 1</xref>.</p>
                  <table-wrap id="bid.65"><label>1</label><caption><title>History of NLM author-indexing policy.</title></caption><table><tr><th align="left" valign="top">Dates</th><th align="left" valign="top">Policy</th></tr><tr><td align="left" valign="top">1966&ndash;1984</td><td align="left" valign="top">MEDLINE did not limit the number of authors.</td></tr><tr><td align="left" valign="top">1984&ndash;1995</td><td align="left" valign="top">The NLM limited the number of authors to 10, with &ldquo;et al.&rdquo; as the eleventh occurrence.</td></tr><tr><td align="left" valign="top">1996&ndash;1999</td><td align="left" valign="top">The NLM increased the limit from 10 to 25. If there were more than 25 authors, the first 24 were listed, the last author was used as the 25th, and the twenty-sixth and beyond became &ldquo;et al.&rdquo;.</td></tr><tr><td align="left" valign="top">2000&ndash;present</td><td align="left" valign="top">MEDLINE does not limit the number of authors.</td></tr></table></table-wrap>
                </sec>
              </sec>
              <sec id="bid.66">
                <title>Search Rules and Field Abbreviations</title>
                <p>It is possible to override PubMed's Automatic Term Mapping by using search rules, syntax, and qualifying terms with search field abbreviations.</p>
                <p>The <xref ref-type="other" rid="bid.1253">Boolean</xref> operators AND, OR, and NOT must be entered in uppercase letters and are processed left to right. Nesting of search terms is possible by enclosing concepts in parentheses. The terms inside the set of parentheses will be processed as a unit and then incorporated into the overall strategy. Terms may be qualified using PubMed's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#SearchFieldDescriptionsandTags">Search Field Descriptions and Tags</ext-link>. Each search term should be followed by the appropriate search field tag, which indicates which field will be searched:</p>
                <p>Search term: o'malley [au] will search only the author field. Specifying the field precludes the Automatic Term Mapping, which would result in the search o'malley[All Fields] if the field were not specified. Similarly, using the search term Cell [Journal] avoids using the MeSH Translation Table, which would interpret Cell as only a text word and MeSH term.</p>
              </sec>
            </sec>
            <sec id="bid.67">
              <title>Using PubMed</title>
              <sec id="bid.68">
                <title>Searching</title>
                <sec id="bid.69">
                  <title>Simple Searching</title>
                  <p>A simple search can be conducted from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/entrez/query.fcgi">PubMed</ext-link> homepage by entering terms in the query box and pressing Enter from the keyboard or clicking on the <named-content content-type="book-spec-emph">Go</named-content> button on the screen.</p>
                  <p>If more than one term is entered in the query box, PubMed will go through the Automatic Term Mapping protocol described in the previous section, first looking for all the terms, as typed, to find an exact match. If the exact phrase is not found, PubMed clips a term off the end and repeats Automatic Term Mapping, again looking for an exact match, but this time to the abbreviated query. This continues until none of the words are found in any one of the translation tables. In this case, PubMed combines terms (with the AND Boolean operator) and applies the Automatic Term Mapping process to each individual word. PubMed ignores <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Stopwords">Stopwords</ext-link>, such as &ldquo;about&rdquo;, &ldquo;of&rdquo;, or &ldquo;what&rdquo;. People can also apply their own Boolean operators (AND, OR, NOT) to multiple search terms; the Boolean operators must be in uppercase.</p>
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: vitamin c common cold</title>
                    <p>Translated as: ((&ldquo;ascorbic acid&rdquo; [MeSH Terms] OR vitamin c [Text Word]) AND (&ldquo;common cold&rdquo; [MeSH Terms] OR common cold [Text Word]))</p>
                  </boxed-text>
                  <boxed-text position="anchor" content-type="website">
                    <title>Search term: single cell separation brain</title>
                    <p>Translated as: (((&ldquo;single person&rdquo; [MeSH Terms] OR single [Text Word]) AND (&ldquo;cell separation&rdquo; [MeSH Terms] OR cell separation [Text Word])) AND (&ldquo;brain&rdquo; [MeSH Terms] OR brain [Text Word]))</p>
                  </boxed-text>
                  <p>If a phrase of more than two terms is not found in any translation table, then the last word of the phrase is dropped, and the remainder of the phrase is sent through the entire process again. This continues, removing one word at a time, until a match is found.</p>
                  <p>If there is no match found during the Automatic Term Mapping process, the individual terms will be combined with AND and searched in All Fields.</p>
                  <p>One can see how PubMed interpreted a search by selecting <named-content content-type="book-spec-emph">Details</named-content> from the Features Bar on the PubMed search pages after completing a search. For more information, see <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Details">Details</ext-link>.</p>
                </sec>
                <sec id="bid.70">
                  <title>Complex Searching</title>
                  <p>There are a variety of ways that PubMed can be searched in a more sophisticated manner than simply typing search terms into the search box and selecting <named-content content-type="book-spec-emph">Go</named-content>. It is possible to construct complex search strategies using Boolean operators and the various functions listed below, provided in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#FeaturesBar">Features Bar</ext-link>:
                    <list list-type="bullet">
                      <list-item id="bid.71">
                        <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Limits">Limits</ext-link> restricts search terms to a specific search field.</p>
                      </list-item>
                      <list-item id="bid.72">
                        <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Index">Preview/Index</ext-link> allows users to view and select terms from search field indexes and to preview the number of search results before displaying citations.</p>
                      </list-item>
                      <list-item id="bid.73">
                        <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#History">History</ext-link> holds previous search strategies and results. The results can be combined to make new searches.</p>
                      </list-item>
                      <list-item id="bid.74">
                        <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Clipboard">Clipboard</ext-link> allows users to save or view selected citations from one search or several searches.</p>
                      </list-item>
                      <list-item id="bid.75">
                        <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Details">Details</ext-link> displays the search strategy as it was translated by PubMed, including error messages.</p>
                      </list-item>
                    </list>
                  </p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.76">
              <title>Additional PubMed Features</title>
              <p>The following resources are available to facilitate effective searches:
                <list list-type="bullet">
                  <list-item id="bid.77">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#MeSHBrowser">MeSH Database</ext-link> allows searching of MeSH, NLM's controlled vocabulary. Users can find MeSH terms appropriate to a search strategy, obtain information about each term, and view the terms within their hierarchical structure.</p>
                  </list-item>
                  <list-item id="bid.78">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#ClinicalQueries">Clinical Queries</ext-link> is a set of search filters developed for clinicians to retrieve clinical studies of the etiology, prognosis, diagnosis, prevention, or treatment of disorders. The Systematic Reviews feature retrieves systematic reviews and meta-analysis studies by topic.</p>
                  </list-item>
                  <list-item id="bid.79">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#JournalBrowser">Journal Database</ext-link> allows searches of journal names, MEDLINE abbreviations, or ISSN numbers for journals that are included in the Entrez system. A list of journals with links to full text is also included.</p>
                  </list-item>
                  <list-item id="bid.80">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#SingleCitationMatcher">Single Citation Matcher</ext-link> is a  &ldquo;fill-in-the-blank&rdquo; form that allows a user to find the PubMed ID (PMID) number for a single article or all citations in a given journal issue by entering partial journal citation information.</p>
                  </list-item>
                  <list-item id="bid.81">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#BatchCitationMatcher">Batch Citation Matcher</ext-link> allows users to find PMID numbers that correspond to their own list of citations. Publishers or other database providers who want to link directly from bibliographic references on their Web sites to entries in PubMed use this service frequently.</p>
                  </list-item>
                  <list-item id="bid.82">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Cubby">Cubby</ext-link> is a place for users to store search strategies, LinkOut preferences, and changes to the default <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#DocumentDeliveryServices">Document Delivery Services</ext-link>.</p>
                  </list-item>
                </list>
              </p>
            </sec>
            <sec id="bid.83">
              <title>Results</title>
              <p>PubMed retrieves and displays search results in the Summary format in the order the record was initially added to PubMed, with the most recent first. (Note that this date can differ widely from the publication date.) Citations can be viewed in several other <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#DisplayingFormats">formats</ext-link> and can be <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Sort">sorted</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Saving">saved</ext-link>, and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Printing">printed</ext-link>, or the full text can be <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#Ordering">ordered</ext-link>.</p>
            </sec>
            <sec id="bid.84">
              <title>Links from PubMed</title>
              <p>A variety of links can be found on PubMed citations including:</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#RelatedPubMedArticlesLink">Related Articles</ext-link>, which retrieves a precalculated set of PubMed citations that are closely related to the selected article. PubMed creates this set by comparing words from the title, abstract, and MeSH terms using a word-weighted <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/computation.html">algorithm</ext-link>.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/">LinkOut</ext-link>, which provides links to publishers, aggregators, libraries, biological databases, sequencing centers, and other Web sites. These link to the provider's site to obtain the full text of articles or related resources, e.g., consumer health information or molecular biology database records. There may be a charge to access the text or information, depending on the policy of the provider.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://web.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Books">Books</ext-link>, which provides links to textbooks so that users can explore unfamiliar concepts found in search results. In collaboration with book publishers, NCBI is adapting textbooks for the Web and linking them to PubMed. The Books link displays a facsimile of the abstract, in which some words or phrases show up as hypertext links to the corresponding terms in the books available at NCBI. Selecting a hyperlinked word or phrase takes you to a list of book entries in which the phrase is found.</p>
              <p>NCBI databases, as well as other resources, may be available from the <named-content content-type="book-spec-emph">Links</named-content> pull-down menu to the right of each citation and from the <named-content content-type="book-spec-emph">Display</named-content> pull-down menu. PubMed will return only the first 500 items when using the <named-content content-type="book-spec-emph">Display</named-content> pull-down menu, from which the following links are available:
                <list list-type="bullet">
                  <list-item id="bid.85">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Protein">Protein</ext-link> &ndash; amino acid (protein) sequences from <xref ref-type="other" rid="bid.1412">SWISS-PROT</xref>, <xref ref-type="other" rid="bid.1374">PIR</xref>, <xref ref-type="other" rid="bid.1380">PRF</xref>, and <xref ref-type="other" rid="bid.1367">PDB</xref> and translated protein sequences from the DNA sequences databases.</p>
                  </list-item>
                  <list-item id="bid.86">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Nucleotide">Nucleotide</ext-link> &ndash; DNA sequences from <xref ref-type="other" rid="bid.1299">GenBank</xref>, <xref ref-type="other" rid="bid.1281">EMBL</xref>, and <xref ref-type="other" rid="bid.1272">DDBJ</xref>.</p>
                  </list-item>
                  <list-item id="bid.87">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Popset">PopSet</ext-link> &ndash; aligned sequences submitted as a set from a population, phylogenetic, or mutation study describing such events as evolution and population variation.</p>
                  </list-item>
                  <list-item id="bid.88">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Structure">Structure</ext-link> &ndash; three-dimensional structures from the Molecular Modeling Database (<xref ref-type="other" rid="bid.1349">MMDB</xref>) that were determined by X-ray crystallography and <xref ref-type="other" rid="bid.1359">NMR</xref> spectroscopy.</p>
                  </list-item>
                  <list-item id="bid.89">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Genome">Genome</ext-link> &ndash; records and graphic displays of entire genomes and chromosomes for megabase-scale sequences.</p>
                  </list-item>
                  <list-item id="bid.90">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=geo">ProbeSet</ext-link> &ndash; gene expression data repository and online resource for the retrieval of gene expression data from any organism or artificial source.</p>
                  </list-item>
                  <list-item id="bid.91">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=OMIM">OMIM</ext-link> &ndash; directory of human genes and genetic disorders.</p>
                  </list-item>
                  <list-item id="bid.1633">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=SNP">SNP</ext-link> &ndash; dbSNP is a database of single nucleotide polymorphisms.</p>
                  </list-item>
                  <list-item id="bid.1634">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=cdd">Domains</ext-link> &ndash; The Domains database is used to identify the conserved domains present in a protein sequence.</p>
                  </list-item>
                  <list-item id="bid.1758">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Domains">3D Domains</ext-link> &ndash; the domains from Entrez Structure.</p>
                  </list-item>
                  <list-item id="bid.1759">
                    <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=pmc">PMC</ext-link> &ndash; PubMed Central.</p>
                  </list-item>
                </list>
              </p>
            </sec>
            <sec id="bid.92">
              <title>How to Create Hyperlinks to PubMed</title>
              <p>The Entrez system provides three distinct ways to create Web <xref ref-type="other" rid="bid.1427">URL</xref> links that search and retrieve items from PubMed and the molecular biology databases: (1) by using the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/eutils_help.html">Entrez Programming Utilities</ext-link>; (2) via the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#SavingaSearchStrategy">URL button</ext-link> on the <named-content content-type="book-spec-emph">Details</named-content> screen; and (3) by constructing URLs <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/linking.html">by hand</ext-link>.</p>
              <p>The Entrez Programming Utilities can be used to create URL links directly to all Entrez data, including PubMed citations and their link information, without using the front-end Entrez query engine. These Utilities provide a fast, efficient way to search and download citation data.</p>
            </sec>
            <sec id="bid.93">
              <title>Customer Support</title>
              <p>If you need more assistance, please contact our <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#ForMoreAssistance">Customer Support</ext-link> services by selecting the <named-content content-type="book-spec-emph">Write to the Help Desk</named-content> link displayed on all PubMed pages or by sending an email to <email>custserv@nlm.nih.gov</email>. You may also contact the NLM Customer Service Desk at 1-888-346-3656 [(1-888)-FINDNLM]. Hours of operation are Monday through Friday from 8:30 a.m. to 8:45 p.m. and Saturday from 9:00 a.m. to 5:00 p.m. (Eastern Time).</p>
              <p>Additional information is also available in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/bsd/pubmed_tutorial/m1001.html">PubMed Tutorial</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/pubs/web_based.html">PubMed Training Manuals</ext-link>, and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/pubs/techbull/tb.html">NLM Technical Bulletin</ext-link>.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/faq.html">FAQs</ext-link> are available on all PubMed pages.</p>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.94" book-part-type="chapter" book-part-number="3">
          <book-part-meta>
            <title-group>
              <title>Macromolecular Structure Databases</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Sayers</surname>
                  <given-names>Eric</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Bryant</surname>
                  <given-names>Steve</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch3d1"/>
            <abstract>
              <title>Summary</title>
              <p>The resources provided by NCBI for studying the three-dimensional (3D) structures of proteins center around two databases: the Molecular Modeling Database (MMDB), which provides structural information about individual proteins; and the Conserved Domain Database (CDD), which provides a directory of sequence and structure alignments representing conserved functional domains within proteins(<xref ref-type="other" rid="bid.1255">CD</xref>s). Together, these two databases allow scientists to retrieve and view structures, find structurally similar proteins to a protein of interest, and identify conserved functional sites.</p>
              <p>To enable scientists to accomplish these tasks, NCBI has integrated MMDB and CDD into the Entrez retrieval system (<xref ref-type="other" rid="bid.588">Chapter 15</xref>). In addition, structures can be found by BLAST, because sequences derived from MMDB structures have been included in the BLAST databases (<xref ref-type="other" rid="bid.610">Chapter 16</xref>). Once a protein structure has been identified, the domains within the protein, as well as domain &ldquo;neighbors&rdquo; (i.e., those with similar structure) can be found. For novel data not yet included in Entrez, there are separate search services available.</p>
              <p>Protein structures can be visualized using Cn3D, an interactive 3D graphic modeling tool. Details of the structure, such as ligand-binding sites, can be scrutinized and highlighted. Cn3D can also display multiple sequence alignments based on sequence and/or structural similarity among related sequences, 3D domains, or members of a CDD family. Cn3D images and alignments can be manipulated easily and exported to other applications for presentation or further analysis.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.95">
              <title>Overview</title>
              <p>The Structure <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure">homepage</ext-link> (<xref ref-type="fig" rid="bid.96">Figure 1</xref>) contains links to the more specialized pages for each of the main tools and databases, introduced below, as well as search facilities for the Molecular Modeling Database (<xref ref-type="other" rid="bid.1349">MMDB</xref>; <xref ref-type="bibr" rid="bid.236">Ref. 1</xref>).<fig id="bid.96"><label>1</label><caption><title>The Structure homepage.</title><p>This page can be found by selecting the <named-content content-type="book-spec-emph">Structure</named-content> link on the tool bar atop many <xref ref-type="other" rid="bid.1353">NCBI</xref> Web pages. Two searches can be performed from this page, an <xref ref-type="other" rid="bid.1282">Entrez</xref> <named-content content-type="book-spec-emph">Structure</named-content> search or a Structure <named-content content-type="book-spec-emph">Summary</named-content> search. Both query the MMDB database. The difference is that the <named-content content-type="book-spec-emph">Entrez Structure</named-content> can take any text as a query (such as a <xref ref-type="other" rid="bid.1367">PDB</xref> code, protein name, text word, author, or journal) and will result initially in a list of one or more document summaries, displayed within the Entrez environment (<xref ref-type="other" rid="bid.588">Chapter 15</xref>), whereas only a PDB code or MMDB ID number can be used for the Structure <named-content content-type="book-spec-emph">Summary</named-content> search, resulting in direct display of the Structure Summary page for that record (<xref ref-type="fig" rid="bid.106">Figure 2</xref>). Announcements about new features or updates can also be found on this page, as well as links to more specialized pages on the various Structure databases and tools.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch3f1" mime-subtype="gif"/></fig></p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/MMDB/mmdb.shtml">MMDB</ext-link> is based on the structures within Protein Data Bank (PDB) and can be queried using the Entrez search engine, as well as via the more direct but less flexible Structure <named-content content-type="book-spec-emph">Summary</named-content> search (see <xref ref-type="fig" rid="bid.96">Figure 1</xref>). Once found, any structure of interest can be viewed using <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/CN3D/cn3d.shtml">Cn3D</ext-link> (<xref ref-type="bibr" rid="bid.237">2</xref>), a piece of software that can be freely downloaded for Mac, PC, and <xref ref-type="other" rid="bid.1426">UNIX</xref> platforms.</p>
              <p>Often used in conjunction with <xref ref-type="other" rid="bid.1264">Cn3D</xref> is the Vector Alignment Search Tool (<xref ref-type="other" rid="bid.1430">VAST</xref>; Refs. <xref ref-type="bibr" rid="bid.238">3</xref>, <xref ref-type="bibr" rid="bid.239">4</xref>). <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/VAST/vast.shtml">VAST</ext-link> is used to precompute &ldquo;structure neighbors&rdquo; or structures similar to each MMDB entry. For people that have a set of 3D coordinates for a protein not yet in MMDB, there is also a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/VAST/vastsearch.html">VAST search service</ext-link>. The output of the precomputed VAST searches is a list of structure records, each representing one of the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/VAST/nrpdb.html">Non-Redundant PDB chain</ext-link> sets (<xref ref-type="other" rid="bid.1360">nr-PDB</xref>), which can also be downloaded. There are four clustered subsets of MMDB that compose nr-PDB, each consisting of clusters having a preset level of sequence similarity.</p>
              <p>The structures within MMDB are now being linked to the NCBI Taxonomy database (<xref ref-type="other" rid="bid.246">Chapter 4</xref>). Known as the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/PDBEAST/pdbeast.shtml">PDBeast</ext-link> project, this effort makes it possible to find: (1) all MMDB structures from a particular organism; and (2) all structures within a node of the taxonomy tree (such as lizards or <italic>Bacillus)</italic>, which launches the Taxonomy Browser showing the number of MMDB records in each node.</p>
              <p>The second database within the <named-content content-type="book-spec-emph">Structure</named-content> resources is the Conserved Domain Database (<xref ref-type="other" rid="bid.1257">CDD</xref>; <xref ref-type="bibr" rid="bid.240">Ref. 5</xref>), originally based largely on <xref ref-type="other" rid="bid.1368">Pfam</xref> and <xref ref-type="other" rid="bid.1403">SMART</xref>, collections of alignments that represent functional domains conserved across evolution. CDD now also contains the alignments of the NCBI COG database, the NCBI Library of Ancient Domains (LOAD) along with new curated alignments assembled at NCBI. CDD can be searched from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd.shtml">CDD</ext-link> page in several ways, including by a domain keyword search. Three tools have been developed to assist in analysis of CDD: (1) the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi">CD-Search</ext-link>, which uses a <xref ref-type="other" rid="bid.1246">BLAST</xref>-based algorithm to search the position-specific scoring matrices (<xref ref-type="other" rid="bid.1386">PSSM</xref>) of CDD alignments; (2) the CD-Browser, which provides a graphic display of domains of interest, along with the sequence alignment; and (3) the Conserved Domain Architecture Retrieval Tool (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/lexington/lexington.cgi?cmd=rps">CDART</ext-link>), which searches for proteins with similar domain architectures.</p>
              <p>All the above databases and tools are discussed in more detail in other parts of this Chapter, including tips on how to make the best use of them.</p>
            </sec>
            <sec id="bid.97">
              <title>Content of the Molecular Modeling Database (MMDB)</title>
              <sec id="bid.98">
                <title>Sources of Primary Data</title>
                <boxed-text id="bid.99" position="margin">
                  <label>1</label>
                  <title>Accession numbers.</title>
                  <p>MMDB records have several types of Accession numbers associated with them, representing the following data types:
  
                    <list list-type="bullet">
                      <list-item id="bid.100">
                        <p> Each MMDB record has at least three Accession numbers: the PDB code of the corresponding PDB record (e.g., 1CYO, 1B8G); a unique MMDB-ID (e.g., 645, 12342); and a gi number for each protein chain. A new MMDB-ID is assigned whenever PDB updates either the sequence or coordinates of a structure record, even if the PDB code is retained.</p>
                      </list-item>
                      <list-item id="bid.1635">
                        <p> If an MMDB record contains more than one polypeptide or nucleotide chain, each chain in the MMDB record is assigned an Accession number in Entrez Protein or Nucleotide consisting of the PDB code followed by the letter designating that chain (e.g., 1B8GA, 3TATB, 1MUHB).</p>
                      </list-item>
                      <list-item id="bid.101">
                        <p>Each 3D Domain identified in an MMDB record is assigned 
                        a unique integer identifier that is appended to the Accession number of the chain to which it belongs (e.g., 1B8G A 2). This new Accession number becomes its identifier in Entrez 3D Domains. New 3D Domain identifiers are assigned whenever a new MMDB-ID is assigned.</p>
                      </list-item>
                      <list-item id="bid.102">
                        <p>For conserved domains, the Accession number is based on 
                        the source database:
                          <preformat>
    Pfam:          pfam00049
    SMART:         smart00078
    LOAD:          LOAD Toprim
    CD:            cd00101
    COG:           COG5641</preformat>
                        </p>
                      </list-item>
                    </list>

                  </p>
                </boxed-text>
                <p>To build MMDB (<xref ref-type="bibr" rid="bid.236">1</xref>), 3D structure data are retrieved from the PDB database (<xref ref-type="bibr" rid="bid.241">6</xref>) administered by the Research Collaboratory for Structural Bioinformatics (<xref ref-type="other" rid="bid.1391">RCSB</xref>). In all cases, the structures in MMDB have been determined by experimental methods, primarily X-ray crystallography and Nuclear Magnetic Resonance (<xref ref-type="other" rid="bid.1359">NMR</xref>) spectroscopy. Theoretical structure models are omitted. The data in each record are then checked for agreement between the atomic coordinates and the primary sequence, and the sequence data are then extracted from the coordinate set. The resulting agreement between sequence and structure allows the record to be linked efficiently into searches and alignment displays involving other NCBI databases.</p>
                <p>The data are converted into <xref ref-type="other" rid="bid.1242">ASN.1</xref> (<xref ref-type="bibr" rid="bid.242">7</xref>), which can be parsed easily and can also accept numerous annotations to the structure data. In contrast to a PDB record, a MMDB record in ASN.1 contains all necessary bonding information in addition to sequence information, allowing consistent display of the 3D structure using Cn3D. The annotations provided in the PDB record by the submitting authors are added, along with uniformly defined secondary structure and domain features. These features support structure-based similarity searches using VAST. Finally, two coordinate subsets are added to the record: one containing only backbone atoms, and one representing a single-conformer model in cases where multiple conformations or structures were present in the PDB record. Both of these additions further simplify viewing both an individual structure and its alignments with structure neighbors in Cn3D. When this process is complete, the record is assigned a unique Accession number, the MMDB-ID (<xref ref-type="boxed-text" rid="bid.99">Box 1</xref>), while also retaining the original four-character PDB code.</p>
              </sec>
              <sec id="bid.103">
                <title>Annotation of 3D Domains</title>
                <p>After initial processing, 3D domains are automatically identified within each MMDB record. 3D domains are annotations on individual MMDB structures that define the boundaries of compact substructures contained within them. In this way, they are similar to secondary structure annotations that define the boundaries of helical or &beta;-strand substructures. Because proteins are often similar at the level of domains, VAST compares each 3D domain to every other one and to complete polypeptide chains. The results are stored in Entrez as a <named-content content-type="book-spec-emph">Related 3D Domain</named-content> link.</p>
                <p>To identify 3D domains within a polypeptide chain, MMDB's domain parser searches for one or more breakpoints in the structure. These breakpoints fall between major secondary structure elements such that the ratio of intra- to interdomain contacts remains above a set threshold. The 3D domains identified in this way provide a means to both increase the sensitivity of structure neighbor calculations and also present 3D superpositions based on compact domains as well as on complete polypeptide chains. They are not intended to represent domains identified by comparative sequence and structure analysis, nor do they represent modules that recur in related proteins, although there is often good agreement between domain boundaries identified by these methods.</p>
              </sec>
              <sec id="bid.104">
                <title>Links to Other NCBI Resources</title>
                <p>After initially processing the PDB record, structure staff add a number of links and other information that further integrate the MMDB record with other NCBI resources. To begin, the sequence information extracted from the PDB record is entered into the Entrez Protein and/or Nucleotide databases as appropriate, providing a means to retrieve the structure information from sequence searches. As with all sequences in Entrez, precomputed BLAST searches are then performed on these sequences, linking them to other molecules of similar sequence. For proteins, these BLAST neighbors may be different than those determined by VAST; whereas VAST uses a conservative significance threshold, the structural similarities it detects often represent remote relationships not detectable by sequence comparison. The literature citations in the PDB record are linked to PubMed so that Entrez searches can allow access to the original descriptions of the structure determinations. Finally, semiautomatic processing of the &ldquo;source&rdquo; field of the PDB record provides links to the NCBI Taxonomy database. Although these links normally follow the genus and species information given, in some cases this information is either absent in the PDB record or refers only to how a sample was obtained. In these cases, the staff manually enters the appropriate taxonomy links.</p>
              </sec>
              <sec id="bid.105">
                <title>The MMDB Record</title>
                <p>The Structure Summary page for each MMDB record summarizes the database content for that record and serves as a starting point for analyzing the record using the NCBI structure tools (<xref ref-type="fig" rid="bid.106">Figure 2</xref>).<fig id="bid.106"><label>2</label><caption><title>The Structure Summary page.</title><p>The page consists of three parts: the header, the view bar, and the graphic display. The header contains basic identifying information about the record: a description of the protein (<italic>Description:</italic>), the author list (<italic>Deposition:</italic>), the species of origin (<italic>Taxonomy:</italic>), literature references (<italic>Reference:</italic>), the MMDB-ID (<italic>MMDB:</italic>), and the PDB code (<italic>PDB:</italic>). Several of these data serve as links to additional information. For example, the species name links to the Taxonomy browser, the literature references link to PubMed, and the PDB code links to the PDB Web site. The view bar allows the user to view the structure record either as a graphic with Cn3D or as a text record in either ASN.1, PDB (RasMol), or Mage formats. The latter can also be downloaded directly from this page. The graphic display contains a variety of information and links to related databases: (<italic>a</italic>) The Chain bar. Each chain of the molecule is displayed as a <italic>dark bar</italic> labeled with residue numbers. To the <italic>left</italic> of this bar is a <named-content content-type="book-spec-emph">Protein</named-content> hyperlink that takes the user to a view of the protein record in Entrez Protein. The bar itself is also a hyperlink and displays the VAST neighbors of the chain. If a structure contains nucleotide sequences, they are displayed in the order contained in the PDB record. A <named-content content-type="book-spec-emph">Nucleotide</named-content> hyperlink to their <italic>left</italic> takes the user to the appropriate record in Entrez Nucleotide. (<italic>b</italic>) The VAST (3D) Domain bar. The <italic>colored bars</italic> immediately below the Chain bar indicate the locations of structural domains found by the original MMDB processing of the protein. In many cases, such a domain contains unconnected sections of the protein sequence, and in such cases, discontinuous pieces making up the domain will have bars of the same color. To the <italic>left</italic> of the Domain bar is a 3D Domains hyperlink (<italic>3d Domains</italic>) that launches the 3D Domains browser in Entrez, where the user can find information about each constituent domain. Selecting a colored segment displays the VAST Structure Neighbors page for that domain. (<italic>c</italic>) The CD bar. Below the VAST Domain bar are <italic>rounded, rectangular bars</italic> representing conserved domains found by a CD-Search. The bars identify the best scoring hits; overlapping hits are shown only if the mutual overlap with hits having better scores is less than 50%. The <italic>CDs</italic> hyperlink to the <italic>left</italic> of the bar displays the CD records in Entrez Domains. Each of the colored bars is also a hyperlink that displays the corresponding CD Summary page configured to show the multiple alignment of the protein sequence with members of the selected CD.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch3f2" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.107">
                <title>VAST Structure Neighbors</title>
                <p>Although VAST itself is not a database, the VAST results computed for each MMDB record are stored with this record and are summarized on a separate page for the whole polypeptide chain as well as for each 3D domain found in the protein (<xref ref-type="fig" rid="bid.108">Figure 3</xref>). These pages can be accessed most easily by clicking on either the chain bar or the 3D Domain bar in the graphic display of the Structure Summary page (<xref ref-type="fig" rid="bid.106">Figure 2</xref>).<fig id="bid.108"><label>3</label><caption><title>VAST Structure Neighbors page.</title><p>The <italic>top portion</italic> of the page contains identifying information about the 3D Domain, along with three functional bars. (<italic>a</italic>) The View bar. This bar allows a user to view a selected alignment either as a graphic using Cn3D or as a sequence alignment in <xref ref-type="other" rid="bid.1312">HTML</xref>, text, or mFASTA format. The user may select which chains to display in the alignment by checking the boxes that appear to the <italic>left</italic> of each neighbor in the <italic>lower portion</italic> of the page. (<italic>b</italic>) The nr-PDB bar. This bar allows a user to either display all matching records in MMDB or to limit the displayed domains to only representatives of the selected nr-PDB set. The user may also select how the matching domains are sorted in the display and whether the results are shown as graphics or as tabulated data. (<italic>c</italic>) The Find bar. This bar allows the user to find specific structural neighbors by entering their PDB or MMDB identifiers. (<italic>d</italic>) The <italic>lower portion</italic> of the page displays a graphical alignment of the various matching domains. The <italic>upper three bars</italic> show summary information about the query sequence: the <italic>top bar</italic> shows the maximum extent of alignment found on all the sequences displayed on the current page (users should note that the appearance of this bar, therefore, depends on which hits are displayed); the <italic>middle bar</italic> represents the query sequence itself that served as input for the VAST search; and the <italic>lower bar</italic> shows any matching CDs and is identical to the CD bar on the Structure Summary page. Listed below these three summary bars are the hits from the VAST search, sorted according to the selection in the nr-PDB bar. Aligned regions are shown in <italic>red</italic>, with gaps indicating unaligned regions. To the <italic>left</italic> of each domain accession is a check box that can be used to select any combination of domains to be displayed either on this page or using Cn3D. Moreover, each of the bars in the display is itself a link, and placing the mouse pointer over any bar reveals both the extent of the alignment by residue number and the data linked to the bar.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch3f3" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.109">
                <title>nr-PDB</title>
                <p>The non-redundant PDB database (nr-PDB) is a collection of four sets of sequence-dissimilar cluster PDB polypeptide chains assembled by NCBI Structure staff. The four sets differ only in their respective levels of non-redundancy. The staff assembles each set by comparing all the chains available from PDB with each other using the BLAST algorithm. The chains are then clustered into groups of similar sequence using a single-linkage clustering procedure. Chains within a sequence-similar group are automatically ranked according to the quality of their structural data. Details of the measures used to determine structure precision and completeness and the methodology of assembling the nr-PDB clusters can be found on the nr-PDB <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/VAST/nrpdb.html">Web page</ext-link>.</p>
              </sec>
            </sec>
            <sec id="bid.110">
              <title>Content of the Conserved Domain Database (CDD)</title>
              <sec id="bid.111">
                <title>What Is a Conserved Domain (CD)?</title>
                <p>CDs are recurring units in polypeptide chains (sequence and structure motifs), the extents of which can be determined by comparative analysis. Molecular evolution uses such domains as building blocks and these may be recombined in different arrangements to make different proteins with different functions. The CDD contains sequence alignments that define the features that are conserved within each domain family. Therefore, the CDD serves as a classification resource that groups proteins based on the presence of these predefined domains. CDD entries often name the domain family and describe the role of conserved residues in binding or catalysis. Conserved domains are displayed in MMDB Structure summaries and link to a sequence alignment showing other proteins in which the domain is conserved, which may provide clues on protein function.</p>
              </sec>
              <sec id="bid.112">
                <title>Sources of Primary Data</title>
                <p>The collections of domain alignments in the CDD are imported either from two databases outside of the NCBI, named Pfam (<xref ref-type="bibr" rid="bid.243">8</xref>) and SMART (<xref ref-type="bibr" rid="bid.244">9</xref>); from the NCBI COB database; from another NCBI collection named LOAD; and from a database curated by the CDD staff. The first task is to identify the underlying sequences in each collection and then link these sequences to the corresponding ones in Entrez. If the CDD staff cannot find the Accession numbers for the sequences in the records from the source databases, they locate appropriate sequences using BLAST. Particular attention is paid to any resulting match that is linked to a structure record in MMDB, and the staff substitute alignment rows with such sequences whenever possible. After the staff imports a collection, they then choose a sequence that best represents the family. Whenever possible, the staff chooses a representative that has a structure record in MMDB.</p>
              </sec>
              <sec id="bid.113">
                <title>The Position-specific Score Matrix (PSSM)</title>
                <p>Once imported and constructed, each domain alignment in CDD is used to calculate a model sequence, called a <xref ref-type="other" rid="bid.1266">consensus sequence</xref>, for each CD. The consensus sequence lists the most frequently found residue in each position in the alignment; however, for a sequence position to be included in the consensus sequence, it must be present in at least 50% of the aligned sequences. Aligned columns covered by the consensus sequence are then used to calculate a <xref ref-type="other" rid="bid.1386">PSSM</xref>, which memorizes the degree to which particular residues are conserved at each position in the sequence. Once calculated, the PSSM is stored with the alignment and becomes part of the CDD. The <xref ref-type="other" rid="bid.1396">RPS-BLAST</xref> tool locates CDs within a query sequence by searching against this database of PSSMs.</p>
              </sec>
              <sec id="bid.114">
                <title>Reverse Position-specific BLAST (RPS-BLAST)</title>
                <p>RPS-BLAST (<xref ref-type="other" rid="bid.610">Chapter 16</xref>) is a variant of the popular Position-specific Iterated BLAST (PSI-BLAST) program. PSI-BLAST finds sequences similar to the query and uses the resulting alignments to build a PSSM for the query. With this PSSM the database is scanned again to draw in more hits and further refine the scoring model. RPS-BLAST uses a query sequence to search a database of precalculated PSSMs and report significant hits in a single pass. The role of the PSSM has changed from &ldquo;query&rdquo; to &ldquo;subject&rdquo;; hence, the term &ldquo;reverse&rdquo; in RPS-BLAST. RPS-BLAST is the search tool used in the CD-Search service.</p>
              </sec>
              <sec id="bid.115">
                <title>The CD Summary</title>
                <p>Analogous to the Structure Summary page, the CD Summary page displays the available information about a given CD and offers various links for either viewing the CD alignment or initiating further searches (<xref ref-type="fig" rid="bid.116">Figure 4</xref>). The CD Summary page can be retrieved by selecting the CD name on any page.<fig id="bid.116"><label>4</label><caption><title>CD summary page.</title><p>The <italic>top</italic> of the page serves as a header and reports a variety of identifying information, including the name and description of the CD, other related CDs with links to their summary pages, as well as the source database, status, and creation date of the CD. A taxonomic node link (<italic>Taxa:</italic>) launches the Taxonomy Browser, whereas a Proteins link (<italic>Proteins:</italic>) uses CDART to show other proteins that contain the CD. <italic>Below</italic> the header is the interface for viewing the CD alignment, which can be done either graphically with Cn3D (if the CD contains a sequence with structural data) or in HTML, text, or mFASTA format. It is also possible to view a selected number of the top-listed sequences, sequences from the most diverse members, or sequences most similar to the query. In addition, users may now select sequences with the NCBI Taxonomy Common  Tree tool. The <italic>lower portion</italic> of the page contains the alignment itself. Members with a structural record in MMDB are listed first, and the identifier of each sequence links to the corresponding record.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch3f4" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.1760">
                <title>CD Records Curated at NCBI</title>
                <p>In 2002, NCBI released the first group of curated CD records, a new and expanding set of annotated protein multiple sequence alignments and corresponding  structure alignments. These new records have Accession numbers beginning with  &ldquo;cd&rdquo; and have been added to the default CD-Search database. Most curated CD records are based on existing family descriptions from SMART and Pfam, but the alignments may have been revised extensively by quantitatively using three-dimensional structures and by re-examining the domain extent. In addition, CDD curators annotate conserved functional residues, ligands, and co-factors contained within the structures. They also record evidence for these sites as pointers to relevant literature or to three-dimensional structures exemplifying their properties. These annotations may be viewed using Cn3D and thus provide a direct way of visualizing functional properties of a protein domain in the context  of its three-dimensional structure. (See <xref ref-type="boxed-text" rid="bid.179">Box 3</xref> and <xref ref-type="fig" rid="bid.184">Figure 7</xref>.)</p>
              </sec>
              <sec id="bid.117">
                <title>The Distinction between 3D Domains and CDs</title>
                <p>The term &ldquo;domain&rdquo; refers in general to a distinct functional and/or structural unit of a protein. Each polypeptide chain in MMDB is analyzed for the presence of two classes of domains, and it is important for users to understand the difference between them. One class, called 3D Domains, is based solely on similar, compact substructures, whereas the second class, called Conserved Domains (CDs), is based solely on conserved sequence motifs. These two classifications often agree, because the compact substructures within a protein often correspond to domains joined by recombination in the evolutionary history of a protein. Note that CD links can be identified even when no 3D structures within a family are known. Moreover, 3D Domain links may also indicate relationships either to structures not included in CDD entries or to structures so distantly related that no significant similarity can be found by sequence comparisons.</p>
              </sec>
            </sec>
            <sec id="bid.118">
              <title>Finding and Viewing Structures</title>
              <boxed-text id="bid.119" position="margin">
                <label>2</label>
                <title>Example query: finding and viewing structural data of a protein.</title>
                <sec id="bid.120">
                  <title>Finding the Structure of a Protein</title>
                  <p>Suppose that we are interested in the biosynthesis of aminocyclopropanes and would like to find structural information on important active site residues in any available aminocyclopropane synthases. To begin, we would go to the <named-content content-type="book-spec-emph">Structure</named-content> main page and enter &ldquo;aminocyclopropane synthase&rdquo; in the <named-content content-type="book-spec-emph">Search</named-content> box. Pressing Enter displays a short list of structures, one of which is 1B8G, 1-aminocyclopropane-1-carboxylate synthase. Perhaps we would like to know the species from which this protein was derived. Selecting the <named-content content-type="book-spec-emph">Taxonomy</named-content> link to the right shows that this protein was derived from <italic>Malux x domestica</italic>, or the common apple tree. Going back to the Entrez results page and selecting the PDB code (1B8G) opens the Structure Summary page for this record. The species is again displayed on this page, along with a link to the <italic>Journal of Molecular Biology</italic> article describing how the structure was determined. We immediately see from this page that this protein appears as a dimer in the structure, with each chain having three 3D domains, as identified by VAST. In addition, CD-Search has identified an &ldquo;aminotran_1_2&rdquo; CD in each chain. Now we are ready to view the 3D structure.</p>
                </sec>
                <sec id="bid.121">
                  <title>Viewing the 3D Structure</title>
                  <p>Once we have found the Structure Summary page, viewing the 3D structure is straightforward. To view the structure in Cn3D, we simply select the <named-content content-type="book-spec-emph">View 3D Structure</named-content> button. The default view is to show helices in green, strands in brown, and loops in blue. This color scheme is also reflected in the Sequence/Alignment Viewer.</p>
                </sec>
                <sec id="bid.122">
                  <title>Locating an Active Site</title>
                  <p>Upon inspecting the structure, we immediately notice that a small molecule is bound to the protein, likely at the active site of the enzyme. How do we find out what that molecule is? One easy way is to return to the Structure Summary page and select the link to the PDB code, which takes us to the PDB Structure Explorer page for 1B8G. Quickly, we see that pyridoxal-5&prime;-phosphate (PLP) is a HET group, or heterogen, in the structure. Our interest piqued, we would now like to know more about the structural domain containing the active site. Returning to Cn3D, we manipulate the structure so that PLP is easily visible and then use the mouse to double-click on any PLP atom. The molecule becomes selected and turns yellow. Now from the <named-content content-type="book-spec-emph">Show/Hide</named-content> menu, we choose <named-content content-type="book-spec-emph">Select by distance</named-content> and <named-content content-type="book-spec-emph">Residues only</named-content> and enter 5 Angstroms for a search radius. Scanning the Sequence/Alignment Viewer, we see that seven residues are now highlighted: 117-119, 230, 268, 270, and 279. Glancing at the 3D Domain display in the Structure Summary page, we note that all of these residues lie in domain 3. We now focus our attention on this domain.</p>
                </sec>
                <sec id="bid.123">
                  <title>Viewing Structure Neighbors of a 3D Domain</title>
                  <p>Given that this enzyme is a dimer, we arbitrarily choose domain 3 in chain A, the accession of which is thus 1B8GA3. By clicking on the 3D Domain bar at a point within domain 3, we are taken to the VAST Structure Neighbors page for this domain, where we find nearly 200 structure neighbors.</p>
                </sec>
                <sec id="bid.124">
                  <title>Restricting the Search by Taxonomy</title>
                  <p>Perhaps we would now like to identify some of the most evolutionarily distant structure neighbors of domain 1B8GA3 as a means of finding conserved residues that may be associated with its binding and/or catalytic function. One powerful way of doing this is to choose structure neighbors from phylogenetically distant organisms. We therefore need to combine our present search with a Taxonomy search. Given that 1B8G is derived from the superkingdom Eukaryota, we would like to find structure neighbors in other superkingdom taxa, such as Eubacteria and Archaea. Returning to the Structure Summary page, select the 3D Domains link in the graphic display to open the list of 3D Domains in Entrez. Finding 1B8GA3 in the list, selecting the <named-content content-type="book-spec-emph">Related 3D Domains</named-content> link shows a list of all the structure neighbors of this domain. From this page, we select <named-content content-type="book-spec-emph">Preview/Index</named-content>, which shows our recent queries. Suppose our set of related 3D Domains is #5. We then perform two searches:
                    <preformat>
    1. #5 AND &ldquo;Archaea&rdquo;[Organism]
    2. #5 AND &ldquo;Eubacteria&rdquo;[Organism]</preformat>
                  </p>
                  <p>Looking at the Archaea results, we find among them 1DJUA3, a domain from an aromatic aminotransferase from <italic>Pyrococcus horikoshii</italic>. Concerning the Eubacteria results, we find among the several hundred matching domains 3TATA2, a tyrosine aminotransferase from <italic>Escherichia coli</italic>.</p>
                </sec>
                <sec id="bid.125">
                  <title>Viewing a 3D Superposition of Active Sites</title>
                  <p>Returning to the VAST Structure Neighbors page for 1B8GA3, we want to select 1DJUA3 and 3TATA2 to display in a structural alignment. One way to do this is to enter these two Accession numbers in the <named-content content-type="book-spec-emph">Find</named-content> box and press <named-content content-type="book-spec-emph">Find</named-content>. We now see only these two neighbors, and we can select <named-content content-type="book-spec-emph">View 3D Structure</named-content> to launch Cn3D.</p>
                  <p>Cn3D again displays the aligned residues in red, and we can highlight these further by selecting <named-content content-type="book-spec-emph">Show aligned residues</named-content> from the <named-content content-type="book-spec-emph">Show/Hide</named-content> menu. The excellent agreement between both the active site structures and the conformations of the bound ligands is readily apparent. Furthermore, by selecting <named-content content-type="book-spec-emph">Style/Coloring Shortcuts/Sequence Conservation/Variety</named-content>, we can easily see that the most highly conserved residues are concentrated near the binding site (<xref ref-type="fig" rid="bid.126">Figure 6</xref>).<fig id="bid.126"><label>6</label><caption><title>VAST structural alignment of 1B8GA3, 3TATA2, and 1DJUA3.</title><p>The backbone atoms of the aligned residues of the three structures are shown colored according to their sequence conservation of each position in the alignment. Highly conserved positions are colored more <italic>red</italic>, whereas poorly conserved positions are colored more <italic>blue</italic>. The bound pyridoxal phosphate ligands are <italic>yellow</italic>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch3f6" mime-subtype="jpg"/></fig></p>
                </sec>
              </boxed-text>
              <p>For an example query on finding and viewing structures, see <xref ref-type="boxed-text" rid="bid.119">Box 2</xref>.</p>
              <sec id="bid.127">
                <title>Why Would I Want to Do This?</title>
                  <list list-type="bullet">
                    <list-item id="bid.128">
                      <p> To determine the overall shape and size of a protein</p>
                    </list-item>
                    <list-item id="bid.129">
                      <p> To locate a residue of interest in the overall structure</p>
                    </list-item>
                    <list-item id="bid.130">
                      <p> To locate residues in close proximity to a residue of interest</p>
                    </list-item>
                    <list-item id="bid.131">
                      <p> To develop or test chemical hypotheses regarding an enzyme mechanism</p>
                    </list-item>
                    <list-item id="bid.132">
                      <p> To locate or predict possible binding sites of a ligand</p>
                    </list-item>
                    <list-item id="bid.133">
                      <p> To interpret mutation studies</p>
                    </list-item>
                    <list-item id="bid.134">
                      <p> To find areas of positive or negative charge on the protein surface</p>
                    </list-item>
                    <list-item id="bid.135">
                      <p> To locate particularly hydrophobic or hydrophilic regions of a protein</p>
                    </list-item>
                    <list-item id="bid.136">
                      <p> To infer the 3D structure and related properties of a protein with unknown structure from the structure of a <xref ref-type="other" rid="bid.1310">homologous</xref> protein</p>
                    </list-item>
                    <list-item id="bid.137">
                      <p> To study evolutionary processes at the level of molecular structure</p>
                    </list-item>
                    <list-item id="bid.138">
                      <p> To study the function of a protein</p>
                    </list-item>
                    <list-item id="bid.139">
                      <p> To study the molecular basis of disease and design novel treatments</p>
                    </list-item>
                  </list>
                <sec id="bid.140">
                  <title>How to Begin</title>
                  <p>The first step to any structural analysis at NCBI is to find the structure records for the protein of interest or for proteins similar to it. One may search MMDB directly by entering search terms such as PDB code, protein name, author, or journal in the Entrez Structure <named-content content-type="book-spec-emph">Search</named-content> box on the Structure <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure">homepage</ext-link>. Alternative points of entry are shown below.</p>
                  <p>By using the full array of Entrez search tools, the resulting list of MMDB records can be honed, ideally, to a workable list from which a record can be selected. Users should note that multiple records may exist for a given protein, reflecting different experimental techniques, conditions, and the presence or absence of various ligands or metal ions. Records may also contain different fragments of the full-length molecule. In addition, many structures of mutant proteins are also available. The PDB record for a given structure generally contains some description of the experimental conditions under which the structure was determined, and this file can be accessed by selecting the PDB code link at the top of the Structure Summary page.</p>
                </sec>
              </sec>
              <sec id="bid.141">
                <title>Alternative Points of Entry</title>
                <p>Structure Summary pages can also be found from the following NCBI databases and tools:
                  <list list-type="bullet">
                    <list-item id="bid.142">
                      <p> Select the Structure <named-content content-type="book-spec-emph">links</named-content> to the right of any Entrez record found; records with Structure links can also be located by choosing <named-content content-type="book-spec-emph">Structure links</named-content> from the <named-content content-type="book-spec-emph">Display</named-content> pull-down menu.</p>
                    </list-item>
                    <list-item id="bid.143">
                      <p> Select the <named-content content-type="book-spec-emph">Related Sequences</named-content> link to the right of an Entrez record to find proteins related by sequence similarity and then select <named-content content-type="book-spec-emph">Structure links</named-content> in the <named-content content-type="book-spec-emph">Display</named-content> pull-down menu.</p>
                    </list-item>
                    <list-item id="bid.144">
                      <p> Choose the PDB database from a <xref ref-type="other" rid="bid.1248">blastp</xref> (protein-protein BLAST) search; only sequences with structure records will be retrieved by BLAST. The <named-content content-type="book-spec-emph">Related Structures</named-content> link provides 3D views in Cn3D.</p>
                    </list-item>
                    <list-item id="bid.145">
                      <p> Select the <named-content content-type="book-spec-emph">3D Structures</named-content> button on any <xref ref-type="other" rid="bid.1250">BLink</xref> report to show those BLAST hits for which structural data are available.</p>
                    </list-item>
                    <list-item id="bid.1761">
                      <p>From the results of any protein BLAST search, click on a red 'S' linkout to view the sequence alignment with a structure record.</p>
                    </list-item>
                  </list>
                </p>
              </sec>
              <sec id="bid.146">
                <title>Viewing 3D Structures</title>
                <sec id="bid.147">
                  <title>3D Domains</title>
                  <p>The 3D domains of a protein are displayed on the Structure Summary page. It is useful to know how many 3D domains a protein contains and whether they are continuous in sequence when viewing the full 3D structure of the molecule.</p>
                </sec>
                <sec id="bid.148">
                  <title>Secondary Structure</title>
                  <p>Knowing the secondary structure of a protein can also be a useful prelude to viewing the 3D structure of the molecule. The secondary structure can be viewed easily by first selecting the <named-content content-type="book-spec-emph">Protein</named-content> link to the left of the desired chain in the graphic display. Finding oneself in Entrez Protein, selecting <named-content content-type="book-spec-emph">Graphics</named-content> in the Display pull-down menu presents secondary structure diagrams for the molecule.</p>
                </sec>
                <sec id="bid.149">
                  <title>Full Protein Structures</title>
                  <p>Cn3D is a software package for displaying 3D structures of proteins. Once it has been <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/CN3D/cn3dinstall.shtml">installed</ext-link> and the Internet browser has been configured correctly, simply selecting the <named-content content-type="book-spec-emph">View 3D Structure</named-content> button on a Structure Summary page launches the application. Once the structure is loaded, a user can manipulate and annotate it using an array of options as described in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/CN3D/cn3dtut.shtml">Cn3D Tutorial</ext-link>. By default, Cn3D colors the structure according to the secondary structure elements. However, another useful view is to color the protein by domain (see <named-content content-type="book-spec-emph">Style</named-content> menu options), using the same color scheme as is shown in the graphic display on the Structure Summary page. These color changes also affect the residues displayed in the Sequence/Alignment Viewer, allowing the identification of domain or secondary structure elements in the primary sequence. In addition to Cn3D, users can also display 3D structures with RasMol or Mage. Structures can also be saved locally as an ASN.1, PDB, or Mage file (depending on the choice of structure viewer) for later display.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.150">
              <title>Finding and Viewing Structure Neighbors</title>
              <p>For an example query on finding and viewing structure neighbors, see <xref ref-type="boxed-text" rid="bid.119">Box 2</xref>.</p>
              <sec id="bid.151">
                <title>Why Would I Want to Do This?</title>
                  <list list-type="bullet">
                    <list-item id="bid.152">
                      <p> To determine structurally conserved regions in a protein family</p>
                    </list-item>
                    <list-item id="bid.153">
                      <p> To locate the structural equivalent of a residue of interest in another related protein</p>
                    </list-item>
                    <list-item id="bid.154">
                      <p> To gain insights into the allowable structural variability in a particular protein family</p>
                    </list-item>
                    <list-item id="bid.155">
                      <p> To develop or test chemical hypotheses regarding an enzyme mechanism</p>
                    </list-item>
                    <list-item id="bid.156">
                      <p> To predict possible binding sites of a ligand from the location of a binding site in a related protein</p>
                    </list-item>
                    <list-item id="bid.157">
                      <p> To identify sites where conformational changes are concentrated</p>
                    </list-item>
                    <list-item id="bid.158">
                      <p> To interpret mutation studies</p>
                    </list-item>
                    <list-item id="bid.159">
                      <p> To find areas of conserved positive or negative charge on the protein surface</p>
                    </list-item>
                    <list-item id="bid.160">
                      <p> To locate conserved hydrophobic or hydrophilic regions of a protein</p>
                    </list-item>
                    <list-item id="bid.161">
                      <p> To identify evolutionary relationships across protein families</p>
                    </list-item>
                    <list-item id="bid.162">
                      <p> To identify functionally equivalent proteins with little or no sequence conservation</p>
                    </list-item>
                  </list>
              </sec>
              <sec id="bid.163">
                <title>How to Begin</title>
                <p>The Vector Alignment Search Tool (VAST) is used to calculate similar structures on each protein contained in the MMDB. The graphic display on each Structure Summary page (<xref ref-type="fig" rid="bid.106">Figure 2</xref>) links directly to the relevant VAST results for both whole proteins and 3D domains:
                  <list list-type="bullet">
                    <list-item id="bid.164">
                      <p> The 3D Domains link transfers the user to Entrez 3D Domains, showing a list of the VAST neighbors.</p>
                    </list-item>
                    <list-item id="bid.165">
                      <p> Selecting the chain bar displays the VAST Structure Neighbors page for the entire chain.</p>
                    </list-item>
                    <list-item id="bid.166">
                      <p> Selecting a 3D Domain bar displays the VAST Structure Neighbors page for the selected domain.</p>
                    </list-item>
                  </list>
                </p>
              </sec>
              <sec id="bid.167">
                <title>Alternative Points of Entry</title>
                  <list list-type="bullet">
                    <list-item id="bid.169">
                      <p> From any Entrez search, select <named-content content-type="book-spec-emph">Related 3D Domains</named-content> to the right of any record found to view the Vast Structure Neighbors page.</p>
                    </list-item>
                  </list>
                <sec id="bid.170">
                  <title>Viewing a 2D Alignment of Structure Neighbors</title>
                  <p>A graphic 2D HTML alignment of VAST neighbors can be viewed as follows:
                    <list list-type="bullet">
                      <list-item id="bid.171">
                        <p> On the lower portion of the VAST Structure Neighbors page (<xref ref-type="fig" rid="bid.108">Figure 3</xref>), select the desired neighbors to view by checking the boxes to their left.</p>
                      </list-item>
                      <list-item id="bid.172">
                        <p> On the <named-content content-type="book-spec-emph">View/Save</named-content> bar, configure the pull-down menus to the right of the <named-content content-type="book-spec-emph">View Alignment</named-content> button.</p>
                      </list-item>
                      <list-item id="bid.173">
                        <p> Select <named-content content-type="book-spec-emph">View Alignment</named-content>.</p>
                      </list-item>
                    </list>
                  </p>
                </sec>
                <sec id="bid.174">
                  <title>Viewing a 3D Alignment of Structure Neighbors</title>
                  <p>Alignments of VAST structure neighbors can be viewed as a 3D image using Cn3D.</p>
                    <list list-type="bullet">
                      <list-item id="bid.175">
                        <p> On the lower portion of the VAST Structure Neighbors page (<xref ref-type="fig" rid="bid.108">Figure 3</xref>), select the desired neighbors to view by checking the boxes to their left.</p>
                      </list-item>
                      <list-item id="bid.176">
                        <p> On the <named-content content-type="book-spec-emph">View/Save</named-content> bar, configure the pull-down menus to the right of the <named-content content-type="book-spec-emph">View 3D Structure</named-content> button.</p>
                      </list-item>
                      <list-item id="bid.177">
                        <p> Select <named-content content-type="book-spec-emph">View 3D Structure</named-content>.</p>
                      </list-item>
                    </list>
                  <p>Cn3D automatically launches and displays the aligned structures. Each displayed chain has a unique color; however, the portions of the structures involved in the alignment are shown in red. These same colors are also reflected in the Sequence/Alignment Viewer. Among the many viewing options provided by Cn3D, of particular use is the <named-content content-type="book-spec-emph">Show/Hide</named-content> menu that allows only the aligned residues to be viewed, only the aligned domains, or all residues of each chain.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.178">
              <title>Finding and Viewing Conserved Domains</title>
              <boxed-text id="bid.179" position="margin">
                <label>3</label>
                <title>Example query: finding and viewing CDs in a protein.</title>
                <sec id="bid.180">
                  <title>Finding CDs in a Protein</title>
                  <p>Suppose that we are interested in topoisomerase enzymes and would like to find human topoisomerases that most closely resemble those found in eubacteria and thus may share a common ancestor. Further suppose that through a colleague, we are aware of a recent and particularly interesting crystal structure of a topoisomerase from <italic>Escherichia coli</italic> with PDB code 1I7D. How can we identify the conserved functional domains in this protein and then find human proteins with the same domains? From the Structure main page, we enter the PDB code 1I7D in the Structure <named-content content-type="book-spec-emph">Summary</named-content> search box and quickly find the Structure Summary page for this record. We see that in this crystal structure, the protein is complexed with a single-stranded oligonucleotide. We also see that the protein has five 3D Domains. Two CDs align to the sequence as well, and they overlap with one another at the N-terminus of the protein in the region corresponding to the first 3D domain.</p>
                </sec>
                <sec id="bid.181">
                  <title>Analyzing CDs Found in a Protein</title>
                  <p>The Struture summary page displays only the CDs that give the best match to the protein sequence. To see all of the matching CDs, we can easily perform a full CD-Search. Select the <named-content content-type="book-spec-emph">Protein</named-content> link to the left of the graphic to reveal the flatfile for the record. Then follow the <named-content content-type="book-spec-emph">Domains</named-content> link in the <named-content content-type="book-spec-emph">Link</named-content> menu on the right to view the results of the CD-Search. Select <named-content content-type="book-spec-emph">Show Details</named-content> to see all CDs matching the query sequence. We find that nine CDs match this sequence, and that the statistics of each match are shown below the alignment graphic. The CD with the best hit is TopA from the COG database, and it is further clear that this domain consists of two smaller domains: TOPRIM (alignments from Pfam, SMART, and curated CD) and a topoisomerase domain (alignments from Pfam and curated CD). We can learn more about these CDs by studying the pairwise alignments at the bottom of the page and by studying their CD Summary pages, reached by selecting the links to their left.</p>
                </sec>
                <sec id="bid.182">
                  <title>Finding Other Proteins with Similar Domain Architecture</title>
                  <p>We now would like to find human proteins that have these same CDs. To perform a CDART search, simply select the <named-content content-type="book-spec-emph">Show Domain Relatives</named-content> button. To limit these results to human proteins, we select the <named-content content-type="book-spec-emph">Subset by Taxonomy</named-content> button. A taxonomic tree is then displayed, and we next check the box for <named-content content-type="book-spec-emph">Mammal</named-content>, the lowest taxa including <italic>Homo sapiens</italic>. Selecting <named-content content-type="book-spec-emph">Choose</named-content> then displays a Common Tree, and by clicking on the appropriate &ldquo;scissor&rdquo; icons, we can cut away all branches except the one leading to <italic>H. sapiens</italic>. We can execute this taxonomic restriction by selecting <named-content content-type="book-spec-emph">Go back</named-content>, and we now find a much shorter list of CDART results. In the most similar group, we find two members, one of which is NP_004609. Selecting the <named-content content-type="book-spec-emph">more&gt;</named-content> link for this record shows the CD-Search results for this human protein. Interestingly, we find that the topoisomerase is very well conserved, whereas only a portion of the TOPRIM domain has been retained.</p>
                </sec>
                <sec id="bid.183">
                  <title>Viewing a CD Alignment with a 3D Structure</title>
                  <p>We now would like to view the alignment of the topoisomerase in the human protein to other members of this CD. On the CD-Search page, select the colored bar of this CD to see a CD-Browser window displaying the alignment. Because this is a curated CD record, we are able to view functional features of the protein domain on a structural template. The rightmost menu in the View Alignment bar shows the available features for this domain, whereas the topmost row in the alignment itself marks the residues involved in this feature with # symbols. The second row of the alignment is the consensus sequence of the CD record, whereas the third row contains the NP_004609 sequence, labeled &ldquo;query&rdquo;. At the bottom of the page, buttons allow Cn3D to be launched with various structural features highlighted. For example, if we are interested in nucleotide binding site II, Cn3D will launch with the view depicted in <xref ref-type="fig" rid="bid.184">Figure 7</xref>, showing the bound nucleotide in orange. Additonal Cn3D windows not shown in Figure 7 allow one to highlight the binding site residues yellow as shown, and these highlights also appear in the sequence window. In this figure, the NP_004609 sequence has been merged into the alignment (bottom row) using tools within Cn3D, and the result shows that this human protein closely conserves these important functional residues.
<fig id="bid.184"><label>7</label><caption><title>Sequence and structure views of the TOP1Ac conserved domain common to type III 
bacterial and eukaryotic DNA topoisomerases.</title><p>The <italic>upper window</italic> displays the structure of the domain with the residues colored according to their sequence conservation, with <italic>red</italic> indicating high 
conservation and <italic>blue</italic> indicating low conservation. The nucleotide bound at site II is shown as an <italic>orange</italic> space-filling model, and the residues involved in this binding site are <italic>yellow</italic>. The <italic>lower window</italic> displays the sequence alignment for the domain with aligned residues shown as colored <italic>capital letters</italic>. Residues aligned to three of the binding site residues are highlighted in <italic>yellow</italic>. The sequence for NP_004609 (gi 10835218) occupies the <italic>bottom row</italic>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch3f7" mime-subtype="jpg"/></fig></p>
                </sec>
              </boxed-text>
              <p>For an example query on finding and viewing conserved domains, see <xref ref-type="boxed-text" rid="bid.179">Box 3</xref>.</p>
              <sec id="bid.185">
                <title>Why Would I Want to Do This?</title>
                  <list list-type="bullet">
                    <list-item id="bid.186">
                      <p> To locate functional domains within a protein</p>
                    </list-item>
                    <list-item id="bid.187">
                      <p> To predict the function of a protein whose function is unknown</p>
                    </list-item>
                    <list-item id="bid.188">
                      <p> To establish evolutionary relationships across protein families</p>
                    </list-item>
                    <list-item id="bid.189">
                      <p> To interpret mutation studies</p>
                    </list-item>
                    <list-item id="bid.190">
                      <p> To predict the structure of a protein of unknown structure</p>
                    </list-item>
                  </list>
              </sec>
              <sec id="bid.191">
                <title>How to Begin</title>
                <p>Following the Domains link for any protein in Entrez, one can find the conserved domains within that protein. The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi">CD-Search</ext-link> (or Protein BLAST, with CD-Search option selected) can be used to find conserved domains (CDs) within a protein. Either the Accession number, gi number, or the <xref ref-type="other" rid="bid.1290">FASTA</xref> sequence can be used as a query.</p>
              </sec>
              <sec id="bid.192">
                <title>Alternative Points of Entry</title>
                <p>Information on the CDs contained within a protein can also be found from these databases and tools:
                  <list list-type="bullet">
                    <list-item id="bid.198">
                      <p> From any Entrez search: select the <named-content content-type="book-spec-emph">Domains</named-content> link to the right of a displayed record.</p>
                    </list-item>
                    <list-item id="bid.193">
                      <p> From the Structure Summary page of a MMDB record: this page displays the CDs within each protein chain immediately below the 3D Domain bar in the graphic display. Selecting the <named-content content-type="book-spec-emph">CDs</named-content> link shows the CD-Search results page.</p>
                    </list-item>
                    <list-item id="bid.194">
                      <p> From an Entrez Domains search: choose <named-content content-type="book-spec-emph">Domains</named-content> from the Entrez <named-content content-type="book-spec-emph">Search</named-content> pull-down menu and enter a search term to retrieve a list of CDs. Clicking on any resulting CD displays the CD Summary page. To find the location of this CD in an aligned protein, select the CD link following a protein name in the bottom portion of this page.</p>
                    </list-item>
                    <list-item id="bid.195">
                      <p> From the CDD page: locate CDs by entering text terms into the search box and proceed as for an Entrez CD search.</p>
                    </list-item>
                    <list-item id="bid.196">
                      <p> From a BLink report: select the <named-content content-type="book-spec-emph">CDD-Search</named-content> button to display the CD-Search results page.</p>
                    </list-item>
                    <list-item id="bid.197">
                      <p> From the BLAST main page: follow the RPS-BLAST link to load the CD-Search page.</p>
                    </list-item>
                  </list>
                </p>
              </sec>
              <sec id="bid.199">
                <title>Viewing Conserved Domains</title>
                <p>Results from a CD search are displayed as colored bars underneath a sequence ruler. Moving the mouse over these bars reveals the identity of each domain; domains are also listed in a format similar to BLAST summary output (<xref ref-type="other" rid="bid.610">Chapter 16</xref>). Pairwise alignments between the matched region of the target protein and the representative sequence of each domain are shown below the bar. Red letters indicate residues identical to those in the representative sequence, whereas blue letters indicate residues with a positive BLOSUM62 score in the BLAST alignment.</p>
              </sec>
              <sec id="bid.200">
                <title>Viewing Multiple Alignments of a Query Protein with Members of a Conserved Domain</title>
                <p>These can be displayed by clicking a CD bar within a MMDB Structure Summary page or from a hyperlinked CD name on a CD-Search results page.</p>
              </sec>
              <sec id="bid.201">
                <title>Viewing CD Alignments in the Context of 3D Structure</title>
                <p>If members of a CD have MMDB records, one of these records can be viewed as a 3D image along with the sequence alignment using Cn3D (launched by selecting the pink dot on a CD-Search results page). As in other alignment views, colored capital letters indicate aligned residues, allowing the sequence of the protein sequence of interest to be mapped onto the available 3D structure.</p>
              </sec>
            </sec>
            <sec id="bid.202">
              <title>Finding and Viewing Proteins with Similar Domain Architectures</title>
              <p>For an example query on finding and viewing proteins with similar domain architectures, see <xref ref-type="boxed-text" rid="bid.179">Box 3</xref>.</p>
              <sec id="bid.203">
                <title>Why Would I Want to Do This?</title>
                  <list list-type="bullet">
                    <list-item id="bid.204">
                      <p> To locate related functional domains in other protein families</p>
                    </list-item>
                    <list-item id="bid.205">
                      <p> To gain insights into how a given CD is situated within a protein relative to other CDs</p>
                    </list-item>
                    <list-item id="bid.206">
                      <p> To explore functional links between different CDs</p>
                    </list-item>
                    <list-item id="bid.207">
                      <p> To predict the function of a protein whose function is unknown</p>
                    </list-item>
                    <list-item id="bid.208">
                      <p> To establish evolutionary relationships across protein families</p>
                    </list-item>
                  </list>
              </sec>
              <sec id="bid.209">
                <title>How to Begin</title>
                <p>Following the <named-content content-type="book-spec-emph">Domain Relatives</named-content> link for any protein in Entrez, one can find other proteins with similar domain architecture. The Conserved Domain Architecture Retrieval Tool (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/lexington/lexington.cgi?cmd=rps">CDART</ext-link>) can take an Accession number or the FASTA sequence as a query to find out the domain architecture of a protein sequence and list other proteins with related domain architectures.</p>
              </sec>
              <sec id="bid.210">
                <title>Alternative Points of Entry</title>
                  <list list-type="bullet">
                    <list-item id="bid.1762">
                      <p>From a CD-Search results page, click <named-content content-type="book-spec-emph">Show Domain Relatives</named-content></p>
                    </list-item>
                    <list-item id="bid.1763">
                      <p>From a CD-Summary page, click the <named-content content-type="book-spec-emph">Proteins</named-content> link</p>
                    </list-item>
                    <list-item id="bid.1764">
                      <p>From an Entrez Domains searc, click the <named-content content-type="book-spec-emph">Proteins</named-content> link in the Links menu</p>
                    </list-item>
                  </list>
              </sec>
              <sec id="bid.211">
                <title>Results of a CDART Search</title>
                <p>These are described in <xref ref-type="fig" rid="bid.212">Figure 5</xref>. The protein &ldquo;hits&rdquo;, which have similar domain architectures to the query sequence, can be further refined by taxonomic group, in which the results can be limited to selected nodes of the taxonomic tree. Furthermore, search results may be limited to those that contain only particular conserved domains.<fig id="bid.212"><label>5</label><caption><title>A CDART results page.</title><p>At the <italic>top</italic> of the CDART results page in a <italic>yellow box</italic>, the query sequence CDs are represented as &ldquo;beads on a string&rdquo;. Each CD had a unique color and shape and is labeled both in the display itself and in a legend located at the <italic>bottom</italic> of the page. The shapes representing CDs are hyperlinked to the corresponding CD summary page. The matching proteins to the query are listed <italic>below</italic> the yellow box, ranked according to the number of non-redundant hits to the domains in the query sequence. Each match is either a single protein, in which case its Accession number is shown, or is a cluster of very similar proteins, in which case the number of members in the cluster is shown. Cluster members can be displayed by selecting the logo to the <italic>left</italic> of its diagram. Selecting any protein Accession number displays the flatfile for that protein. To the <italic>right</italic> of any drawing for a single protein (either on the main results page or after expanding a protein cluster) is a <named-content content-type="book-spec-emph">more&gt;</named-content> link, which displays the CD-Search results page for the selected protein so that the sequence alignment, e.g., of a CDART hit with a CD contained in the original protein of interest, can be examined.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch3f5" mime-subtype="gif"/></fig></p>
              </sec>
            </sec>
            <sec id="bid.213">
              <title>Links Between Structure and Other Resources</title>
              <sec id="bid.214">
                <title>Integration with Other NCBI Resources</title>
                <p>As illustrated in the sections above, there are numerous connections between the Structure resources and other databases and tools available at the NCBI. What follows is a listing of major tools that support connections.</p>
                <sec id="bid.215">
                  <title>Entrez</title>
                  <p>Because Entrez is an integrated database system (<xref ref-type="other" rid="bid.588">Chapter 15</xref>), the links attached to each structure give immediate access to PubMed, Protein, Nucleotide, 3D Domain, CDD, or Taxonomy records.</p>
                </sec>
                <sec id="bid.216">
                  <title>BLAST</title>
                  <p>Although the BLAST service is designed to find matches based solely on sequence, the sequences of Structure records are included in the BLAST databases, and by selecting the PDB search database, BLAST searches only the protein sequences provided by MMDB records. A new <named-content content-type="book-spec-emph">Related Structure</named-content> link provides 3D views for sequences with structure data identified in a BLAST search.</p>
                </sec>
                <sec id="bid.217">
                  <title>BLink</title>
                  <p>The BLink report represents a precomputed list of similar proteins for many proteins (see, for example, links from LocusLink records; <xref ref-type="other" rid="bid.788">Chapter 19</xref>). The <named-content content-type="book-spec-emph">3D Structures</named-content> option on any BLink report shows the BLAST hits that have 3D structure data in MMDB, whereas the <named-content content-type="book-spec-emph">CDD-Search</named-content> button displays the CD-Search results page for the query protein.</p>
                </sec>
                <sec id="bid.218">
                  <title>Microbial Genomes</title>
                  <p>A particularly useful interface with the structural databases is provided on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/micr.html">Microbial Genomes page</ext-link> (<xref ref-type="bibr" rid="bid.245">10</xref>). To the left of the list of genomes are several hyperlinks, two of which offer users direct access to structural information. The red <named-content content-type="book-spec-emph">[D]</named-content> link displays a listing of every protein in the genome, each with a link to a BLink page showing the results of a BLAST pdb search for that protein. The <named-content content-type="book-spec-emph">[S]</named-content> link displays a similar protein list for the selected genome, but now with a listing of the conserved domains found in each protein by a CD-Search.</p>
                </sec>
              </sec>
              <sec id="bid.219">
                <title>Links to Non-NCBI Resources</title>
                <sec id="bid.220">
                  <title>The Protein Data Bank (PDB)</title>
                  <p>As stated elsewhere, all records in the MMDB are obtained originally from the Protein Data Bank (PDB) (<xref ref-type="bibr" rid="bid.241">6</xref>). Links to the original PDB records are located on the Structure Summary page of each MMDB record. Updates of the MMDB with new PDB records occur once a month.</p>
                </sec>
                <sec id="bid.221">
                  <title>Pfam and SMART</title>
                  <p>The CDD staff imports CD collections from both the Pfam and SMART databases. Links to the original records in these databases are located on the appropriate CD Summary page. Both Pfam and SMART are updated several times per year in roughly bimonthly intervals, and the CDD staff update CDD accordingly.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.222">
              <title>Saving Output from Database Searches</title>
              <sec id="bid.223">
                <title>Exporting Graphics Files from Cn3D</title>
                <p>Structures displayed in Cn3D can be exported as a Portable Network Graphics (<xref ref-type="other" rid="bid.1377">PNG</xref>) file from within Cn3D (the Export PNG command in the <named-content content-type="book-spec-emph">File</named-content> menu). The structure file itself, in the orientation currently being viewed, can also be saved for later launching in Cn3D.</p>
              </sec>
              <sec id="bid.224">
                <title>Saving Individual MMDB Records</title>
                <p>Individual MMDB records can be saved/downloaded to a local computer directly from the Structure Summary page for that record. <named-content content-type="book-spec-emph">Save File</named-content> in the <named-content content-type="book-spec-emph">View</named-content> bar downloads the file in a choice of three formats: ASN.1 (select <named-content content-type="book-spec-emph">Cn3D</named-content>); PDB (select <named-content content-type="book-spec-emph">RasMol</named-content>); or Mage (select <named-content content-type="book-spec-emph">Mage</named-content>).</p>
              </sec>
              <sec id="bid.225">
                <title>Saving VAST Alignments</title>
                <p>Alignments of VAST neighbors can be saved/downloaded from the VAST Structure Neighbors page of any MMDB record. By selecting options in the <named-content content-type="book-spec-emph">View Alignment</named-content> pull-down menu, the alignment data can be saved, formatted as HTML, text, or <xref ref-type="other" rid="bid.1343">mFASTA</xref>, and then saved.</p>
              </sec>
            </sec>
            <sec id="bid.226">
              <title>FTP</title>
              <sec id="bid.1765">
                <title>MMDB</title>
                <p>Users can download the NCBI Structure databases from the NCBI FTP site: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/mmdb">ftp://ftp.ncbi.nih.gov/mmdb</ext-link>. A Readme file contains descriptions of the contents and information about recent updates. Within the mmdb directory are four subdirectories that contain the following data:
                  <list list-type="bullet">
                    <list-item id="bid.227">
                      <p> mmdbdata: the current MMDB database (NOTE: these files can not be read directly by Cn3D).</p>
                    </list-item>
                    <list-item id="bid.228">
                      <p> vastdata: the current set of VAST neighbor annotations to MMDB records</p>
                    </list-item>
                    <list-item id="bid.229">
                      <p> nrtable: the current non-redundant PDB database</p>
                    </list-item>
                    <list-item id="bid.230">
                      <p> pdbeast: table listing the taxonomic classification of MMDB records</p>
                    </list-item>
                  </list>
                </p>
              </sec>
              <sec id="bid.1766">
                <title>CDD</title>
                <p>CDD data can be downloaded from <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/pub/mmdb/cdd">ftp://ftp.ncbi.nih.gov/pub/mmdb/cdd</ext-link>. A Readme file contains descriptions of the data archives. Users can download the PSSMs for each CD record, the sequence alignments in mFASTA format, or a text file containing the accessions and descriptions of all CD records. </p>
              </sec>
            </sec>
            <sec id="bid.231">
              <title>Frequently Asked Questions</title>
                 <list list-type="bullet">
                  <list-item id="bid.232">
                    <p>
                      <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/Structure/CN3D/cn3dfaq.shtml">Cn3D</ext-link>
                    </p>
                  </list-item>
                  <list-item id="bid.233">
                    <p>
                      <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/Structure/VAST/vastsearch_faq.html">VAST searches</ext-link>
                    </p>
                  </list-item>
                  <list-item id="bid.234">
                    <p>
                      <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/cdd/cdd_help.shtml">CDD</ext-link>
                    </p>
                  </list-item>
                </list>
            </sec>
            </body>
            <back>
              <ref-list>
              <title>References</title>
                <ref id="bid.236">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Wang</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Anderson</surname>
                        <given-names>JB</given-names>
                      </name>
                      <name>
                        <surname>Chen</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Geer</surname>
                        <given-names>LY</given-names>
                      </name>
                      <name>
                        <surname>He</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Hurwitz</surname>
                        <given-names>DI</given-names>
                      </name>
                      <name>
                        <surname>Liebert</surname>
                        <given-names>CA</given-names>
                      </name>
                      <name>
                        <surname>Madej</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Marchler</surname>
                        <given-names>GH</given-names>
                      </name>
                      <name>
                        <surname>Marchler-Bauer</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Panchenko</surname>
                        <given-names>AR</given-names>
                      </name>
                      <name>
                        <surname>Shoemaker</surname>
                        <given-names>BA</given-names>
                      </name>
                      <name>
                        <surname>Song</surname>
                        <given-names>JS</given-names>
                      </name>
                      <name>
                        <surname>Thiessen</surname>
                        <given-names>PA</given-names>
                      </name>
                      <name>
                        <surname>Yamashita</surname>
                        <given-names>RA</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>SH</given-names>
                      </name>
                    </person-group>
                    <article-title>MMDB: Entrez's 3D-structure database</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>249</fpage>
                    <lpage>252</lpage>
                    <pub-id pub-id-type="pmid">11752307</pub-id>
                    <pub-id pub-id-type="other">99072</pub-id>
                  </citation>
                </ref>
                <ref id="bid.237">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Wang</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Geer</surname>
                        <given-names>LY</given-names>
                      </name>
                      <name>
                        <surname>Chappey</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Kans</surname>
                        <given-names>JA</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>SH</given-names>
                      </name>
                    </person-group>
                    <article-title>Cn3D: sequence and structure views for Entrez</article-title>
                    <source>Trends Biochem Sci</source>
                    <year>2000</year>
                    <volume>25</volume>
                    <fpage>300</fpage>
                    <lpage>302</lpage>
                    <pub-id pub-id-type="pmid">10838572</pub-id>
                  </citation>
                </ref>
                <ref id="bid.238">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Madej</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Gibrat</surname>
                        <given-names>J-F</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>SH</given-names>
                      </name>
                    </person-group>
                    <article-title>Threading a database of protein cores</article-title>
                    <source>Proteins</source>
                    <year>1995</year>
                    <volume>23</volume>
                    <fpage>356</fpage>
                    <lpage>369</lpage>
                    <pub-id pub-id-type="pmid">8710828</pub-id>
                  </citation>
                </ref>
                <ref id="bid.239">
                  <label>4</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Gibrat</surname>
                        <given-names>J-F</given-names>
                      </name>
                      <name>
                        <surname>Madej</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>SH</given-names>
                      </name>
                    </person-group>
                    <article-title>Surprising similarities in structure comparison</article-title>
                    <source>Curr Opin Struct Biol</source>
                    <year>1996</year>
                    <volume>6</volume>
                    <fpage>377</fpage>
                    <lpage>385</lpage>
                    <pub-id pub-id-type="pmid">8804824</pub-id>
                  </citation>
                </ref>
                <ref id="bid.240">
                  <label>5</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Marchler-Bauer</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Panchenko</surname>
                        <given-names>AR</given-names>
                      </name>
                      <name>
                        <surname>Shoemaker</surname>
                        <given-names>BA</given-names>
                      </name>
                      <name>
                        <surname>Thiessen</surname>
                        <given-names>PA</given-names>
                      </name>
                      <name>
                        <surname>Geer</surname>
                        <given-names>LY</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>SH</given-names>
                      </name>
                    </person-group>
                    <article-title>CDD: a database of conserved domain alignments with links to domain three-dimensional structure</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>281</fpage>
                    <lpage>283</lpage>
                    <pub-id pub-id-type="pmid">11752315</pub-id>
                    <pub-id pub-id-type="other">99109</pub-id>
                  </citation>
                </ref>
                <ref id="bid.241">
                  <label>6</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Westbrook</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Feng</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Jain</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Bhat</surname>
                        <given-names>TN</given-names>
                      </name>
                      <name>
                        <surname>Thanki</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Ravichandran</surname>
                        <given-names>V</given-names>
                      </name>
                      <name>
                        <surname>Gilliland</surname>
                        <given-names>GL</given-names>
                      </name>
                      <name>
                        <surname>Bluhm</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Weissig</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Greer</surname>
                        <given-names>DS</given-names>
                      </name>
                      <name>
                        <surname>Bourne</surname>
                        <given-names>PE</given-names>
                      </name>
                      <name>
                        <surname>Berman</surname>
                        <given-names>HM</given-names>
                      </name>
                    </person-group>
                    <article-title>The Protein Data Bank: unifying the archive</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>245</fpage>
                    <lpage>248</lpage>
                    <pub-id pub-id-type="pmid">11752306</pub-id>
                    <pub-id pub-id-type="other">99110</pub-id>
                  </citation>
                </ref>
                <ref id="bid.242">
                  <label>7</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Ohkawa</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Ostell</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>S</given-names>
                      </name>
                    </person-group>
                    <article-title>MMDB: an ASN.1 specification for macromolecular structure</article-title>
                    <source>Proc Int Conf Intell Syst Mol Biol</source>
                    <year>1995</year>
                    <volume>3</volume>
                    <fpage>259</fpage>
                    <lpage>267</lpage>
                    <pub-id pub-id-type="pmid">7584445</pub-id>
                  </citation>
                </ref>
                <ref id="bid.243">
                  <label>8</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Bateman</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Birney</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Cerruti</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Durbin</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Etwiller</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Eddy</surname>
                        <given-names>SR</given-names>
                      </name>
                      <name>
                        <surname>Griffiths-Jones</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Howe</surname>
                        <given-names>KL</given-names>
                      </name>
                      <name>
                        <surname>Marshall</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Sonnhammer</surname>
                        <given-names>ELL</given-names>
                      </name>
                    </person-group>
                    <article-title>The Pfam proteins family database</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>276</fpage>
                    <lpage>280</lpage>
                    <pub-id pub-id-type="pmid">11752314</pub-id>
                    <pub-id pub-id-type="other">99071</pub-id>
                  </citation>
                </ref>
                <ref id="bid.244">
                  <label>9</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Letunic</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Goodstadt</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Dickens</surname>
                        <given-names>NJ</given-names>
                      </name>
                      <name>
                        <surname>Doerks</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Schultz</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Mott</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Ciccarelli</surname>
                        <given-names>F</given-names>
                      </name>
                      <name>
                        <surname>Copley</surname>
                        <given-names>RR</given-names>
                      </name>
                      <name>
                        <surname>Ponting</surname>
                        <given-names>CP</given-names>
                      </name>
                      <name>
                        <surname>Bork</surname>
                        <given-names>P</given-names>
                      </name>
                    </person-group>
                    <article-title>SMART: a Web-based tool for the study of genetically mobile domains. Recent improvements to the SMART domain-based sequence annotation resource</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>242</fpage>
                    <lpage>244</lpage>
                    <pub-id pub-id-type="pmid">11752305</pub-id>
                    <pub-id pub-id-type="other">99073</pub-id>
                  </citation>
                </ref>
                <ref id="bid.245">
                  <label>10</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Wang</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Tatusov</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Tatusova</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>Links from genome proteins to known 3D structures</article-title>
                    <source>Genome Res</source>
                    <year>2000</year>
                    <volume>10</volume>
                    <fpage>1643</fpage>
                    <lpage>1647</lpage>
                    <pub-id pub-id-type="pmid">11042161</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part><?mtl hellip?>
<book-part id="bid.246" book-part-type="chapter" book-part-number="4">
<book-part-meta>
<title-group>
<title>The Taxonomy Project</title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Federhen</surname>
<given-names>Scott</given-names></name>
</contrib>
</contrib-group>
<history>
<date date-type="created">
<day>9</day>
<month>10</month>
<year>2002</year>
</date>
<date date-type="updated">
<day>13</day>
<month>8</month>
<year>2003</year>
</date>
</history>
<alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" 
alternate-form-type="pdf" xlink:href="ch4d1"/>
<abstract>
<title>Summary</title>
<p>The NCBI Taxonomy database is a curated set of names and 
classifications for all of the organisms that are represented 
in GenBank. When new sequences are submitted to GenBank, the 
submission is checked for new organism names, which are then 
classified and added to the Taxonomy database. As of April 
2003, there were 176,890 total taxa represented.</p>
<p>There are two main tools for viewing the information in the 
Taxonomy database: the Taxonomy Browser, and Taxonomy Entrez. 
Both systems allow searching of the Taxonomy database for names, 
and both link to the relevant sequence data. However, the 
Taxonomy Browser provides a hierarchical view of the classification 
(the best display for most casual users interested in exploring 
our classification), whereas Entrez Taxonomy provides a uniform 
indexing, search, and retrieval engine with a common mechanism 
for linking between the Taxonomy and other relevant Entrez 
databases.</p>
</abstract>
</book-part-meta>
<body>
<sec id="bid.247">
<title>Introduction</title>
<boxed-text id="bid.249" position="margin">
<label>1</label>
<title>History of the Taxonomy Project.</title>
<p>By the time the NCBI was created in 1988, the nucleotide 
sequence databases (GenBank, <abbrev>EMBL</abbrev>, and 
<abbrev>DDBJ</abbrev>) each maintained their own taxonomic 
classifications. All three classifications derived from 
the one developed at the Los Alamos National Lab 
(<abbrev>LANL</abbrev>) but had diverged considerably. 
Furthermore, the protein sequence databases (<xref 
ref-type="other" rid="bid.1412">SWISS-PROT</xref> and 
<xref ref-type="other" rid="bid.1374">PIR</xref>) each 
developed their own taxonomic classifications that were 
very different from each other and from the nucleotide 
database taxonomies. To add to the mix, in 1990 the NCBI 
and the <xref ref-type="other" rid="bid.1358">NLM</xref> 
initiated a journal-scanning program to capture and annotate 
sequences reported in the literature that had not been 
submitted to any of the sequence databases. We, of course, 
began to assign our own taxonomic classifications for these 
records.</p>
<p>The Taxonomy Project started in 1991, in association with 
the launch of Entrez (<xref ref-type="other" 
rid="bid.588">Chapter 15</xref>). The goal was to combine 
the many taxonomies that existed at the time into a single 
classification that would span all of the organisms represented 
in any of the GenBank sources databases (<xref ref-type="other" 
rid="bid.2">Chapter 1</xref>).</p>
<p>To represent, manipulate, and store versions of each of 
the different database taxonomies, we wrote a stand-alone, 
tree-structured database manager, TaxMan. This also allowed 
us to merge the taxonomies into a single composite classification. 
The resulting hybrid was, at first, a bigger mess than any of 
the pieces had been, but it gave us a starting point that spanned 
all of the names in all of the sequence databases. For many years, 
we cleaned up and maintained the NCBI Taxonomy database with 
TaxMan.</p>
<p>After the initial unification and clean-up of the taxonomy for 
Entrez was complete, Mitch Sogin organized a workshop to give us 
advice on the clean-up and recommendations for the long-term 
maintenance of the taxonomy. This was held at the NCBI in 1993 
and included: Mitch Sogin (protists), David Hillis (chordates), 
John Taylor (fungi), S.C. Jong (fungi), John Gunderson (protists), 
Russell Chapman (algae), Gary Olsen (bacteria), Michael Donoghue 
(plants), Ward Wheeler (invertebrates), Rodney Honeycutt 
(invertebrates), Jack Holt (bacteria), Eugene Koonin (viruses), 
Andrzej Elzanowski (PIR taxonomy), Lois Blaine (ATCC), and Scott 
Federhen (NCBI). Many of these attendees went on to serve as 
curators for different branches of the classification. In particular, 
David Hillis, John Taylor, and Gary Olsen put in long hours to help 
the project move along.</p>
<p>In 1995, as more demands were made on the Taxonomy database, 
the system was moved to a SyBase relational database 
(<abbrev>TAXON</abbrev>), originally developed by Tim Clark. 
Hierarchical organism indexing was added to the Nucleotide and 
Protein domains of Entrez, and the Taxonomy browser made its 
first appearance on the Web.</p>
<p>In 1997, the <abbrev>EMBL</abbrev> and <abbrev>DDBJ</abbrev> 
databases agreed to adopt the NCBI taxonomy as the standard 
classification for the nucleotide sequence databases. Before that, 
we would see new organism names from the <abbrev>EMBL</abbrev> and 
<abbrev>DDBJ</abbrev> only after their entries were released to 
the public, and any corrections (in spelling, or nomenclature, or 
classification) would have to be made after the fact. We now receive 
taxonomy consults on new names from the <abbrev>EMBL</abbrev> and 
<abbrev>DDBJ</abbrev> before the release of their entries, just as 
we do from our own GenBank indexers. <abbrev>SWISS-PROT</abbrev> 
has also recently (2001) agreed to use our Taxonomy database and 
send us taxonomy consults.</p>
</boxed-text>
<p>Organismal taxonomy is a powerful organizing principle in the 
study of biological systems. Inheritance, homology by common descent, 
and the conservation of sequence and structure in the determination 
of function are all central ideas in biology that are directly related 
to the evolutionary history of any group of organisms. Because of this, 
taxonomy plays an important cross-linking role in many of the 
<xref ref-type="other" rid="bid.1353">NCBI</xref> tools and databases.</p>
<p>The NCBI Taxonomy database is a curated set of names and classifications 
for all of the organisms that are represented in <xref ref-type="other" 
rid="bid.1299">GenBank</xref>. When new sequences are submitted to GenBank, 
the submission is checked for new organism names, which are then classified 
and added to the taxonomy database. As of April 1, 2003, there were 
4,653 families, 26,427 genera, 130,207 species, and 176,890 total taxa 
represented.</p>
<p>Of the several different ways to build a taxonomy, our group maintains 
a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/primer/phylo.html">phylogenetic 
taxonomy</ext-link>. In a phylogenetic classification scheme, the structure 
of the taxonomic tree approximates the evolutionary relationships among the 
organisms included in the classification (the &ldquo;tree of life&rdquo;; 
see <xref ref-type="fig" rid="bid.248">Figure 1</xref>).
<fig id="bid.248">
<label>1</label>
<caption>
<title>A phylogenetic classification scheme.</title>
<p>If two organisms (A and B) are listed more closely together in the taxonomy 
than either is to organism C, the assertion is that C diverged from the 
lineage leading to A+B earlier in evolutionary history, and that A and B 
share a common ancestor that is not in the direct line of evolutionary 
descent to species C. For example, the current consensus it that the closest 
living relatives of the birds are the crocodiles; therefore, our classification 
does not include the familiar taxon Reptilia (turtles, lizards and snakes, 
and crocodiles), which excludes the birds, and would break the phylogenetic 
principle outlined above.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" 
xlink:href="ch4f1" mime-subtype="gif"/>
</fig>
</p>
<p>Our classification represents an assimilation of information from many 
different sources (see <xref ref-type="boxed-text" rid="bid.249">Box 1</xref>). 
Much of the success of the project is attributable to the flood of new 
molecular data that has revolutionized our understanding of the phylogeny 
of many groups, especially of previous poorly understood groups such as 
Bacteria, Archea, and Fungi. Users should be aware that some parts of the 
classification are better developed than others and that the primary 
<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/taxonomyhome.html/
index.cgi?chapter=resources">systematic and phylogenetic literature</ext-link> 
is the most reliable information source.</p>
<p>We do not rely on sequence data alone to build our classification, and 
we do not perform phylogenetic analysis ourselves as part of the taxonomy 
project. Most of the organisms in GenBank are represented by only a 
snippet of sequence; therefore, sequence information alone is not enough 
to build a robust phylogeny. The vast majority of species are not there 
at all, although about 50% of the birds and the mammals are represented. 
We therefore also rely on analyses from morphological studies; the 
challenge of modern systematics is to unify molecular and morphological 
data to elucidate the evolutionary history of life on earth.</p>
</sec>
<sec id="bid.250">
<title>Adding to the Taxonomy Database</title>
<p>Currently, more than 100 new species are added to the database 
daily, and the rate is accelerating as sequence analysis becomes an 
ever more common component of systematic research and the taxonomic 
description of new species.</p>
<sec id="bid.251">
<title>Sources of New Names</title>
<p>The <xref ref-type="other" rid="bid.1281">EMBL</xref> and 
<xref ref-type="other" rid="bid.1272">DDBJ</xref> databases, as well 
as GenBank, now use the NCBI Taxonomy as the standard classification 
for nucleotide sequences (see <xref ref-type="boxed-text" 
rid="bid.249">Box 1</xref>). Nearly all of the new species found in 
the Taxonomy database are via sequences submitted to one of these 
databases from species that are not yet represented. In these cases, 
the NCBI taxonomy group is consulted, and any problems with the 
nomenclature and classifications are resolved before the sequence 
entries are released to the public. We also receive consults for 
submissions that are not identified to the species level (e.g., 
&ldquo;Hantavirus&rdquo; or &ldquo;Bacillus sp.&rdquo;) and for 
Anything that looks confusing, incorrect, or incomplete to the 
database indexers. All consults include information on the problem 
Organism names, source features, and publication titles (if any). 
The email addresses for the submitters are also included in case 
we need to contact them about the nomenclature, classification, 
or annotation of their entries.</p>
<p>The number and complexity of organisms in a submission can vary 
enormously. Many contain a single new name, others may include 100 
species, all from the same familiar genus, whereas others may include 
100 names (only half identified at the species level) from 100 genera 
(all of which are new to the Taxonomy database) without any other 
identifying information at all.</p>
<p>Some new organism names are found by software when the protein 
sequence databases (<abbrev>SWISS-PROT</abbrev>, <abbrev>PIR</abbrev>, 
and the <xref ref-type="other" rid="bid.1380">PRF</xref>) are added to 
Entrez; because most of the entries in the protein databases have been 
derived from entries in the nucleotide database, this is a small number. 
The NCBI structure group may also find new names in the <xref 
ref-type="other" rid="bid.1367">PDB</xref> protein structure database. 
Finally, because we made the Taxonomy database publicly accessible on 
the Web, we have had a steady stream of comments and corrections to our 
spellings and classification from outside users.</p>
</sec>
<sec id="bid.252">
<title>More on Submission</title>
<p>We often receive consults on submissions with explicitly new species 
names that will be published as part of the description of a new species. 
These sequence entries (like any other) may be designated &ldquo;hold 
until published&rdquo; (<abbrev>HUP</abbrev>) and will not be released 
until the corresponding journal article has been published. These species 
names will not appear on any of our taxonomy Web sites until the 
corresponding sequence entries have been released.</p>
<p>Occasionally, the same new genus name is proposed simultaneously 
for different taxa; in one case, two papers with conflicting new names 
had been submitted to the same journal, and both had gone through one 
round of review and revision without detection of the duplication. Although 
these duplications would have been discovered in time, the increasingly 
common practice of including some sequence analysis in the description of 
a new species can lead to earlier detection of these problems. In many 
cases, the new species name proposed in the submitted manuscript is 
changed during the editorial review process, and a different name appears 
in the publication. Submitters are encouraged to inform us when their new 
descriptions have been published, particularly if the proposed names have 
been changed.</p>
<p>We strongly encourage the submission of strain names for cultured 
bacteria, algae, and fungi and for sequences from laboratory animals in 
biochemical and genetic studies; of cultivar names for sequences from 
cultivated plants; and of specimen vouchers (something that definitively 
ties the sequence to its source) for sequences from phylogenetic studies. 
There are many other kinds of useful information that may be contained 
within the sequence submission, but these data are the bare minimum 
necessary to maintain a reliable link between an entry in the sequence 
database and the biological source material.</p>
</sec>
</sec>
<sec id="bid.253">
<title>Using the Taxonomy Browser</title>
<p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" 
ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/
taxonomyhome.html/">Taxonomy Browser</ext-link> (TaxBrowser) provides 
a hierarchical view of the classification from any particular place 
in the taxonomy. This is probably the display of choice for most casual 
users (browsers) of the taxonomy who are interested in exploring our 
classification. The TaxBrowser displays only the subset of taxa from 
the taxonomy database that is linked to public sequence entries. About 
15% of the full Taxonomy database is not displayed on the public Web 
pages because the names are from sequence entries that have not yet been 
released.</p>
<p>TaxBrowser is updated continuously. New species will appear on a daily 
basis as the new names appear in sequence entries indexed during the daily 
release cycle of the Entrez databases. New taxa in the classification 
appear in TaxBrowser on an ongoing basis, as sections of the taxonomy 
already linked to public sequence entries are revised.</p>
<sec id="bid.254">
<title>The Hierarchical Display</title>
<p>The browser produces two different kinds of Web pages: 
<list list-type="arabic">
<list-item><label>(1)</label>
<p>hierarchy pages, which present a familiar indented <!--<glossref>-->
flatfile<!--</glossref>--> view of the taxonomic classification, 
centered on a particular taxon in the database; and</p>
</list-item>
<list-item><label>(2)</label>
<p>taxon-specific pages, which summarize all of the information that 
we associate with any particular taxonomic entry in the database. For 
example, &ldquo;hominidae&rdquo; as a search term from the <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/
taxonomyhome.html/">TaxBrowser homepage</ext-link> finds our human 
family (<xref ref-type="fig" rid="bid.255">Figure 2</xref>).
<fig id="bid.255">
<label>2</label>
<caption>
<title>The Taxonomy Browser hierarchical display for the family 
Hominidae.</title>
<p>(<italic>a</italic>) There are four genera listed in this family 
(Gorilla, Homo, Pan, and Pongo) with six species-level names 
(<italic>Gorilla gorilla</italic>, <italic>Homo sapiens</italic>, 
<italic>Pan paniscus</italic>, <italic>Pan troglodytes</italic>, 
<italic>Pongo pygmaeus</italic>, and <italic>Pongo</italic> sp.) 
and 2 subspecies. Common names are shown in <italic>parentheses</italic> 
if they are available in the Taxonomy database. The lineage above 
Hominidae is shown in the <italic>line</italic> at the 
<italic>top</italic> of the display; selecting the word 
<named-content content-type="book-spec-emph">Lineage</named-content> 
will toggle back and forth between the abbreviated lineage (the display 
used in GenBank flatfiles) and the full lineage (as it appears in the 
Taxonomy database). Selecting any of the taxa <italic>above</italic> 
Hominidae (in the lineage) or <italic>below</italic> Hominidae (in 
the hierarchical display) will refocus the browser on that taxon 
instead of the Hominidae. Selecting Hominidae itself, however, will 
display the taxon-specific page for the Hominidae. (<italic>b</italic>) 
The default setting displays three levels of the classification on 
the hierarchy pages. To change this, enter a different number in the 
<named-content content-type="book-spec-emph">Display levels</named-content> 
box and select <named-content content-type="book-spec-emph">Accept</named-content>. 
If any of the check boxes to the <italic>right</italic> of the <named-content 
content-type="book-spec-emph">Display levels</named-content> box are 
selected (i.e., <italic>Nucleotide, Protein, Structures, ...</italic>), 
the numbers of records in the corresponding Entrez database that are 
associated with that taxon will appear as hyperlinks. Selecting a link 
retrieves those records.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" 
xlink:href="ch4f2a" mime-subtype="gif"/>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" 
xlink:href="ch4f2b" mime-subtype="gif"/>
</fig>
</p>
</list-item>
</list>
</p>
</sec>
<sec id="bid.256">
<title>The Taxon-specific Display</title>
<p>The taxon-specific browser display page shows all of the information 
that is associated with a particular taxon in the Taxonomy database and 
some information collected through links with related databases 
(<xref ref-type="fig" rid="bid.257">Figure 3</xref>).
<fig id="bid.257">
<label>3</label>
<caption>
<title>The Taxonomy Browser taxon-specific display.</title>
<p>The name, taxid, rank, <xref ref-type="other" rid="bid.1300">genetic 
code</xref>s, and other names (if any) associated with this taxon are all 
listed. The full lineage is shown; selecting the word <named-content 
content-type="book-spec-emph">Lineage</named-content> toggles between 
the abbreviated and full versions. There may also be citation information 
and comments hyperlinked to the appropriate sources. The numbers of Entrez 
database records that link to this Taxonomy record are displayed and can 
be retrieved via the hotlinked entry counts. LinkOut links to external 
resources appear at the bottom of the page.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" 
xlink:href="ch4f3" mime-subtype="gif"/>
</fig>
</p>
<p>There are two sets of links to Entrez records from the Taxonomy 
Browser. The &ldquo;subtree links&rdquo; are accumulated up the tree 
in a hierarchical fashion; for example, there are 16 million nucleotide 
records and a half million protein records associated with the Chordata 
(<xref ref-type="fig" rid="bid.257">Figure 3a</xref>). These are all 
linked into the taxonomy at or below the species leves and can be 
retrieved en masse via the subtree hotlink.</p>
<p>&ldquo;Direct links&rdquo; will retrieve Entrez records that are 
linked directly to this particular node in the taxonomy database. Many 
of the Entrez domains (e.g., sequences and structures) are linked into 
the taxonomy at or below the species level; it is a data error when a 
sequence entry is directly linked into the taxonomy at a taxon somewhere 
above the species level. For other Entrez domains (e.g., literature and 
phylogenetic sets), this is not the case. A journal article may talk 
about several different species but may also refer directly to the 
<italic>phylum Chordata</italic>. We have searched the full text of the 
articles in the PubMed Central archive with the scientific names from the 
taxonomy database. Twenty-seven articles in PubMed Central refer directly 
to the <italic>phylum Chordata</italic>; 9,299 articles are linked into the 
taxonomy somewhere in the <italic>Chordata</italic> subtree. The PopSet 
domain contains population studies, phylogenetic sets, and alignments. We 
have recently changed the way that we index phylogenetic sets in Entrez. 
The five &ldquo;Direct links&rdquo; at the <italic>Chordata</italic> 
will retrieve the phylogenetic sets that explicitly span the 
<italic>Chordata</italic>; the &ldquo;Subtree links&rdquo; will also 
include phylogenetic sets that are completely contained within the 
<italic>Chordata</italic>.</p>
<p>The taxon-specific browser pages now also show the NCBI LinkOut links 
to external resources. These include links to a broad range of different 
kinds of resources and are provided for the convenience of our users; the 
NCBI does not vouch for the content of these resources, although we do 
make an effort to ensure that they are of good scientific quality. A 
complete list of external resources can be found <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/entrez/journals/
loprovlink.html">here</ext-link>. Groups interested in participating in 
the LinkOut program should visit the <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/
entrez/linkout/">LinkOut homepage</ext-link>.</p>
</sec>
<sec id="bid.258">
<title>Search Options</title>
<boxed-text id="bid.262" position="margin">
<label>2</label>
<title>TaxBrowser results from using the wild card search, C* elegans.</title>
<!--TONYA, ASK DEBBIE IF THIS SHOULD BE A LIST. SURE LOOKS 
LIKE A SIMPLE LIST TO ME.-->
<p><italic>Cunninghamella elegans</italic></p>
<p><italic>Caenorhabditis elegans</italic></p>
<p><italic>Codonanthe elegans</italic></p>
<p><italic>Cyclamen coum subsp. elegans</italic></p>
<p><italic>Cestrum elegans</italic></p>
<p> <italic>Chaerophyllum elegans</italic></p>
<p><italic>Chalara elegans</italic></p>
<p><italic>Chrysemys scripta elegans</italic></p>
<p><italic>Ceuthophilus elegans</italic></p>
<p><italic>Carpolepis elegans</italic></p>
<p><italic>Cylindrocladiella elegans</italic></p>
<p><italic>Coluria elegans</italic></p>
<p><italic>Cymbidium elegans</italic></p>
<p><italic>Coronilla elegans</italic></p>
<p><italic>Gymnothamnion elegans</italic> (synonym: 
<italic>Callithamnion elegans</italic>)</p>
<p><italic>Centruroides elegans</italic></p>
</boxed-text>
<p>There are several different ways to search for names in the 
Taxonomy database. If the search results in a terminal node in our 
taxonomy, the taxon-specific browser page is displayed; if the search 
returns with an internal (non-terminal) node, the hierarchical 
classification page is displayed.</p>
<!--TONYA, SEE IF DEBBIE AGREES THAT THIS SHOULD BE DEF-LIST.-->
<p><bold>Complete Name.</bold> By default, TaxBrowser looks for the 
complete name when a term is typed into the search box. It looks for 
a case-insensitive, full-length string match to all of the nametypes 
stored in the Taxonomy database. For example, <italic>Homo sapiens, 
Escherichia, Tetrapoda,</italic> and <italic>Embryophyta</italic> would 
all retrieve results.</p>
<p>Names can be duplicated in the Taxonomy database, but the taxonomy 
browser can only be focused on a single taxon at any one time. If a 
complete name search retrieves more than one entry from the taxonomy, 
an intermediate name selection screen appears (<xref ref-type="fig" 
rid="bid.260">Figure 4</xref>). Each duplicated name includes a manually 
curated suffix that differentiates between the duplicated names.
<fig id="bid.260">
<label>4</label>
<caption>
<title>Examples of search results for duplicated names.</title>
<p>Searches for <italic>Bacillus</italic> (<italic>a</italic>), 
<italic>Proboscidea</italic> (<italic>b</italic>), <italic>Anopheles</italic> 
(<italic>c</italic>), and reptiles (<italic>d</italic>) result in the 
options shown above. If the duplicated name is not the primary (scientific) 
name of the node [as in (<italic>d</italic>)], the primary name is given 
first, followed by the nametype and the duplicated name in <italic>square 
brackets</italic>.</p></caption><graphic 
xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch4f4" 
mime-subtype="gif"/></fig></p>
<p><bold>Wild Card.</bold> This is a regular expression search, *, with 
wild cards. It is useful when the correct spelling of a scientific name 
is uncertain or to find ambiguous combinations for abbreviated species 
names. For example, <italic>C* elegans</italic> results in a list of 16 
species and subspecies (<xref ref-type="boxed-text" rid="bid.262">Box 2</xref>). 
Note: there is still only one <italic>H. sapiens</italic>.</p>
<p><bold>Token Set.</bold> This treats the search string as an unordered set 
of tokens, each of which must be found in one of the names associated with a 
particular node. For example, &ldquo;sapiens&rdquo; retrieves:
<preformat>
Homo sapiens
Homo sapiens neanderthalensis
</preformat>
</p>
<p><bold>Phonetic Name.</bold> This search qualifier can be used when the 
user has exhausted all other search options to find the organism of interest. 
The results using this function can be patchy, however. For example, 
&ldquo;drozofila&rdquo; and &ldquo;kaynohrhabdietees&rdquo; retrieve 
respectable results; however, &ldquo;seenohrabdietees&rdquo; and 
&ldquo;eshereesheeya&rdquo; are not found.</p>
<p><bold>Taxonomy ID.</bold> This allows searching by the numerical 
unique identifier (taxid) of the NCBI Taxonomy database, e.g., 
<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/
wwwtax.cgi?id=9606">9606</ext-link> or 
<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/
wwwtax.cgi?id=666">666</ext-link>.</p>
</sec>
<sec id="bid.266">
<title>How to Link to the TaxBrowser</title>
<p>There is a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" 
ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/
taxonomyhome.html/index.cgi?chapter=howlink">help page</ext-link> 
that describes how to make hyperlinks to the Taxonomy Browser pages.</p>
</sec>
</sec>
<sec id="bid.267">
<title>The Taxonomy Database: TAXON</title>
<p>The NCBI Taxonomy database is stored as a SyBase relational database, 
called TAXON. The NCBI taxonomy group maintains the database with a 
customized software tool, the Taxonomy Editor. Each entry in the database 
is a &ldquo;taxon&rdquo;, also referred to as a &ldquo;node&rdquo; in the 
database. The &ldquo;root node&rdquo; (taxid1) is at the top of the 
hierarchy. The path from the root node to any other particular taxon in 
the database is called its &ldquo;lineage&rdquo;; the collection of all 
of the nodes beneath any particular taxon is called its &ldquo;subtree&rdquo;. 
Each node in the database may be associated with several names, of several 
different nametypes. For indexing and retrieval purposes, the nametypes are 
essentially equivalent.</p>
<p>The Taxonomy database is populated with species names that have appeared 
in a sequence record from one of the nucleotide or protein databases. If a 
name has ever appeared in a sequence record at any time (even if it is not 
found in the current version of the record), we try to keep it in the 
Taxonomy database for tracking purposes (as a synonym, a misspelling, or 
other nametype), unless there are good reasons for removing it completely 
(for example, if it might cause a future submission to map to the wrong 
place in the taxonomy).</p>
<sec id="bid.268">
<title>Taxids</title>
<table-wrap id="bid.269">
<label>1</label>
<caption>
<title>Files on the taxonomy FTP site.</title>
</caption>
<table>
<thead>
<tr>
<th align="left" valign="top">File</th>
<th align="left" valign="top">Uncompresses to</th>
<th align="left" valign="top">Description</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">taxdump.tar.Z<xref ref-type="fn" 
rid="mul2"><sup><italic>a</italic></sup></xref></td>
<td align="left" valign="top">readme.txt</td>
<td align="left" valign="top">A terse description of the dmp files</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="left" valign="top">nodes.dmp</td>
<td align="left" valign="top">Structure of the database; lists each 
taxid with its parent taxid, rank, and other values associated with 
each node (genetic codes, etc.)</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="left" valign="top">names.dmp</td>
<td align="left" valign="top">Lists all the names associated with 
each taxid</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="left" valign="top">delnodes.dmp</td>
<td align="left" valign="top">Deleted taxid list</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="left" valign="top">merged.dmp</td>
<td align="left" valign="top">Merged nodes file</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="left" valign="top">division.dmp</td>
<td align="left" valign="top">GenBank division files</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="left" valign="top">gencode.dmp</td>
<td align="left" valign="top">Genetic codes files</td>
</tr>
<tr>
<td align="left" valign="top"/>
<td align="left" valign="top">gc.prt</td>
<td align="left" valign="top">Print version of genetic codes</td>
</tr>
<tr>
<td align="left" valign="top">gi_taxid_nucl.dmp.gz</td>
<td align="left" valign="top">gi_taxid_nucl.dmp</td>
<td align="left" valign="top">A list of gi_taxid pairs for every 
live gi-identified sequence in the nucleotide sequence database</td>
</tr>
<tr>
<td align="left" valign="top">gi_taxid_prot.dmp.gz</td>
<td align="left" valign="top">gi_taxid_prot.dmp</td>
<td align="left" valign="top">A list of gi_taxid pairs for every 
live gi-identified sequence in the protein sequence database</td>
</tr>
<tr>
<td align="left" valign="top">gi_taxid_nucl_diff.dmp</td>
<td align="left" valign="top">gi_taxid_nucl_diff</td>
<td align="left" valign="top">List of differences between latest 
gi_taxid_nucl and previous listing</td>
</tr>
<tr>
<td align="left" valign="top">gi_taxid_prot_diff.dmp</td>
<td align="left" valign="top">gi_taxid_prot_diff</td>
<td align="left" valign="top">List of differences between latest 
gi_taxid_prot and previous listing</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn symbol="a" id="mul2">
<p>For non-UNIX users, the file taxdmp.zip includes the same 
(zip compressed) data.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Each taxon in the database has a unique identifier, its taxid. 
Taxids are assigned sequentially. When a taxon is deleted, its taxid 
disappears and is not reassigned (<xref ref-type="table" 
rid="bid.269">Table 1</xref>; see the <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" 
xlink:href="ftp://ncbi.nlm.nih.gov/pub/taxonomy/">FTP</ext-link> 
for a list of deleted taxids). When one taxon is merged with another 
taxon (e.g., if the names were determined to be synonyms or one was 
a misspelling), the taxid of the node that has disappeared is listed 
as a &ldquo;secondary taxid&rdquo; to the taxid of the node that 
remains (see the merged taxid file on the <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" 
xlink:href="ftp://ncbi.nlm.nih.gov/pub/taxonomy/">FTP</ext-link> site). 
In either case, the taxid that has disappeared will never be assigned 
to a new entry in the database.</p>
</sec>
</sec>
<sec id="bid.270">
<title>Nomenclature Issues</title>
<sec id="bid.271">
<title>TAXON Nametypes</title>
<p>There are many possible types of names that can be associated 
with an organism taxid in TAXON. To track and display the names 
correctly, the various names associated with a taxid are tagged 
with a nametype, for example &ldquo;scientific name&rdquo;, 
&ldquo;synonym&rdquo;, or &ldquo;common name&rdquo;. Each taxid 
<bold>must have one</bold> (and only one) scientific name but may 
have zero or many other names (for example, several synonyms, 
several common names, along with only one &ldquo;GenBank common 
name&rdquo;).</p>
<p>When sequences are submitted to GenBank, usually only a scientific 
name is included; most other names are added by NCBI taxonomists at 
the time of submission or later, when further information is discovered. 
For a complete description of each nametype used in TAXON, see 
<xref ref-type="app" rid="bid.301">Appendix 1</xref>.</p>
</sec>
<sec id="bid.272">
<title>Classes of TAXON Scientific Names</title>
<p>Scientific names, the only required nametype for a taxid, can be 
further qualified into different classes. Not all &ldquo;scientific 
names&rdquo; that accompany sequence submissions are true Linnaean 
Latin binomial names; if the taxon is not identified to the species 
level, it is not possible to assign a binomial name to it. For 
indexing and retrieval purposes, TAXON needs to know whether the 
scientific name is a Latin binomial name, or otherwise. A full 
listing of the classes of TAXON scientific names can be viewed in 
<xref ref-type="app" rid="bid.317">Appendix 2</xref>.</p>
</sec>
<sec id="bid.273">
<title>Duplicated Names</title>
<p>The treatment of duplicated names was discussed briefly in the 
section on the Taxonomy browser. For our purposes, there are four 
main classes of duplicated scientific names: (1) real duplicate names, 
(2) structural duplicates, (3) polyphyletic genera, and (4) other 
duplicate names.</p>
<sec id="bid.274">
<title>Real Duplicate Names</title>
<p>There are several main codes of nomenclature for living organisms: 
the Zoological Code (International Code of Zoological Nomenclature, 
ICZN; for animals), the Botanical Code (International Code of Botanical 
Nomenclature, ICBN; for plants), the Bacteriological Code (International 
Code of Nomenclature of Bacteria, ICNB; for prokaryotes), and the Viral 
Code (International Code of Virus Classification and Nomenclature, ICVCN; 
for viruses). Within each code, names are required to be unique. When 
duplicate names are discovered within a code, one of them is changed 
(generally, the newer duplicate name). However, the codes are complex, 
and not all names are subject to these restrictions. For example, 
<italic>Polyphaga</italic> is both a genus of cockroaches and a suborder 
of beetles, and the damselfly genus <italic>Lestoidea</italic> is listed 
within the superfamily Lestoidea.</p>
<p>There is no real effort to make the scientific names of taxa unique 
among Codes, and among the relatively small set of names represented in 
the NCBI taxonomy database (20,000 genera), there are approximately 200 
duplicate names (or about 1%), mostly at the genus level.</p>
<p>Early in 2002, the first duplicate species name was recorded in the 
Taxonomy database. <italic>Agathis montana</italic> is both a wasp and 
a conifer. In this case, we have used the full species names (with 
authorities) to provide unambiguous scientific names for the sequence 
entries. (The conifer is listed as <italic>Agathis montana</italic> 
de Laub; the wasp, <italic>Agathis montana</italic> Shest).</p>
</sec>
<sec id="bid.1636">
<title>Structural Duplicates</title>
<p>In the Zoological and Bacteriological Codes, the subgenus that 
Includes the type species is required to have the same name as the genus. 
This is a systematic source of duplicate names. For these duplicates, 
we use the associated rank in the unique name, e.g., 
<italic>Drosophila</italic> &lt;genus&gt; and <italic>Drosophila</italic> 
&lt;subgenus&gt;. Duplicated genera/subgenera also occur in the Botanical 
Codes, e.g. <italic>Pinus</italic> &lt;genus&gt; and <italic>Pinus</italic> 
&lt;subgenus&gt;.</p>
</sec>
<sec id="bid.1637">
<title>Polyphyletic Genera</title>
<p>Certain genera, especially among the asexual forms of Ascomycota 
and Basidiomycota, are polyphyletic, i.e., they do not share a common 
ancestor. Pending taxonomic revisions that will transfer species assigned 
to &ldquo;form&rdquo; genera such as <italic>Cryptococcus</italic> to 
more natural genera, we have chosen to duplicate such polyphyletic genera 
in different branches of the Taxonomy database. This will maintain a 
phylogenetic classification and ensure that all species assigned to a 
polyphyletic genus can be retrieved when searching on the genus. Therefore, 
for example, the basidiomycete genus <italic>Sporobolomyces</italic> is 
represented in three different branches of the Basidiomycota: 
<italic>Sporobolomyces</italic> &lt;Sporidiobolaceae&gt;, 
<italic>Sporobolomyces</italic> &lt;Agaricostilbomycetidae&gt;, and 
<italic>Sporobolomyces</italic> &lt;Erythrobasidium clade&gt;.</p>
</sec>
<sec id="bid.1638">
<title>Other Duplicate Names</title>
<p>We list many duplicate names in other nametypes (apart from our 
preferred &ldquo;scientific name&rdquo; for each taxon). Most of these 
are included for retrieval purposes, common names or the names of 
familiar paraphyletic taxa that we have not included in our 
classification, e.g. Osteichthyes, Coelenterata, and reptiles.</p>
</sec>
</sec>
<sec id="bid.1639">
<title>Other TAXON Data Types</title>
<p>Aside from names, there are several optional types of information 
that may be associated with a taxid. These are (1) rank, such as 
species, genus or family; (2) genetic code, for translating proteins; 
(3) GenBank division; (4) literature citations; and (5) abbreviated 
lineage, for display in GenBank flat files. For more details on these 
data types, see <xref ref-type="app" rid="bid.331">Appendix 3</xref>.</p>
</sec>
</sec>
<sec id="bid.1640">
<title>Taxonomy in Entrez: A Quick Tour</title>
<p>The TAXON database is a node within the Entrez integrated retrieval 
system (<xref ref-type="other" rid="bid.588">Chapter 15</xref>) that 
provides an important organizing principle for other Entrez databases. 
Taxonomy provides an alternative view of <abbrev>TAXON</abbrev> to that 
of TaxBrowser. Entrez adds some very powerful capabilities (for example, 
<xref ref-type="other" rid="bid.1253">Boolean</xref> queries, search 
history, and both internal and external links) to <abbrev>TAXON</abbrev>, 
but in many ways it is an unnatural way to represent such hierarchical 
data in Taxonomy. (TaxBrowser is the way to view the taxonomy 
hierarchically.)</p>
<p>Taxonomy was the first Entrez database to have an internal 
hierarchical structure. Because Entrez deals with unordered sets 
of objects in a given domain, an alternative way to represent these 
hierarchical relationships in Entrez was required (see the section 
<xref rid="bid.288" ref-type="sec">Hierarchy Fields</xref>, below).</p>
<p>The main focus of the Entrez Taxonomy <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/entrez/
query.fcgi?db=Taxonomy">homepage</ext-link> is the search bar but 
also worth noting are the <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/
help/helpdoc.html">Help</ext-link> and <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/Browser/
wwwtax.cgi?mode=Root">TaxBrowser</ext-link> hotlinks that lead to 
Entrez generic help documentation and the Taxonomy browser, respectively.</p>
<p>The default Entrez search is case insensitive and can be for any of 
the names that can be found in the Taxonomy database. Thus, any of the 
following search terms, Homo sapiens, homo sapiens, human, or Man, will 
retrieve the node for <italic>Homo sapiens</italic>.</p>
<p>As for other Entrez databases, Taxonomy supports Boolean searching, 
a <named-content content-type="book-spec-emph">History</named-content> 
function, and searches limited by field. The Taxonomy fields can be 
browsed under <named-content content-type="book-spec-emph">Preview/
Index</named-content>, some are specific to Taxonomy (such as 
<named-content content-type="book-spec-emph">Lineage</named-content> 
or <named-content content-type="book-spec-emph">Rank</named-content>), 
and others are found in all Entrez databases (such as Entrez 
<named-content content-type="book-spec-emph">Date</named-content>).</p>
<p>Each search result, listed in document summary (DocSum) format, may 
have several links associated with it. For example, for the search 
result Homo sapiens, the <named-content 
content-type="book-spec-emph">Nucleotide</named-content> link will 
retrieve all the human sequences from the nucleotide databases, 
and the <named-content content-type="book-spec-emph">Genome</named-content> 
link will retrieve the human genome from the Genomes database.</p>
<sec id="bid.280">
<title>Search Tips and Tricks</title>
<p>A helpful list follows:
<list list-type="arabic">
<list-item>
<label>1</label>
<p>A search for Hominidae retrieves a single, hyperlinked entry. 
Selecting the link shows the structure of the taxon. On the other 
hand, a search for Hominidae[subtree] will retrieve a 
nonhierarchical list of all of the taxa listed within the 
Hominidae.</p>
</list-item>
<list-item>
<label>2</label>
<p>A search for species[rank] yields a list of all species in 
the Taxonomy database (108,020 in May 2002).</p>
</list-item>
<list-item>
<label>3</label>
<p>Find the Taxonomy update frequency by selecting Entrez 
<named-content content-type="book-spec-emph">Date</named-content> 
from the pull-down menu under <named-content 
content-type="book-spec-emph">Preview/Index</named-content>, 
typing &ldquo;2002/02&rdquo; in the box and selecting 
<named-content content-type="book-spec-emph">Index</named-content>. 
The result:
<preformat>
2002/01 (5176)
2002/01/03 (478)
2002/01/08 (2)
2002/01/10 (2260)
2002/01/14 (7)
2002/01/16 (239)</preformat>
shows that in January 2002, 5,176 new taxa were added, the bulk 
of which appeared in Entrez for the first time on January 10, 2002. 
These taxa can be retrieved by selecting 2002/01/10, then selecting 
the <named-content content-type="book-spec-emph">AND</named-content> 
button above the window, followed by <named-content 
content-type="book-spec-emph">Go</named-content>.</p>
</list-item>
<list-item>
<label>4</label>
<p>An overview of the distribution of taxa in the DocSum list can be 
seen if <named-content content-type="book-spec-emph">Summary</named-content> 
is changed to <named-content content-type="book-spec-emph">Common 
Tree</named-content>, followed by selecting <named-content 
content-type="book-spec-emph">Display</named-content>.</p>
</list-item>
<list-item>
<label>5</label>
<p>To filter out less interesting names from a DocSum list, add some 
terms to the query, e.g., 2002/01/10[date] NOT uncultured[prop] 
NOT unspecified[prop].</p></list-item>
</list>
</p>
</sec>
<sec id="bid.281">
<title>Displays in Taxonomy Entrez</title>
<p>There are a variety of choices regarding how search results can be 
displayed in Taxonomy Entrez.</p>
<!--TONYA, SEE IF DEBBIE AGREES THIS SHOULD BE DEF-LIST-->
<p><bold>Summary.</bold> This is the default display view. There are 
as many as four pieces of information in this display, if they are all 
present in the Taxonomy database: (1) scientific name of the taxon; 
(2) common name, if one is available; (3) taxonomic rank, if one is 
assigned; and (4) <xref ref-type="other" rid="bid.1246">BLAST</xref> 
name, inherited from the taxonomy, e.g., Homo sapiens (human), 
species, mammals.</p>
<p><bold>Brief.</bold> Shows only the scientific names of the taxa. 
This view can be used to download lists of species names from Entrez.</p>
<p><bold>Tax ID List.</bold> Shows only the taxids of the taxa. 
This view can be used to download taxid lists from Entrez.</p>
<p><bold>Info.</bold> Shows a summary of most of the information 
associated with each taxon in the Taxonomy database (similar to the 
TaxBrowser taxon-specific display; <xref ref-type="fig" 
rid="bid.257">Figure 3</xref>). This can be downloaded as a text file; 
an <xref ref-type="other" rid="bid.1435">XML</xref> representation 
of these data is under development.</p>
<p><bold>Common Tree.</bold> A special display that shows a skeleton 
view of the relationships among the selected set of taxa and is 
described in the section below.</p>
<p><bold>LinkOut.</bold> Displays a list of the linkout links (if any) 
for each of the selected sets of taxa (see <xref ref-type="other" 
rid="bid.650">Chapter 17</xref>).</p>
<p><bold>Entrez Links.</bold> The remaining views follow Entrez 
links from the selected set of taxa to the other Entrez databases 
(Nucleotide, Protein, Genome, etc.) The <named-content 
content-type="book-spec-emph">Display</named-content> view 
allows all links for a whole set of taxa to be viewed at once.</p>
</sec>
</sec>
<sec id="bid.282">
<title>The Common Tree Viewer</title>
<p>The Common Tree view shows an abbreviated view of the taxonomic 
hierarchy and is designed to highlight the relationships between a 
selected set of organisms. <xref ref-type="fig" rid="bid.283">Figure 
5</xref> shows the Common Tree view for a familiar set of model 
organisms.
<fig id="bid.283">
<label>5</label>
<caption><title>The Common Tree view for some model organisms.</title>
<p>The ten species shown in <italic>bold</italic> are the ones that 
were selected as input to the Common Tree display. The other taxa 
displayed show the taxonomic relationships between the selected taxa. 
For example, <italic>Eutheria</italic> is included because it is 
the smallest taxonomic group in our classification that includes 
both <italic>Homo sapiens</italic> and <italic>Mus musculus</italic>. 
A &ldquo;+&rdquo; box in the tree indicates that part of the 
taxonomic classification has been suppressed in this abbreviated 
view; selecting the &ldquo;+&rdquo; will fill in the missing lineage 
(and change the &ldquo;+&rdquo; to a &ldquo;-&rdquo;). The 
<named-content content-type="book-spec-emph">Expand All</named-content> 
and <named-content content-type="book-spec-emph">Collapse All</named-content> 
buttons at the <italic>top</italic> of the display will do this globally. 
The <named-content content-type="book-spec-emph">Search for</named-content> 
box at the <italic>top</italic> of the display can be used to add taxa to 
the Common Tree display; taxa can be removed using the list at the 
<italic>bottom</italic> of the page.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" 
xlink:href="ch4f5" mime-subtype="gif"/>
</fig>
</p>
<p>If there are more than few dozen taxa selected for the common 
tree view, the display becomes visually complex and generally less 
useful. When a large list of taxa is sent to the Common Tree display, 
a summary screen is displayed first. For example, we currently list 
727 families in the Viridiplantae (plants and green algae) 
(<xref ref-type="fig" rid="bid.284">Figure 6</xref>).
<fig id="bid.284">
<label>6</label>
<caption>
<title>The Common Tree summary page for the plants and green 
algae.</title>
<p>The taxa are aggregated at the predetermined set of nodes in the 
Taxonomy database that have been assigned &ldquo;BLAST names&rdquo;. 
This serves an informal, very abbreviated, vernacular classification 
that gives a convenient overview. The BLAST names will often not 
provide complete coverage for all species at all levels in the tree. 
Here, not all of our <italic>flowering plants</italic> are flagged 
as <italic>eudicots</italic> or <italic>monocots</italic>. The Common 
Tree summary display recognizes cases such as these and lists the 
remaining taxa as <italic>other flowering plants</italic>. The full 
Common Tree for some or all of the taxa can be seen by selecting the 
check box next to <italic>monocots</italic> on the summary page and 
then <named-content content-type="book-spec-emph">Choose</named-content>. 
This will display the full Common Tree view for 108 families of 
monocots.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" 
xlink:href="ch4f6" mime-subtype="gif"/>
</fig>
</p>
<p>There are several formatting options for saving the common tree 
display to a text file: text tree, phylip tree, and taxid list.</p>
<p>Hyperlinks to a common tree display can be made in two ways: 
<list list-type="arabic">
<list-item><label>(1)</label>
<p>by specifying the common tree view in an Entrez query 
<xref ref-type="other" rid="bid.1427">URL</xref> (for example, 
<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/entrez/
query.fcgi?cmd=search&amp;db=taxonomy&amp;
term=loprovbni[filter]&amp;doptcmdl=TxTree">this link</ext-link>, 
which displays the common tree view of all of the taxonomy nodes 
with LinkOut links to the Butterfly Net International Web site); or</p>
</list-item>
<list-item><label>(2)</label>
<p>by providing a list of taxids directly to the common tree 
cgi function (for example, <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/CommonTree/
wwwcmt.cgi?id=7227&amp;id=6239&amp;id=9606&amp;id=31033&amp;
id=7955&amp;id=10090&amp;id=4932&amp;id=4896&amp;id=4577&amp;
id=3702">this link</ext-link>, which will display a live version 
of <xref ref-type="fig" rid="bid.283">Figure 5</xref>).</p>
</list-item>
</list>
</p>
<sec id="bid.285">
<title>Using Batch Taxonomy Entrez</title>
<p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" 
ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/
batchentrez.cgi?db=Taxonomy">Batch Entrez</ext-link> page allows you 
to upload a file of taxids or taxon names into Taxonomy Entrez.</p>
</sec>
</sec>
<sec id="bid.286">
<title>Indexing Taxonomy in Entrez</title>
<p>As for any Entrez database, the contents are indexed by creating 
term lists for each field of each database record (or taxid). For 
<abbrev>TAXON</abbrev>, the types of fields include name fields, 
hierarchy fields, inherited fields, and generic Entrez fields.</p>
<sec id="bid.287">
<title>Name Fields</title>
<p>There are five different index fields for names in Taxonomy Entrez.</p>
<!--TONYA, SEE IF DEBBIE AGREES THIS IS A DEF-LIST-->
<p><bold>All names</bold>, [name] in an Entrez search &ndash; this 
is the default search field in Taxonomy Entrez. This is different 
from most Entrez databases, where the default search field is the
composite [All Fields].</p>
<p><bold>Scientific name</bold>, [sname] &ndash; using [sname] as a 
qualifier in a search restricts it to the nametype &ldquo;scientific 
name&rdquo;, the single preferred name for each taxon.</p>
<p><bold>Common name</bold>, [cname] &ndash; restricts the search 
to common names.</p>
<p><bold>Synonym</bold>, [synonym] &ndash; restricts the search to 
the &ldquo;synonym&rdquo; nametype.</p>
<p><bold>Taxid</bold>, [uid] &ndash; restricts the search to taxonomy 
IDs, the unique numerical identifiers for taxa in the database. Taxids 
are not indexed in the other Entrez name fields.</p>
</sec>
<sec id="bid.288">
<title>Hierarchy Fields</title>
<boxed-text id="bid.289" position="margin">
<label>3</label>
<title>Examples of combining the subtree and lineage field limits with 
Boolean operators for searching Taxonomy Entrez.</title>
<!--TONYA, TALK TO DEBBIE ABOUT THIS.  I'M NOT SURE WHETHER
THIS SHOULD BE SOME TYPE OF LIST OR NOT-->
<p><bold>(1) Mammalia[subtree] AND Mammalia[lineage]</bold> returns 
the taxon Mammalia.</p>
<p><bold>(2) Mammalia[subtree] OR Mammalia[lineage]</bold> returns 
all of the taxa in a direct parent&ndash;child relationship with the 
taxon Mammalia.</p>
<p><bold>(3) root[subtree] NOT (Mammalia[subtree] OR Mammalia[lineage])</bold> 
returns all of the taxa not in a direct parent&ndash;child relationship 
with the taxon Mammalia.</p>
<p><bold>(4) Sauropsida[subtree] NOT Aves[subtree]</bold> will 
retrieve the members of the classical taxon Reptilia, excluding the 
birds.</p>
</boxed-text>
<p>The [lineage] and [subtree] index fields are a way to superimpose 
the hierarchical relationships represented in the taxonomy on top of 
the Entrez data model. For an example of how to use these field 
limits for searching Taxonomy Entrez, see <xref ref-type="boxed-text" 
rid="bid.289">Box 3</xref>.</p>
<!--TONYA, TALK TO DEBBIE ABOUT WHETHER THIS IS A DEF-LIST-->
<p><bold>Lineage</bold>. For each node, the [lineage] index field 
retrieves all of the taxa listed at or above that node in the taxonomy. 
For example, the query Mammalia[lineage] retrieves 18 taxa from Entrez.</p>
<p><bold>Subtree</bold>. For each node, the [subtree] index field 
retrieves all of the taxa listed at or below that node in the taxonomy. 
For example, the query Mammalia[lineage] retrieves 4,021 taxa from 
Entrez (as of March 9, 2002).</p>
<p><bold>Next level</bold>. Returns all of the direct children of a 
given taxon.</p>
<p><bold>Rank</bold>. Returns all of the taxa of a given Linnaean rank. 
The query Aves[subtree] AND species[rank] retrieves all of the species 
of birds with public sequence entries (there are 2,459, approximately 
half of the currently described species of extant birds).</p>
</sec>
<sec id="bid.290">
<title>Inherited Fields</title>
<p>The genetic code [gc], mitochondrial genetic code [mgc], and GenBank 
division [division] fields are all inherited within the taxonomy. The 
information in these fields refers to the genetic code used by a taxon 
or in which GenBank division it resides. Because whole families or branches 
may use the same code or reside in the same GenBank division, this property 
is usually indexed with a taxon high in the taxonomic tree, and the 
information is inherited by all those taxa below it. If there is no [gc] 
field associated with a taxon in the database, it is assumed that the 
standard genetic code is used. A genetic code may be referred to by either 
name or translation table number. For example, the two equivalent queries, 
standard[gc] and translation table 1[gc], each retrieves the set of 
organisms that use the standard genetic code for translating genomic 
sequences. Likewise, these two queries echinoderm mitochondrial[mgc] and 
translation table 9[mgc] will each retrieve the set of organisms that 
use the echinoderm mitochondrial genetic code for translating their 
mitochondrial sequences.</p>
</sec>
<sec id="bid.291">
<title>Generic Entrez Fields</title>
<boxed-text id="bid.292" position="margin">
<label>4</label>
<title>The properties [prop] field of Taxonomy in Entrez.</title>
<p>There are several useful terms and phrases indexed in the [prop] 
field. Possible search strategies that specify the prop field are 
explained below.</p>
<!--TONYA, ASK IF THIS SHOULD BE A LIST. "LABELED" PARA'S
WOULD SEEM TO INDICATE THESE ARE ITEMS, AND PRESUMABLY THE 
ADDITIONAL PARA'S ARE JUST PART OF THE LIST-ITEM.-->
<p><bold>(1) Using functional nametypes and classifications</bold></p>
<p>unspecified [prop] not identified at the species level</p>
<p>uncultured [prop] environmental sample sequences</p>
<p>unclassified [prop] listed in an &ldquo;unclassified&rdquo; bin</p>
<p>incertae sedis [prop] listed in an &ldquo;incertae sedis&rdquo; bin</p>
<p>We do not explicitly flag names as &ldquo;unspecified&rdquo; in 
<abbrev>TAXON</abbrev>; rather, we rely on heuristics to index names as 
&ldquo;unspecified&rdquo; in the properties field. Many are missed. 
Taxa are indexed as &ldquo;uncultured&rdquo; if they are listed within 
an environmental samples bin or if their scientific names begin with 
the word &ldquo;uncultured&rdquo;.</p>
<p><bold>(2) Using rank level of taxon</bold></p>
<p>All of these search strategies below are valid. Taxonomy 
Eentrez displays only taxa that are linked to public sequence entries, 
and because sequence entries are supposed to correspond to the Taxonomy 
database at or below the species level, the Entrez query: terminal 
[prop] NOT &ldquo;at or below species level&rdquo; [prop] should only 
retrieve problem cases.</p>
<p>above genus level [prop]</p>
<p>above species level [prop]</p>
<p>&ldquo;at or below species level&rdquo; [prop] (needs explicit 
quotes)</p>
<p>below species level [prop]</p>
<p>terminal [prop]</p>
<p>non terminal [prop]</p>
<p>
<bold>(3) Inherited value assignment points</bold></p>
<p>genetic code [prop]</p>
<p>mitochondrial genetic code [prop]</p>
<p>standard [prop] invertebrate mitochondrial [prop]</p>
<p>translation table 5 [prop]</p>
<p>The query &ldquo;genetic code [prop]&rdquo; retrieves all of 
the taxa at which one of the genomic genetic codes is explicitly set. 
The second query retrieves all of the taxa at which one of the 
mitochondrial genetic codes is explicitly set, and so on.</p>
<p>division [prop]</p>
<p>INV [prop]</p>
<p>invertebrates [prop]</p>
<p>The above terms index the assignments of the GenBank division 
codes, which are divided along crude taxonomic categories (see 
<xref ref-type="other" rid="bid.2">Chapter 1</xref>). We have placed 
the division flags in the database so as to preserve the original 
assignment of species to GenBank divisions.</p>
</boxed-text>
<p>The remaining index fields are common to most or all Entrez 
domains, although some have special features in the taxonomy domain. 
For example, the field text word, [word], indexes words from the 
Taxonomy Entrez name indexes. Most punctuation is ignored, and the 
index is searched one word at a time; therefore, the search &ldquo;homo 
sapiens[word]&rdquo; will retrieve nothing.</p>
<p>Several useful terms are indexed in the properties field, [prop], 
including functional nametypes and classifications, the rank level of 
a taxon, and inherited values. See <xref ref-type="boxed-text" 
rid="bid.292">Box 4</xref> for a detailed discussion of searches 
using the [prop] field.</p>
<p>More information on using the generic Entrez fields can be found 
in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" 
ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/
query/static/help/helpdoc.html">Entrez Help</ext-link> documents.</p>
</sec>
<sec id="bid.293">
<title>Taxonomy Fields in Other Entrez Databases</title>
<p>Many of the Entrez databases (Nucleotide, Protein, Genome, etc.) 
include an <named-content content-type="book-spec-emph">Organism</named-content> 
field, [orgn], that indexes entries in that database by taxonomic group. 
All of the names associated with a taxon (scientific name, synonyms, 
common names, and so on) are indexed in the <named-content 
content-type="book-spec-emph">Organism</named-content> field and will 
retrieve the same set of entries. The <named-content 
content-type="book-spec-emph">Organism</named-content> field will 
retrieve all of the entries below the term and any of their children.</p>
<p>To not retrieve such &ldquo;exploded&rdquo; terms, the unexploded 
indexes should be used. This query will only retrieve the entries that 
are linked directly to <italic>Homo sapiens</italic>: Homo sapiens[orgn:noexp]. 
This query will not retrieve entries that are linked to the subordinate 
node <italic>Homo sapiens neanderthalensis</italic>.</p>
<p>Taxids are indexed with the prefix txid: txid9606 [orgn].</p>
<p>Source organism modifiers are indexed in the [properties] field, 
and such queries would be in the form: src strain[prop], src variety[prop], 
or src specimen voucher[prop]. These queries will retrieve all entries with 
a strain qualifier, a variety qualifier, or a specimen_voucher qualifier, 
respectively.</p>
<p>All of the organism source feature modifiers (/clone, /serovar, /variety, 
etc.) are indexed in the text word field, [text word]. For example, one 
could query GenBank for: &ldquo;strain k-12&rdquo; [text word]. Because 
strain information is inconsistent in the sequence databases (as in the 
literature), a better query would be: &ldquo;strain k 12&rdquo;[word] OR 
&ldquo;strain k12&rdquo;[word]. Note: explicit double-quotes may be necessary 
for some of these queries.</p>
</sec>
</sec>
<sec id="bid.294">
<title>The Taxonomy Statistics Page</title>
<p>The Taxonomy <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" 
ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/
txstat.cgi">Statistics page</ext-link> displays tables of counts of the 
number of taxa in the public subset of the Taxonomy database. The numbers 
displayed are hyperlinks that will retrieve the corresponding set of entries. 
The table can be configured to display data based on three criteria: Entrez 
release date, rank, and taxa. The default setting shows the counts by rank 
for a pre-selected set of taxa (across all dates).</p>
<p>The checkboxes <named-content 
content-type="book-spec-emph">unclassified</named-content>, 
<named-content content-type="book-spec-emph">uncultured</named-content>, 
and <named-content content-type="book-spec-emph">unspecified</named-content> 
will exclude the corresponding sets of taxa from the count. These work by 
appending the terms &ldquo;NOT unclassified[prop]&rdquo;, etc., to the 
statistics query. Checking 
<named-content content-type="book-spec-emph">uncultured</named-content> 
and <named-content content-type="book-spec-emph">unspecified</named-content> 
removes about 20% of taxa in the database and gives a much better count of 
the number of formally described species. As of April 1, 2003, the count 
was as follows:
<preformat>
Archaea: 82 genera, 364 species
Bacteria: 1163 genera, 9927 species
Eukaryota: 24939 genera, 74832 species</preformat>
</p>
<p>Selecting one of the <named-content 
content-type="book-spec-emph">rank</named-content> categories (e.g., 
species) loads a new table that shows, in this example, the number of 
new species added to each taxon each year, starting with 1993. The 
<named-content content-type="book-spec-emph">Interval</named-content> 
pull-down menu shows release statistics in finer detail. The list of 
taxa in the display can be customized through the <named-content 
content-type="book-spec-emph">Customize</named-content> link.</p>
</sec>
<sec id="bid.295">
<title>Other Relevant References</title>
<sec id="bid.296">
<title>Taxonomy FTP</title>
<p>A complete copy of the public NCBI taxonomy database is deposited 
several times a day on our <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" 
xlink:href="ftp://ftp.ncbi.nih.gov/pub/taxonomy/">FTP site</ext-link>. 
See <xref ref-type="table" rid="bid.269">Table 1</xref> for details.</p>
</sec>
<sec id="bid.297">
<title>Tax BLAST</title>
<p>Taxonomy BLAST reports (<abbrev>Tax BLAST</abbrev>) are available 
from the <abbrev>BLAST</abbrev> results page and from the 
<xref ref-type="other" rid="bid.1250">BLink</xref> pages. <abbrev>Tax 
BLAST</abbrev> post-processes the <abbrev>BLAST</abbrev> output results 
according to the source organisms of the sequences in the 
<abbrev>BLAST</abbrev> results page. A 
<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/blast/taxblasthelp.html">help 
page</ext-link> is available that describes the three different views 
presented on the <abbrev>Tax BLAST</abbrev> page (Lineage Report, 
Organism Report, and Taxonomy Report).</p>
</sec>
<sec id="bid.298">
<title>Toolkit Function Libraries</title>
<p>The function library for the taxonomy application software in the 
NCBI Toolkit is ncbitxc2.a (or libncbitxc2.a). The source code can be 
found in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" 
ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/IEB/
ToolBox/SB/hbr.html">NCBI Toolkit Source Browser</ext-link> and can 
be downloaded from the toolbox directory on the <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" 
xlink:href="ftp://ncbi.nlm.nih.gov/toolbox/ncbi_tools/">FTP site</ext-link>.
</p>
</sec>
</sec>
<sec id="bid.299">
<title>NCBI Taxonomists</title>
<p>In the early years of the project, Scott Federhen did all of 
the software and database development. In recent years, Vladimir 
Soussov and his group have been responsible for software and database 
development.</p>
<list list-type="simple">
<list-item><p>Scott Federhen (1990&ndash;present)</p></list-item>
<list-item><p>Andrzej Elzanowski (1994&ndash;1997)</p></list-item>
<list-item><p>Detlef Leipe (1994&ndash;present)</p></list-item>
<list-item><p>Mark Hershkovitz (1996&ndash;1997)</p></list-item>
<list-item><p>Carol Hotton (1997&ndash;present)</p></list-item>
<list-item><p>Mimi Harrington (1999&ndash;2000)</p></list-item>
<list-item><p>Ian Harrison (1999&ndash;2002)</p></list-item>
<list-item><p>Sean Turner (2000&ndash;present)</p></list-item>
<list-item><p>Rick Sternberg (2001&ndash;present)</p></list-item>
</list>
</sec>
<sec id="bid.300">
<title>Contact Us</title>
<p>If you have a comment or correction to our Taxonomy database, 
perhaps a misspelling or classification or if something looks wrong, 
please send a message to <email>info@ncbi.nlm.nih.gov</email>.</p>
</sec>
</body>
<back>
<app-group>
<app id="bid.301">
<title>Appendix 1. <abbrev>TAXON</abbrev> nametypes.</title>
<sec id="bid.302">
<title>Scientific Name</title>
<p>Every node in the database is required to have exactly one 
&ldquo;scientific name&rdquo;. Wherever possible, this is a 
validly published name with respect to the relevant code of 
nomenclature. Formal names that are subject to a code of 
nomenclature and are associated with a validly published 
description of the taxon will be Latinized uninomials above 
the species level, binomials (e.g. <italic>Homo sapiens</italic>) 
at the species level, and trinomials for the formally described 
infraspecific categories (e.g., <italic>Homo sapiens 
neanderthalensis</italic>). For many of our taxa, it is not 
possible to find an appropriate formal scientific name; these 
nodes are given an informal &ldquo;scientific name&rdquo;. 
The different classes of informal names are discussed in 
<xref ref-type="app" rid="bid.317">Appendix 2</xref>. Functional 
Classes of <abbrev>TAXON</abbrev> Scientific Names.</p>
<p>The scientific name is the one that will be used in all of 
the sequence entries that map to this node in the Taxonomy database. 
Entries that are submitted with any of the other names associated 
with this node will be replaced with this name. When we change 
the scientific name of a node in the Taxonomy database, the 
corresponding entries in the sequence databases will be updated 
to reflect the change. For example, we list <italic>Homo 
neanderthalensis</italic> as a synonym for <italic>Homo sapiens 
neanderthalensis</italic>. Both are in common use in the 
literature. We try to impose consistent usage on the entries 
in the sequence databases, and resolving the nomenclatural 
disputes that inevitably arise between submitters is one of the 
most difficult challenges that we face.</p>
</sec>
<sec id="bid.303">
<title>Synonym</title>
<p>The &ldquo;synonym&rdquo; nametype is applied to both synonyms 
in the formal nomenclatural sense (objective, nomenclatural, 
homotypic <italic>versus</italic> subjective, heterotypic) and 
more loosely to include orthographic variants and a host of names 
that have found their way into the taxonomy database over the years, 
because they were found in sequence entries and later merged into 
the same taxon in the Taxonomy database.</p>
</sec>
<sec id="bid.304">
<title>Acronym</title>
<p>The &ldquo;acronym&rdquo; nametype is used primarily for the 
viruses. The International Committee on Taxonomy of Viruses 
(<abbrev>ICTV</abbrev>) maintains an official list of acronyms for 
viral species, but the literature is often full of common variants, 
and it is convenient to list these as well. For example, we list 
HIV, LAV-1, HIV1, and HIV-1 as acronyms for the human immunodeficiency 
virus type 1.</p>
</sec>
<sec id="bid.305">
<title>Anamorph</title>
<p>The term &ldquo;anamorph&rdquo; is reserved for names applied 
to asexual forms of fungi, which present some special nomenclatural 
challenges. Many fungi are known to undergo both sexual and asexual 
reproduction at different points in their life cycle (so-called 
&ldquo;perfect&rdquo; fungi); for many others, however, only the 
asexually reproducing (anamorphic or mitosporic) form is known (in 
some, perhaps many, asexual species, the sexual cycle may have been 
lost altogether). These anamorphs, often with simple and not especially 
diagnostic morphology, were given Linnaean binomial names. A number of 
named anamorphic species have subsequently been found to be associated 
with sexual forms (teleomorphs) with a different name (for example, 
<italic>Aspergillus nidulans</italic> is the name given to the asexual 
stage of the teleomorphic species <italic>Emericella nidulans</italic>). 
In these cases, the teleomorphic name is given precedence in the 
GenBank Taxonomy database as the &ldquo;scientific&rdquo; name, and 
the anamorphic name is listed as an &ldquo;anamorph&rdquo; nametype.</p>
</sec>
<sec id="bid.306">
<title>Misspelling</title>
<p>The &ldquo;misspelling&rdquo; nametype is for simple misspellings. 
Some of these are included because the misspelling is present in the 
literature, but most of them are there because they were once found in 
a sequence entry (which has since been corrected). We keep them in the 
database for tracking purposes, because copies of the original sequence 
entry can still be retrieved. Misspellings are not listed on the TaxBrowser 
pages nor on the Taxonomy Entrez Info display views, but they are indexed 
in the Entrez search fields (so that searches and Entrez queries with 
the misspelling will find the appropriate node).</p>
</sec>
<sec id="bid.307">
<title>Misnomer</title>
<p>&ldquo;Misnomer&rdquo; is a rarely used nametype. It is used for 
names that might otherwise be listed as &ldquo;misspellings&rdquo; but 
which we want to appear on the browser and Entrez display pages.</p>
</sec>
<sec id="bid.308">
<title>Common Name</title>
<p>The &ldquo;common name&rdquo; nametype is used for vernacular 
names associated with a particular taxon. These may be found at any 
level in the hierarchy; for example, &ldquo;human&rdquo;, 
&ldquo;reptiles&rdquo;, and &ldquo;pale devil's-claw&rdquo; are all 
used. Common names should be in lowercase letters, except where part 
of the name is derived from a proper noun, for example, &ldquo;American 
butterfish&rdquo; and &ldquo;Robert's arboreal rice rat&rdquo;.</p>
<p>The use of common names is inherently variable, regional, and often 
inconsistent. There is generally no authoritative reference that 
regulates the use of common names, and there is often not perfect 
correspondence between common names and formally described scientific 
taxa; therefore, there are some caveats to their use. For scientific 
discourse, there is no substitute for formal scientific names. 
Nevertheless, common names are invaluable for many indexing, retrieval, 
and display purposes. The combination &ldquo;<italic>Oecomys roberti</italic> 
(Robert's arboreal rice rat)&rdquo; conveys much more information 
than either name by itself. Issues raised by the variable, regional, 
and inexact use of common names are partly addressed by the &ldquo;genbank 
common name&rdquo; nametype (below) and the ability to customize names 
in the GenBank flatfile.</p>
</sec>
<sec id="bid.309">
<title><abbrev>BLAST</abbrev> Name</title>
<p>The &ldquo;<abbrev>BLAST</abbrev> names&rdquo; are a specially 
designated set of common names selected from the Taxonomy database. 
These were chosen to provide a pool of familiar names for large 
groups of organisms (such as &ldquo;insects&rdquo;, &ldquo;mammals&rdquo;, 
&ldquo;fungi&rdquo;, and others) so that any particular species (which 
may not have an informative common name of its own) could inherit a 
meaningful collective common name from the Taxonomy database. This was 
originally developed for <abbrev>BLAST</abbrev>, because a list of 
<abbrev>BLAST</abbrev> results will typically include entries from many 
species identified by Latin binomials, which may not be familiar to 
all users. <abbrev>BLAST</abbrev> names may be nested; for example, 
&ldquo;eukaryotes&rdquo;, &ldquo;animals&rdquo;, &ldquo;chordates&rdquo;, 
&ldquo;mammals&rdquo;, and &ldquo;primates&rdquo; are all flagged as 
&ldquo;blast names&rdquo;.</p>
<p><abbrev>BLAST</abbrev> names are now used in several other applications, 
for example the <xref ref-type="other" rid="bid.1415">Tax BLAST</xref> 
displays, the Summary view in Entrez Taxonomy, and in the Summary 
display of the Common Tree format.</p>
</sec>
<sec id="bid.310">
<title>In-part</title>
<p>The &ldquo;in-part&rdquo; nametype is included for retrieval 
terms that have a broader range of application than the taxon or 
taxa at which they appear. For example, we list reptiles and Reptilia 
as in-part nametypes at our nodes <italic>Testudines</italic> (the 
turtles), <italic>Lepidosauria</italic> (the lizards and snakes), and 
<italic>Crocodylidae</italic> (the crocodilians).</p>
</sec>
<sec id="bid.311">
<title>Includes</title>
<p>The &ldquo;includes&rdquo; nametype is the opposite of the in-part 
nametype and is included for retrieval terms that have a narrower 
scope of application than the taxon at which they appear. For example, 
we could list &ldquo;reptiles&rdquo; as an &ldquo;includes&rdquo; 
nametype for the <italic>Amniota</italic> (or at any higher node in the 
lineage).</p>
</sec>
<sec id="bid.312">
<title>Equivalent Name</title>
<p>The &ldquo;equivalent name&rdquo; nametype is a catch-all category, 
used for names that we would like to associate with a particular node 
in the database (for indexing or tracking purposes) but which do not 
seem to fit well into any of the other existing nametypes.</p>
</sec>
<sec id="bid.313">
<title>GenBank Common Name</title>
<p>The &ldquo;genbank common name&rdquo; was introduced to provide a 
mechanism by which, when there is more than one common name associated 
with a particular node in the taxonomy, one of them could be designated 
to be the common name that should be used by default in the GenBank flatfiles 
and other applications that are trying to find a common name to use for 
display (or other) purposes. This is not intended to confer any special 
status or blessing on this particular common name over any of the other 
common names that might be associated with the same node, and we have 
developed mechanisms to override this choice for a common name on a 
case-by-case basis if another name is more appropriate or desirable for 
a particular sequence entry. Each node may have at most one &ldquo;genbank 
common name&rdquo;.</p>
</sec>
<sec id="bid.314">
<title>GenBank Acronym</title>
<p>There may be more than one acronym associated with a particular node 
in the Taxonomy database (particularly if several virus names have been 
synonymized in a single species). Just as with the &ldquo;genbank common 
name&rdquo;, the &ldquo;genbank acronym&rdquo; provides a mechanism to 
designate one of them to be the acronym that should be used for display 
(or other) purposes. Each node may have at most one &ldquo;genbank 
acronym&rdquo;.</p>
</sec>
<sec id="bid.315">
<title>GenBank Synonym</title>
<p>The &ldquo;genbank synonym&rdquo; nametype is intended for those 
special cases in which there is more than one name commonly used in 
the literature for a particular species, and it is informative to have 
both names displayed prominently in the corresponding sequence record. 
Each node may have at most one &ldquo;genbank synonym&rdquo;. For example,
<preformat>
SOURCE Takifugu rubripes (Fugu rubripes)
ORGANISM Takifugu rubripes
</preformat>
</p>
</sec>
<sec id="bid.316">
<title>GenBank Anamorph</title>
<p>Although the use of either the anamorph or teleomorph name is 
formally correct under the International Code of Botanical Nomenclature, 
we prefer to give precedence to the telemorphic name as the 
&ldquo;scientific name&rdquo; in the Taxonomy database, both to 
emphasize their commonality and to avoid having two (or more) taxids 
that effectively apply to the same organism. However, in many cases, 
the anamorphic name is much more commonly used in the literature, 
especially when sequences are normally derived from the asexual form 
of the species. In these cases, the &ldquo;genbank anamorph&rdquo; 
nametype can be used to annotate the corresponding sequence records 
with both names. Each node may have at most one &ldquo;genbank 
anamorph&rdquo;. For example:
<preformat>
SOURCE Emericella nidulans (anamorph: Aspergillus nidulans)
ORGANISM Emericella nidulans
</preformat>
</p>
</sec>
</app>
<app id="bid.317">
<title>Appendix 2. Functional classes of <abbrev>TAXON</abbrev> 
scientific names.</title>
<sec id="bid.318">
<title>Formal Names</title>
<p>Whenever possible, formal scientific names are used for taxa. 
There are several codes of nomenclature that regulate the description 
and use of names in different branches of the tree of life. These are: 
the International Code of Zoological Nomenclature (<abbrev>ICZN</abbrev>), 
the International Code of Botanical Nomenclature (<abbrev>ICBN</abbrev>), 
the International Code of Nomenclature for Cultivated Plants 
(<abbrev>ICNCP</abbrev>), the International Code of Nomenclature of 
Bacteria (<abbrev>ICNB</abbrev>), and the International Code of Virus 
Classification and Nomenclature (<abbrev>ICVCN</abbrev>).</p>
<p>The viral code is less well developed than the others, but it includes 
an official classification for the viruses as well as a list of approved 
species names. Viral names are not Latin binomials (as required by the 
other codes), although there are some instances (e.g., <italic>Herpesvirus 
papio</italic> or <italic>Herpesvirus sylvilagus</italic>). When possible, 
we try to use <abbrev>ICTV</abbrev>-approved names for viral taxa, but 
new viral species names appear in the literature (and therefore in the 
sequence databases) much faster than they are approved into the 
<abbrev>ICTV</abbrev> lists. We are working to set up taxonomy LinkOut 
links (see <xref ref-type="other" rid="bid.650">Chapter 17</xref>) 
to the <abbrev>ICTV</abbrev> database, which will make the subset of 
<abbrev>ICTV</abbrev>-approved names explicit.</p>
<p>The zoological, botanical, and bacteriological codes mandate Latin 
binomials for species names. They do not describe an official 
classification (such as the <abbrev>ICTV</abbrev>), with the exception 
that the binomial species nomenclature itself makes the classification 
to the genus level explicit. If a genus is found to be polyphyletic, 
the classification cannot be corrected without formally renaming at least 
some of the species in the genus. (This is somewhat reminiscent of the 
&ldquo;smart identifier&rdquo; problem in computer science.)</p>
<p>The fungi are subject to the botanical code. The cyanobacteria 
(blue-green algae) have been subject to both the botanical and the 
bacteriological codes, and the issue is still controversial.</p>
<sec id="bid.1641">
<title>Authorities</title>
<p>&ldquo;Authorities&rdquo; appear at the end of the formal species 
name and include at least the name or standard abbreviation of the 
taxonomist who first described that name in the scientific literature. 
Other information may appear in the authority as well, often the year 
of description, and can become quite complicated if the taxon has been 
transferred or amended by other taxonomists over the years. We do not 
use authorities in our taxon names, although many are included in the 
database listed as synonyms. We have made an exception to this rule in 
the case of our first duplicated species name in the database, 
<italic>Agathis montana</italic>, to provide unambiguous names for the 
corresponding sequence entries.</p>
</sec>
<sec id="bid.1642">
<title>Subspecies</title>
<p>All three of the codes of nomenclature for cellular organisms 
provide for names at the subspecies level. The botanical and 
bacteriological codes include the string &ldquo;subsp.&rdquo; in the 
formal name; the zoological code does not, e.g., <italic>Homo sapiens 
neanderthalensis</italic>, <italic>Zea mays subsp. mays</italic>, and 
<italic>Klebsiella pneumoniae subsp. ozaenae</italic>.</p>
</sec>
<sec id="bid.321">
<title>Varietas and Forma</title>
<p>The botanical code (but none of the others) provides for two 
additional formal ranks beneath the subspecies level, varietas and 
forma. These names will include the strings &ldquo;var.&rdquo; and 
&ldquo;f.&rdquo;, respectively, e.g., <italic>Marchantia paleacea 
var. diptera, Penicillium aurantiogriseum var. neoechinulatum, Salix 
babylonica f. rokkaku</italic>, or <italic>Fragaria vesca subsp. 
Vesca f. alba</italic></p>
</sec>
</sec>
<sec id="bid.1643">
<title>Other Subspecific Names</title>
<p>We list taxa with other subspecific names where it seems useful 
and appropriate and where it is necessary to find places for names 
in the sequence databases. For indexing purposes in the Genomes 
division of Entrez, it is convenient to have strain-level nodes for 
bacterial species with a complete genome sequence, particularly when 
there are two or more complete genome sequences available for 
different strains of the same species, e.g., <italic>Escherichia 
coli</italic> K12, <italic>Escherichia coli</italic> O157:H7, 
<italic>Escherichia coli</italic> O157:H7 EDL933, <italic>Mycobacterium 
tuberculosis</italic> CDC1551, and <italic>Mycobacterium 
tuberculosis</italic> H37Rv.</p>
<p>Several other classes of subspecific groups do not have formal 
standing in the nomenclature but represent well-characterized and 
biologically meaningful groups, e.g., serovar, pathovar, forma 
specialis, and others. In many cases, these may eventually be promoted 
to a species; therefore, it is convenient to represent them 
independently from the outset, e.g., <italic>Xanthomonas campestris</italic> 
pv. <italic>campestris, Xanthomonas campestris</italic> pv. 
<italic>vesicatoria</italic>, <italic>Pneumocystis carinii</italic> f. sp. 
<italic>hominis, Pneumocystis carinii</italic> f. sp. <italic>mustelae, 
Salmonella enterica</italic> subsp. <italic>enterica</italic> serovar 
Dublin, and <italic>Salmonella enterica</italic> subsp. 
<italic>enterica</italic> serovar Panama.</p>
<p>Many other names below the species level have been added to the 
Taxonomy database to accommodate SWISS-PROT entries, where strain 
(and other) information is annotated with the organism name for some 
species.</p>
</sec>
<sec id="bid.323">
<title>Informal Names</title>
<p>In general, we try to avoid unqualified species names such as 
<italic>Bacillus</italic> sp., although many of them exist in the 
Taxonomy database because of earlier sequence entries. 
<italic>Bacillus</italic> sp. is a particularly egregious example, 
because <italic>Bacillus</italic> is a duplicated genus name and 
could refer to either a bacterium or an insect. In our database, 
<italic>Bacillus</italic> sp. is assumed to be a bacterium, but 
<italic>Bacillus</italic> sp. P-4-N, on the other hand, is 
classified with the insects.</p>
<p>When entries are not identified at the species level, multiple 
sequences can be from the same unidentified species. Sequences 
from multiple different unidentified species in the same genus 
are also possible. To keep track of this, we add unique informal 
names to the Taxonomy database, e.g., a meaningful identifier 
from the submitters could be used. This could be a strain name, 
a culture collection accession, a voucher specimen, an isolate 
name or location&mdash;anything that could tie the entry to the 
literature (or even to the lab notebook). If nothing else is 
available, we may construct a unique name using a default formula 
such as the submitter's initials and year of submission. This way, 
if a formal name is ever determined or described for any of these 
organisms, we can synonymize the informal name with the formal 
one in the Taxonomy database, and the corresponding entries in 
the sequence databases will be updated automatically. For example, 
AJ302786 was originally submitted (in November 2000) as 
<italic>Agathis</italic> sp. and was added to the Taxonomy database 
as <italic>Agathis</italic> sp. RDB-2000. In January 2002, this wasp 
was identified as belonging to the species <italic>Agathis montana</italic>, 
and the node was renamed; the informal name <italic>Agathis</italic> 
sp. RDB-2000 was listed as a synonym. A separate member of the genus, 
<italic>Agathis sp.</italic> DMA-1998, is still listed with an informal 
name.</p>
<p>Here are some examples of informal names in the Taxonomy database:
<preformat>
Anabaena sp. PCC 7108
Anabaena sp. M14-2
Calophyllum sp. 'Fay et al. 1997'
Scutellospora sp. Rav1/RBv2/RCv3
Ehrlichia-like sp. 'Schotti variant'
Gilia sp. Porter and Heil 7991
Camponotus n. sp. BGW-2001
Camponotus sp. nr. gasseri BGW-2001
Drosophila sp. 'white tip scutellum'
Chrysoperla sp. 'C.c.2 slow motorboat'
Saranthe aff. eichleri Chase 3915
Agabus cf. nitidus IR-2001
Simulium damnosum s.l. 'Kagera'
Amoebophrya sp. ex Karlodinium micrum
</preformat>
</p>
<p>We use single quotes when it seems appropriate to group a phrase 
into a single lexical unit. Some of these names include abbreviations 
with special meanings.</p>
<p>&ldquo;n. sp.&rdquo; indicates that this is a new, undescribed 
species and not simply an unidentified species. &ldquo;sp. nr.&rdquo; 
indicates &ldquo;species near&rdquo;. In the example above, this 
indicates that this is similar to Camponotus gasseri. &ldquo;aff.&rdquo;, 
affinis, related to but not identical to the species given. 
&ldquo;cf.&rdquo;, confer; literally, &ldquo;compare with&rdquo; 
conveys resemblance to a given species but is not necessarily related 
to it. &ldquo;s.l.&rdquo;, sensu lato; literally, &ldquo;in the broad 
sense&rdquo;. &ldquo;ex&rdquo;, &ldquo;from&rdquo; or &ldquo;out of&rdquo; 
the biological host of the specimen.</p>
<p>Note that names with cf., aff., nr., and n. sp. are not unique and 
should have unique identifiers appended to the name.</p>
<p>Cultured bacterial strains and other specimens that have not been 
identified to the genus level are given informal names as well, e.g., 
Desulfurococcaceae str. SRI-465; crenarchaeote OlA-6.</p>
<p>Names such as Camponotus sp. 1 are avoided, because different 
submitters might easily use the same name to refer to different species. 
See <xref ref-type="boxed-text" rid="bid.292">Box 4</xref> for how to 
retrieve these names in Taxonomy Entrez.</p>
</sec>
<sec id="bid.1644">
<title>Uncultured Names</title>
<p>Sequences from environmental samples are given &ldquo;uncultured&rdquo; 
names. In these studies, nucleotide sequences are cloned directly from 
the environment and come from varied sources, such as Antarctic sea ice, 
activated sewer sludge, and dental plaque. Apart from the sequence itself, 
there is no way to identify the source organisms or to recover them for 
further studies. These studies are particularly important in bacterial 
systematics work, which shows that the vast majority of environmental 
bacteria are not closely related to laboratory cultured strains (as 
measured by 16S rRNA sequences). Many of the deepest-branching groups 
in our bacterial classification are defined only by anonymous sequences 
from these environmental samples studies, e.g., candidate division OP5, 
candidate division Termite group 1, candidate subdivision kps59rc, 
phosphorous removal reactor sludge group, and marine archael group 1.</p>
<p>These samples vary widely in length and in quality, from short 
single-read sequences of a few hundred base pairs to high-quality, 
full-length 16S sequences. We now give all of these samples anonymous 
names, which may indicate the phylogenetic affiliation of the sequence, 
as far is it may be determined, e.g., uncultured archaeon, uncultured 
crenarchaeote, uncultured gamma proteobacerium, or uncultured enterobacterium. 
See <xref ref-type="boxed-text" rid="bid.292">Box 4</xref> for how to retrieve 
these names in Taxonomy Entrez.</p>
</sec>
<sec id="bid.325">
<title>Candidatus Names</title>
<p>Some groups of bacteria have never been cultured but can be 
characterized and reliably recovered from the environment by other 
means. These include endosymbiotic bacteria and organisms similar to 
the phytoplasmas, which can be identified by the plant diseases that 
they cause. We do not give these &ldquo;uncultured&rdquo; names, as 
above. These represent a special challenge for bacterial nomenclature, 
because a formal species description requires the designation of a 
cultured type strain. The bacteriological code has a special provision 
for names of this sort, Candidatus, e.g., Candidatus Endobugula or 
Candidatus Endobugula sertula; Candidatus Phlomobacter or Candidatus 
Phlomobacter fragariae. These often appear in the literature without 
the Candidatus prefix; therefore, we list the unqualified names as 
synonyms for retrieval purposes.</p>
</sec>
<sec id="bid.326">
<title>Informal Names above the Species Level</title>
<p>We allow informal names for unranked nodes above the species 
level as well. These should all be phylogenetically meaningful 
groups, e.g., the Fungi/Metazoa group, eudicotyledons, 
Erythrobasidium clade, RTA clade, and core jakobids. In addition, 
there are several other classes of nodes and names above the species 
level that explicitly do not represent phylogenetically meaningful 
groups. These are outlined below.</p>
<sec id="bid.327">
<title>Unclassified Bins</title>
<p>We are expected to add new species names to the database in a 
timely manner, preferably within a day or two. If we are able to 
find only a partial classification for a new taxon in the database, 
we place it as deeply as we can and list it in an explicit 
&ldquo;unclassified&rdquo; bin. As more information becomes available, 
these bins are emptied, and we give full classifications to the taxa 
listed there. In general, we suppress the names of the unclassified 
bins themselves so they do not appear in the abbreviated lineages 
that appear in the GenBank flatfiles, e.g., unclassified Salticidae, 
unclassified Bacteria, and unclassified Myxozoa.</p>
</sec>
<sec id="bid.1645">
<title>Incertae Sedis Bins</title>
<p>If the best taxonomic opinion available is that the position of 
a particular taxon is uncertain, then we will list it in an 
&ldquo;incertae sedis&rdquo; bin. This is a more permanent assignment 
than for taxa that are listed in unclassified bins, e.g., 
<italic>Neoptera incertae sedis, Chlorophyceae incertae sedis,</italic> 
and <italic>Riodininae incertae sedis.</italic></p>
</sec>
<sec id="bid.329">
<title>Mitosporic Bins</title>
<p>Fungi that were known only in the asexual (mitosporic, anamorphic) 
state were placed formerly in a separate, highly polyphyletic category 
of &ldquo;imperfect&rdquo; fungi, the Deuteromycota. Spurred especially 
by the development of molecular phylogenetics, current mycological 
practice is to classify anamorphic species as close to their sexual 
relatives as available information will support. Mitosporic categories 
can occur at any rank, e.g., <italic>mitosporic Ascomycota, mitosporic 
Hymenomycetes, mitosporic Hypocreales,</italic> and <italic>mitosporic 
Coniochaetaceae</italic>. The ultimate goal is to fully incorporate 
anamorphs into the natural phylogenetic classification.</p>
</sec>
</sec>
<sec id="bid.330">
<title>Other Names</title>
<p>The requirement that the Taxonomy database includes names 
from all of the entries in the sequence database introduces a 
number of names that are not typically treated in a taxonomic 
database. These are listed in the top-level group &ldquo;Other&rdquo;. 
Plasmids are typically annotated with their host organism, using 
the /plasmid source organism qualifier. Broad-host-range plasmids 
that are not associated with any single species are listed in their 
own bin. Plasmid and transposon names from very old sequence entries 
are listed in separate bins here as well. Plasmids that have been 
artificially engineered are listed in the &ldquo;vectors&rdquo; bin.</p>
</sec>
</app>
<app id="bid.331">
<title>Appendix 3. Other <abbrev>TAXON</abbrev> data types.</title>
<sec id="bid.332">
<title>Ranks</title>
<p>We do not require that Linnaean ranks be assigned to all of 
our taxa, but we do include a standard rank table that allows us 
to assign formal ranks where it seems appropriate. We do not require 
that sibling taxa all have the same rank, but we do not allow taxa 
of higher rank to be listed beneath taxa of lower rank. We allow 
unranked nodes to be placed at any point in our classification.</p>
<p>The one rank that we particularly care about is &ldquo;species&rdquo;. 
We try to ensure that all of the sequence entries map into the Taxonomy 
database at or below a species-level node.</p>
</sec>
<sec id="bid.333">
<title>Genetic Codes</title>
<p>The genetic codes and mitochondrial genetic codes that are 
appropriate for translating protein sequences in different branches 
of the tree of life are assigned at nodes in the Taxonomy database 
and inherited by species at the terminal branches of the tree. Plastid 
sequences are all translated with the standard genetic code, but many 
of the mRNAs undergo extensive RNA editing, making it difficult or 
impossible to translate sequences from the plastid genome directly. 
The genetic codes are listed on our <ext-link 
xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" 
xlink:href="http://www.ncbi.nlm.nih.gov/Taxonomy/Utils/
wprintgc?mode=c">Web site</ext-link>.</p>
</sec>
<sec id="bid.334">
<title>GenBank Divisions</title>
<p>GenBank taxonomic division assignments are made in the Taxonomy 
database and inherited by species at the terminal branches of the 
tree, just as with the genetic codes.</p>
</sec>
<sec id="bid.335">
<title>References</title>
<p>The Taxonomy database allows us to store comments and references 
at any taxon. These may include hotlinks to abstracts in <xref 
ref-type="other" rid="bid.1387">PubMed</xref>, as well as links to 
external addresses on the Web.</p>
</sec>
<sec id="bid.336">
<title>Abbreviated Lineage</title>
<p>Some branches of our taxonomy are many levels deep, e.g., the 
bony fish (as we moved to a phylogenetic classification) and the 
drosophilids (a model taxon for evolutionary studies). In many cases, 
the classification lines in the GenBank flatfiles became longer than 
the sequences themselves. This became a storage and update issue, and 
the classification lines themselves became less helpful as generally 
familiar taxa names became buried within less recognizable taxa.</p>
<p>To address this problem, the Taxonomy database allows us to flag 
taxa that should (or should not) appear in the abbreviated classification 
line in the GenBank flatfiles. The full lineages are indexed in Entrez 
and displayed in the Taxonomy Browser.</p>
</sec>
</app>
</app-group>
</back>
</book-part>
<?mtl hellip?><book-part id="bid.1143" book-part-type="chapter" book-part-number="5">
          <book-part-meta>
            <title-group>
              <title>The Single Nucleotide Polymorphism Database (dbSNP) of Nucleotide Sequence Variation</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Kitts</surname>
                  <given-names>Adrienne</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Sherry</surname>
                  <given-names>Stephen</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch5d1"/>
            <abstract>
              <title>Summary</title>
              <p>Sequence variations exist at defined positions within genomes and are responsible for individual 

phenotypic characteristics, including a person's propensity toward complex disorders such as heart disease and cancer. As 

tools for understanding human variation and molecular genetics, sequence variations can be used for gene mapping, definition 

of population structure, and performance of functional studies.</p>
              <p>The Single Nucleotide Polymorphism database (dbSNP) is a public-domain archive for a broad collection of 

simple genetic <xref ref-type="other" rid="bid.1379">polymorphism</xref>s. This collection of polymorphisms includes 

single-base nucleotide substitutions (also known as single nucleotide polymorphisms or SNPs), small-scale multi-base 

deletions or insertions (also called deletion insertion polymorphisms or DIPs), and retroposable element insertions and 

microsatellite repeat variations (also called short tandem repeats or STRs). Please note that in this chapter, you can 

substitute any class of variation for the term SNP. Each dbSNP entry includes the sequence context of the polymorphism (i.e., 

the surrounding sequence), the occurrence frequency of the polymorphism (by population or individual), and the experimental 

method(s), protocols, and conditions used to assay the variation.</p>
              <p>The dbSNP accepts submissions for variations in any species and from any part of a genome. This document 

will provide you with options for finding SNPs in dbSNP, discuss dbSNP content and organization, and furnish instructions to 

help you create your own (local) copy of dbSNP.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.1144">
              <title>Introduction</title>
              <p>The db<xref ref-type="other" rid="bid.1405">SNP</xref> has been designed to support submissions 

and research into a broad range of biological problems. These include physical mapping, functional analysis and 

pharmacogenomics, association studies, and evolutionary studies. Because <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/SNP/">dbSNP</ext-link> was developed to complement <xref ref-type="other" rid="bid.1299">GenBank</xref>, it may contain nucleotide sequences (<xref ref-type="fig" rid="bid.1157">Figure 1</xref>) 

from any organism; currently, the majority of the data is for human and mouse.<fig id="bid.1157"><label>1</label><caption><title>The structure of the flanking sequence.</title><p>The structure of the flanking sequence in dbSNP is a composite of bases 

either assayed for variation or included from published sequence. We make the distinction to distinguish regions of sequence 

that have been experimentally surveyed for variation (<italic>assay</italic>) from those regions that have not been surveyed 

(<italic>flank</italic>). The minimum sequence length for a variation definition (SNPassay) is 25 bp for both the 5&prime; and 

3&prime; flanks and 100 bp overall to ensure an adequate sequence for accurate mapping of the variation on reference genome 

sequence. (<italic>a</italic>) Flanking sequence completely surveyed for variation. Both 5&prime; and 3&prime; flanking 

sequence can be defined with 5&prime;_assay and 3&prime;_assay fields, respectively, when all flanking sequence was examined 

for variation. This can occur in both experimental contexts (e.g., denaturing high-pressure liquid chromatography or DNA 

sequencing) and computational contexts (e.g., analysis of <xref ref-type="other" rid="bid.1243">BAC</xref> overlap 

sequence). (<italic>b</italic>) Partial survey of flanking sequence can occur when detection methods examine only a region of 

sequence surrounding the variation that is shorter than either the 25 bp per flank rule or the 100 bp overall length rule. In 

these experimental designs (e.g., chip hybridization, enzymatic cleavage), we designate the experimental sequence 

5&prime;_assay or 3&prime;_assay, and you can insert published sequence (usually from a gene reference sequence) as 

5&prime;_flank or 3&prime;_flank to construct a sequence definition that will satisfy the length rules. (<italic>c</italic>) 

Unknown or no survey of flanking sequence can occur when variations are captured from published literature without an 

indication of survey conditions. In these cases, the entire flanking sequence is defined as 5&prime;_flank and 

3&prime;_flank.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch5f1" mime-subtype="gif"/></fig></p>
              <sec id="bid.1646">
                <title>Physical Mapping</title>
                <p>In the physical mapping of nucleotide sequences, variations are used as positional markers. When mapped to 

a unique location in a genome, variation markers work with the same logic as Sequence Tagged Sites 
(<xref ref-type="other" rid="bid.1410">STS</xref>s) or framework <xref ref-type="other" rid="bid.1347">microsatellite</xref> 

markers. As is the case for STSs, the position of a variation is defined by its unique flanking sequence, and hence, 

variations can serve as stable landmarks in the genome, even if the variation is fixed for one <xref ref-type="other" rid="bid.1240">allele</xref> in a sample. When multiple alleles are observed in a sample pedigree, pedigree members can 

be tested for variation <xref ref-type="other" rid="bid.1302">genotype</xref>s as in traditional physical mapping 

studies.</p>
              </sec>
              <sec id="bid.1647">
                <title>Functional Analysis</title>
                <p>Variations that occur in functional regions of genes or in conserved non-coding regions can cause 

significant changes in the complement of transcribed sequences, leading to changes in protein expression that can affect 

aspects of the <xref ref-type="other" rid="bid.1369">phenotype</xref> such as metabolism or cell signaling. We note 

possible functional implications of DNA sequence variations in dbSNP in terms of how the variation alters mRNA 

transcripts.</p>
              </sec>
              <sec id="bid.1648">
                <title>Association Studies</title>
                <p>The associations between variations and complex genetic traits are more ambiguous than simple, single-gene 

mutations that lead to a phenotypic change. When multiple genes are involved in a trait, then the identification of the 

genetic causes of the trait requires the identification of the chromosomal segment combinations, or haplotypes, that carry 

the putative gene variants.</p>
              </sec>
              <sec id="bid.1649">
                <title>Evolutionary Studies</title>
                <p>The variations in dbSNP currently represent an uneven but large sampling of genome diversity. The human 

data in dbSNP include submissions from the SNP Consortium, variations mined from genome sequence as part of the human genome 

project, and individual lab contributions of variations in specific genes, mRNAs, <xref ref-type="other" rid="bid.1283">EST</xref>s, or genomic regions.</p>
              </sec>
              <sec id="bid.1650">
                <title>Null Results Are Important</title>
                <p>Systematic surveys of sequence variation will undoubtedly reveal sequences that are invariant in the 

sample. These observations can be submitted to dbSNP as NoVariation records that record the sequence, the population, and the 

sample size that were used in the survey.</p>
              </sec>
            </sec>
            <sec id="bid.1145">
              <title>Searching dbSNP</title>
              <p>The SNP database can be queried from the dbSNP homepage <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/SNP/">http://www.ncbi.nlm.nih.gov/SNP/</ext-link> (<xref ref-type="fig" rid="bid.1146">Figure 

2</xref>) by using the <xref ref-type="other" rid="bid.1282">Entrez</xref> SNP <named-content content-type="book-spec-emph">Search</named-content> box at the top of the 

page or by using the links to eight basic dbSNP search options (located just below the Entrez SNP <named-content content-type="book-spec-emph">Search</named-content> box). Each 

search option is described below.<fig id="bid.1146"><label>2</label><caption><title>The dbSNP homepage.</title><p>We organized the dbSNP homepage with links to documentation, FTP, and sub-query 

pages on the <italic>left sidebar</italic> and a selection of query modules on the <italic>right sidebar</italic>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch5f2" mime-subtype="gif"/></fig></p>
              <sec id="bid.1147">
                <title>Entrez SNP</title>
                <p>The dbSNP is now a part of the Entrez integrated information retrieval system (<xref ref-type="other" rid="bid.588">Chapter 15</xref>) and may be searched using either qualifiers (aliases) or a combination of 28 

different search fields. A complete list of the qualifiers and search fields can be found on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=snp">Entrez SNP site</ext-link>.</p>
              </sec>
              <sec id="bid.1148">
                <title>Single Record (Search by IDs) Query in dbSNP</title>
                <p>Use this query module to select SNPs based on dbSNP record identifiers. These include reference 

SNP (refSNP) cluster ID numbers (rs#), submitted SNP Accession numbers (ss#), and local (or submitter) IDs for the same 

variations.</p>
              </sec>
              <sec id="bid.1149">
                <title>SNP Submission Information Queries</title>
                <table-wrap id="bid.1161">
                  <label>1</label>
                  <caption>
                    <title>Method classes organize submissions by a general methodological or experimental 

approach to assaying for variation in the DNA sequence.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="middle">Method class</th>
                      <th align="center" valign="middle">Class code in 

Sybase, ASN.1, and XML</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Denaturing high pressure 

liquid chromatography (DHPLC)</td>
                      <td align="center" valign="top">1</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">DNA 

hybridization</td>
                      <td align="center" valign="top">2</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Computational 

analysis</td>
                      <td align="center" valign="top">3</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Single-stranded 

conformational polymorphism (SSCP)</td>
                      <td align="center" valign="top">5</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Other</td>
                      <td align="center" valign="top">6</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Unknown</td>
                      <td align="center" valign="top">7</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Restriction fragment 

length polymorphism (<xref ref-type="other" rid="bid.1393">RFLP</xref>)</td>
                      <td align="center" valign="top">8</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Direct DNA 

sequencing</td>
                      <td align="center" valign="top">9</td>
                    </tr>
                  </table>
                </table-wrap>
                <table-wrap id="bid.1163">
                  <label>2</label>
                  <caption>
                    <title>Population classes organize population samples by geographic region.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Population class</th>
                      <th align="left" valign="top">Description</th>
                      <th align="center" valign="top">Population class in 

Sybase, ASN.1, and XML</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Central Asia</td>
                      <td align="left" valign="top">Samples from Russia and 

its satellite Republics and from nations bordering the Indian Ocean between East Asia and the Persian Gulf regions.</td>
                      <td align="center" valign="top">8</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Central/South 

Africa</td>
                      <td align="left" valign="top">Samples from nations 

south of the Equator, Madagascar, and neighboring island nations.</td>
                      <td align="center" valign="top">4</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Central/South 

America</td>
                      <td align="left" valign="top">Samples from Mainland 

Central and South America and island nations of the western Atlantic, Gulf of Mexico, and Eastern Pacific.</td>
                      <td align="center" valign="top">10</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">East Asia</td>
                      <td align="left" valign="top">Samples from eastern and 

south eastern Mainland Asia and from Northern Pacific island nations.</td>
                      <td align="center" valign="top">6</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Europe</td>
                      <td align="left" valign="top">Samples from Europe 

north and west of Caucasus Mountains, Scandinavia, and Atlantic islands.</td>
                      <td align="center" valign="top">5</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Multi-National</td>
                      <td align="left" valign="top">Samples that were 

designated to maximize measures of heterogeneity or sample human diversity in a global fashion. Examples include 

OEFNER|GLOBAL and the CEPH repository.</td>
                      <td align="center" valign="top">1</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">North America</td>
                      <td align="left" valign="top">All samples north of the 

Tropic of Cancer, including defined samples of United States Caucasians, African Americans, Hispanic Americans, and the NHGRI 

polymorphism discovery resource (NCBI|NIHPDR).</td>
                      <td align="center" valign="top">9</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">North/East Africa and 

Middle East</td>
                      <td align="left" valign="top">Samples collected from 

North Africa (including the Sahara desert), East Africa (south to the Equator), Levant, and the Persian Gulf.</td>
                      <td align="center" valign="top">2</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Pacific</td>
                      <td align="left" valign="top">Samples from Australia, 

New Zealand, Central and Southern Pacific Islands, and Southeast Asian peninsular/island nations.</td>
                      <td align="center" valign="top">7</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Unknown</td>
                      <td align="left" valign="top">Samples with unknown 

geographic provinces that are not global in nature.</td>
                      <td align="center" valign="top">11</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">West Africa</td>
                      <td align="left" valign="top">Sub-Saharan nations 

bordering the Atlantic north of the Congo River and central/southern Atlantic island nations.</td>
                      <td align="center" valign="top">3</td>
                    </tr>
                  </table>
                </table-wrap>
                <p>Use this module to construct a query that will select SNPs based on submission records by 

laboratory (submitter), new data (called &ldquo;new batches&rdquo;, this query limitation is more recent than a 

user-specified date), methods used to assay for variation (<xref ref-type="table" rid="bid.1161">Table 1</xref>), 

populations of interest (<xref ref-type="table" rid="bid.1163">Table 2</xref>), or publication information.</p>
              </sec>
              <sec id="bid.1150">
                <title>dbSNP Batch Query</title>
                <p>Use sets of variation IDs collected from other queries to generate a variety of SNP reports.</p>
              </sec>
              <sec id="bid.1151">
                <title>Locus Information Query</title>
                <p>The links in this section all point to the <xref ref-type="other" rid="bid.1353">NCBI</xref><xref ref-type="other" rid="bid.1334">LocusLink</xref> resource. This resource permits 

users to query for records by gene name, gene symbol, gene product, gene ontology, or Accession number. Records in LocusLink 

will have a pointer back to dbSNP if you have associated one or more variations with a gene feature.</p>
              </sec>
              <sec id="bid.1152">
                <title>Free Form (Entrez-like) and Easy Form Queries</title>
                <p>Free Form is the most flexible query structure in dbSNP. Modeled on the NCBI Entrez retrieval 

system, queries can be conducted using multiple database field values to restrict a query to specific subsets of data. The 

Easy Form query is identical to the Free Form query, with the exception that the Easy Form query has a series of pull-down 

menus from which a value can be selected for the most popular query fields.</p>
              </sec>
              <sec id="bid.1153">
                <title>Between-Markers Positional Query</title>
                <p>Use this query approach if you are interested in retrieving variations that have been mapped to a 

specific region of the genome bounded by two STS markers. Other map-based queries are available through the NCBI 

<xref ref-type="other" rid="bid.1336">Map Viewer</xref> tool.</p>
              </sec>
              <sec id="bid.1154">
                <title>ADA Section 508-compliant Link</title>
                <p>All links located on the left sidebar of the dbSNP homepage are also provided in text format at 

the bottom of the page to support browsing by text-based Web browsers. Suggestions for improving database access by disabled 

persons should be sent to: <email>snp-admin@ncbi.nlm.nih.gov</email>.</p>
              </sec>
            </sec>
            <sec id="bid.1155">
              <title>Submitted Content</title>
              <p>The SNP database has two major classes of content: the first class is submitted data, i.e., original 

observations of sequence variation (<xref ref-type="fig" rid="bid.1169">Figure 3</xref>); and the second class is computed 

content, i.e., content generated during the dbSNP &ldquo;build&rdquo; cycle by computation on original submitted data. 

Computed content consists of refSNPs, other computed data, and links that increase the utility of dbSNP.<fig id="bid.1169"><label>3</label><caption><title>dbSNP submission details for an assay report.</title><p>The major sections of the report are described in the <italic>main panel</italic>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch5f3" mime-subtype="gif"/></fig></p>
              <p>A complete copy of the SNP database is publicly available and can be downloaded from the SNP <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/snp/">FTP</ext-link> site (see the section <italic>How to Create a Local Copy of 

dbSNP</italic>). dbSNP accepts submissions from public laboratories and private organizations. (There are online <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/SNP/get_html.cgi?whichHtml=how_to_submit">instructions</ext-link> for preparing a 

submission to dbSNP.) A short tag or abbreviation called Submitter HANDLE uniquely defines each submitting laboratory and 

groups the submissions within the database. The 10 major data elements of a submission follow.</p>
              <sec id="bid.1156">
                <title>Flanking Sequence Context DNA or cDNA</title>
                <p>The essential component of a submission to dbSNP is the nucleotide sequence itself. dbSNP accepts 

submissions as either genomic DNA or cDNA (i.e., sequenced mRNA transcript) sequence. Sequence submissions have a minimum 

length requirement to maximize the specificity of the sequence in larger contexts, such as a reference genome sequence. We 

also structure submissions so that the user can distinguish regions of sequence actually surveyed for variation from regions 

of sequence that are cut and pasted from a published reference sequence to satisfy the minimum-length requirements. <xref ref-type="fig" rid="bid.1157">Figure 1</xref> shows the details of flanking sequence structure.</p>
              </sec>
              <sec id="bid.1158">
                <title>Alleles</title>
                <table-wrap id="bid.1159">
                  <label>3</label>
                  <caption>
                    <title>Allele definitions define the class of the variation in dbSNP.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">dbSNP variation 

class<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref>, <xref ref-type="fn" rid="mul4"><sup><italic>b</italic></sup></xref></th>
                      <th align="center" valign="top">Rules for assigning 

allele classes</th>
                      <th align="center" valign="top">Sample allele 

definition</th>
                      <th align="center" valign="top">Class code in Sybase, 

ASN.1, and XML<xref ref-type="fn" rid="mul5"><sup><italic>c</italic></sup></xref></th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Single Nucleotide 

Polymorphisms (SNPs)<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref></td>
                      <td align="left" valign="top">Strictly defined as 

single base substitutions involving A, T, C, or G.</td>
                      <td align="left" valign="top">A/T</td>
                      <td align="center" valign="top">1</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Deletion/Insertion 

Polymorphisms (DIPs)<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref></td>
                      <td align="left" valign="top">Designated using the 

full sequence of the insertion as one allele, and either a fully defined string for the variant allele or a &ldquo;-&rdquo; 

character to specify the deleted allele. This class will be assigned to a variation if the variation alleles are of different 

lengths or if one of the alleles is deleted (&ldquo;-&rdquo;).</td>
                      <td align="left" valign="top">T/-/CCTA/G</td>
                      <td align="center" valign="top">2</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Heterozygous 

sequence<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref></td>
                      <td align="left" valign="top">The term heterozygous is 

used to specify a region detected by certain methods that do not resolve the polymorphism into a specific sequence motif. In 

these cases, a unique flanking sequence must be provided to define a sequence context for the variation.</td>
                      <td align="left" valign="top">(heterozygous)</td>
                      <td align="center" valign="top">3</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Microsatellite or short 

tandem repeat (STR)<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref></td>
                      <td align="left" valign="top">Alleles are designated 

by providing the repeat motif and the copy number for each allele. Expansion of the allele repeat motif designated in dbSNP 

into full-length sequence will be only an approximation of the true genomic sequence because many microsatellite markers are 

not fully sequenced and are resolved as size variants only.</td>
                      <td align="left" valign="top">(CAC)8/9/10/11</td>
                      <td align="center" valign="top">4</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Named variant<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref></td>
                      <td align="left" valign="top">Applies to 

insertion/deletion polymorphisms of longer sequence features, such as retroposon dimorphism for <xref ref-type="other" rid="bid.1239">Alu</xref> or line elements. These variations frequently include a deletion &ldquo;-&rdquo; indicator for 

the absent allele.</td>
                      <td align="left" valign="top">(alu) / -</td>
                      <td align="center" valign="top">5</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">No-variation<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref></td>
                      <td align="left" valign="top">Reports may be submitted 

for segments of sequence that are assayed and determined to be invariant in the sample.</td>
                      <td align="left" valign="top">(NoVariation)</td>
                      <td align="center" valign="top">6</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Mixed<xref ref-type="fn" rid="mul4"><sup><italic>b</italic></sup></xref></td>
                      <td align="left" valign="top"/>
                      <td align="left" valign="top">Mix of other 

classes</td>
                      <td align="center" valign="top">7</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Multi-Nucleotide 

Polymorphism (MNP)<xref ref-type="fn" rid="mul3"><sup><italic>a</italic></sup></xref></td>
                      <td align="left" valign="top">Assigned to variations 

that are multi-base variations of a single, common length.</td>
                      <td align="left" valign="top">GGA/AGT</td>
                      <td align="center" valign="top">8</td>
                    </tr>
                  </table>
                  <table-wrap-foot>
                    <fn symbol="a" id="mul3">
                      <p>Seven of the classes apply to both submissions of 

variations (submitted SNP assay, or ss#) and the non-redundant refSNP clusters (rs#'s) created in dbSNP.</p>
                    </fn>
                    <fn symbol="b" id="mul4">
                      <p>The &ldquo;Mixed&rdquo; class is assigned to refSNP 

clusters that group submissions from different variation classes.</p>
                    </fn>
                    <fn symbol="c" id="mul5">
                      <p>Class codes have a numeric representation in the 

database itself and in the export versions of the data (<xref ref-type="other" rid="bid.1242">ASN.1</xref> and 

<xref ref-type="other" rid="bid.1435">XML</xref>).</p>
                    </fn>
                  </table-wrap-foot>
                </table-wrap>
                <p>Alleles define variation class (<xref ref-type="table" rid="bid.1159">Table 3</xref>). In the 

dbSNP submission scheme, we define single-nucleotide variants as G, A, T, or C. We do not permit ambiguous IUPAC codes, such 

as N, in the allele definition of a variation. In cases where variants occur in close proximity to one another, we permit 

IUPAC codes such as N, and in the flanking sequence of a variation, we actually encourage them. See <xref ref-type="table" rid="bid.1159">Table 3</xref> for the rules that guide dbSNP post-submission processing in assigning allele classes to each 

variation.</p>
              </sec>
              <sec id="bid.1160">
                <title>Method</title>
                <p>Each submitter defines the methods in their submission as either the techniques used to assay 

variation or the techniques used to estimate allele frequencies. We group methods by method class (<xref ref-type="table" rid="bid.1161">Table 1</xref>) to facilitate queries using general experimental technique as a query field. The submitter 

provides all other details of the techniques in a free-text description of the method. Submitters can also use the 

<named-content content-type="book-spec-emph">METHOD EXCEPTION</named-content> field to describe changes to a general protocol for particular sets of data (batch-specific 

details). Submitters generally define methods only once in the database.</p>
              </sec>
              <sec id="bid.1162">
                <title>Population</title>
                <p>Each submitter defines population samples either as the group used to initially identify 

variations or as the group used to identify population-specific measures of allele frequencies. These populations may be one 

and the same in some experimental designs. We assign populations a population class (<xref ref-type="table" rid="bid.1163">Table 

2</xref>) based on the geographic origin of the sample. These broad categories provide a general framework for organizing 

the approximately 700 (as of this writing) sample descriptions in dbSNP. Similar to method descriptions, population 

descriptions minimally require the submitter to provide a Population ID and a free-text description of the sample.</p>
              </sec>
              <sec id="bid.1164">
                <title>Sample Size</title>
                <p>There are two sample-size fields in dbSNP. One field is called the <named-content content-type="book-spec-emph">SNPASSAY SAMPLE SIZE</named-content>, 

and it reports the number of chromosomes in the sample used to initially ascertain or discover the variation. The other 

sample size field is called <named-content content-type="book-spec-emph">SNPPOPUSE SAMPLE SIZE</named-content>, and it reports the number of chromosomes used as the denominator 

in computing estimates of allele frequencies. These two measures need not be the same.</p>
              </sec>
              <sec id="bid.1165">
                <title>Population-specific Allele Frequencies</title>
                <table-wrap id="bid.1171">
                  <label>4</label>
                  <caption>
                    <title>Validation status codes summarize the available validation data in assay reports and 

refSNP clusters.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Validation 

evidence</th>
                      <th align="center" valign="top">Description</th>
                      <th align="center" valign="top">Code in database for 

ss#</th>
                      <th align="center" valign="top">Code in FTP dumps for 

ss#</th>
                      <th align="center" valign="top">Code in database for 

rs#</th>
                      <th align="center" valign="top">Code in FTP dumps for 

rs#</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Not validated</td>
                      <td align="left" valign="top">For ss#, no batch update 

or validation data, no frequency data (or frequency is 0 or 1). rs# status code is OR'd from the ss# codes.</td>
                      <td align="center" valign="top">0</td>
                      <td align="center" valign="top">Not present</td>
                      <td align="center" valign="top">0<xref ref-type="fn" rid="mul6"><sup><italic>a</italic></sup></xref></td>
                      <td align="center" valign="top">Not present</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Multiple 

reporting</td>
                      <td align="left" valign="top">Status = 1 for an rs# 

with at least two ss# numbers; having at least one ss# is validated by a non-computational method. For a ss#, status = 1 if 

the method is non-computational.</td>
                      <td align="center" valign="top">1</td>
                      <td align="center" valign="top">1<xref ref-type="fn" rid="mul7"><sup><italic>b</italic></sup></xref></td>
                      <td align="center" valign="top">1,0<xref ref-type="fn" rid="mul7"><sup><italic>b</italic></sup></xref></td>
                      <td align="center" valign="top">1</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">With frequency</td>
                      <td align="left" valign="top">Frequency data is 

present with a value between 0 and 1.</td>
                      <td align="center" valign="top">2</td>
                      <td align="center" valign="top">2</td>
                      <td align="center" valign="top">2</td>
                      <td align="center" valign="top">2</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Both frequency</td>
                      <td align="left" valign="top">For ss#, the method is 

non-computational and frequency data is present. If the ss# is a single cluster member, then the rs# code is set to 

2.</td>
                      <td align="center" valign="top">3</td>
                      <td align="center" valign="top">3</td>
                      <td align="center" valign="top">3/2</td>
                      <td align="center" valign="top">3</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Submitter 

validation</td>
                      <td align="left" valign="top">Submission of a batch 

update or validation section that reports a second validation method on the assay.</td>
                      <td align="center" valign="top">4</td>
                      <td align="center" valign="top">4</td>
                      <td align="center" valign="top">4</td>
                      <td align="center" valign="top">4</td>
                    </tr>
                  </table>
                  <table-wrap-foot>
                    <fn symbol="a" id="mul6">
                    <p>If the rs# has a single ss# with code 1, then rs# is set to code 0.</p>
                    </fn>
                    <fn symbol="b" id="mul7">
                      <p>For a single member rs where the ss# validation status = 1, the rs# validation status is 

set to 0.</p>
                    </fn>
                  </table-wrap-foot>
                </table-wrap>
                <p>Alleles typically exist at different frequencies in different populations; a very common allele in 

one population may be quite rare in another population. Also, allelic variants can emerge as <xref ref-type="other" rid="bid.1381">private polymorphisms</xref> when particular populations have been reproductively isolated from 

neighboring groups, as is the case with religious isolates or island populations. Frequency data are submitted to dbSNP as 

allele counts or binned frequency intervals, depending on the precision of the experimental method used to make the 

measurement. dbSNP contains records of allele frequencies for specific population samples defined by each submitter 

(<xref ref-type="table" rid="bid.1171">Table 4</xref>).</p>
              </sec>
              <sec id="bid.1166">
                <title>Population-specific Genotype Frequencies</title>
                <p>Similar to alleles, <xref ref-type="other" rid="bid.1402">genotypes</xref> have 

frequencies in populations that can be submitted to dbSNP.</p>
              </sec>
              <sec id="bid.1167">
                <title>Population-specific Heterozygosity Estimates</title>
                <p>Some methods for detection of variation (e.g., denaturing high-pressure liquid chromatography or 

DHPLC) can recognize when DNA fragments contain a variation without resolving the precise nature of the sequence change. 

These data define an empirical measure of <xref ref-type="other" rid="bid.1306">heterozygosity</xref> when submitted 

to dbSNP.</p>
              </sec>
              <sec id="bid.1168">
                <title>Individual Genotypes</title>
                <p>dbSNP accepts individual genotypes for samples from publicly available repositories such as 

<xref ref-type="other" rid="bid.1724">CEPH</xref> or <xref ref-type="other" rid="bid.1726">Coriell</xref>. 

Genotypes reported in dbSNP contain links to population and method descriptions as shown in <xref ref-type="fig" rid="bid.1169">Figure 3</xref>. General genotype data provide the foundation for individual haplotype definitions and are 

useful for selecting positive and negative control reagents in new experiments.</p>
              </sec>
              <sec id="bid.1170">
                <title>Validation Information</title>
                <p>dbSNP accepts individual assay records (ss# numbers) without validation evidence. When possible, 

however, we try to distinguish high-quality validated data from unconfirmed (usually computational) variation reports. Assays 

validated directly by the submitter through the <named-content content-type="book-spec-emph">VALIDATION</named-content> section show the type of evidence used to confirm the 

variation. Additionally, dbSNP will flag an assay as validated (<xref ref-type="table" rid="bid.1171">Table 4</xref>) when 

we observe frequency or genotype data for the record.</p>
              </sec>
            </sec>
            <sec id="bid.1172">
              <title>Computed Content (The dbSNP Build Cycle)</title>
              <p>We release the content of dbSNP to the public in periodic &ldquo;builds&rdquo; that we synchronize with 

the release of new genome assemblies (<xref ref-type="other" rid="bid.1440">Chapter 14</xref>). During each build, we 

cluster the data submitted since the last build into existing refSNPs and form new refSNPs when necessary. The following 12 

tasks define the sequence of steps in the dbSNP build cycle (<xref ref-type="fig" rid="bid.1173">Figure 4</xref>).<fig id="bid.1173"><label>4</label><caption><title>The dbSNP build cycle.</title><p>The dbSNP build cycle starts with close of data for new submissions. We map all 

data, including existing refSNP clusters and new submissions, to reference genome sequence if available for the organism. 

Otherwise, we map them to non-redundant DNA sequences from GenBank. We then use map data on co-occurrence of hit locations to 

either merge submissions into existing clusters or to create new clusters. We then annotate the new non-redundant refSNP (rs) 

set on reference sequences and dump the contents of dbSNP in a variety of comprehensive and denormalized formats on the dbSNP 

FTP site for release with the online build of the database.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch5f4" mime-subtype="gif"/></fig></p>
              <sec id="bid.1174">
                <title>Mapping and Reclustering New Submissions</title>
                <p>Each build starts with a &ldquo;close of data&rdquo; that defines the set of new submissions that 

will be mapped to genome sequence by <xref ref-type="other" rid="bid.1340">MegaBLAST</xref> for subsequent 

reclustering and annotation. The set of new data entering each build typically includes all submissions received since the 

close of data in the previous build.</p>
              </sec>
              <sec id="bid.1175">
                <title>Resource Integration</title>
                <p>We annotate the non-redundant set of variations (refSNP cluster set) on reference genome sequence 

<xref ref-type="other" rid="bid.1267">contig</xref>s, chromosomes, mRNAs, and proteins as part of the NCBI 

<xref ref-type="other" rid="bid.1392">RefSeq</xref> project (<xref ref-type="other" rid="bid.681">Chapter 

18</xref>). We compute summary properties for each refSNP cluster, which we then use to build fresh indexes for dbSNP 

in Entrez and to update the variation map in the NCBI Map Viewer. Finally, we update links between dbSNP and <!--DO NOT 

UNCOMMENT<glossref>-->dbMHC<!--</glossref>DO NOT UNCOMMENT-->, <xref ref-type="other" rid="bid.1425">UniSTS</xref>, 

<xref ref-type="other" rid="bid.1334">LocusLink</xref>, <xref ref-type="other" rid="bid.1387">PubMed</xref>, 

and <xref ref-type="other" rid="bid.1424">UniGene</xref>.</p>
              </sec>
              <sec id="bid.1176">
                <title>Public Release</title>
                <p>Public release of a new build involves an update to the public database and the production of a 

new set of files on the dbSNP FTP site. We make an announcement to the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mailman/listinfo/dbsnp-announce">dbsnp-announce</ext-link> mailing list when the new build 

is publicly available.</p>
              </sec>
              <sec id="bid.1177">
                <title>refSNP Cluster Assignment for Non-Redundant Datasets</title>
                <p>Data submitted to dbSNP are clustered and provide a non-redundant set of variations for each 

organism in the database. We maintain these clusters as refSNPs in dbSNP in parallel to the underlying submitted data. We 

distinguish refSNPs from assay submissions by using an <bold>rs</bold>-prefixed (<bold>r</bold>ef<bold>S</bold>NP) 

Accession number instead of the <bold>ss</bold>-prefixed (<bold>s</bold>ubmitted <bold>S</bold>NP) Accession number 

assigned to individual submissions.</p>
                <p>refSNPs are compact sets of identifiers that are used to annotate variations on other NCBI 

resources. A refSNP has a number of summary properties that are computed over all cluster members (<xref ref-type="fig" rid="bid.1178">Figure 5</xref>). We export the entire refSNP set in many report formats on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/snp/">FTP</ext-link> site and as sets of results through a dbSNP batch query. We maintain both refSNPs 

and submitted SNPs as <xref ref-type="other" rid="bid.1290">FASTA</xref> databases for <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/SNP/snpblastByChr.html">BLAST searches</ext-link> of dbSNP.<fig id="bid.1178"><label>5</label><caption><title>Schematic view of refSNP rs7412 and its connections to underlying submissions 

ss9266, ss870165, and ss1542565.</title><p>rs7412 has an average heterozygosity of 12.7% based on the frequency data 

provided by the three submissions, and the cluster as a whole is validated because one of the underlying submissions has been 

experimentally validated. rs7412 is annotated as a variation feature on RefSeq contigs, mRNAs, and proteins. Pointers in the 

refSNP summary record direct the user to additional information on the three submitter Web sites, through the linkout 

<xref ref-type="other" rid="bid.1427">URL</xref>s supplied in each submission. These Web sites may contain additional 

data that were used in the initial variation call, or it may be additional phenotype or molecular data that indicate the 

function of the variation.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch5f5" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.1179">
                <title>Summary Data Measures</title>
                <p>We compute summary measures for each refSNP to integrate data provided by each independent 

submitter.</p>
                <sec id="bid.1180">
                  <title>The refSNP Clustering Process</title>
                  <p>Submitters can arbitrarily define variations on either strand of DNA sequence; therefore, 

submissions in a refSNP cluster can be reported on the forward or reverse strand. The orientation of the refSNP and, hence, 

its sequence and allele string, is set by a cluster exemplar. By convention, the clustering process <xref ref-type="fig" rid="bid.1181">(Figure 6</xref>) picks a cluster exemplar as that member of a cluster with the longest sequence. In subsequent 

builds, this sequence may be in reverse orientation to the current orientation of the refSNP. When this occurs, we try to 

preserve the orientation of the refSNP, if possible, by using the reverse complement of the cluster exemplar to set the 

orientation of the refSNP sequence.<fig id="bid.1181"><label>6</label><caption><title>The dbSNP reclustering process.</title><p>We define clusters on shared locations (refSNPs) when we BLAST all 

existing refSNPs against contig sequence. In cases where contig sequence is not available or the variation is defined in a 

mRNA flanking sequence that will not map to a contig, we compute the refSNP set based on hits to the RefSeq or GenBank 

sequences for the organism. We rank map hits as either LO or HI quality and parse the hits to assign a <xref ref-type="other" rid="bid.1431">weight</xref> to each refSNP. We make the set of all hits unique by dropping contig location 

and retaining only the first occurrence of each rs#, ss#, ID, and the cluster in which each ID number appears. The resulting 

data include all refSNPs for the current build that have at least one hit to contig, RefSeq, or GenBank sequence. We then 

compare the set of new submissions that fail to hit these sequence sets against each other using BLAST. This removes any 

potential redundancy in the incoming data, and unmapped refSNPs are instantiated for these data as well. This final merged 

set of data constitutes the refSNP set for the current build.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch5f6" mime-subtype="gif"/></fig></p>
                  <p>Once the clustering process determines the orientation of all member sequences in a 

cluster, it will gather a comprehensive set of alleles for a refSNP cluster.</p>
                  <boxed-text position="anchor" content-type="website">
                    <title>Hint:</title>
                    <p>When the alleles of a submission appear to be different from the alleles of its 

parent refSNP, check the orientation of the submission for reverse orientation.</p>
                  </boxed-text>
                </sec>
                <sec id="bid.1182">
                  <title>Summary Measures of Variation</title>
                  <p>The best single measure of a variation's diversity in different populations is its average 

<xref ref-type="other" rid="bid.1306">heterozygosity</xref>. This measure serves as the general probability that both 

alleles are in a diploid individual or in a sample of two chromosomes. Estimates of average heterozygosity have an 

accompanying standard error based on the sample sizes of the underlying data, which reflects the overall uncertainty of the 

estimate.</p>
                  <p>Additional summary measures of variation include counts of populations and individuals 

sampled for this variation.</p>
                </sec>
              </sec>
              <sec id="bid.1183">
                <title>Mapping to Reference Genome Sequence</title>
                <p>When reference genome assemblies are available, we use them as anchor sequence to place refSNP 

clusters into a genomic context. We clean dbSNP flanking sequence with <xref ref-type="other" rid="bid.1734">RepeatMasker</xref> and then re-map them to the most current build of each genome using MegaBLAST. The 

mapping results then define a new non-redundant set of variations for the genome.</p>
              </sec>
              <sec id="bid.1184">
                <title>Reclustering</title>
                <p>We define a refSNP operationally as a variation at a location on an ideal reference chromosome. 

Such reference chromosomes are the goal of genome assemblies. However, work is still in progress in cases such as the human 

genome project; therefore, we must currently define a refSNP as a variation in the interim reference contig sequence. Every 

time there is a genomic assembly update, the interim reference contig sequence changes, and refSNPs must be updated or 

reclustered.</p>
                <p>The reclustering process begins when NCBI updates the genomic assembly. We BLAST all existing 

refSNPs as well as any newly submitted SNPs (not yet bound to a refSNP cluster) against the genome assembly. Then, we cluster 

SNPs that co-locate at the same place on the genome into a single refSNP. Usually, new clusters are composed entirely of new 

submitted SNPs, or else the newly submitted SNPs cluster to an already existing refSNP. When newly submitted SNPs cluster 

among themselves, they are assigned to a new refSNP ID#, and when they cluster with an already existing refSNP, they are 

assigned to the cluster for that refSNP.</p>
                <p>Sometimes a refSNP will co-locate with another existing refSNP. In this case, the refSNP with a 

higher ID number is retired, and all the submitted SNPs in its cluster are reassigned to the refSNP with the lower ID 

number.</p>
                <p>Once the clusters are formed, the variation of a refSNP is the union of all possible alleles 

defined in the set of submitted SNPs that composed the cluster. <xref ref-type="fig" rid="bid.1181">Figure 6</xref> is a 

detailed flow chart of the reclustering process.</p>
              </sec>
              <sec id="bid.1185">
                <title>NCBI Contig Annotation</title>
                <p>We annotate <xref ref-type="other" rid="bid.1431">weight</xref> 1 and 2 refSNP variations 

on NCBI RefSeq chromosomes, contig sequences, mRNAs, and proteins as variation features with multiple allele qualifiers (one 

per allele). Weight 2 records receive an additional warning note to indicate the ambiguous nature of the mapping result.</p>
                <boxed-text position="anchor" content-type="website">
                  <title>Hint:</title>
                  <p>The two hits that define a weight 2 variation may not reflect <xref ref-type="other" rid="bid.1365">paralogy</xref> in the genome. Sequence assemblies are imperfect, and some regions of unique genome 

sequence are potentially reflected in two or more contig sequence fragments. Because we are currently unable to distinguish 

such cases from true paralogy, we annotate the variation in both locations with a warning and leave the assessment of the 

flanking sequence to the user.</p>
                </boxed-text>
                <p>We do not believe that weight 3 and weight 10 variations have sufficient utility to warrant their 

annotation, but the mapping results for these variations are still available in dbSNP.</p>
                <p>We annotate NoVariation records on NCBI RefSeq chromosomes, contig sequences, mRNAs, and proteins 

as a miscellaneous feature, or misc_feat. All dbSNP annotations also include a db_Xref cross-reference pointer back to dbSNP 

that uses the refSNP ID number.</p>
              </sec>
              <sec id="bid.1186">
                <title>Annotating GenBank and Other RefSeq Records</title>
                <p>GenBank records can be annotated only by their original authors. Therefore, when we find 

high-quality hits of refSNP records to the <xref ref-type="other" rid="bid.1311">HTGS</xref> and non-redundant 

divisions of GenBank, we connect them using LinkOut (<xref ref-type="other" rid="bid.650">Chapter 17</xref>).</p>
                <p>We annotate RefSeq mRNAs with variation features when the refSNP has a high-quality hit to the 

mRNA sequence. If the variation is in the coding region of the transcript and has a non-synonymous allele that changes the 

protein sequence, we also annotate the variation on the protein translation of the mRNA. The alleles in protein annotations 

are the amino acid translations of the affected codons.</p>
              </sec>
              <sec id="bid.1187">
                <title>NCBI Map Viewer Variation and Linkage Maps</title>
                <p>The Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>) can show multiple 

maps of sequence features in common chromosome coordinates. The variation map shows all variation features that we annotate 

on the current genome assembly. There are two ways to see the variation data. The default graphic mode shows the data as tick 

marks on the vertical coordinate scale. When <named-content content-type="book-spec-emph">variation</named-content> is selected as the master map, a summary of map quality, 

quality warning, functional relationships to genes, average heterozygosity with standard error, and validation information 

are provided. If genotype, haplotype, or LinkOut data are available, the master map will also contain links to this 

information.</p>
                <boxed-text position="anchor" content-type="website">
                  <title>Hint:</title>
                  <p>The summary values can be viewed or downloaded directly as a tab-delimited table if you 

select the <named-content content-type="book-spec-emph">Show Data as Table</named-content> option from the left sidebar.</p>
                </boxed-text>
              </sec>
              <sec id="bid.1188">
                <title>Functional Analysis</title>
                <sec id="bid.1189">
                  <title>Variation Functional Class</title>
                  <table-wrap id="bid.1190">
                    <label>5</label>
                    <caption>
                      <title>Function codes for refSNPs in gene features. <xref ref-type="fn" rid="mul8"><sup><italic>a</italic></sup></xref></title>
                    </caption>
                    <table>
                      <tr>
                        <th align="left" valign="top">Functional 

class</th>
                        <th align="left" valign="top">Description</th>
                        <th align="center" valign="top">Database 

code</th>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Locus 

region</td>
                        <td align="left" valign="top">Variation is 

within 2 Kb 5&prime; or 500 bp 3&prime; of a gene feature (on either strand), but the variation is not in the transcript for 

the gene. This class is indicated with an L in graphical summaries.</td>
                        <td align="center" valign="top">1</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Coding</td>
                        <td align="left" valign="top">Variation is in 

the coding region of the gene. This class is assigned if the allele-specific class is unknown. This class is indicated with a 

C in graphical summaries.</td>
                        <td align="center" valign="top">2</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Coding-synon</td>
                        <td align="left" valign="top">The variation 

allele is synonymous with the contig codon in a gene. An allele receives this class when substitution and translation of the 

allele into the codon makes no change to the amino acid specified by the reference sequence. A variation is a synonymous 

substitution if all alleles are classified as contig reference or coding-synon. This class is indicated with a C in graphical 

summaries.</td>
                        <td align="center" valign="top">3</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Coding-nonsynon</td>
                        <td align="left" valign="top">The variation 

allele is nonsynonymous for the contig codon in a gene. An allele receives this class when substitution and translation of 

the allele into the codon changes the amino acid specified by the reference sequence. A variation is a nonsynonymous 

substitution if any alleles are classified as coding-nonsynon. This class is indicated with a C or N in graphical 

summaries.</td>
                        <td align="center" valign="top">4</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">mRNA-UTR</td>
                        <td align="left" valign="top">The variation is 

in the transcript of a gene but not in the coding region of the transcript. This class is indicated by a T in graphical 

summaries.</td>
                        <td align="center" valign="top">5</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Intron</td>
                        <td align="left" valign="top">The variation is 

in the intron of a gene but not in the first two or last two bases of the intron. This class is indicated by an L in 

graphical summaries.</td>
                        <td align="center" valign="top">6</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Splice-site</td>
                        <td align="left" valign="top">The variation is 

in the first two or last two bases of the intron. This class is indicated by a T in graphical summaries.</td>
                        <td align="center" valign="top">7</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Contig-reference</td>
                        <td align="left" valign="top">The variation 

allele is identical to the contig nucleotide. Typically, one allele of a variation is the same as the reference genome. The 

letter used to indicate the variation is a C or N, depending on the state of the alternative allele for the 

variation.</td>
                        <td align="center" valign="top">8</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Coding-exception</td>
                        <td align="left" valign="top">The variation is 

in the coding region of a gene, but the precise location cannot be resolved because of an error in the alignment of the exon. 

The class is indicated by a C in graphical summaries.</td>
                        <td align="center" valign="top">9</td>
                      </tr>
                    </table>
                    <table-wrap-foot>
                      <fn symbol="a" id="mul8">
                        <p>Most gene features are defined by the location of the variation with respect to 

transcript <xref ref-type="other" rid="bid.1287">exon</xref> boundaries. Variations in coding regions, however, have a 

functional class assigned to each allele for the variation because these classes depend on allele sequence.</p>
                      </fn>
                    </table-wrap-foot>
                  </table-wrap>
                  <p>We compute a functional context for sequence variations by inspecting the flanking 

sequence for gene features during the contig annotation process. We are also currently developing a method to do the same 

analysis on RefSeq/GenBank mRNAs.</p>
                  <p><xref ref-type="table" rid="bid.1190">Table 5</xref> defines variation functional 

classes. We base class on the relationship between a variation and any local gene features. When a variation is near a 

transcript or in a transcript interval but not in the coding region, then we define the functional class by the position of 

the variation relative to the structure of the aligned transcript. In other words, a variation may be near a gene 

(<xref ref-type="other" rid="bid.1332">locus</xref> region), in a <xref ref-type="other" rid="bid.1429">UTR</xref> (mrna-utr), in an <xref ref-type="other" rid="bid.1323">intron</xref> (intron), or in a 

splice site (splice site). If the variation is in a coding region, then the functional class of the variation depends on how 

each allele may affect the translated peptide sequence.</p>
                  <p>Typically, one allele of a variation will be the same as the contig (contig reference), 

and the other allele will be either a synonymous change or a nonsynonymous change. In some cases, one allele will be a 

synonymous change, and the other allele will be a nonsynonymous change. If any allele is a nonsynonymous change, then the 

variation is classified as a nonsynonymous variation. Otherwise, the variation is classified as a synonymous variation.</p>
                    <list list-type="bullet">
                      <title>The Four Basic Outcomes When a Variation Is in Coding Sequence</title>
                      <list-item id="bid.1191">
                        <p> The allele is the same as the contig (contig reference) and hence 

causes no change to the translated sequence.</p>
                      </list-item>
                      <list-item id="bid.1192">
                        <p> The allele, when substituted for the reference sequence, yields a 

new codon that encodes the same amino acid. This is termed a synonymous substitution.</p>
                      </list-item>
                      <list-item id="bid.1193">
                        <p> The allele, when substituted for the reference sequence, yields a 

new codon that encodes a different amino acid. This is termed a nonsynonymous substitution.</p>
                      </list-item>
                      <list-item id="bid.1194">
                        <p> A problem with the annotated coding region feature prohibits 

conceptual translation. In this case, we note the variation class as coding, based solely on position.</p>
                      </list-item>
                    </list>
                  <p>Because functional classification is defined by positional and sequence parameters, two 

facts emerge: (<italic>a</italic>) if a gene has multiple transcripts because of alternative splicing, then a variation may 

have several different functional relationships to the gene; and (<italic>b</italic>) if multiple genes are densely packed in a 

contig region, then a variation at a single location in the genome may have multiple, potentially different, relationships to 

its local gene neighbors.</p>
                </sec>
                <sec id="bid.1195">
                  <title>SNP Position in 3D Structure</title>
                  <p>When a SNP results in amino acid sequence change, knowing where that amino acid lies in 

the protein structure is valuable. We provide this information using the following procedure. To find the location of a SNP 

within a particular protein, we attempt to identify similar proteins whose structure is known by comparing the protein 

sequence against proteins from the <xref ref-type="other" rid="bid.1367">PDB</xref> database of known protein 

structures using BLAST. Then, if we find matches, we use the BLAST alignment to identify the amino acid in the protein of 

known structure that corresponds to the amino acid containing the SNP. We store the position of the amino acid on the 3D 

structure that corresponds to the amino acid containing the SNP in the dbSNP table SNP3D.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.1196">
              <title>Resource Integration</title>
              <sec id="bid.1197">
                <title>Links from SNP Records to Submitter Websites</title>
                <p>The SNP database supports and encourages connections between assay records (ss#'s) and 

supplementary data on the submitter's Web site. This connection is made using the <named-content content-type="book-spec-emph">LINKOUT</named-content> field in the SNPassay batch 

header. <xref ref-type="other" rid="bid.1331">LinkOut</xref> URLs are base URLs to which dbSNP can append the 

submitter's ID for the variation to construct a complete URL to the specific data for the record. We provide LinkOut pointers 

in the batch header section of SNP detail reports and in the refSNP report cluster membership section.</p>
              </sec>
              <sec id="bid.1198">
                <title>Links within NCBI</title>
                <p>We make the following connections between refSNP clusters and other NCBI resources during the 

contig annotation process:</p>
                <sec id="bid.1199">
                  <title>LocusLink</title>
                  <p>There are two methods by which we localize variations to known genes: (<italic>a</italic>) 

if a variation is mapped to the genome, we note the variation/gene relationship (<xref ref-type="table" rid="bid.1190">Table 

5</xref>) during functional classification and store the locus_id of the gene in the dbSNP table SNPContigLocusId; and 

(<italic>b</italic>) if the variation does not map to the genome, we look for high-quality blast hits for the variation against 

mRNA sequence. We note these hits with the protein_ID (PID) of the protein (the conceptual translation of the mRNA 

transcript). LocusLink scans this table nightly and updates the table MapLinkPID with the locus_id for the gene when the 

protein is a known product of a gene.</p>
                </sec>
                <sec id="bid.1200">
                  <title>UniSTS</title>
                  <p>When an original submitted SNP record shows a relationship between a SNP and a STS, we 

share the data with dbSTS and establish a link between the SNP and the STS record. We also examine refSNPs for proximity to 

STS features during contig annotation. When we determine that a variation needs to be placed within an STS feature, we note 

the relationship in the dbSNP table SnpInSts.</p>
                </sec>
                <sec id="bid.1201">
                  <title>UniGene</title>
                  <p>The contig annotation pipeline relates refSNPs to UniGene EST clusters based on shared 

chromosomal location. We store Variation/<xref ref-type="other" rid="bid.1424">UniGene cluster</xref> relationships 

in the dbSNP table UnigeneSnp.</p>
                </sec>
                <sec id="bid.1202">
                  <title>PubMed</title>
                  <p>We connect individual submissions to PubMed record(s) of publications cited at the time of 

submission. If you want to view links from PubMed to dbSNP, select <named-content content-type="book-spec-emph">LinkOuts</named-content> as a PubMed query result.</p>
                </sec>
                <sec id="bid.1203">
                  <title>dbMHC</title>
                  <p>dbSNP stores the underlying variation data that define HLA alleles at the nucleotide 

level. The combinations of alleles that define specific HLA alleles are stored in dbMHC. dbSNP points to dbMHC at the 

haplotype level, and dbMHC points to dbSNP at both the haplotype and variation level.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.1204">
              <title>How to Create a Local Copy of dbSNP</title>
              <p>dbSNP is a relational database with about 100 tables. NCBI deploys dbSNP in both MSSQL and Sybase 

environments, and the public can download the full contents of the database from the dbSNP <xref ref-type="other" rid="bid.1295">FTP</xref> site. The following sections will guide you in this process.</p>
              <sec id="bid.1205">
                <title>Schema: The dbSNP Physical Model</title>
                <p>A schema is a necessary part of constructing your own copy of dbSNP because it is a visual 

representation of dbSNP and shows the logical relationship between data in dbSNP. It is available as a printable PDF <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/snp/Sybase/schema/PDM_dbSNP_work.pdf">file</ext-link> from the dbSNP FTP site.</p>
                <p>Data in dbSNP are organized into &ldquo;zones&rdquo; or boxes, depending on the nature of the 

data. Each zone is color coded to allow the viewer to find the data more easily. The current color groupings are <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/outreach/gettingstarted/snpftp/sybase.html">available</ext-link> online. 

The data dictionary currently includes a description of all the tables in dbSNP, tables of columns and their properties, and 

tables of foreign keys in the conceptual model. Foreign keys are not enforced in the physical model because they make it 

harder to load table data asynchronously. In the future, we will add descriptions of individual columns. The data dictionary 

described above is available in rich text format in the file dbSNPdataDictionary.rtf.gz. The data are also available online 

from the dbSNP Web site.</p>
                <sec id="bid.1206">
                  <title>Resources Required for Creating a Local Copy of dbSNP</title>
                  <sec id="bid.1207">
                    <title>Software:</title>
                    
                      <list list-type="bullet">
                        <list-item id="bid.1208">
                          <p><bold>Relational database software</bold>. If you are 

planning to create a local copy of dbSNP, you must first have a relational database server, such as Sybase, Microsoft SQL 

server, or Oracle. dbSNP at NCBI runs on both Sybase and MSSQL servers, but we know of users who have successfully created 

their local copy of dbSNP on Oracle.</p>
                        </list-item>
                        <list-item id="bid.1209">
                          <p><bold>Data loading tool</bold>. Loading data from the 

dbSNP FTP site into a database requires a bulk data-loading tool, which usually comes with a database installation. For 

example, we use the bcp (bulk-copy) utility that comes with Sybase.</p>
                        </list-item>
                        <list-item id="bid.1210">
                          <p><bold>winzip/gzip to decompress FTP files</bold>. 

Complete instructions on how to uncompress *.gz and *.Z files can be found <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Ftp/uncompress.html">here</ext-link>.</p>
                        </list-item>
                        <list-item id="bid.1211">
                          <p><bold>Perl binaries (optional)</bold>. There is a 

sample Perl script that validates whether all rows in the bcp file are loaded successfully into the table. See <named-content content-type="book-spec-emph">validating 

data loading</named-content> in the <italic>Stepwise Procedure for Creating a Copy of dbSNP</italic> for more details.</p>
                        </list-item>
                      </list>
                  </sec>
                  <sec id="bid.1212">
                    <title>Hardware:</title>
                    
                      <list list-type="bullet">
                        <list-item id="bid.1213">
                          <p><bold>Computer platforms/OS</bold>. Databases can be 

maintained on any PC, Mac, or <xref ref-type="other" rid="bid.1426">UNIX</xref> with an Internet connection.</p>
                        </list-item>
                        <list-item id="bid.1214">
                          <p><bold>Disk space</bold>. Currently, a complete dbSNP 

database needs 20 GB. You need to first create a database of at least 20 GB in size. Aside from the database device size of 

20 GB, you will also need another 15 GB to store downloaded data files before loading them into your database. Allowing space 

for other miscellaneous uses, we recommend free disk space of 40 GB.</p>
                        </list-item>
                        <list-item id="bid.1215">
                          <p><bold>Internet connection</bold>. We recommend a 

high-speed connection to download such large database files.</p>
                        </list-item>
                      </list>
                    
                  </sec>
                  <sec id="bid.1216">
                    <title>dbSNP Data Location</title>
                    <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nlm.nih.gov/snp/Sybase/">FTP</ext-link> 

directory contains the schema, data, and SQL statements to create the tables and indices for dbSNP. The /data subdirectory 

contains all data files of dbSNP tables, organized as one file per table. The file name convention is: 

&lt;tablename&gt;.bcp.gz. We use the Sybase bcp tool to generate the data files was one line per table row. Columns of data 

in each line are tab delimited.</p>
                    <p>The /schema subdirectory contains the files dbSNP_table.atx.gz and 

dbSNP_index.atx.gz. These files use standard SQL DDL language to create tables and indexes.</p>
                    <p>There are many utilities available to generate table/index creation statements 

from a database. We use a tool called <italic>atxtract</italic> to generate the files.</p>
                  </sec>
                </sec>
              </sec>
              <sec id="bid.1217">
                <title>Stepwise Procedure for Creating a Local Copy of dbSNP</title>
                <list list-type="arabic">
                <list-item>
                <label><bold>1</bold></label><p><bold> Prepare the local area.</bold> (check available space, etc.)</p></list-item>
                <list-item><label><bold>2</bold></label><p><bold>Download schema files.</bold> <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/snp/Sybase/schema/dbSNP_table.atx.gz">dbSNP_table.atx.gz</ext-link> and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/snp/Sybase/schema/dbSNP_index.atx.gz">dbSNP_index.atx.gz</ext-link>. Save the file in your local directory and decompress the files.
                <boxed-text position="anchor" content-type="website">
                  <title>Hint:</title>
                  <p>For example, on UNIX operating systems, use gunzip to decompress the files: gunzip 
                  dbSNP_table.atx.gz gunzip dbSNP_index.atx.gz to get the files dbSNP_table.atx and dbSNP_index.atx.</p>
                </boxed-text></p></list-item>
                <list-item><label><bold>3</bold></label><p><bold>Create the tables.</bold> Open dbSNP_table.atx file with a text editor. Replace 'snp_sherry' with the name of the database you have created on your server.
                <boxed-text position="anchor" content-type="website">
                  <title>Hint:</title>
                  <p>If you are using Sybase and the SQL server, use the command <named-content content-type="book-spec-emph">
isql -S &lt;servername&gt; -U username -P password -i dbSNP_table.atx</named-content></p>
                </boxed-text></p></list-item>
                <list-item>
                <label><bold>4</bold></label><p><bold>Download the data.</bold> Table data files can be found at: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/snp/Sybase/data">ftp://ftp.ncbi.nih.gov/snp/Sybase/data</ext-link>. There are many pieces of FTP client software available that will facilitate downloading large number of files.
                <boxed-text position="anchor" content-type="website">
                  <title>Hint:</title>
                  <p>Use &ldquo;ftp -i&rdquo; to turn off interactive prompting during multiple file transfers 

to avoid having to hit &ldquo;yes&rdquo; to confirm transfer hundreds of times.</p>
                </boxed-text></p></list-item>
                <list-item>
                <label><bold>5</bold></label><p><bold>Sample FTP protocol.</bold> Type <named-content content-type="book-spec-emph">ftp -i ftp.ncbi.nih.gov</named-content> (use 

&ldquo;anonymous&rdquo; as the user name and your email as a password.) Type <named-content content-type="book-spec-emph">cd snp/Sybase/data</named-content>. Type <named-content content-type="book-spec-emph">ls</named-content> (if you see a list of &lt;tablename&gt;.bcp.gz files, you are in the right directory). Type <named-content content-type="book-spec-emph">binary</named-content> (to set binary transfer mode). Type <named-content content-type="book-spec-emph">mget *.gz</named-content> (to initiate transfer. Depending on the speed 

of connection, this may take more than 1 hour because the total transfer size currently is 1.5 GB and growing). To decompress 

the *.gz files, type <named-content content-type="book-spec-emph">gunzip *.gz</named-content> (currently, the total size of the uncompressed bcp files is more than 10 

GB).</p></list-item>
                <list-item><label><bold>6</bold></label><p><bold>Load data.</bold> Use the data-loading tool of your database server (e.g., bcp for 

Sybase).
                <boxed-text position="anchor" content-type="website">
                  <title>Hint:</title>
                  <p>Because of the sheer volume of data, our tests to load data and then create indexes on a 

Sun Sparc Sybase server using &ldquo;fast bcp&rdquo; took about 8 hours.</p>
                </boxed-text></p></list-item>
                <list-item><label><bold>7</bold></label><p><bold>Create indices.</bold> Open the dbSNP_index.atx file with a text editor and 

replace &ldquo;snp_sherry&rdquo; with the local name of the database as you have created it on your server. If you are using 

Sybase and the SQL server, use the command <named-content content-type="book-spec-emph">isql -S servername -U username -P password -i dbSNP_index.atx</named-content></p></list-item>
                <list-item><label><bold>8</bold></label><p><bold>Use scripts to automate data loading.</bold> In the dbSNP FTP snp/Sybase area, 

there is a sample UNIX C shell script cmd.loaddata that will combine tasks of creating tables, loading data, and creating 

indexes.</p></list-item>
                <list-item><label><bold>9</bold></label><p><bold>Validate data loading.</bold> In the dbSNP FTP snp/Sybase area, there is a sample 

Perl script that validates whether all data are loaded successfully. For each table, it compares the number of rows in the 

bcp file with the number of rows in the newly loaded table and reports discrepancies.</p></list-item>
                <list-item><label><bold>10</bold></label><p><bold>Data integrity (creating a partial local copy of dbSNP).</bold> dbSNP is a 

relational database. Each table has either a unique index or a primary key. Foreign keys are not reinforced. There are 

advantages and a disadvantage to this approach. The advantages are: it makes it easy to drop and re-create the table using 

dbSNP_table.atx, which makes it possible to create a partial local copy of dbSNP. For example, if you are interested only in 

the original submitted SNP and their population frequencies and not in their map locations on NCBI genome contigs or GenBank 

Accession numbers (both are huge tables), then these tables can be skipped (i.e., SNPContigLoc and MapLink). Of course, to 

select tables for a particular query, the contents of each table and the dbSNP entity relationship (ER) diagram need to be 

understood. The disadvantage of unreinforced references is that either the stored procedures or the external code needs to be 

written to ensure the referential integrity.</p></list-item>
</list>
              </sec>
            </sec>
            <sec id="bid.2001">
              <title>Appendix 1. dbSNP report formats.</title>
              <sec id="bid.2002">
                <title>ASN.1 Brief and Full Versions</title>
                <p>
	The file docsum.asn is the ASN structure definition file for both formats of the data. The 00readme file, located in the /specs subdirectory of the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/snp/">dbSNP FTP</ext-link> site, includes comments defining the sections common to both the full and the brief formats and the sections included in only the full version. ASN.1 text or binary output can be converted into one or more of the following formats: flatfile, FASTA, docsum, chromosome report, RS/SS, and XML. To do this, request ASN.1 output from either the dbSNP FTP site or the dbSNP batch query pages, and use dstool (located in the &ldquo;bin&rdquo; directory of the dSNP FTP site) to locally convert the output into as many alternative formats as needed.
	</p>
              </sec>
              <sec id="bid.2003">
                <title>XML Brief and Full Versions</title>
                <p>The XML DTD comprises three files located in the /specs subdirectory of the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/snp/">SNP FTP</ext-link> site: NSE.dtd, NSE.mod, and NCBI_Entity.mod (type definitions). The file NSE.dtd is simply a header file that bundles the two module files NCBI_Entity and NSE (NCBI SNP Exchange). The real structure definitions for the SNP data are in NSE.mod. This file is a direct conversion of the ASN.1 definition file (docsum.asn) into XML.  All three XML files described above must be included in your download to parse the dbSNP XML data dumps successfully. The 00readme file located in the /specs subdirectory includes comments defining the sections common to both the full and the brief formats and the sections included in only the full version.</p>
              </sec>
              <sec id="bid.2004">
                <title>FASTA: ss and rs</title>
                <p>The FASTA report format provides the flanking sequence for each report of variation in dbSNP, as well as all submitted sequences that have no variation. ss FASTA contains all submitted SNP sequences in FASTA format, whereas rs FASTA contains all the reference SNP sequences in FASTA format.  The FASTA data format is typically used for sequence comparisons using BLAST. BLAST SNP is useful for conducting a few sequence comparisons in the FASTA format, wherease multiple FASTA sequence comparisons will require the construction of a local BLAST database of FASTA formatted data and the installation of a local stand-alone version of BLAST. </p>
              </sec>
              <sec id="bid.2005">
                <title>rs docsum Flatfile</title>
                <p>The rs docsum flatfile report is generated from the ASN.1 datafiles and is provided in the 
	files &ldquo;/ASN1_flat/ds_flat_chXX.flat&rdquo;. As with all of the large report dumps, files are generated per chromosome (chXX in file name). Because flatfile reports are compact, they will not provide you with as much information as the ASN.1 binary report, but they are useful for scanning human SNP data manually because they provide detailed information at a glance. A full description of the information provided in the rs docsum flatfile format is available in the 00readme file, located in the SNP directory of the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/snp/">SNP FTP</ext-link> site.</p>
              </sec>
              <sec id="bid.2006">
                <title>Chromosome Reports</title>
                <p>The chromosome reports format provides an ordered list of RefSNPs in approximate chromosome coordinates.  Chromosome reports is a small file to download but contains a great deal of information that might be helpful in identifying SNPs useful as markers on maps or contigs because the coordinate system used in this format is the same as that used for the NCBI genome Map Viewer.  It should also be mentioned that the chromosome reports directory might contain the multi/ file and/or the noton/ files.  These files are lists (in chromosome report format) of SNPs that hit multiple chromosomes in the genome and those that did not hit any chromosomes in the genome, respectively. A full description of the information provided in the chromosome reports format is available in the 00readme file, located in the SNP directory of the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/snp/">SNP FTP</ext-link> site.</p>
              </sec>
            </sec>
            <sec id="bid.2007">
              <title>Appendix 2. Rules and methodology for mapping.</title>
              <p>A cycle of flank sequence masking and MegaBLAST alignment to the NCBI genome assembly of an organism is initiated either by the appearance of FASTA-formatted genome sequence for a new build of the assembly or by the significant accrual of newly submitted SNP data for that organism. </p>
              <sec id="bid.2008">
                <title>Flank Sequence Masking</title>
                <p>We prepare FASTA-formatted sequence for unclustered ss assays of a particular organism and run RepeatMasker (with the mixed case option set) on the sequence, using the appropriate ALU-repeat library. For human, we use <preformat>repmask -q -xsmall</preformat> and for mouse we use <preformat>repmask -m -q -xsmall</preformat><named-content content-type="book-spec-emph">dust</named-content> is used if no library of repetitive elements is available for the organism. The mixed case sequence is then loaded into the database, replacing the original submitted flank.  The SubSNPRepMask table is updated with repeat masker output that describes the nature and location of the repeat, and a set of mixed-case FASTA is prepared for the entire refSNP set.</p>
              </sec>
              <sec id="bid.2009">
                <title>Organism-specific Genome Mapping</title>
                <p>Both refSNP and subSNP FASTA sets are aligned to the genome assembly using MegaBLAST, and the SNP position is computed from the list of ungapped alignments returned with the <named-content content-type="book-spec-emph">-D1</named-content> option. We filter the alignments into either a high-threshold category (95 % of SNP flank aligning with fewer than sixth mismatches) or a low-threshold category (75% of the alignment length with fewer than 3% mismatches).  The rest of the MegaBLAST alignments are dropped.  MegaBLAST parameters are set to a default position with the exception of <named-content content-type="book-spec-emph">-F m,</named-content> which suppresses seeding hits but allows extension through regions of lowercase sequence.</p>
                <p>We also perform tailored BLAST algorithms to specialized subsets of the database that are known to fail or perform poorly in the standard pathway. Short sequences of 75 bases or fewer are BLASTed with word 
size = 22.  To capture hits in cases where the variation length might break the alignment, each flank of microsatellites or indels are BLASTed independently and paired up in post processing to provide the SNP position. We perform high-stringency BLASTs for many of the heavily masked SNPs using wordsize = 40 and 99% quality in any alignment.  This is an attempt to capture low-multiplicity hits of repetitive sequence without overburdening the BLAST algorithm with a large number of seeded alignments.</p>
                <p>To reduce hit multiplicity in cases where SNPs align several times on the genome, we examine the relative hit quality of multiple hits. Hits are discarded when the mismatch count is greater than the minimum mismatch count. The mismatch count parameter is currently set to 3.  We BLAST refSNPs and subSNPs against GenBank mRNA, RefSeq mRNA, and GenBank clone accessions by a similar procedure.</p>
                <p>We send the output from the genome alignment (or clone alignment where a genome is not available) through a clustering procedure that assigns rs IDs to unclustered new submissions.  In some cases, this clustering procedure will re-cluster existing refSNP clusters based on colocation on the genome. After clustering, we update BLAST output to synchronize the ss IDs and the re-clustered rs IDs to the associated refSNP cluster defined in the database. Location data are loaded into the database, wherease cumulative hit data for each mapped SNP are computed and loaded to the database. Finally, we update SNPFlankStatus.

</p>
              </sec>
            </sec>
            <sec id="bid.2010">
              <title>Appendix 3. 3D structure neighbor analysis.</title>
              <p>When a protein is known to have a structure neighbor, dbSNP projects the RefSNPs located in that protein sequence onto sequence structures. </p>
              <p>First, contig annotation results provide the SNP ID (snp_id), protein accession (protein_acc), contig and SNP amino acid residue (residue), as well as the amino acid position (aa_position) for a particular RefSNP.  These data can be found in the dbSNP table, SNPContigLocusId.  FASTA sequence is then obtained for each protein accession using the program idfetch, with the command line parameters set to: <preformat>-t 5 -dp -c 1 -q</preformat></p>
              <p>We BLAST these sequences against the PDB database using blastall with the command line parameters set to: 
<preformat>-p blastp -d pdb -i protein.fasta -o result.blast -e 0.0001 -m 3 -I T -v 1 -b 1</preformat></p>
              <p>Each SNP position in the protein sequence is used to determine its corresponding amino acid and amino acid position in the 3D structure from the BLAST result. These data are stored in the SNP3D table.</p>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.337" book-part-type="chapter" book-part-number="6">
          <book-part-meta>
            <title-group>
              <title>The Gene Expression Omnibus (GEO): A Gene Expression and Hybridization Repository</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Edgar</surname>
                  <given-names>Ron</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Lash</surname>
                  <given-names>Alex</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch6d1"/>
            <abstract>
              <title>Summary</title>
              <p>The Gene Expression Omnibus (<abbrev>GEO</abbrev>) project was 
	                    initiated at NCBI in 1999 in response to the growing demand for a 
	                    public repository for data generated from high-throughput microarray 
	                    experiments. <abbrev>GEO</abbrev> has a flexible and open design 
	                    that allows the submission, storage, and retrieval of many types 
	                    of data sets, such as those from high-throughput gene expression, 
	                    genomic hybridization, and antibody array experiments. 
	                    <abbrev>GEO</abbrev> was never intended to replace lab-specific 
	                    gene expression databases or laboratory information management 
	                    systems (<abbrev>LIMS</abbrev>), both of which usually cater 
	                    to a particular type of data set and analytical method. Rather, 
	                    <abbrev>GEO</abbrev> complements these resources by acting as 
	                    a central, molecular abundance&ndash;data distribution hub. 
	                    <abbrev>GEO</abbrev> is available on the World Wide Web at 
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/geo">http://www.ncbi.nih.gov/geo</ext-link>.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.338">
              <title>Site Description</title>
              <p>High-throughput hybridization array- and sequencing-based experiments have become increasingly common in molecular biology laboratories in recent years (<xref ref-type="bibr" rid="bid.398">1</xref>&ndash;<xref ref-type="bibr" rid="bid.401">4</xref>). These techniques are used to measure the molecular abundance of mRNA, genomic DNA, and proteins in absolute or relative terms.  The main attraction of these techniques is their highly parallel nature; large numbers of simultaneous molecular sampling events are performed under very similar conditions. This means that time and resources are saved, and complex biological systems can be represented in a more holistic manner. Furthermore, the development of tissue arrays means that it is possible to analyze, in parallel, the gene expression of large numbers of tumor tissue samples from patients at different stages of cancer development (<xref ref-type="bibr" rid="bid.402">5</xref>).</p>
              <p>Because of the plethora of measuring techniques for molecular abundance in use, our primary goal in creating the Gene Expression Omnibus (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo">GEO</ext-link>) was to cover the broadest possible spectrum of these techniques and remain flexible and responsive to future trends, rather than choosing only one of these techniques or setting rigid requirements and standards for entry. In taking this approach, however, we recognize that there are obvious, inherent limitations to functionality and analysis that can be provided on such heterogeneous data sets.</p>
              <p>This chapter is both more current and more detailed than the previous literature report on GEO (<xref ref-type="bibr" rid="bid.403">6</xref>). However, more detailed descriptions, tools, and news releases are available on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo">GEO Web site</ext-link>.</p>
            </sec>
            <sec id="bid.339">
              <title>Design and Implementation</title>
              <fig id="bid.340">
                <label>1</label>
                <caption>
                  <title>GEO design.</title>
                  <p>The entity&ndash;relationship diagram for GEO.</p>
                </caption>
                <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6f1" mime-subtype="gif"/>
              </fig>
              <fig id="bid.342">
                <label>2</label>
                <caption>
                  <title>GEO implementation example.</title>
                  <p>An actual example of three samples referencing one platform and contained in a single series.</p>
                </caption>
                <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6f2" mime-subtype="gif"/>
              </fig>
              <table-wrap id="bid.341">
                <label>1</label>
                <caption>
                  <title>Entity prefixes, types, and subtypes in the GEO database.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Accession prefix</th>
                    <th align="left" valign="top">Entity type</th>
                    <th align="left" valign="top">Subtype</th>
                    <th align="left" valign="top">Description</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GPL</td>
                    <td align="left" valign="top">Platform</td>
                    <td align="left" valign="top">Commercial nucleotide array</td>
                    <td align="left" valign="top">Commercially available nucleotide hybridization array</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Commercial tissue array</td>
                    <td align="left" valign="top">Commercially available tissue array</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Commercial antibody array</td>
                    <td align="left" valign="top">Commercially available antibody array</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Non-commercial nucleotide array</td>
                    <td align="left" valign="top">Nucleotide array that is not commercially available</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Non-commercial tissue array</td>
                    <td align="left" valign="top">Tissue array that is not commercially available</td>
                  </tr>
                  <tr content-type="rowsep-1">
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Non-commercial antibody array</td>
                    <td align="left" valign="top">Antibody array that is not commercially available</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GSM</td>
                    <td align="left" valign="top">Sample</td>
                    <td align="left" valign="top">Dual channel</td>
                    <td align="left" valign="top">Dual mRNA target sample hybridization</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Single channel</td>
                    <td align="left" valign="top">Single mRNA target sample hybridization</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Dual channel genomic</td>
                    <td align="left" valign="top">Dual DNA target sample hybridization, e.g., array CGH</td>
                  </tr>
                  <tr content-type="rowsep-1">
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">SAGE</td>
                    <td align="left" valign="top">Serial analysis of gene expression</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GSE</td>
                    <td align="left" valign="top">Series</td>
                    <td align="left" valign="top">Time&ndash;course</td>
                    <td align="left" valign="top">Time&ndash;course experiment, e.g., yeast cell cycle</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Dose&ndash;response</td>
                    <td align="left" valign="top">Dose&ndash;response experiment, e.g., response to drug dosage</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Other ordered</td>
                    <td align="left" valign="top">Ordered, but unspecified</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">Other</td>
                    <td align="left" valign="top">Unordered</td>
                  </tr>
                </table>
              </table-wrap>
              <p>The three principle components (or entities) of GEO are modeled after the three organizational units common to high-throughput gene expression and array-based methodologies. These entities are called <italic>platforms</italic>, <italic>samples</italic> and <italic>series</italic> (<xref ref-type="fig" rid="bid.340">Figure 1</xref>; <xref ref-type="table" rid="bid.341">Table 1</xref>).  A <italic>platform</italic> is, essentially, a list of probes that defines what set of molecules may be detected in any experiment using that platform. A <italic>sample</italic> describes the set of molecules that are being probed and references a single platform used to generate molecular abundance data. Each sample has one, and only one, parent platform that must be defined previously. A <italic>series</italic> organizes samples into the meaningful data sets that make up an experiment and are bound together by a common attribute.</p>
              <p>The GEO repository is a relational database, which required that some fundamental implementation decisions were made:</p>
              <p>(<italic>a</italic>) GEO does not store raw hybridization-array image data, although &ldquo;reference&rdquo; images of less than 100 Kb may be stored. This decision was based on an assertion that most users of the data within the GEO repository would not be equipped to use raw image data (<xref ref-type="bibr" rid="bid.404">7</xref>); although some may disagree, this means that repository storage requirements are reduced roughly by a factor of 20.</p>
              <p>(<italic>b</italic>) We decided to use a different storage mechanism for data and metadata. Within the GEO repository, metadata are stored in designated fields within the database table. However, data from the entire set of probe attributes (for each platform) and molecular abundance measurements (for each sample) are stored as a single, text-compressed <xref ref-type="other" rid="bid.1251">BLOB</xref>. This mode of data storage allows great flexibility in the amount and type of information stored in this BLOB. It allows any number of supplementary attributes or measurements to be provided by the submitter, including optional or submitter-defined information. For example, a microarray (the platform) consisting of several thousand spots (the probes) would have a set of probe attributes, some of which are defined by GEO. The GEO-defined attributes include, for each probe, the position within the array and biological reagent contents of each probe such as a <xref ref-type="other" rid="bid.1299">GenBank</xref> Accession number, open reading frame (ORF) name, and clone identifier, as well as any number of submitter-defined columns. As another example, the set of probe-target measurements given in the data from a sample may contain the final, relevant abundance value of the probe defined in its platform, as well as any other GEO-defined (e.g., raw signal, background signal) and submitter-defined data.</p>
              <p>Once a platform, sample, or series is defined by a submitter, an Accession number (i.e., a unique, stable identifier) is assigned (<xref ref-type="fig" rid="bid.342">Figure 2</xref>). Whether a GEO Accession number refers to a platform, sample, or series can be understood by the Accession number &ldquo;prefix&rdquo;. Platforms have the prefix GPL, samples have the prefix GSM, and series have the prefix GSE.</p>
            </sec>
            <sec id="bid.343">
              <title>Retrieving Data</title>
              <fig id="bid.355">
                <label>3</label>
                <caption>
                  <title>GEO retrieval statistics.</title>
                  <p>Daily usage statistics evaluated over a 4-week period January 24 to February 20, 2002. Web server <italic>GET</italic> (<italic>blue</italic>) and <italic>POST</italic> (<italic>magenta</italic>) calls are evaluated for URL <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo/query/acc.cgi">http://www.ncbi.nlm.nih.gov/geo/query/acc.cgi</ext-link>. <italic>GET</italic> calls correspond roughly to links being followed from other Web pages, most likely following Entrez ProbeSet queries. <italic>POST</italic> calls roughly correspond to direct queries by Accession number. The spike of activity seen from January 29 to January 31 represents retrievals by one IP address and most likely represent an automated &ldquo;Web crawler&rdquo; pull.</p>
                </caption>
                <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6f3" mime-subtype="gif"/>
              </fig>
              <boxed-text id="bid.344" position="margin">
                <label>1</label>
                <title>GEO Web site Accession Display tool.</title>
                <p>It is very easy to use the <named-content content-type="book-spec-emph">Accession Display</named-content> tool:
                <list list-type="arabic">
                  <list-item id="bid.345">
                    <label>1</label>
                    <p>Type in a valid public or private<xref ref-type="fn" rid="bid.351"><sup><italic>a</italic></sup></xref> GEO Accession number in the top <named-content content-type="book-spec-emph">GEO accession</named-content> box.</p>
                  </list-item>
                  <list-item id="bid.346">
                    <label>2</label>
                    <p>Select desired display options.</p>
                  </list-item>
                  <list-item id="bid.347">
                    <label>3</label>
                    <p>Press the <named-content content-type="book-spec-emph">Go</named-content> button.</p>
                  </list-item>
                </list>
                </p>
                <p>Three types of display options are currently available:
                <list list-type="bullet">
                  <list-item id="bid.348">
                    <p><bold>Scope</bold> allows you to display the GEO accession(s) that you want to target for display. You may display the GEO accession, which is typed into the <named-content content-type="book-spec-emph">GEO accession</named-content> box itself (<named-content content-type="book-spec-emph">Self</named-content>), or any (<named-content content-type="book-spec-emph">Platform</named-content>, <named-content content-type="book-spec-emph">Samples</named-content>, or <named-content content-type="book-spec-emph">Series</named-content>) or all (<named-content content-type="book-spec-emph">Family</named-content>) of the accessions related to an accession. GEO platforms (GPL prefix) may have related samples and, through those related samples, related series. GEO samples (GSM prefix) will always have one related platform and may have multiple, related series. GEO series (GSE prefix) will have at least one related sample and, through those related samples, will have at least one related platform. The <named-content content-type="book-spec-emph">Family</named-content> setting will retrieve all accessions (of different types) related to self (including self).</p>
                  </list-item>
                  <list-item id="bid.349">
                    <p><named-content content-type="book-spec-emph">Format</named-content> allows you to display the GEO accession in human-readable, linked HTML form or in machine-readable, SOFT form (<xref ref-type="boxed-text" rid="bid.354">Box 2</xref>).</p>
                  </list-item>
                  <list-item id="bid.350">
                    <p><named-content content-type="book-spec-emph">Amount</named-content> allows you to control the amount of data that you will see displayed. <named-content content-type="book-spec-emph">Brief</named-content> displays the accession's metadata only. <named-content content-type="book-spec-emph">Quick</named-content> displays the accession's metadata and the first 20 rows of its data set. <named-content content-type="book-spec-emph">Full</named-content> displays the accession's metadata and the full data set. <named-content content-type="book-spec-emph">Data</named-content> omits the accession's metadata, showing only the links to other accessions as well as the full data set.</p>
                  </list-item>
                </list>
                </p>
                <fn-group>
                  <fn id="bid.351" symbol="a">
                    <p>To view one's own private, currently unreleased accessions, login with username and password at the bottom <named-content content-type="book-spec-emph">login</named-content> bar.</p>
                  </fn>
                </fn-group>
              </boxed-text>
              <p>A GEO Accession number is required to retrieve data from the GEO repository database (<xref ref-type="fig" rid="bid.355">Figure 3</xref>). An Accession number may be acquired in any number of ways, including direct reference, such as from a publication citing data deposited to GEO, or through a query interface, such as through <xref ref-type="other" rid="bid.1353">NCBI</xref>'s <xref ref-type="other" rid="bid.1282">Entrez</xref> ProbeSet interface (covered below).</p>
              <p>Given a valid GEO Accession number, the Accession Display tool available on the GEO Web site provides a number of options for the retrieval and display of repository contents (see <xref ref-type="boxed-text" rid="bid.344">Box 1</xref>).</p>
            </sec>
            <sec id="bid.353">
              <title>Depositing Data</title>
              <fig id="bid.352">
                <label>4</label>
                <caption>
                  <title>GEO submission statistics.</title>
                  <p>Cumulative individual sample measurements submitted to GEO are shown. Data are presented by quarter since operations began on July 25, 2000.</p>
                </caption>
                <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6f4" mime-subtype="gif"/>
              </fig>
              <boxed-text id="bid.354" position="margin">
                <label>2</label>
                <title>SOFT.</title>
                <p>Simple Omnibus Format in Text (SOFT) is a line-based, ASCII text format that allows for the representation of multiple GEO platforms, samples, and series in one file. In SOFT, metadata appear as label-value pairs and are associated with the tab-delimited text tables of platforms and samples. SOFT has been designed for easy manipulation by readily available line-scanning software and may be quite readily produced from, and imported into, spreadsheet, database, and analysis software. More information about SOFT and the submission process is available from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo">GEO Web site</ext-link>.</p>
              </boxed-text>
              <p>There are several formats in which data can be deposited and retrieved from GEO. For deposit: (1) a file containing an ASCII-encoded text table of data can be uploaded, and metadata fields can be interactively entered through a series of Web forms; or (2) both data and metadata for one or more platforms, samples, or series can be uploaded directly in a format we call Simple Omnibus Format in Text, or SOFT (<xref ref-type="boxed-text" rid="bid.354">Box 2</xref>).</p>
              <p>Interactive and direct modes of communication are available for new data submissions and updating data submissions. The interactive Web form route is straightforward and most suited for occasional submissions of a relatively small number of samples. Bulk submissions of large data sets may be rapidly incorporated into GEO via direct deposit of SOFT formatted data.</p>
              <p>Submissions may be held private for a maximum of 6 months; this policy allows data release concordant with manuscript publication. Such submissions are given a final Accession number at the time of submission, which may be quoted in a publication.</p>
              <p>Currently, submissions are validated according to a limited set of criteria (see the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo">GEO Web site</ext-link> for more details). Submissions are scanned by our staff to assure that the submissions are organized correctly and include meaningful information. It is entirely up to the submitter to make the data useful to others.</p>
              <p>A quarterly, cumulative graph of the number of individual molecular abundance measurements in public submissions made through the first quarter of 2002 is shown in <xref ref-type="fig" rid="bid.355">Figure 4</xref>.</p>
            </sec>
            <sec id="bid.356">
              <title>Search and Integration</title>
              <table-wrap id="bid.358">
                <label>2</label>
                <caption>
                  <title>Entrez ProbeSet fields.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Field name</th>
                    <th align="left" valign="top">Description</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Accession</td>
                    <td align="left" valign="top">GEO accession identifier</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Author</td>
                    <td align="left" valign="top">Author of GEO sample</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">CloneID</td>
                    <td align="left" valign="top">Clone identifier of GEO sample's platform</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Country</td>
                    <td align="left" valign="top">Country of GEO sample's submitter</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Email</td>
                    <td align="left" valign="top">email of submitter</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GBAcc</td>
                    <td align="left" valign="top">GenBank Accession of GEO sample's platform</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Institute</td>
                    <td align="left" valign="top">Institute of GEO sample's submitter</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Keyword</td>
                    <td align="left" valign="top">Keyword of GEO sample</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">ORF</td>
                    <td align="left" valign="top">Open reading frame (ORF) designation of GEO sample's platform</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Organism</td>
                    <td align="left" valign="top">Organism of GEO sample and its parent taxonomic nodes</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">RefSeq</td>
                    <td align="left" valign="top">RefSeq accession of GEO sample's platform</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">SAGEtag</td>
                    <td align="left" valign="top">Serial analysis of gene expression (SAGE) 10-bp tag of GEO sample</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Subtype</td>
                    <td align="left" valign="top">Subtype of GEO sample</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Target ref</td>
                    <td align="left" valign="top">Target reference of GEO sample</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Target src</td>
                    <td align="left" valign="top">Target source of GEO sample</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Text Word</td>
                    <td align="left" valign="top">Word from description of GEO sample or sample's platform, and word from the titles of sample and its platform</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Title</td>
                    <td align="left" valign="top">Titles of GEO sample and its platform</td>
                  </tr>
                </table>
              </table-wrap>
              <boxed-text id="bid.357" position="margin">
                <label>3</label>
                <title>Entrez ProbeSet indexing and linking process.</title>
                <p>The basic unit (defined by a unique identifier, or UID, in Entrez parlance) in Entrez ProbeSet is the GEO sample, fused with its affiliated platform and series information. The indexing process iterates through all platforms in the GEO database, extracting metadata and the data table and fishing for any sequence-based identifiers such as GenBank Accession, ORFs, Clone IDs, or SAGE tags. Each sample belonging to that platform is in turn assigned a new UID and indexed with the above platform information plus any related series metadata (<xref ref-type="table" rid="bid.358">Table 2</xref>).</p>
                <p>GenBank Accessions, PubMed references, and taxonomy information are also linked to the appropriate Entrez databases for cross-reference and appear in the <named-content content-type="book-spec-emph">Links</named-content> section of the display. Neighbors (related intra-Entrez database links) are generated for UIDs sharing the same GEO platform or series.</p>
              </boxed-text>
              <p>Extensive indexing and linking on the data in GEO are performed periodically and can be queried through Entrez ProbeSet (<xref ref-type="boxed-text" rid="bid.357">Box 3</xref>). Many users of Entrez will recognize this interface as similar to that of other popular NCBI resources such as <xref ref-type="other" rid="bid.1387">PubMed</xref> and GenBank. As with any Entrez database, a <xref ref-type="other" rid="bid.1253">Boolean</xref> phrase may be entered and restricted to any number of supported attribute fields (<xref ref-type="table" rid="bid.358">Table 2</xref>). Matches are linked to the full GEO entry as well as to other Entrez databases, currently Nucleotide, Taxonomy, and PubMed, as well as related Entrez ProbeSet entries. (See <xref ref-type="other" rid="bid.588">Chapter 15</xref> for more details.) Entrez ProbeSet is accessible through the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=geo">Entrez Web site</ext-link> as one of the drop-down menu selections.</p>
            </sec>
            <sec id="bid.359">
              <title>Example of Retrieving Data</title>
              <p>Because samples are oftentimes organized into meaningful data sets within series, an example of retrieving a series and all the data of its associated samples and platform(s) is illustrative of the retrieval capabilities of the GEO Web site. For this example, to select a series of interest, we scan down a list of series in the GEO repository. However, to arrive at our series of interest, we could have just as well performed an Entrez ProbeSet query and followed GEO accession links to a sample and then to its related series, or followed links from PubMed to Entrez ProbeSet, and then to GEO. A step-by-step example of selecting a series of data and retrieving the data for this series from the GEO repository follows:
              <list list-type="arabic">
                <list-item id="bid.360">
                  <label>1</label>
                  <p>Select the linked number of public series from the table of Repository Contents given on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/geo">GEO homepage</ext-link>: <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6df1" mime-subtype="gif"/></p>
                </list-item>
                <list-item id="bid.361">
                  <label>2</label>
                  <p>Scan down the list of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo/query/browse.cgi?view=series">public series</ext-link> in the GEO repository and select GSE27, on sporulation in yeast: <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6df2" mime-subtype="gif"/></p>
                </list-item>
                <list-item id="bid.362">
                  <label>3</label>
                  <p>The description of GSE27 on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE27">Accession Display</ext-link> allows a summary assessment of the data. The data set can be downloaded in SOFT format: <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6df3" mime-subtype="gif"/></p>
                </list-item>
                <list-item id="bid.363">
                  <label>4</label>
                  <p>In the <named-content content-type="book-spec-emph">Accession Display</named-content> options, select <named-content content-type="book-spec-emph">Scope:</named-content>Family, <named-content content-type="book-spec-emph">Format:</named-content>SOFT, and <named-content content-type="book-spec-emph">Amount:</named-content>Full and then press the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE27&amp;targ=all &amp;form=text&amp;view=full">go</ext-link> button: <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6df4" mime-subtype="gif"/></p>
                </list-item>
                <list-item id="bid.364">
                  <label>5</label>
                  <p>A browser dialog states that it took 19 seconds to download the 5 MB SOFT file of data and metadata for one series (GSE27), seven samples (GSM992 to GSM1000), and one platform (GPL67). <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6df5" mime-subtype="gif"/></p>
                </list-item>
              </list>
              </p>
            </sec>
            <sec id="bid.365">
              <title>Future Directions</title>
              <fig id="bid.366">
                <label>5</label>
                <caption>
                  <title>Constellation of NCBI gene expression resources.</title>
                  <p>Anticipated development of gene expression resources at NCBI is shown. <italic>Blue spheres</italic> represent Web sites, <italic>orange cylinders</italic> represent primary NCBI databases, <italic>green cylinders</italic> represent secondary databases, and <italic>yellow cylinders</italic> represent tertiary NCBI interface databases. <italic>Arrows</italic> represent data flow, and <italic>lines</italic> represent Web site links.</p>
                </caption>
                <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch6f5" mime-subtype="gif"/>
              </fig>
              <p>The GEO resource is under constant development and aims to improve its indexing, linking, searching, and display capabilities to allow vigorous data mining. Because the data sets stored within GEO are from heterogeneous techniques and sources, they are not necessarily comparable. For this reason, we have defined a ProbeSet to be a collection of GEO samples that contains comparable data. The selection of GEO samples into ProbeSets is necessary before integrating data in the GEO repository into other NCBI resources (see <xref ref-type="other" rid="bid.588">Chapter 15</xref>, <xref ref-type="other" rid="bid.610">Chapter 16</xref>, and <xref ref-type="other" rid="bid.1562">Chapter 20</xref>), as well as for developing useful display tools for these data (<xref ref-type="fig" rid="bid.366">Figure 5</xref>).</p>
            </sec>
            <sec id="bid.367">
              <title>Frequently Asked Questions</title>
              <list list-type="simple">
                <list-item id="bid.368">
                  <label>1</label>
                  <p>How do I submit my data?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.369">
                  <p>To submit data, an identity within the GEO resource must first be established. On first login, authentication and contact information must be provided. Authentication information (username and password) is used to identify users making submissions and updates to submissions. Contact information is displayed when repository contents are retrieved by others. This information is entered only once and can be updated at any time.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.370">
                  <label>2</label>
                  <p>Is there a &ldquo;hold until date&rdquo; feature in GEO?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.371">
                  <p>Yes. This feature allows a submitter to submit data to GEO and receive a GEO Accession number before the data become public. There is currently a 6-month limit to this hold period. All private data are publicly released eventually.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.372">
                  <label>3</label>
                  <p>What kinds of data will GEO accept?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.373">
                  <p>GEO was designed around the common features of most of the high-throughput gene expression and array-based measuring technologies in use today. These technologies include hybridization filter, spotted microarray, high-density oligonucleotide array, serial analysis of gene expression, and Comparative Genomic Hybridization (CGH) and protein (antibody) arrays but may be expanded in the future.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.374">
                  <label>4</label>
                  <p>Does GEO archive raw data images?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.375">
                  <p>No. However, a reference image will be optionally accepted (limited to 100 Kb in size in JPEG format). In combination with optional references to horizontal and vertical coordinates, this image can be used to provide the user of the data with a qualitative assessment of the data.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.376">
                  <label>5</label>
                  <p>Are there any Quality Assurance (QA) measurements that are required by GEO?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.377">
                  <p>Not at this time. These requirements may be added in the future.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.378">
                  <label>6</label>
                  <p>How can I submit QA measurements to GEO?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.379">
                  <p>QA measurements are currently optional. If QA measurements are performed at the image-analysis step, these can be submitted as additional sample data.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.380">
                  <label>7</label>
                  <p>How can I make corrections to data that I have already submitted?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.381">
                  <p>By logging in with a username and password, an option to update a previous submission or your contact information is given. Accession updates can also be made through a link from the Accession Display after logging in. Updating the data of an already existing and valid GEO Accession number will cause a new version of that data element to be created. Alterations of metadata will not create a new version. All of the various versions of a data element will remain in the database.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.382">
                  <label>8</label>
                  <p>How are submitters authenticated?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.383">
                  <p>In their first submission to GEO, submitters will be asked to select a username and password. This username and password can be used to submit additional data in the future without reentering contact information, as well as to authenticate the submitter when updating or resubmitting data elements under an existing GEO Accession number.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.384">
                  <label>9</label>
                  <p>How do I get data from GEO?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.385">
                  <p>You need not login to retrieve data. All the data are available for downloading. NCBI places no restrictions on the use of data whatsoever but does not guarantee that no restrictions exist from others. You should carefully read NCBI's data disclaimer, available on the GEO Web site.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.386">
                  <label>10</label>
                  <p>What kind of queries and retrievals will be possible in GEO?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.387">
                  <p>Currently, there are three ways to retrieve submissions. One way is by entering a valid GEO Accession number into the query box on the header bar of this page; this will take you to the Accession Display. Another is to use the platform, sample, and series lists, located on the GEO Statistics page. Sophisticated queries of GEO data and linking to other Entrez databases can be accomplished by using Entrez ProbeSet.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.388">
                  <label>11</label>
                  <p>What does <named-content content-type="book-spec-emph">Scope</named-content> mean in the Accession Display?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.389">
                  <p>GEO platforms (GPL prefix) may have related samples and, through those related samples, related series. GEO samples (GSM prefix) will always have one related platform and may have multiple, related series. GEO series (GSE prefix) will have at least one related sample and, through those related samples, will have at least one related platform. The <named-content content-type="book-spec-emph">Family</named-content> setting will retrieve all accessions (of different types) related to self (including Self). Please see <xref ref-type="boxed-text" rid="bid.344">Box 1</xref> for more details.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.390">
                  <label>12</label>
                  <p>What is SOFT?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.391">
                  <p>SOFT stands for Simple Omnibus Format in Text. SOFT is an ASCII text format that was designed to be a machine-readable representation of data retrieved from, or submitted to, GEO. SOFT output is obtained by using the Accession Display, and SOFT can be used to submit data to GEO. Please see Box 2 for more details.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.392">
                  <label>13</label>
                  <p>What does the word &ldquo;taxon&rdquo; mean?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.393">
                  <p>The NCBI's Taxonomy group has constructed and maintains a taxonomic hierarchy based upon the most recent information, which is described in <xref ref-type="other" rid="bid.246">Chapter 4</xref> of this Handbook.</p>
                </list-item>
              </list>
            </sec>
            <sec id="bid.394">
              <title>Acknowledgments</title>
              <p>We gratefully acknowledge the work of Vladimir Soussov, as well as the entire NCBI Entrez team, especially Grisha Starchenko, Vladimir Sirotinin, Alexey Iskhakov, Anton Golikov, and Pramod Paranthaman. We thank Jim Ostell for guidance, Lou Staudt for discussions during our initial planning for GEO, and the extreme patience shown by Brian Oliver, Wolfgang Huber, and Gavin Sherlock when making the first data submissions. Admirable patience was also exhibited by Al Zhong during the development of the direct deposit validator. Special thanks go to Manish Inala and Wataru Fujibuchi for their continuing work on future features and tools.</p>
            </sec>
            <sec id="bid.395">
              <title>Contributors</title>
              <table-wrap id="bid.396">
                <label>3</label>
                <caption>
                  <title>Selective data set survey.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Source</th>
                    <th align="left" valign="top">Accessions</th>
                    <th align="left" valign="top">Description</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NHGRI melanoma study</td>
                    <td align="left" valign="top">GSE1</td>
                    <td align="left" valign="top">This series represents a group of cutaneous malignant melanomas and unrelated controls that were clustered based on correlation coefficients calculated through a comparison of gene expression profiles.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Stanford Microarray Database</td>
                    <td align="left" valign="top">GSE4 to GSE9, and GSE18 to GSE29</td>
                    <td align="left" valign="top">These series represent microarray studies from the public collection of the Stanford Microarray Database (<xref ref-type="other" rid="bid.1404">SMD</xref>).</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Cancer Genome Anatomy Project</td>
                    <td align="left" valign="top">GSE14</td>
                    <td align="left" valign="top">This series represents the Cancer Genome Anatomy Project SAGE library collection. Libraries contained herein were either produced through CGAP funding or donated to CGAP.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Affymetrix Gene Chips&trade;</td>
                    <td align="left" valign="top">GPL71 to GPL101</td>
                    <td align="left" valign="top">These platforms represent the latest probe attributes of the commercially available Affymetrix Gene Chips&trade; high density oligonucleotide arrays.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">National Children's Medical Center Microarray Center</td>
                    <td align="left" valign="top">GSM1131 to GSM1345</td>
                    <td align="left" valign="top">These samples represent direct deposits of data derived from Affymetrix Gene Chip&trade; arrays and come from the Microarray Center database at the National Children's Medical Center.</td>
                  </tr>
                </table>
              </table-wrap>
              <p><xref ref-type="table" rid="bid.396">Table 3</xref> shows a collection of data sets from various sources. Ron Edgar, Michael Domrachev, Tugba Suzek, Tanya Barrett, and Alex E. Lash contributed to this NCBI resource.</p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.398">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Schena</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Shalon</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Davis</surname>
                        <given-names>RW</given-names>
                      </name>
                      <name>
                        <surname>Brown</surname>
                        <given-names>PO</given-names>
                      </name>
                    </person-group>
                    <article-title>Quantitative monitoring of gene expression patterns with a complementary DNA microarray</article-title>
                    <source>Science</source>
                    <year>1995</year>
                    <volume>270</volume>
                    <fpage>467</fpage>
                    <lpage>470</lpage>
                    <pub-id pub-id-type="pmid">7569999</pub-id>
                  </citation>
                </ref>
                <ref id="bid.399">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Lipshutz</surname>
                        <given-names>RJ</given-names>
                      </name>
                      <name>
                        <surname>Morris</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Chee</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Hubbell</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Kozal</surname>
                        <given-names>MJ</given-names>
                      </name>
                      <name>
                        <surname>Shah</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Shen</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Yang</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Fodor</surname>
                        <given-names>SP</given-names>
                      </name>
                    </person-group>
                    <article-title>Using oligonucleotide probe arrays to access genetic diversity</article-title>
                    <source>Biotechniques</source>
                    <year>1995</year>
                    <volume>19</volume>
                    <fpage>442</fpage>
                    <lpage>447</lpage>
                    <pub-id pub-id-type="pmid">7495558</pub-id>
                  </citation>
                </ref>
                <ref id="bid.400">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Velculescu</surname>
                        <given-names>VE</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Vogelstein</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Kinzler</surname>
                        <given-names>KW</given-names>
                      </name>
                    </person-group>
                    <article-title>Serial analysis of gene expression</article-title>
                    <source>Science</source>
                    <year>1995</year>
                    <volume>270</volume>
                    <fpage>484</fpage>
                    <lpage>487</lpage>
                    <pub-id pub-id-type="pmid">7570003</pub-id>
                  </citation>
                </ref>
                <ref id="bid.401">
                  <label>4</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Emili</surname>
                        <given-names>AQ</given-names>
                      </name>
                      <name>
                        <surname>Cagney</surname>
                        <given-names>G</given-names>
                      </name>
                    </person-group>
                    <article-title>Large-scale functional analysis using peptide or protein arrays</article-title>
                    <source>Nat Biotechnol</source>
                    <year>2000</year>
                    <volume>18</volume>
                    <fpage>393</fpage>
                    <lpage>397</lpage>
                    <pub-id pub-id-type="pmid">10748518</pub-id>
                  </citation>
                </ref>
                <ref id="bid.402">
                  <label>5</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Leighton</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Torhorst</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Mihatsch</surname>
                        <given-names>MJ</given-names>
                      </name>
                      <name>
                        <surname>Sauter</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Kallioniemi</surname>
                        <given-names>OP</given-names>
                      </name>
                    </person-group>
                    <article-title>Tissue microarrays for high-throughput molecular profiling of tumor specimens</article-title>
                    <source>Nat Med</source>
                    <year>1998</year>
                    <volume>4</volume>
                    <fpage>844</fpage>
                    <lpage>847</lpage>
                    <pub-id pub-id-type="pmid">9662379</pub-id>
                  </citation>
                </ref>
                <ref id="bid.403">
                  <label>6</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Edgar</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Domrachev</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Lash</surname>
                        <given-names>AE</given-names>
                      </name>
                    </person-group>
                    <article-title>Gene Expression Omnibus: NCBI gene expression and hybridization array data repository</article-title>
                    <source>Nucleic Acids Research</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>207</fpage>
                    <lpage>210</lpage>
                    <pub-id pub-id-type="pmid">11752295</pub-id>
                    <pub-id pub-id-type="other">99122</pub-id>
                  </citation>
                </ref>
                <ref id="bid.404">
                  <label>7</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Brazma</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Hingamp</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Quackenbush</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Sherlock</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Spellman</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Stoeckert</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Aach</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Ansorge</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Ball</surname>
                        <given-names>CA</given-names>
                      </name>
                      <name>
                        <surname>Causton</surname>
                        <given-names>HC</given-names>
                      </name>
                      <name>
                        <surname>Gaasterland</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Glenisson</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Holstege</surname>
                        <given-names>FC</given-names>
                      </name>
                      <name>
                        <surname>Kim</surname>
                        <given-names>IF</given-names>
                      </name>
                      <name>
                        <surname>Markowitz</surname>
                        <given-names>V</given-names>
                      </name>
                      <name>
                        <surname>Matese</surname>
                        <given-names>JC</given-names>
                      </name>
                      <name>
                        <surname>Parkinson</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Robinson</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Sarkans</surname>
                        <given-names>U</given-names>
                      </name>
                      <name>
                        <surname>Schulze-Kremer</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Stewart</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Taylor</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Vilo</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Vingron</surname>
                        <given-names>M</given-names>
                      </name>
                    </person-group>
                    <article-title>Minimum information about a microarray experiment (MIAME)&mdash;toward standards for microarray data</article-title>
                    <source>Nat Genet</source>
                    <year>2001</year>
                    <volume>29</volume>
                    <fpage>365</fpage>
                    <lpage>371</lpage>
                    <pub-id pub-id-type="pmid">11726920</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.405" book-part-type="chapter" book-part-number="7">
          <book-part-meta>
            <title-group>
              <title>Online Mendelian Inheritance in Man (OMIM): A Directory of Human Genes and Genetic Disorders</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Maglott</surname>
                  <given-names>Donna</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>S. Amberger</surname>
                  <given-names>Joanna</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Hamosh</surname>
                  <given-names>Ada</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch7d1"/>
            <abstract>
              <title>Summary</title>
              <p>Online Mendelian Inheritance in Man (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=OMIM">OMIM</ext-link><sup>TM<xref ref-type="fn" rid="bid.406">*</xref></sup>) is a timely, authoritative compendium of bibliographic material and observations on inherited disorders and human genes. It is the continuously updated electronic version of <italic>Mendelian 
Inheritance in Man</italic> (MIM). MIM was last published in 1998 (<xref ref-type="bibr" rid="bid.1651">1</xref>) and is authored and edited by Dr. Victor A. McKusick and a team of science writers, editors, scientists, and physicians at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.jhu.edu">The Johns Hopkins University</ext-link> and 
around the world (<xref ref-type="bibr" rid="bid.1652">2</xref>). 
Curation of the database and editorial decisions take place at The Johns Hopkins University School of Medicine.
OMIM provides authoritative free text overviews of genetic disorders and gene loci that can be used by clinicians, researchers, students, and educators. In addition, OMIM has many rich connections to relevant primary data resources such as bibliographic, sequence, and map information.</p>
              <fn-group>
                <fn id="bid.406" symbol="*">
                  <p>Trademark status. OMIM<sup>TM</sup> and Online Mendelian Inheritance in Man<sup>TM</sup> are trademarks of The Johns Hopkins University.</p>
                </fn>
              </fn-group>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.407">
              <title>Content and Access</title>
              <sec id="bid.408">
                <title>OMIM Entries</title>
                <table-wrap id="bid.409">
                  <label>1</label>
                  <caption>
                    <title>The OMIM numbering system.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">MIM number range<xref ref-type="fn" rid="mul10"><sup><italic>a</italic></sup></xref></th>
                      <th align="left" valign="top">Explanation</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">100000-199999</td>
                      <td align="left" valign="top">Autosomal loci or 
phenotypes (entries created before May 15, 1994)</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">200000-299999</td>
                      <td align="left" valign="top">Autosomal loci or 
phenotypes (entries created before May 15, 1994)</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">300000-399999</td>
                      <td align="left" valign="top">X-linked loci or phenotypes</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">400000-499999</td>
                      <td align="left" valign="top">Y-linked loci or phenotypes</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">500000-599999</td>
                      <td align="left" valign="top">Mitochondrial loci or phenotypes</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">600000-</td>
                      <td align="left" valign="top">Autosomal loci or phenotypes 
(entries created after May 15, 1994)</td>
                    </tr>
                  </table>
                  <table-wrap-foot>
                    <fn symbol="a" id="mul10">
                      <p>MIM numbers are frequently preceded by a symbol. 
An asterisk (*) before a MIM number indicates that the entry describes a distinct gene or phenotype and that the mode of inheritance of the
phenotype has been proved (in the judgment of the authors and editors)
and that the phenotype described is not known to be determined by a 
gene represented by other asterisked entries in MIM. </p>
                      <p>A number sign (#) before a MIM 
number describing a phenotype indicates that the phenotype is caused by mutation in a gene represented by another entry and usually in any of two or more genes represented by other entries.  The number sign is also used for phenotypes that result from specific chromosomal aberrations, such as Down syndrome, and for contiguous gene syndromes, such as Langer-Giedion syndrome.  Whenever a number sign is used, the reason is stated at the outset of the entry.</p>
                      <p>The absence of an asterisk (or other
sign) preceding the number indicates that the distinctness of the
phenotype as a mendelizing entity or the characterization of the gene in the human is not established.</p>
                    </fn>
                  </table-wrap-foot>
                </table-wrap>
                <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=OMIM">OMIM</ext-link> comprises descriptive, full-text MIM entries, a tabular Synopsis of the (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/htbin-post/Omim/getmap">Human Gene Map</ext-link>) that includes the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/htbin-post/Omim/getmorbid">Morbid Map</ext-link>, clinical synopses, and mini-MIMs.</p>
                <p>OMIM entries are authored and edited by 
experts in the field and by the OMIM staff, based on information in 
the published literature. All entries are assigned a unique, stable, 
six-digit ID number and provide names and symbols used for the disorder and/or gene, a literature-based description, citations, contributor information, and creation and editing dates.
Because MIM is derived from the primary literature, the text is replete with citations and links to PubMed. OMIM authors create entries for each unique gene or Mendelian disorder for which sufficient information exists and do not wittingly create more
than one entry for each gene <xref ref-type="other" rid="bid.1332">locus</xref>.</p>
                <p>MIM is organized into autosomal, X-linked,
Y-linked, and mitochondrial catalogs, and MIM numbers are assigned
sequentially within each catalog (<xref ref-type="table" rid="bid.409">Table 1</xref>).  The kinds of information that may be included in MIM entries are approved name and symbol (obtained from the HUGO Nomenclature Committee), alternative names and symbols in common use, and a text description of the disease or gene.
Many of the longest entries and most new entries in MIM have headings
within the text that may include Clinical Features,
Inheritance, Population Genetics, Heterogeneity, Genotype/Phenotype
Correlations, Cloning, Gene Structure, Gene Function, Mapping, and more. Information on selected disease-causing mutations is contained in the Allelic Variant section of the entry describing the gene.</p>
                <p>With the increasing complexity of 
biological information, OMIM makes a critical contribution by 
distilling what is known about a gene or disease into a single, 
searchable entry. The rich text of the OMIM entry, along with the 
source reference citations, make it easy to retrieve data of 
interest.  The OMIM entry can then serve as a gateway to other sources
of related information via the many curated and computed links within
each entry.</p>
              </sec>
              <sec id="bid.410">
                <title>The OMIM Gene Map</title>
                <p>The OMIM <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/htbin-post/Omim/getmap">Synopsis of the Human Gene Map</ext-link> is a tabular listing of the genes and 
loci represented in MIM ordered pter to qter from chromosome 1&ndash;22, X, and Y. The information in the map includes the cytogenetic location, symbol, title, MIM number, method of mapping, comments, associated disorders and their MIM numbers, and the map location of the mouse <xref ref-type="other" rid="bid.1362">ortholog</xref>. Links are provided from the cytogenetic location to the human <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview">Map
Viewer</ext-link>, from MIM numbers to OMIM entries, and from mouse map 
locations to the Mouse Genome Database (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.informatics.jax.org/">MGD</ext-link>).</p>
                <p>The Synopsis of the human gene map has also been sorted
alphabetically by disorder and is referred to as the Morbid Map.</p>
              </sec>
              <sec id="bid.411">
                <title>Access</title>
                <p>OMIM can be found either by direct query 
via <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=OMIM">Entrez</ext-link>, from other data resources within NCBI that connect to OMIM directly (for example, LocusLink or UniGene), or through Entrez cross-indexing (for example, from a <xref ref-type="other" rid="bid.1387">PubMed</xref> abstract of an article cited in an OMIM 
entry). OMIM is indexed for retrieval in Entrez using a weighted 
system so that if the query term(s) appears in the title of a MIM entry, it will appear at the top of the retrieval list.  
<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/omimhelp.html#SearchFields">Field restrictions</ext-link> are supported for some types of information, or one can use the <named-content content-type="book-spec-emph">Limits</named-content> page to restrict a search (<xref ref-type="fig" rid="bid.412">Figure 1</xref>).  There are also several format options for viewing a retrieval set that may affect which entries are displayed (<xref ref-type="boxed-text" rid="bid.415">Box 1</xref>).<fig id="bid.412"><label>1</label><caption><title>The <named-content content-type="book-spec-emph">Limits</named-content>
options for searching OMIM in Entrez.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch7f1" mime-subtype="gif"/></fig></p>
                <p>Queries can also be entered in the query 
box shown on all Entrez database pages, selecting OMIM as the 
database in the <named-content content-type="book-spec-emph">Search</named-content> option. The <named-content content-type="book-spec-emph">Preview/Index,</named-content> <named-content content-type="book-spec-emph">History</named-content>, <named-content content-type="book-spec-emph">Clipboard</named-content>, or <xref ref-type="other" rid="bid.1270">Cubby</xref> functions can also be used for OMIM, as 
for the rest of Entrez.</p>
                <p>The Gene Map can be accessed from an 
individual OMIM entry via the cytogenetic location displayed when
appropriate under the entry titles. The Gene Map may also be queried <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/htbin-post/Omim/getmap">directly</ext-link>. When queried directly, the first entry that matches the query is shown in the top row of the table, followed by 19 entries ordered by cytogenetic location. The <named-content content-type="book-spec-emph">Find Next</named-content> button can be used to find additional gene map entries that match the query.</p>
              </sec>
            </sec>
            <sec id="bid.413">
              <title>Guide to OMIM Pages</title>
              <sec id="bid.414">
                <title>Query Bar</title>
                <boxed-text id="bid.415" position="margin">
                  <label>1</label>
                  <title><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/omimhelp.html#DisplayFormats">Display options</ext-link> for viewing OMIM query results.</title>
                  <p>Title (default)</p>
                  <p>Details</p>
                  <p>Clinical Synopsis</p>
                  <p>Allelic Variants</p>
                  <p>mini-MIM</p>
                  <p>ASN.1</p>
                  <p>LinkOut</p>
                  <p>Related Entries</p>
                  <p>Genome Links</p>
                  <p>Nucleotide Links</p>
                  <p>Protein Links</p>
                  <p>PubMed Links</p>
                  <p>SNP Links</p>
                  <p>Structure Links</p>
                  <p>UniSTS Links</p>
                  <p>
                    <bold>Obtaining multiple 
views of a query result:</bold>
                  </p>
                  <p>1. Enter query term or terms 
(example: renal failure hypertension).</p>
                  <p>2. Default display is <named-content content-type="book-spec-emph">Titles</named-content>.
				</p>
                  <p>3. Select <named-content content-type="book-spec-emph">Clinical 
Synopsis</named-content> and click on <named-content content-type="book-spec-emph">Display</named-content> at the left to see the 
<named-content content-type="book-spec-emph">Clinical Synopsis</named-content> section of all entries that have them.</p>
                  <p>4. Similarly, select 
<named-content content-type="book-spec-emph">mini-MIM</named-content> or <named-content content-type="book-spec-emph">Allelic Variants</named-content>.
				</p>
                  <p>NOTE: In the same bar, the number 
of entries to display and the format in which to display them can be 
configured by use of the <named-content content-type="book-spec-emph">Show</named-content> and <named-content content-type="book-spec-emph">Text</named-content> buttons, 
respectively.</p>
                </boxed-text>
                <p>OMIM is queried via a standard Entrez 
query bar. The mechanics of selecting entries to display, how to 
display them, and identifying related entries either within Entrez or 
from external resources is also according to the Entrez/LinkOut 
standard. The display options (<xref ref-type="boxed-text" rid="bid.415">Box 
1</xref>) allow the user to format results in several ways. A 
useful function particular to OMIM is the option to display 
<named-content content-type="book-spec-emph">Allelic Variants</named-content>, <named-content content-type="book-spec-emph">Clinical Synopsis</named-content>, or <named-content content-type="book-spec-emph">mini-MIM</named-content> views for a retrieval set.</p>
              </sec>
              <sec id="bid.416">
                <title>OMIM Navigation</title>
                <p>The OMIM <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=OMIM">homepage</ext-link> and the search results pages share the same navigational links to the 
<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=OMIM&amp;cmd=Limits">advanced query</ext-link> page (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/htbin-post/Omim/getmap">Gene 
Map</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/htbin-post/Omim/getmorbid">Morbid 
Map</ext-link>), <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/omimhelp.html">Help</ext-link> documents, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/omimfaq.html">FAQs</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/dispupdates.html">statistics</ext-link>, 
and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/allresources.html">related resources</ext-link>.</p>
                <p>When viewing the text of an OMIM entry, 
however, the navigational links serve as an electronic table of 
contents. The section headings within the entry are listed similar to a table of contents, and selecting one moves the display to that section. Within an entry, selecting the <named-content content-type="book-spec-emph">MIM #</named-content> link takes you back to the top of the entry.</p>
                <p>OMIM staff actively contributes to the 
curation of data in LocusLink. Thus, if the MIM number is represented in LocusLink, a reciprocal LocusLink link is provided in the section to the left. Other links provided by the LocusLink collaboration may also be listed in this section, e.g., links to Nomenclature, Reference Sequences, or UniGene clusters that are specific to the subject OMIM entry.</p>
                <p>The sequence links in the LocusLink 
section may be different from the Entrez indexing links available via
the <named-content content-type="book-spec-emph">Links</named-content> link at the top right of an entry. The Entrez 
indexing links result indirectly from the references in the OMIM entry and may include related sequences in other species, for example. 
Thus, OMIM pages allow two levels of sequence connection: the specific ones in the left section under LocusLink and the indirect but still informative ones through Entrez indexing link at the upper right. More information on OMIM link types can be found in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/omimhelp.html#LinkTypesOverview">Help</ext-link> documents.</p>
                <p>OMIM entries may also contain a link to 
LinkOut for resources external to NCBI (see <xref ref-type="other" rid="bid.650">Chapter 17</xref>). Some of these external resouces are 
curated by OMIM staff, in which case they are displayed by name. 
Others can be seen either by selecting <named-content content-type="book-spec-emph">LinkOut</named-content> in the 
<named-content content-type="book-spec-emph">Links</named-content> pull-down menu or by selecting the <named-content content-type="book-spec-emph">LinkOut</named-content> 
display option in the query bar.</p>
              </sec>
              <sec id="bid.417">
                <title>Entrez Links</title>
                <p>At the top of any report page, or 
associated with each entry in the query result page, are the links to 
related data generated from Entrez (<xref ref-type="other" rid="bid.588">Chapter 15</xref>). Here, PubMed links to the PubMed 
abstracts of the reference citations in the entry. <named-content content-type="book-spec-emph">Related 
Entries</named-content> are to all other OMIM entries referenced in the subject 
entry. <named-content content-type="book-spec-emph">Nucleotide</named-content>, <named-content content-type="book-spec-emph">Protein</named-content>, and <named-content content-type="book-spec-emph">LinkOut</named-content> 
connections are as documented in the previous section.</p>
              </sec>
              <sec id="bid.418">
                <title>The OMIM Entry</title>
                <p>Each OMIM entry has a unique number and is
given a primary title and symbol. This is the title that is displayed
in the document retrieval list. Alternative designations are listed
below the primary title. Some entries contain information that is 
related but not synonymous to the primary topic and is not addressed in another entry (e.g., splice variants, phenotypic variants, etc.).  
This information is set off by the word &ldquo;included&rdquo; in the 
title. The first &ldquo;included&rdquo; title is displayed in the 
document retrieval list. The cytogenetic map location when known is given for each entry. When a disease shows genetic heterogeneity, more than one map location may be given. The &ldquo;light bulb&rdquo; icon at the end of text paragraphs links to related articles in PubMed. References within the text are linked to the complete citation at the end of the entry. There, the PubMed ID is linked to the PubMed abstract.</p>
                <p>Some entries contain an <named-content content-type="book-spec-emph">Allelic 
Variants</named-content> section, which lists noteworthy mutations for the gene. 
Allelic variants are given a 10 digit number: the 6-digit number of the parent locus, followed by a decimal point and a unique 4-digit variant number. Criteria for inclusion include the first mutation to be discovered, high population frequency, distinctive phenotype, historic significance, unusual mechanism of mutation, unusual pathogenetic mechanism, and distinctive inheritance (e.g., dominant with some mutations, recessive with other mutations in the same gene).  Most of the allelic variants represent disease-producing mutations. A few polymorphisms are included, many of which show a statistically significant association with specific common disorders.</p>
              </sec>
            </sec>
            <sec id="bid.419">
              <title>FTP</title>
              <p>The OMIM data are available for <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/Omim/omimfaq.html#download">bulk transfer</ext-link>, but it should be noted that there are <bold>restrictions on use</bold>.</p>
              <p>The OMIM<sup>TM</sup> database, including the 
collective data contained therein, is the property of The Johns 
Hopkins University, which holds the copyright thereto. The OMIM 
database is made available to the general public subject to certain 
restrictions. You may use the OMIM database and data obtained from 
this site for your personal use, for educational or scholarly use, or 
for research purposes only. The OMIM database may not be copied, 
distributed, transmitted, duplicated, reduced, or altered in any way 
for commercial purposes or for the purpose of redistribution without 
a license from The Johns Hopkins University. Requests for information 
regarding a license for commercial use or redistribution of the OMIM 
database may be sent via email to <email>techlicense@jhmi.edu</email>.</p>
            </sec>
            <sec id="bid.1756">
              <title>Legal Statement</title>
              <p>OMIM is funded by a contract from the National Library of Medicine and the National Human Genome Research Institute and by
licensing fees paid to the Johns Hopkins University by commercial
entities for adaptations of the database. The terms of these licenses
are being managed by the Johns Hopkins University in accordance with its
conflict of interest policies.</p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.1651">
                  <label>1</label>
                  <citation>
		  <person-group>
		  <name><surname>McKusick</surname>
		  <given-names>VA</given-names></name>
		  </person-group>
		  <etal>et al.</etal>
		  <source>Mendelian Inheritance in Man</source>
		  <edition>12th ed.</edition>
		  <publisher-loc>Baltimore</publisher-loc> 
		  <publisher-name>Johns Hopkins University Press</publisher-name> 
		  <year>1998</year>
                  </citation>
                </ref>
                <ref id="bid.1652">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Hamosh</surname>
                        <given-names>A</given-names>
                      </name>
                    </person-group>
                    <article-title>Online Mendelian Inheritance in Man 
(OMIM), a knowledgebase of human genes and genetic disorders</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>52</fpage>
                    <lpage>55</lpage>
                    <pub-id pub-id-type="pmid">11752252</pub-id>
                    <pub-id pub-id-type="other">99152</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.423" book-part-type="chapter" book-part-number="8">
          <book-part-meta>
            <title-group>
              <title>The NCBI BookShelf: Searchable Biomedical Books</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Trawick</surname>
                  <given-names>Bart</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Beck</surname>
                  <given-names>Jeff</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>McEntyre</surname>
                  <given-names>Jo</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch8d1"/>
            <abstract>
              <title>Summary</title>
              <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Books">BookShelf</ext-link> is a collection of biomedical books that can be searched directly in Entrez or found via keyword links in PubMed abstracts. Books have been added to the BookShelf in collaboration with authors and publishers, and the complete content (including all figures and tables) is free to use for anyone with an Internet connection.</p>
              <p>The online books are displayed one section at a time, with navigation provided to other parts of the current chapter or to other chapters within the book. Many of the books on the BookShelf can be browsed without any restriction at all; others have less flexibility for navigating the complete content. The publisher (or the owner of the content) defines the rules for access.</p>
              <p>The books are linked to PubMed through research papers citations within the text. In the future, more links may be established between the BookShelf and other resources at NCBI, such as gene and protein sequences, genomes, and macromolecular structures.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.424">
              <title>Content Acquisition</title>
              <sec id="bid.425">
                <title>Basis for Inclusion</title>
                <p>The BookShelf provides a venue through which publishers and authors can make the full text of biomedical books available to the scientific community. As the BookShelf grows, we welcome proposals from authors, editors, or publishers of any &ldquo;in-scope&rdquo; texts, from undergraduate textbooks to more specialized publications, including collections of review articles or workshop proceedings. The scope of the BookShelf is broadly biomedical, including clinical works and those concerning basic biological and chemical sciences.</p>
              </sec>
              <sec id="bid.426">
                <title>Information for Authors, Editors, and Publishers</title>
                <p>All books made available on the BookShelf have been provided in electronic format to <xref ref-type="other" rid="bid.1353">NCBI</xref>. The publisher must provide the content <italic>in full</italic>. Because the complete content is required for display and indexing purposes, the authors, editors, and publisher of a book should all be a part of the decision to participate in the BookShelf. Interested parties can contact us at <email>books@ncbi.nlm.nih.gov</email> for more information or to make a proposal for the inclusion of a book. There is a simple contract that specifies the terms of use of the content. A sample contract can also be obtained by request at the above email address.</p>
                <p>The complete contents of each book will be converted into <xref ref-type="other" rid="bid.1435">XML</xref> according to the NCBI Book Document Type Definition (<xref ref-type="other" rid="bid.1277">DTD</xref>), a public domain DTD developed at NCBI for this project (see below). Books may be submitted to conform to this DTD, or NCBI will convert the source data to validate against the Book DTD. If a conversion needs to be done, the content must be in a format robust enough to meet the needs of the BookShelf publishing system and DTD.</p>
                <p>Any book that was printed from <xref ref-type="other" rid="bid.1400">SGML</xref> or XML should allow for a straightforward conversion. We have had success converting books from Word, XYWrite, PDF, and Quark Express formats, and we anticipate that we would also be able to convert from other desktop publishing packages. <xref ref-type="other" rid="bid.1312">HTML</xref> and PDF formats are less desirable because the data formats are less detailed.</p>
                <p>Figures should be supplied in TIFF format, although GIF and JPEG formats may be accepted. The submitted text files are converted into XML according to the NCBI Book DTD; graphic files are converted into GIF and JPEG formats. Three hard copies of the book are also required, along with the electronic files.</p>
                <p>The XML files are stored in a database. When a reader requests a book, chapter, or section, the XML is retrieved from the database and converted into HTML on the fly using Extensible Stylesheet Language Transformations (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.w3.org/Style/XSL/">XSLT</ext-link>) and Cascading Style Sheets (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.w3.org/Style/CSS/">CSS</ext-link>).</p>
              </sec>
            </sec>
            <sec id="bid.427">
              <title>How to Use the Books</title>
              <p>There are three ways to access the content in BookShelf:</p>
              <p>1. Through hyperlinked terms in PubMed abstracts</p>
              <p>2. By a direct search using search terms or phrases (in the same way as the bibliographic database of PubMed is searched)</p>
              <p>3. Through the Table of Contents of the book (note: some publishers restrict browsing through the entire book by disabling hyperlinks in the Table of Contents)</p>
              <sec id="bid.428">
                <title>Links from PubMed</title>
                <p>The BookShelf can be accessed from all PubMed abstract pages. When viewing a full PubMed abstract, select the <named-content content-type="book-spec-emph">Books</named-content> hyperlink in the upper right-hand corner. This generates a version of the abstract in which certain phrases and terms appear as hypertext links (see <xref ref-type="fig" rid="bid.429">Figure 1</xref>). The linked term may be one or more words in length. If a word or a phrase is linked, it means that the exact phrase also appears in at least one book. Selecting a linked phrase retrieves a list of books that contain that term.<fig id="bid.429"><label>1</label><caption><title>A PubMed abstract showing terms linked to books.</title><p>This view was generated by selecting <named-content content-type="book-spec-emph">Links</named-content> found in the <italic>top right corner</italic> of the abstract and then selecting <named-content content-type="book-spec-emph">Books</named-content> from the drop-down menu that appears.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch8f1" mime-subtype="gif"/></fig></p>
                <p>A statistical weighting system based on the frequency of each phrase in a book section, relative to the rest of the book, is used to identify &ldquo;good&rdquo; phrases. A phrase that appears repeatedly in only a few sections and rarely in other parts of the book indicates a definitive phrase for those few sections; therefore, it ranks highly. Furthermore, the appearance of a phrase in the title, for example, has a greater value in the weighting system than one appearing solely in the text.</p>
                <p>Each PubMed abstract can thus be linked to the appropriate book pages. This method allows two very dissimilar types of text&mdash;the dense, focused PubMed abstracts and the more descriptive book text&mdash;to find common ground.</p>
              </sec>
              <sec id="bid.430">
                <title>Direct Search of Books</title>
                <table-wrap id="bid.431">
                  <label>1</label>
                  <caption>
                    <title>Field limits for use in the BookShelf.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Field<xref ref-type="fn" rid="mul11"><sup><italic>a</italic></sup></xref></th>
                      <th align="left" valign="top">Use</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">[Author]</td>
                      <td align="left" valign="top">Search for the authors of books or chapters.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">[Book]</td>
                      <td align="left" valign="top">Typically used with a Boolean expression to limit a search to a particular book.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">[PmId]</td>
                      <td align="left" valign="top">Locate a journal article citation in a book by its PubMed ID.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">[Rid]</td>
                      <td align="left" valign="top">Locate a particular book element (such as a figure or table) by its reference ID.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">[Secondary Text]</td>
                      <td align="left" valign="top">Search for secondary text, e.g., units (mg/l, etc.)</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">[Title]</td>
                      <td align="left" valign="top">Search for words used in any title (book, chapter, section, subsection, figure, etc.).</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">[Type]</td>
                      <td align="left" valign="top">Locate a division of a book such as a section, chapter, or figure group.</td>
                    </tr>
                  </table>
                  <table-wrap-foot>
                    <fn symbol="a" id="mul11">
                      <p>Filters are applied immediately following a search term, with no separating spaces, e.g., watson[author] AND cmed[book].</p>
                    </fn>
                  </table-wrap-foot>
                </table-wrap>
                <p>Book contents may be searched directly from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Books">BookShelf homepage</ext-link> by using the search boxes located in the Table of Contents and navigation bars of books or by selecting <named-content content-type="book-spec-emph">Books</named-content> from the pull-down menu in any Entrez database search bar (see <xref ref-type="other" rid="bid.588">Chapter 15</xref>). Search terms may be combined using <xref ref-type="other" rid="bid.1253">Boolean</xref> operators that conform to PubMed syntax (see <xref ref-type="other" rid="bid.42">Chapter 2</xref>). The BookShelf also allows search fields to be specified. A complete list of BookShelf database fields can be found in <xref ref-type="table" rid="bid.431">Table 1</xref>.</p>
              </sec>
              <sec id="bid.432">
                <title>Interpreting Results from a Search or PubMed Link</title>
                <p>Results are shown as a list of books in which the term is found, along with the number of sections, figures, and tables that contain the term (<xref ref-type="fig" rid="bid.433">Figure 2</xref>). The book that contains the most hits appears at the top of the list. Choosing the hyperlinked number of items that is associated with a particular book will then display a document summary list of the individual sections, figures, and tables found. (When a term or phrase is found fewer than 20 times within the BookShelf, the document summary page is shown directly, without first displaying the results clustered by book.)<fig id="bid.433"><label>2</label><caption><title>A results page from a BookShelf search.</title><p>If more than 20 sections, tables, and/or figures are found that contain the query term, a summary page, such as the one above, is displayed.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch8f2" mime-subtype="gif"/></fig></p>
                <p>The document summary list is sorted with the most relevant documents shown at the top of the page. The sorting makes use of scores allocated to phrases as a measure of how relevant they are to a given section (a part of the statistical weighting system also used for linking PubMed abstracts to the books). For each book section found, the title, along with some context regarding the hierarchy of the section location (e.g., the chapter and book), is given. An icon is used to distinguish figure legends and tables from text sections (<xref ref-type="fig" rid="bid.434">Figure 3</xref>). Selecting a hyperlinked section title displays the part of the book that contains the search term. From this point, the user may be able to navigate further throughout the book content, according to the policy of the publisher (see <italic>Navigating Book Content</italic>).<fig id="bid.434"><label>3</label><caption><title>A document summary list of book sections.</title><p>The most relevant sections to the query appear at the top of the list. Note the icons that designate figure and table hits. The list may also be displayed in a brief format that lists only the section names that contain the term by choosing <named-content content-type="book-spec-emph">Brief</named-content> from the drop-down menu to the immediate <italic>right</italic> of the <named-content content-type="book-spec-emph">Display</named-content> button.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch8f3" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.435">
                <title>Navigating Book Content</title>
                <p>Each HTML page of content seen in a Web browser represents one section of one chapter of a book, i.e., all of the content (including subsections and so on) within the first-level heading of a chapter. The amount of content this represents varies according to the structure of the original book. Some books have very long sections, some short, some a mixture; although on the whole, most chapters are divided into 3-10 sections.</p>
                <p>The top of every page contains links to both short and detailed Tables of Contents and a description of the current location within the book (<xref ref-type="fig" rid="bid.436">Figure 4</xref>). The hierarchal elements that describe the current location are hyperlinked and may be used to travel up the organizational levels of a book. Additionally, a navigation sidebar shows the current section among its peers and lists the figures and tables found within the current section (<xref ref-type="fig" rid="bid.436">Figure 4</xref>). Reference citations in the text are linked where possible to PubMed abstracts by the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/citmatch.html">Citation Matcher</ext-link>. References internal to the book, e.g., to other chapters or sections, figures, tables, and boxes, are also hyperlinked. Further navigation from the current page to other parts of the book depends on the access policy of the publisher.<fig id="bid.436"><label>4</label><caption><title>Navigation elements that appear on every book HTML page.</title><p>Hyperlinks to various sections within a chapter appear within a navagation bar to the <italic>left</italic> of each page. Hyperlinks may be disabled within some books at the request of the publisher. A <named-content content-type="book-spec-emph">Search</named-content> box is located <italic>below</italic> the navigation bar. At the <italic>top</italic> of each page is a hyperlinked, hierarchical tree that illustrates a page's relative position within a book.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch8f4" mime-subtype="gif"/></fig></p>
              </sec>
            </sec>
            <sec id="bid.437">
              <title>Technology</title>
              <sec id="bid.438">
                <title>Text Conversion and XML</title>
                <p>All book content submitted to the BookShelf is converted to XML according to the public NCBI Book DTD.</p>
                <p>Any files submitted in SGML or XML are converted to the Book DTD using <xref ref-type="other" rid="bid.1437">XSLT</xref>. Books submitted in desktop publishing or word processing formats are converted by a contractor using proprietary technologies. Once the book XML is valid against the DTD, the book is ready to be loaded into the database.</p>
                <p>The XML for a book is generally a set of files. Each chapter and appendix is an independent file, as is the frontmatter. The book is pulled together by the book.xml file, which defines the structure of the book. For example, a book with two chapters and a bibliography at the end of the book would be structured as follows:
                  <preformat>
&lt;!DOCTYPE book SYSTEM "ncbi-book.dtd" [
&lt;!-- graphics -->
&lt;!ENTITY % Graphics SYSTEM "graphics.xml">
%Graphics;
&lt;!ENTITY frontmatter SYSTEM "fm.xml">
&lt;!ENTITY chapter1 SYSTEM "ch1.xml">
&lt;!ENTITY chapter2 SYSTEM "ch2.xml">
&lt;!--back-->;
&lt;!ENTITY biblist SYSTEM "biblist.xml">
]>
&lt;book>
&amp;frontmatter;
&lt;body>
&amp;chapter1;
&amp;chapter2;
&lt;/body>
&lt;back>
&amp;biblist;
&lt;/back>
&lt;/book></preformat>
                </p>
                <p>The book.xml file is composed of two parts. The first part defines all of the components that are required to build the book. These definitions occur within the &lt;!DOCTYPE [ ] &gt; tag. The second part builds the structure of the book. The root element is &lt;book&gt;. &lt;book&gt; contains whatever is in &amp;frontmatter;, &lt;body&gt;, and &lt;back&gt;.</p>
                <p>The book.xml example above refers to five external files: fm.xml, which contains all of the frontmatter of the book; ch1.xml, which contains chapter 1; ch2.xml, which contains chapter 2; biblist.xml, which contains the bibliography for the book; and graphics.xml, which defines the images. If any of these files is not valid according to the DTD or if the files are not found where they are defined in their &lt;!ENTITY &gt; declaration, then the book will not be valid.</p>
              </sec>
              <sec id="bid.439">
                <title>Images</title>
                <p>All of the images, including figures, icons, and book-specific character graphics (see Special Characters below), are called out in the text as entities. The entities are defined in the graphics.xml file.</p>
                <p>&lt;!ENTITY ch2fu6 SYSTEM "data/mga/pictures/ch2/ch2fu6.gif" NDATA GIF&gt;</p>
                <p>&lt;!ENTITY ch2fu7 SYSTEM "data/mga/pictures/ch2/ch2fu7.jpg" NDATA JPG&gt;</p>
                <p>&lt;!ENTITY ch2fu8 SYSTEM "data/mga/pictures/ch2/ch2fu8.gif" NDATA GIF&gt;</p>
                <p>&lt;!ENTITY ch2fu9 SYSTEM "data/mga/pictures/ch2/ch2fu9.gif" NDATA GIF&gt;</p>
                <p>&lt;!ENTITY ch2fu10 SYSTEM "data/mga/pictures/ch2/ch2fu10.gif" NDATA GIF&gt;</p>
                <p>&lt;!ENTITY ch2e1 SYSTEM "data/mga/pictures/ch2/ch2e1.gif" NDATA GIF&gt;</p>
                <p>&lt;!ENTITY ch2e2 SYSTEM "data/mga/pictures/ch2/ch2e2.gif" NDATA GIF&gt;</p>
                <p>Graphic files are converted into GIF and JPEG formats and optimized for display on the Web. The images are not loaded into the database; they are retrieved from a file server when called by the HTML page.</p>
              </sec>
              <sec id="bid.440">
                <title>Math and Formulae</title>
                <p>Math expressions and chemical formulae and structures are handled as images.</p>
              </sec>
              <sec id="bid.441">
                <title>Special Characters</title>
                <p>The BookShelf uses the same character sets that PubMed Central uses (see <xref ref-type="other" rid="bid.457">Chapter 9</xref>). These include a number of standard <xref ref-type="other" rid="bid.1326">ISO</xref> character sets (8879 and 9573), along with a set of characters that has been defined to accommodate characters not in the standard set. The ISO Standard Character sets referenced are listed in Box 1 of <xref ref-type="other" rid="bid.457">Chapter 9</xref>. Special characters are converted to the BookShelf/PubMed Central (<xref ref-type="other" rid="bid.1375">PMC</xref>) character set during conversion into XML. Characters created for one book (book-specific characters) are called out in the XML as images. To provide for the most flexibility in displaying characters across platforms, BookShelf uses <xref ref-type="other" rid="bid.1428">UTF-8</xref> encoding whenever possible. Because not all browsers support the same subset of UTF-8 characters and some characters cannot be represented in UTF-8, the BookShelf displays characters as a combination of GIFs and UTF-8 characters, depending on the Browser/OS combination and the character to be displayed.</p>
              </sec>
            </sec>
            <sec id="bid.442">
              <title>The BookShelf Data Flow</title>
              <p>The XML files for each book are stored in a <xref ref-type="other" rid="bid.1413">Sybase</xref> database. To show the best representation to each reader, the reader's browser and operating system are noted and passed to the rendering software. When a reader requests a book, chapter, or section, the XML and the character images or UTF-8 characters appropriate to the reader's system are retrieved from the database.</p>
              <p>The XML is converted to HTML using XSLT stylesheets. The look of the HTML pages is controlled further using CSS, which allow manipulation of colors, fonts, and typefaces (<xref ref-type="fig" rid="bid.443">Figure 5</xref>).<fig id="bid.443"><label>5</label><caption><title>Processing and data flow of books.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch8f5" mime-subtype="gif"/></fig></p>
              <p>The Table of Contents for a book is created from actual elements within the book content, rather than from the Table of Contents given in the book frontmatter of the hard copy. This ensures that the Table of Contents represents the content accurately as it is organized on BookShelf.</p>
            </sec>
            <sec id="bid.444">
              <title>NCBI Book DTD</title>
              <sec id="bid.445">
                <title>History</title>
                <p>In the first version of the NCBI BookShelf project, Quark files were converted directly to HTML for display online. The result was effective, illustrating the value of having a textbook online and linked to PubMed; however, it was labor intensive, limiting, and not scaleable.</p>
                <p>To simplify the delivery of books online and to allow for the expansion of linking within the Entrez system, NCBI decided to convert all content into a centralized XML format. The normalized XML content is easier to render, allows added value such as the addition of links to other NCBI databases, and simplifies the addition of new volumes.</p>
                <p>PMC created a new DTD for the BookShelf project, which was based on the ISO-12803 DTD. As more books were converted to the NCBI Book DTD, changes had to be made to accommodate the data.</p>
                <p>The NCBI Book DTD is a public DTD available on request from <email>books@ncbi.nlm.nih.gov</email>.</p>
              </sec>
            </sec>
            <sec id="bid.446">
              <title>Frequently Asked Questions</title>
              <list list-type="simple">
                <list-item id="bid.447">
                  <label>1</label>
                  <p>How do I access the books at NCBI?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.448">
                  <p>The online books can be accessed by direct searching in Entrez or through PubMed abstracts. After performing a general PubMed search, click on the author name of one of the search results to view the abstract. A hypertext link called <named-content content-type="book-spec-emph">Links</named-content> is displayed to the right of the abstract title. This link contains a drop-down menu consisting of various choices, depending upon the specific abstract. Choosing <named-content content-type="book-spec-emph">Books</named-content> from this drop-down menu will highlight keywords in the abstract that, when selected, initiate a search of all BookShelf content for that particular term.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.449">
                  <label>2</label>
                  <p>Which books are available at NCBI?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.450">
                  <p>The book list is updated on a regular basis and can be viewed on the BookShelf homepage: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Books">http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Books</ext-link>.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.451">
                  <label>3</label>
                  <p>Can I search the books at NCBI?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.452">
                  <p>Yes, the books can be searched either as a complete collection or as a single, selected book (restricted using search options found under <named-content content-type="book-spec-emph">Limits</named-content>).</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.453">
                  <label>4</label>
                  <p>Can I browse the whole book?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.454">
                  <p>The system has been designed so that the user is delivered to the most relevant book sections for a particular term or concept. Although navigation is possible in the immediate vicinity of the page to which you are delivered, it may not be possible to browse the complete book on BookShelf. The range of navigation for each book is determined on a case-by-case basis, in agreement with the publisher.</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.455">
                  <label>5</label>
                  <p>I am the publisher/author/editor of a book. How can I participate?</p>
                </list-item>
              </list>
              <list list-type="simple">
                <list-item id="bid.456">
                  <p>Please email <email>books@ncbi.nlm.nih.gov</email> to discuss potential projects.</p>
                </list-item>
              </list>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.457" book-part-type="chapter" book-part-number="9">
          <book-part-meta>
            <title-group>
              <title>PubMed Central (PMC): An Archive for Literature from Life Sciences Journals</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Beck</surname>
                  <given-names>Jeff</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Sequeira</surname>
                  <given-names>Ed</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch9d1"/>
            <abstract>
              <title>Summary</title>
              <p>PubMed Central (PMC) is the National Library of Medicine's digital archive of full-text journal literature. Journals deposit material in PMC on a voluntary basis. Articles in PMC may be retrieved either by browsing a table of contents for a specific journal or by searching the database. Certain journals allow the full text of their articles to be viewed directly in PMC. These are always free, although there may be a time lag of a few weeks to a year or more between publication of a journal issue and when it is available in PMC. Other journals require that PMC direct users to the journal's own Web site to see the full text of an article. In this case, the material will always be available free to any user no more than 1 year after publication but will usually be available only to the journal's subscribers for the first 6 months to 1 year.</p>
              <p>To increase the functionality of the database, a variety of links are added to the articles in PMC: between an article correction and the original article; from an article to other articles in PMC that cite it; from a citation in the references section to the corresponding abstract in PubMed and to its full text in PMC; and from an article to related records in other Entrez databases such as Reference Sequences, OMIM, and Books.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.458">
              <title>A PubMed Central (PMC) Site Guide</title>
              <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pubmedcentral.gov/">PMC homepage</ext-link> has a list of all journals available in PMC and the earliest available issue for each journal (<xref ref-type="fig" rid="bid.459">Figure 1</xref>). From here, a table of contents for the latest available issue of a journal or a list of all issues of the journal available through PMC can be viewed.<fig id="bid.459"><label>1</label><caption><title>The PMC journal list.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f1" mime-subtype="gif"/></fig></p>
              <p>Every article citation in a table of contents includes one or more links (<xref ref-type="fig" rid="bid.460">Figure 2</xref>). Articles for which the full text is available directly in PMC generally have links to an Abstract view, a Full Text view, and a PDF (printable view). Where applicable, they also have links to Corrections and to supplementary data that may be available for the article. In cases where the full text is available only at the journal publisher's site, there is only one link, to a PubLink page (described below).<fig id="bid.460"><label>2</label><caption><title>PMC table of contents.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f2" mime-subtype="gif"/></fig></p>
              <p>In addition to the header information for the article itself, the upper part of a Full Text or PubLink page contains a variety of links, including links to other forms of the article, to related information in PubMed and other <xref ref-type="other" rid="bid.1282">Entrez</xref> databases, and to corrections or &ldquo;cited-in&rdquo; lists where these apply (<xref ref-type="fig" rid="bid.461">Figure 3</xref>). The sidebar in the body of a Full Text page (<xref ref-type="fig" rid="bid.462">Figure 4</xref>) has links to tables and thumbnail images of any figures in the article, which when selected will display the full figure. Figures and tables may also be opened directly from the point in the text where they are referenced. Citations in the References section of an article frequently include a link to the corresponding PubMed abstract and sometimes also have a link to the full text of the referenced article in PMC (<xref ref-type="fig" rid="bid.463">Figure 5</xref>).<fig id="bid.461"><label>3</label><caption><title>Header of Abstract and Full Text pages.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f3" mime-subtype="gif"/></fig><fig id="bid.462"><label>4</label><caption><title>Body of Full Text page.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f4" mime-subtype="gif"/></fig><fig id="bid.463"><label>5</label><caption><title>References.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f5" mime-subtype="gif"/></fig></p>
              <p>An Abstract page is identical to a Full Text page that has been cut off at the end of the abstract.</p>
              <p>A PMC PubLink page (<xref ref-type="fig" rid="bid.464">Figure 6</xref>) is similar to an Abstract page, except that it does not have links to alternate forms (Full Text or PDF) of the article in PMC. Instead, it contains a link to the full text at the publisher's site and information about when it will be freely available.<fig id="bid.464"><label>6</label><caption><title>PubLink page.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f6" mime-subtype="gif"/></fig></p>
              <p>When an article has been cited by other articles in PMC, a &ldquo;cited-in&rdquo; link displays just under the article header information on both the Abstract and Full Text pages. Selecting this cited-in link gives you a list of the articles that have referenced the subject article (<xref ref-type="fig" rid="bid.465">Figure 7</xref>).<fig id="bid.465"><label>7</label><caption><title>Cited-in list.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f7" mime-subtype="gif"/></fig></p>
              <p>PMC article citations may also be retrieved by doing a search in PMC or through a PubMed search. (In PubMed, use the <named-content content-type="book-spec-emph">subsets</named-content> limit if you want to find only articles that are available in PMC.)</p>
            </sec>
            <sec id="bid.466">
              <title>Participation in PMC</title>
              <p>Participation by publishers in PMC is voluntary, although participating journals must meet certain editorial standards. A participating journal is expected to include all of its peer-reviewed primary research articles in PMC. Journals are encouraged to also deposit other content such as review articles, essays, and editorials. Review journals, and similar publications that have no primary research articles, are also invited to include their contents in PMC. However, primary research papers without peer review are not accepted.</p>
              <p>Journals that deposit material in PMC may make the full text viewable directly in PMC or may require that PMC link to the journal site for viewing the complete article. In the latter case, the full text must be freely available at the journal site no more than 1 year after publication. In the case of full text that is viewable directly in PMC, which by definition is free, the journal may delay the release of its material for more than 1 year after publication, although all current journals have delays of 1 year or less.</p>
              <p>In either case, the journal must provide <xref ref-type="other" rid="bid.1400">SGML</xref> or <xref ref-type="other" rid="bid.1435">XML</xref> for the full text, along with any related high-resolution image files. All data must meet PMC standards for syntactically correct and complete data.</p>
              <p>The rationale behind the insistence on free access is that continued use of the material, which is encouraged by open access, serves as the best test of the durability and utility of the archive as technology changes over time. PMC does not claim copyright on any material deposited in the archive. Copyright remains with the journal publisher or with individual authors, whichever is applicable.</p>
              <p>Refer to <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pubmedcentral.gov/about/pubinfo.html">Information for Publishers</ext-link> to learn more about participating in PMC.</p>
            </sec>
            <sec id="bid.467">
              <title>Links to Other NCBI Resources</title>
              <p>From Abstract and Full Text pages in PMC are links to related articles in PubMed and to related records in other Entrez databases, such as Nucleotides or Books. These are identical to the links between databases that you can find in any Entrez record.</p>
            </sec>
            <sec id="bid.468">
              <title>PMC Architecture</title>
              <p>PubMed Central is an XML-based publishing system for full-text journal articles. All journal content in the archive was either supplied in, or has been converted to, a Document Type Definition (<xref ref-type="other" rid="bid.1277">DTD</xref>) written at <xref ref-type="other" rid="bid.1353">NCBI</xref> for the publication and storage of full-text articles.</p>
              <p>The content is displayed dynamically on the PMC site by journal, volume, and issue (if applicable). XML, Web graphics, PDFs, and supplemental data are stored in a <xref ref-type="other" rid="bid.1413">Sybase</xref> database. When a reader requests an article, the XML is retrieved from the database, and it is converted to <xref ref-type="other" rid="bid.1312">HTML</xref> using <xref ref-type="other" rid="bid.1437">XSLT</xref> stylesheets. The look of the HTML pages is controlled further by using Cascading Style Sheets (<xref ref-type="other" rid="bid.1269">CSS</xref>), which allow manipulation of colors, fonts, and typefaces.</p>
            </sec>
            <sec id="bid.469">
              <title>Data Flow: 1. SGML/XML Processing</title>
              <p>We receive journal content either directly from publishers or from publishers' vendors. This content includes:
                <list list-type="bullet">
                  <list-item id="bid.470">
                    <p> SGML or XML of the articles to be deposited</p>
                  </list-item>
                  <list-item id="bid.471">
                    <p> High-resolution images</p>
                  </list-item>
                  <list-item id="bid.472">
                    <p> Supplemental data associated with the articles</p>
                  </list-item>
                  <list-item id="bid.473">
                    <p> PDF versions of the articles</p>
                  </list-item>
                </list>
              </p>
              <p>All of the text is converted to a central DTD, the PMC DTD, and the images are converted to Web format (GIF and JPEG). These files, along with any supplemental data or PDFs, are loaded into the database for linking, indexing, and retrieval (<xref ref-type="fig" rid="bid.474">Figure 8</xref>). The source text in SGML or XML format is parsed against the source DTD. If the source files do not conform to the DTD, they are returned to the publisher or vendor for correction.<fig id="bid.474"><label>8</label><caption><title>Data flow.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f8" mime-subtype="gif"/></fig></p>
              <p>Once all of the files in an issue or batch have been validated, they are converted to PMC XML (referred to as a PMC XML file or PXML) using XSLT (<xref ref-type="fig" rid="bid.475">Figure 9</xref>). Because XSLT is an XML conversion tool, SGML source files must first be converted to XML. This is done using SX, which is available from <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.jclark.com/sp/index.htm">James Clark</ext-link>. The transformation will close any empty elements and insert ending tags for any element that is not closed.<fig id="bid.475"><label>9</label><caption><title>Text conversion flow.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9f9" mime-subtype="gif"/></fig></p>
              <p>For each publisher that submits SGML, an XML version of the DTD is created. This is used to parse the output of SX before the XML conversion is started. The XML version of the publisher's DTD is not used to validate the source data (because this has already been done using the original DTD).</p>
              <p>XSLT requires that the input file be valid XML. If a DTD is not available for validation, the parser will check the syntax; it will also replace all of the character entities with the appropriate <xref ref-type="other" rid="bid.1428">UTF-8</xref> representation. This can cause a problem because the relationship between the characters in the input file and their UTF-8 representations may not be one to one. This means that characters translated to UTF-8 might not translate back to the original character entity accurately.</p>
              <p>After the XSLT conversion, the original character entities are converted to character entities that are valid under the PMC DTD. Character translation tables for each source DTD regulate this conversion.</p>
              <p>The resulting XML is then validated against the PMC DTD.</p>
              <p>Several other items are created along with the PXML file. These are:</p>
              <p>1. An entity file (articlename.ent). This file lists all of the character entities (from the PMC DTD-defined entity sets) that are in the article. One entity file is created for each article. This information is loaded into the database and is used to prepare the final HTML file for display. A sample is below:
                <preformat>
    agr
    deg
    Dgr
    ldquo
    lsqb
    lt
    pmc811
    mdash</preformat>
              </p>
              <p>2. A PubMedID file (articlename.pmid). This file includes a set of reference citations in the format
                  <preformat>
Journal_Title|year|volume|first_page|AuthorName|Refno|pubmedid</preformat>
</p>
              <p>When the article is converted, this information is collected from each journal citation in the bibliography and sent in a query to PubMed (using the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/getids.cgi">Citation Matcher</ext-link> utility). If a value is returned, it is written in the last field. If no value is returned, an error message is written in this field. This information is saved so that if the article ever needs to be reconverted, the PubMed IDs will not need to be looked up again. A sample is below:
                  <preformat>
Nucleic Acids Res|1992|20|2673|Murray JM|13:2:493:16|1319571
J Biol Chem|1999|274|35297|Naumann M|13:2:493:17|10585392
EMBO J|2000|19|3475|Osaka F|13:2:493:18|10880460
Nature|2000|405|662|Osterlund MT|13:2:493:19|0
FASEB J|1998|12|469|Seeger M|13:2:493:20|9535219</preformat>
</p>
              <p>3. A source node file (articlename.src). When an article is passed through the XSLT conversion, a list is made of each node, or named piece of information, that is included in the file. As the conversion is running, each node that is being processed is recorded. When the conversion is complete, the processed node list is compared with the list of nodes in the source file, and any piece of information that was not processed is reported in a conversion log. A sample is below:
                  <preformat>
/ART
/ART/@AID
/ART/@DATE
/ART/@ISS
/ART/BM/FN/P/EXREF
/ART/BM/FN/P/EXREF/@ACCESS
/ART/BM/FN/P/EXREF/@TYPE
/ART/FM
/ART/FM/ABS
/ART/FM/ABS/P
/ART/FM/ABS/P/EMPH
/ART/FM/ACC
/ART/FM/ATL
/ART/FM/ATL/EMPH
/ART/FM/AUG
/ART/FM/PUBFRONT/CPYRT/CPYRTNME/COLLAB
/ART/FM/PUBFRONT/CPYRT/DATE
/ART/FM/PUBFRONT/CPYRT/DATE/YEAR
/ART/FM/PUBFRONT/DOI
/ART/FM/PUBFRONT/EXTENT
/ART/FM/PUBFRONT/FPAGE
/ART/FM/PUBFRONT/ISSN
/ART/FM/PUBFRONT/LPAGE
/ART/FM/RE
/ART/FM/RV</preformat>
</p>
              <sec id="bid.476">
                <title>Image Processing</title>
                <p>To accommodate the archiving requirements of the PubMed Central project, it is important that figures be submitted in the greatest resolution possible, in TIFF or EPS format. Figures in these formats will be available for data migration when formats change in the future, and PubMed Central will be able to keep all of the figures current.</p>
                <p>For display on the PMC site, two copies of each figure are made: a GIF thumbnail (100 pixels wide) and a JPEG file that will be displayed with the figure caption when the figure is requested.</p>
              </sec>
              <sec id="bid.477">
                <title>Supplemental Data Processing</title>
                <p>Supplemental data include any supporting information that accompanies the article but is not part of the article. They may be text files, Word document files, spreadsheet files, executables, video, and others.</p>
                <p>Sometimes a journal has a Web site where all of this supplemental information is stored. In this case, PubMed Central establishes links from the article to the supplemental information on the publisher's site.</p>
                <p>In other cases, the supplemental data files are submitted with the article to be loaded into the PMC database. Either way, the information concerning this supplemental data is collected in a Supplemental Data file, which includes the location of the supplemental file(s), the type of information that is available, and how the link should be built from the article. PMC does not validate any of the supplemental data files that are supplied.</p>
              </sec>
              <sec id="bid.478">
                <title>Mathematics</title>
                <p>Mathematical symbols and notations can be difficult to display in HTML because of built-up expressions and unusual characters. For the most part, expressions that are simple enough to display using HTML are not handled as math unless they are tagged as math specifically. Publishers that supply content to PMC handle math expressions in one of two ways: supplied images or encoding in SGML.</p>
                <sec id="bid.479">
                  <title>1. Math Images</title>
                  <p>Any expression that cannot be tagged by the source DTD is supplied as an image. In this case, PMC will pass the image callout through to the PXML file and display the supplied image in the HTML file.</p>
                </sec>
                <sec id="bid.480">
                  <title>2. Math in SGML</title>
                  <p>Several of the Source DTDs used by publishers to submit data to PMC are robust enough to allow coding of almost any mathematical expression in SGML. Most of these were derived from the Elsevier DTD; therefore, many of the elements are similar.</p>
                  <p>During article conversion, any items that are recognized as math are translated into Te&chi;. This would include any expression tagged specifically as a &ldquo;formula&rdquo; or &ldquo;display formula,&rdquo; as well as any free-standing expression that cannot be represented in HTML. These expressions include radicals, fractions, and anything with an overbar (other than accented characters). For example:
                  <disp-quote>
                    <p>&ldquo;<italic>x</italic> + <italic>y</italic> = 2<italic>z</italic>&rdquo; would not be recognized as a math expression, but &ldquo;&lt;formula&gt;<italic>x</italic> + <italic>y</italic> = 2<italic>z</italic>&lt;/formula&gt;&rdquo; would be.</p>
                  </disp-quote>
                  <disp-quote>
                    <p>&ldquo;1/2&rdquo; would not be recognized as a math expression, but &lt;fraction&gt;&lt;numerator&gt;1&lt;/numerator&gt;&lt;denominator&gt;2 &lt;/denominator&gt;&lt;/fraction&gt; would be.</p>
                  </disp-quote>
                  <disp-quote>
                    <p>&ldquo;&lt;radical&gt;2x&lt;/radical&gt;&rdquo; would be recognized as a math expression, as would &ldquo;&lt;overbar&gt;47X&lt;/overbar&gt;.&rdquo;</p>
                  </disp-quote>
                  </p>
                  <p>The SGML:
						<preformat>
&lt;FD ID="E2"&gt;I&lt;SUP&gt;&lt;UP&gt;o&lt;/UP&gt;
&lt;/SUP&gt;&lt;INF&gt;&lt;UP&gt;f&lt;/UP&gt;&lt;/INF&gt;&amp;cjs1134;I&lt;INF&gt;&lt;UP&gt;f&lt;/UP&gt;&lt;/INF&gt;=&amp;phgr;
&lt;SUP&gt;&lt;UP&gt;o&lt;/UP&gt;&lt;/SUP&gt;&lt;INF&gt;&lt;UP&gt;f&lt;/UP&gt;&lt;/INF&gt;&amp;cjs1134;&amp;phgr;&lt;INF&gt;
&lt;UP&gt;f&lt;/UP&gt;&lt;/INF&gt;=1&amp;plus;K&lt;INF&gt;&lt;UP&gt;sv&lt;/UP&gt;&lt;/INF&gt;&amp;lsqb;&lt;UP&gt;Q
&lt;/UP&gt;&amp;rsqb;&lt;/FD&gt;</preformat></p>
                  <p>Converts to:
						<preformat>
&lt;fd id="E2"&gt;&lt;math mathtype="tex" 
id="M2"&gt;&bsol;documentclass[12pt]&lcub;minimal&rcub;
&bsol;usepackage&lcub;wasysym&rcub;
&bsol;usepackage[substack]&lcub;amsmath&rcub;
&bsol;usepackage&lcub;amsfonts&rcub;
&bsol;usepackage&lcub;amssymb&rcub;
&bsol;usepackage&lcub;amsbsy&rcub;
&bsol;usepackage[mathscr]&lcub;eucal&rcub;
&bsol;usepackage&lcub;mathrsfs&rcub;
&bsol;DeclareFontFamily&lcub;T1&rcub;&lcub;linotext&rcub;&lcub;&rcub;
&bsol;DeclareFontShape&lcub;T1&rcub;&lcub;linotext&rcub;&lcub;m&rcub;&lcub;n&rcub; &lcub; &lt;-&gt; linotext &rcub;&lcub;&rcub;
&bsol;DeclareSymbolFont&lcub;linotext&rcub;&lcub;T1&rcub;&lcub;linotext&rcub;&lcub;m&rcub;&lcub;n&rcub;
&bsol;DeclareSymbolFontAlphabet&lcub;&bsol;mathLINOTEXT&rcub;&lcub;linotext&rcub;
&bsol;begin&lcub;document&rcub;
&bsol;[
I^&lcub;o&rcub;_&lcub;f&rcub;/I_&lcub;f&rcub;=&lcub;&bsol;phi&rcub;^&lcub;o&rcub;_&lcub;f&rcub;/&lcub;&bsol;phi&rcub;_&lcub;f&rcub;=1+K_&lcub;sv&rcub;[Q]
&bsol;]
&bsol;end&lcub;document&rcub;
&lt;/math&gt;&lt;/fd&gt;</preformat></p>
                  <p>When the articles are loaded into the database, the equation markup is written into an Equation table. This table will also include the equation image, which will be created from the Te&chi; markup.</p>
                  <p>The image for the equation shown above in SGML and PXML (with Te&chi;) is:</p>
                  <p>
                    <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch9df2" mime-subtype="gif"/>
                  </p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.482">
              <title>Data Flow: 2. Loading the Database</title>
              <p>Because all of the content in PubMed Central is in the same format&mdash;PMC DTD&mdash;loading articles into the database is relatively straightforward. Once an article is loaded into the production (public) database, it will retain its ArticleId (article ID number) in perpetuity. On loading, each article is validated against the PMC DTD. Also, any external files that are referenced by the XML are checked. If any file, such as a figure, is missing, the loading will be aborted.</p>
              <p>The database loading software and daily maintenance programs perform several other tests to ensure the accuracy and vitality of the archive:
                <list list-type="arabic">
                  <list-item id="bid.483">
                    <label>1</label>
                    <p>Journal identity. The Journal title being loaded is verified against the <xref ref-type="other" rid="bid.1327">ISSN</xref> number in the PXML to verify that the journal identity is correct.</p>
                  </list-item>
                  <list-item id="bid.484">
                    <label>2</label>
                    <p>Duplicate articles. An article may not be loaded more than once. Any changes to the article must be submitted as a replacement article, which will use the same ArticleId.</p>
                  </list-item>
                  <list-item id="bid.485">
                    <label>3</label>
                    <p>Publication date/delay. Rules for delay of publication embargo can be set up in the database to ensure that an issue will not be released to the public before a certain amount of time has passed since the publisher made the issue available.</p>
                  </list-item>
                  <list-item id="bid.486">
                    <label>4</label>
                    <p>PubMed IDs. PubMedIDs for the article being loaded or any bibliographic citation in the article that are not defined in the PXML are looked up upon loading.</p>
                  </list-item>
                  <list-item id="bid.487">
                    <label>5</label>
                    <p>Link updates. Links between related articles and from articles to external sources are updated daily.</p>
                  </list-item>
                </list>
              </p>
              <p>The database has been designed to allow multiple versions of articles. In addition to article information, the database also stores information on content suppliers and publishers and journal-specific information.</p>
            </sec>
            <sec id="bid.488">
              <title>Special Characters</title>
              <boxed-text id="bid.489" position="margin">
                <label>1</label>
                <title>ISO Standard Character sets used by PMC.</title>
                
                  <preformat>
&lt;!ENTITY % ISOlat1 PUBLIC "ISO 8879-1986//ENTITIES Added Latin 1//EN"&gt;
&lt;!ENTITY % ISOlat2 PUBLIC "ISO 8879-1986//ENTITIES Added Latin 2//EN"&gt;
&lt;!ENTITY % ISOnum PUBLIC "ISO 8879-1986//ENTITIES Numeric and Special Graphic//EN"&gt;
&lt;!ENTITY % ISOpub PUBLIC "ISO 8879-1986//ENTITIES Publishing//EN"&gt;
&lt;!ENTITY % ISOgrk1 PUBLIC "ISO 8879-1986//ENTITIES Greek Letters//EN"&gt;
&lt;!ENTITY % ISOgrk2 PUBLIC "ISO 8879-1986//ENTITIES Monotoniko Greek//EN"&gt;
&lt;!ENTITY % ISOtech PUBLIC "ISO 8879-1986//ENTITIES General Technical//EN"&gt;
&lt;!ENTITY % ISOdia PUBLIC "ISO 8879-1986//ENTITIES Diacritical Marks//EN"&gt;
&lt;!ENTITY % ISOAMSO PUBLIC "ISO 9573-13:1991//ENTITIES Added Math Symbols: Ordinary //EN"&gt;
&lt;!ENTITY % ISOAMSR PUBLIC "ISO 9573-13:1991//ENTITIES Added Math Symbols: Relations //EN"&gt;
&lt;!ENTITY % ISOamsr PUBLIC "ISO 8879-1986//ENTITIES Added Math Symbols: Relations//EN"&gt;
&lt;!ENTITY % ISOamsn PUBLIC "ISO 8879-1986//ENTITIES Added Math Symbols: Negated Relations//EN"&gt;
&lt;!ENTITY % ISOAMSA PUBLIC "ISO 9573-13:1991//ENTITIES Added Math Symbols: Arrow Relations //EN"&gt;
&lt;!ENTITY % ISOAMSB PUBLIC "ISO 9573-13:1991//ENTITIES Added Math Symbols: Binary Operators //EN"&gt;
&lt;!ENTITY % ISOamsc PUBLIC "ISO 8879-1986//ENTITIES Added Math Symbols: Delimiters//EN"&gt;
&lt;!ENTITY % ISOmopf PUBLIC "ISO 9573-13:1991//ENTITIES Math Alphabets: Open Face//EN"&gt;
&lt;!ENTITY % ISOmscr PUBLIC "ISO 9573-13:1991//ENTITIES Math Alphabets: Script//EN"&gt;
&lt;!ENTITY % ISOmfrk PUBLIC "ISO 9573-13:1991//ENTITIES Math Alphabets: Fraktur//EN"&gt;
&lt;!ENTITY % ISObox PUBLIC "ISO 8879:1986//ENTITIES Box and Line Drawing//EN"&gt;
&lt;!ENTITY % ISOcyr1 PUBLIC "ISO 8879:1986//ENTITIES Russian Cyrillic//EN"&gt;
&lt;!ENTITY % ISOcyr2 PUBLIC "ISO 8879:1986//ENTITIES Non-Russian Cyrillic//EN"&gt;
&lt;!ENTITY % ISOGRK3 PUBLIC "ISO 9573-13:1991//ENTITIES Greek Symbols //EN"&gt;
&lt;!ENTITY % ISOGRK4 PUBLIC "ISO 9573-13:1991//ENTITIES Alternative Greek Symbols //EN"&gt;
&lt;!ENTITY % ISOTECH PUBLIC "ISO 9573-13:1991//ENTITIES General Technical //EN"&gt;</preformat>
                
              </boxed-text>
              <p>PubMed Central uses a number of standard <xref ref-type="other" rid="bid.1326">ISO</xref> character sets (8879 and 9573), along with a set of characters that has been defined to accommodate characters not in the standard set. The ISO Standard Character sets referenced are listed in <xref ref-type="boxed-text" rid="bid.489">Box 1</xref>.
		</p>
              <p>Each publisher DTD defines a set of characters that may be used in their articles. Generally, these publisher DTDs use the same standard ISO character sets listed in <xref ref-type="boxed-text" rid="bid.489">Box 1</xref>. Any character that cannot be represented by the standard ISO sets is defined in a publisher-specific character set. These publisher-specified characters are converted into characters in the PMC entity list during conversion (see <italic>SGML/XML Processing</italic>). The PMC <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pubmedcentral.nih.gov/pmcdoc/dtd/pmc.ent">entity list</ext-link> is publicly available.</p>
              <p>The supplied data also include groups of entities that are to be combined in the final document. Sometimes these are grouped in a tag such as:
				<preformat>    &lt;A&gt;&lt;AC&gt;&amp;alpha;&lt;/AC&gt;&lt;AC&gt;&amp;acute;&lt;/AC&gt;&lt;/A&gt;</preformat></p>
              <p>and sometimes they are just positioned next to each other in the text. These combined entities must be mapped either to an ISO character or to a character in the PMC character set.</p>
              <p>For the most flexibility in displaying characters across platforms, PMC uses UTF-8 encoding whenever possible. Because not all browsers support the same subset of UTF-8 characters and some characters cannot be represented in UTF-8, PMC displays characters as a combination of GIFs and UTF-8 characters, depending on the Browser/OS combination and the character to be displayed.</p>
            </sec>
            <sec id="bid.490">
              <title>PMC DTD</title>
              <sec id="bid.491">
                <title>History</title>
                <p>In the first version of the PMC project, the SGML and XML were loaded into a database in its native format. The HTML rendering software was then required to convert content from different sources into normalized HTML on the fly when a reader requested an article.</p>
                <p>This was slow and cumbersome on the rendering side and was not scaleable. At that time, PMC was receiving content for about five journals in two DTDs, the keton.dtd from HighWire Press and the article.dtd from BioMed Central. The set-up for a new journal was difficult, and it soon became obvious that this solution would not scale easily.</p>
                <p>To satisfy the archiving requirement for the PMC project and to simplify the delivery of articles online, PubMed Central decided to convert all content into a centralized format. The normalized content is easier to render, allows enhanced value such as links to other NCBI databases to be added, and simplifies content archiving.</p>
                <p>PMC created a new DTD, which was strongly influenced by the BioMedCentral article.dtd and the keton.dtd. The original emphasis was on simplicity. As more and more articles from more and more journals were converted to the PMC DTD, changes had to be made to accommodate the data. The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pubmedcentral.nih.gov/pmcdoc/dtd/pmc-1.dtd">PMC DTD</ext-link> is publicly available.</p>
              </sec>
              <sec id="bid.492">
                <title>Review and Revision of the PMC DTD</title>
                <p>Because the PMC DTD grew rapidly, it was feared that the original &ldquo;simplicity&rdquo; of its design would lead to confusing data structures. With more and more publishers inquiring about submitting content directly in the PMC DTD, PubMed Central decided that an independent review was necessary. <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.mulberrytech.com">Mulberry Technologies, Inc.</ext-link>, an electronic publishing consultancy specializing in SGML- and XML-based systems, reviewed the DTD and created a modified version.</p>
                <p>At approximately the same time, under the auspices of a Mellon Grant to explore ejournal archiving, Harvard University Library contracted with <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.inera.com">Inera, Inc.</ext-link> to review a variety of DTDs from selected publishers, PMC included. The study focused on two key questions:
                  <list list-type="arabic">
                    <list-item id="bid.493">
                      <label>1</label>
                      <p>Can a common DTD be designed and developed into which publishers' proprietary SGML files can be transformed to meet the requirements of an archiving institution?</p>
                    </list-item>
                    <list-item id="bid.494">
                      <label>2</label>
                      <p>If such a structure can be developed, what are the issues that will be encountered when transforming publishers' SGML files into the archive structure for deposit into the archive?</p>
                    </list-item>
                  </list>
                </p>
                <p>The requirement of the archival article DTD was defined as the ability to represent the intellectual content of journal articles. This <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.diglib.org/preserve/hadtdfs.pdf">study</ext-link> is available and suggestions from the study were used in the NLM Archiving DTD Suite. </p>
                <p>The NLM Archiving DTD will not be backwards-compatible with the pmc-1.dtd. It should be publicly available by the end of 2002, along with complete documentation for publishers and authors. A draft version is available (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pubmedcentral.nih.gov/pmcdoc/dtd/nlm_lib/0.1/documentation/HTML/index.html">http://www.pubmedcentral.nih.gov/pmcdoc/dtd/nlm_lib/0.1/documentation/HTML/index.html</ext-link>), along with a draft version of the documentation.</p>
              </sec>
            </sec>
            <sec id="bid.495">
              <title>Frequently Asked Questions</title>
              <p>Please refer to the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pubmedcentral.gov/about/faq.html">PMC site</ext-link> for answers to frequently asked questions.</p>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.496" book-part-type="chapter" book-part-number="10">
          <book-part-meta>
            <title-group>
              <title>The SKY/CGH Database for Spectral Karyotyping and Comparative Genomic Hybridization Data</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author" rid="bid.ch10-au1">
                <name>
                  <surname>Knutsen</surname>
                  <given-names>Turid</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author" rid="bid.ch10-au2">
                <name>
                  <surname>Gobu</surname>
                  <given-names>Vasuki</given-names>
                </name>
               </contrib>
              <contrib contrib-type="author" rid="bid.ch10-au2">
                <name>
                  <surname>Knaus</surname>
                  <given-names>Rodger</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author" rid="bid.ch10-au1">
                <name>
                  <surname>Ried</surname>
                  <given-names>Thomas</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author" rid="bid.ch10-au2">
                <name>
                  <surname>Sirotkin</surname>
                  <given-names>Karl</given-names>
                </name>
              </contrib>
            </contrib-group>
            <aff id="bid.ch10-au1">
              <institution>National Cancer Institute (NCI)</institution>
            </aff>
            <aff id="bid.ch10-au2">
              <institution>National Center for Biotechnology Information (NCBI)</institution>
            </aff>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch10d1"/>
            <abstract>
              <title>Summary</title>
              <fig id="bid.499">
                <label>1</label>
                <caption>
                  <title>A metaphase spread after SKY hybridization.</title>
                  <p>The RGB image demonstrates cytogenetic abnormalities in a cell from a secondary leukemia cell line. <italic>Arrows</italic> indicate some of the many chromosomal translocations in this cell line.</p>
                </caption>
                <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch10f1" mime-subtype="jpg"/>
              </fig>
              <p>Spectral Karyotyping (<xref ref-type="other" rid="bid.1401">SKY</xref>) (<xref ref-type="bibr" rid="bid.524">1</xref>&ndash;<xref ref-type="bibr" rid="bid.530">7</xref>) and Comparative Genomic Hybidization (<xref ref-type="other" rid="bid.1261">CGH</xref>) (<xref ref-type="bibr" rid="bid.531">8</xref>&ndash;<xref ref-type="bibr" rid="bid.534">11</xref>) are complementary fluorescent molecular cytogenetic techniques that have revolutionized the detection of chromosomal abnormalities. SKY permits the simultaneous visualization of all human or mouse chromosomes in a different color, facilitating the detection of chromosomal translocations and rearrangements (<xref ref-type="fig" rid="bid.499">Figure 1</xref>). CGH uses the hybridization of differentially labeled tumor and reference DNA to generate a map of DNA copy number changes in tumor genomes.</p>
              <p>The goal of the SKY/CGH database is to allow investigators to submit and analyze both clinical and research (e.g., cell lines) SKY and CGH data. The database is growing and currently has a total of about 700 datasets, some of which are being held private until published. Several hundred labs around the world use this technique, with many more looking at the data they generate. Submitters can enter data from their own cases in either of two formats, public or private; the public data is generally that which has already been published, whereas the private data can be viewed only by the submitters, who can transfer it to the public format at their discretion. The results are stored under the name of the submitter and are listed according to case number. The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/sky/">homepage</ext-link> includes a basic description of SKY and CGH techniques and provides links to a more detailed explanation and relevant literature.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.500">
              <title>Database Content</title>
              <p>Detailed information on how to submit data either to the SKY or CGH sectors of the database can be found through links on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/sky/">homepage</ext-link>. What follows is a brief outline.</p>
              <sec id="bid.501">
                <title>Spectral Karyotyping</title>
                <p>The submitter enters the written <xref ref-type="other" rid="bid.1328">karyotype</xref>, the number of normal and abnormal copies for each chromosome, and the number of cells for each clone. Each abnormal chromosome segment is then described by typing in the beginning and ending bands, starting from the top of the chromosome (<xref ref-type="fig" rid="bid.502">Figure 2</xref>); the computer then builds a colored <xref ref-type="other" rid="bid.1730">ideogram</xref> of this chromosome and eventually a full karyotype (<xref ref-type="other" rid="bid.1402">SKYGRAM</xref>) with each normal and abnormal chromosome displayed in its unique SKY classification color, with band overlay (<xref ref-type="fig" rid="bid.503">Figure 3</xref>). Each breakpoint submitted is automatically linked by a button marked <xref ref-type="other" rid="bid.1293">FISH</xref> to the human Map Viewer (<xref ref-type="fig" rid="bid.504">Figure 4</xref>; <xref ref-type="other" rid="bid.1562">Chapter 20</xref>), which provides a list of genes at that site and available FISH clones for that breakpoint.<fig id="bid.502"><label>2</label><caption><title>SKY data entry form for two different abnormal chromosomes, built segment by segment, for the SKYGRAM image.</title><p>Chromosome images on the <italic>left</italic> are the result of entering the start (<italic>top</italic>) and stop (<italic>bottom</italic>) band for each segment.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch10f2" mime-subtype="gif"/></fig><fig id="bid.503"><label>3</label><caption><title>A SKYGRAM from the SKY/CGH database demonstrating cytogenetic abnormalities in a transitional cell carcinoma of the bladder cell line.</title><p>The written karyotype and a one-line case summary are also provided.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch10f3" mime-subtype="gif"/></fig><fig id="bid.504"><label>4</label><caption><title>Map Viewer image depicting the information on genes, clones, FISH clones, map sequences, and STSs for a specific chromosomal breakpoint (5q13) identified in a SKYGRAM image.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch10f4" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.505">
                <title>Comparative Genomic Hybridization</title>
                <p>The CGH database displays gains, losses, and amplification of chromosomes and chromosome segments. The data are entered either by hand or automatically from various CGH software programs. In the manual format, the submitter enters the band information for each affected chromosome, describing the start band and stop band for each gain or loss, and the computer program displays the final karyotype with vertical bars to the left or right of each chromosome, indicating loss or gain, respectively (<xref ref-type="fig" rid="bid.506">Figure 5</xref>).<fig id="bid.506"><label>5</label><caption><title>A CGH profile from the SKY/CGH database demonstrating copy number gains and losses in a tumor cell line established from a metastatic lymph node in von Hippel-Lindau renal cell carcinoma.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch10f5" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.507">
                <title>Case Information</title>
                <p>Clinical data submitted include case identification, World Health Organization (WHO) disease classification code, diagnosis, organ, tumor type, and disease stage. To obtain the correct classification code, a link is provided to the NCI's <xref ref-type="other" rid="bid.1342">Metathesaurus</xref><sup>TM</sup><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ncievs.nci.nih.gov/NCI-Metaphrase.html">site</ext-link>, which includes, among its many systems, the codes developed by the WHO and NCI, and published as the International Classification of Diseases, 3rd edition (ICD-O-3).</p>
              </sec>
              <sec id="bid.508">
                <title>Reference Information</title>
                <p>The references for the published cases are entered into the Case Information page and are linked to their abstracts in PubMed.</p>
              </sec>
              <sec id="bid.509">
                <title>Tools for Data Entry</title>
                <sec id="bid.510">
                  <title>SKYIN</title>
                  <p>A colored karyotype with band overlay is presented to the submitter, who then builds each aberrant chromosome by cutting and pasting (by clicking with the mouse at appropriate breakpoints) (<xref ref-type="fig" rid="bid.511">Figure 6</xref>). Each aberrant chromosome is then inserted into the full karyotype, which is displayed as a SKYGRAM.<fig id="bid.511"><label>6</label><caption><title>SKYIN format.</title><p>Clicking on a chromosome brings up that chromosome with band overlay. Using the cursor, the operator cuts and pastes together each abnormal chromosome. The abnormal chromosome shown is a combination of chromosomes 10, 16, and 22. <italic>Inset,</italic> an example of the clinical information entered for a case.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch10f6" mime-subtype="gif"/></fig></p>
                </sec>
                <sec id="bid.512">
                  <title>Karyotype Parser</title>
                  <p>To speed up the entry of cytogenetic data into the database, NCBI has built a computer program to automatically read short-form karyotypes, extract the information therein, and insert it into the SKY database (<xref ref-type="fig" rid="bid.513">Figure 7</xref>). Karyotypes are written according to specific rules described in <italic>An International System for Human Cytogenetic Nomenclature (1995)</italic> (<xref ref-type="bibr" rid="bid.535">12</xref>). Using these rules, the parser (1) breaks the karyotype into small syntactic components, (2) assembles information from these components into an information structure in computer memory, (3) transforms this information into the formats required for an application, and (4) uses the information in the application, i.e., inserts it into the database. To accomplish this, the syntactic parser first extracts the information out of each piece of the input; the pieces are then put directly into a tree structure that represents karyotype semantics. For insertion into the SKY database, the karyotype information is transformed into ASN.1 structures that reflect the design of the database.<fig id="bid.513"><label>7</label><caption><title>Karyotype Parser.</title><p>The short-form written karyotype, entered in the karyotype field in (<italic>a</italic>), has been converted into a modified long-form karyotype (<italic>b</italic>), which describes each abnormal chromosome from <italic>top</italic> to <italic>bottom</italic>. Both short and long terms use standardized symbols and abbreviations specified by <xref ref-type="other" rid="bid.1325">ISCN</xref>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch10f7" mime-subtype="gif"/></fig></p>
                </sec>
                <sec id="bid.514">
                  <title>NCI Metathesaurus</title>
                  <p>Data submitters must use the same terminology for diagnosis (morphology) and organ site (topography) to permit comparison or combination of the data in the SKY/CGH database. From the many different disease classification systems, the <italic>International Classification of Diseases for Oncology</italic>, 3rd edition (ICD-O-3)(<xref ref-type="bibr" rid="bid.536">13</xref>) was selected as the database's standard. It contains a morphology tree and a topography tree. In most cases, the submitter must select one term from each tree to fully classify a case. To find and select the correct ICD-O-3 morphology and topography terms, the user is referred to <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ncievs.nci.nih.gov/NCI-Metaphrase.html">NCI Metathesaurus</ext-link>, a comprehensive biomedical terminology database, produced by the NCI Center for Bioinformatics Enterprise Vocabulary Service. This tool facilitates mapping concepts from one vocabulary to other standard vocabularies.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.515">
              <title>Data Analysis: Query Tools</title>
              <sec id="bid.516">
                <title>Quick Search Format</title>
                <p>Quick Search can be found at the top of the SKY/CGH <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/sky/">homepage</ext-link> and can be used for several types of information in the database; these are defined in Searchable Topics in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/sky/help.cgi">Help</ext-link> section. Topics include cytogenetic information (whole chromosome, chromosome arm, or chromosome breakpoint), submitter name, case name, cell line by name, diagnosis, site of disease, treatment, hereditary disorders, mouse strain, and genotype. One or more terms can be entered, and there are options to search SKY alone, CGH alone, SKY AND CGH, or the default, SKY OR CGH.</p>
                <p>The query results page displays information on all relevant cases, clones, and cells, along with details of SKY and/or CGH studies and clinical information for each case.</p>
              </sec>
              <sec id="bid.517">
                <title>Advanced Search Format</title>
                <p>All of the public clinical and cytogenetic information can be searched. This format is currently under development.</p>
              </sec>
              <sec id="bid.518">
                <title>The CGH Case Comparison Tool</title>
                <p>This <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/sky/webcghdrawing.cgi">tool</ext-link> compares and summarizes the CGH profiles from multiple cases on one ideogram. There are numerous criteria that can be used for comparison, such as diagnosis, tumor site, mouse strain, and gain or loss of specific chromosomes, chromosome arms, or chromosome bands.</p>
              </sec>
            </sec>
            <sec id="bid.519">
              <title>Data Integration</title>
              <sec id="bid.520">
                <title>Integration with the NCBI Map Viewer</title>
                <p>All chromosomal bands, including breakpoints involved in chromosomal abnormalities, are linked to the Map Viewer database (<xref ref-type="fig" rid="bid.504">Figure 4</xref> ; see also <xref ref-type="other" rid="bid.1562">Chapter 20</xref>) The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview">Map Viewer</ext-link> provides graphical displays of features on NCBI's assembly of human genomic sequence data as well as cytogenetic, genetic, physical, and radiation hybrid maps. Map features that can be seen along the sequence include NCBI contigs, the BAC <xref ref-type="other" rid="bid.1419">tiling path</xref>, and the location of genes, STSs, FISH mapped clones, ESTs, <xref ref-type="other" rid="bid.1301">GenomeScan</xref> models, and variation (SNPs; see <xref ref-type="other" rid="bid.1143">Chapter 5</xref>).</p>
              </sec>
              <sec id="bid.521">
                <title>SKY/CGH Database Links</title>
                <p>Links are provided to related Web sites including: chromosome databases (e.g., the Mitelman database); other NCI (e.g., <xref ref-type="other" rid="bid.1260">CGAP</xref> and <xref ref-type="other" rid="bid.1254">CCAP</xref>) and NCBI [e.g., the Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>), LocusLink (<xref ref-type="other" rid="bid.788">Chapter 19</xref>) resources; and PubMed (<xref ref-type="other" rid="bid.42">Chapter 2</xref>)] sites; The Jackson Laboratory; and several other CGH sites.</p>
              </sec>
            </sec>
            <sec id="bid.522">
              <title>Contributors</title>
              <p/>
              <p><bold>NCBI:</bold> Karl Sirotkin, Vasuki Gobu, Rodger Knaus, Joel Plotkin, Carolyn Shenmen, and Jim Ostell</p>
              <p><bold>NCI:</bold> Turid Knutsen, Hesed Padilla-Nash, Meena Augustus, Evelin Schr&ouml;ck, Ilan R. Kirsch, Susan Greenhut, James Kriebel, and Thomas Ried</p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.524">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Schr&ouml;ck</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>du Manoir</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Veldman</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Schoell</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Wienberg</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Ferguson-Smith</surname>
                        <given-names>MA</given-names>
                      </name>
                      <name>
                        <surname>Ning</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Ledbetter</surname>
                        <given-names>DH</given-names>
                      </name>
                      <name>
                        <surname>Bar-Am</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Soenksen</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Garini</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>Multicolor spectral karyotyping of human chromosomes</article-title>
                    <source>Science</source>
                    <volume>273</volume>
                    <fpage>494</fpage>
                    <lpage>497</lpage>
                    <year>1996</year>
                    <pub-id pub-id-type="pmid">8662537</pub-id>
                  </citation>
                </ref>
                <ref id="bid.525">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Liyanage</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Coleman</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>du Manoir</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Veldman</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>McCormack</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Dickson</surname>
                        <given-names>RB</given-names>
                      </name>
                      <name>
                        <surname>Barlow</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Wynshaw-Boris</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Janz</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Wienberg</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Ferguson-Smith</surname>
                        <given-names>MA</given-names>
                      </name>
                      <name>
                        <surname>Schr&ouml;ck</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>Multicolour spectral karyotyping of mouse chromosomes</article-title>
                    <source>Nat Genet</source>
                    <volume>14</volume>
                    <fpage>312</fpage>
                    <lpage>315</lpage>
                    <year>1996</year>
                    <pub-id pub-id-type="pmid">8896561</pub-id>
                  </citation>
                </ref>
                <ref id="bid.526">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Liyanage</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>du Manoir</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Heselmeyer</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Auer</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Macville</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Schr&ouml;ck</surname>
                        <given-names>E</given-names>
                      </name>
                    </person-group>
                    <article-title>Tumor cytogenetics revisited: comparative genomic hybridization and spectral karyotyping</article-title>
                    <source>J Mol Med</source>
                    <volume>75</volume>
                    <fpage>801</fpage>
                    <lpage>814</lpage>
                    <year>1997</year>
                    <pub-id pub-id-type="pmid">9428610</pub-id>
                  </citation>
                </ref>
                <ref id="bid.527">
                  <label>4</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Weaver</surname>
                        <given-names>ZA</given-names>
                      </name>
                      <name>
                        <surname>McCormack</surname>
                        <given-names>SJ</given-names>
                      </name>
                      <name>
                        <surname>Liyanage</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>du Manoir</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Coleman</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Schr&ouml;ck</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Dickson</surname>
                        <given-names>RB</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>A recurring pattern of chromosomal aberrations in mammary gland tumors of MMTV&ndash;c-myc transgenic mice</article-title>
                    <source>Genes Chromosomes Cancer</source>
                    <volume>25</volume>
                    <fpage>251</fpage>
                    <lpage>260</lpage>
                    <year>1999</year>
                    <pub-id pub-id-type="pmid">10379871</pub-id>
                  </citation>
                </ref>
                <ref id="bid.528">
                  <label>5</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Knutsen</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>SKY: a comprehensive diagnostic and research tool. A review of the first 300 published cases</article-title>
                    <source>J Assoc Genet Technol</source>
                    <volume>26</volume>
                    <fpage>3</fpage>
                    <lpage>15</lpage>
                    <year>2000</year>
                  </citation>
                </ref>
                <ref id="bid.529">
                  <label>6</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Padilla-Nash</surname>
                        <given-names>HM</given-names>
                      </name>
                      <name>
                        <surname>Heselmeyer-Haddad</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Wangsa</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Ghadimi</surname>
                        <given-names>BM</given-names>
                      </name>
                      <name>
                        <surname>Macville</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Augustus</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Schr&ouml;ck</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Hilgenfeld</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>Jumping translocations are common in solid tumor cell lines and result in recurrent fusions of whole chromosome arms</article-title>
                    <source>Genes Chromosomes Cancer</source>
                    <volume>30</volume>
                    <fpage>349</fpage>
                    <lpage>363</lpage>
                    <year>2001</year>
                    <pub-id pub-id-type="pmid">11241788</pub-id>
                  </citation>
                </ref>
                <ref id="bid.530">
                  <label>7</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Phillips</surname>
                        <given-names>JL</given-names>
                      </name>
                      <name>
                        <surname>Ghadimi</surname>
                        <given-names>BM</given-names>
                      </name>
                      <name>
                        <surname>Wangsa</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Padilla-Nash</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Worrell</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Hewitt</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Walther</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Linehan</surname>
                        <given-names>WM</given-names>
                      </name>
                      <name>
                        <surname>Klausner</surname>
                        <given-names>RD</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>Molecular cytogenetic characterization of early and late renal cell carcinomas in Von Hippel-Lindau (VHL) disease</article-title>
                    <source>Genes Chromosomes Cancer</source>
                    <volume>31</volume>
                    <fpage>1</fpage>
                    <lpage>9</lpage>
                    <year>2001</year>
                    <pub-id pub-id-type="pmid">11284029</pub-id>
                  </citation>
                </ref>
                <ref id="bid.531">
                  <label>8</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Kallioniemi</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Kallioniemi</surname>
                        <given-names>OP</given-names>
                      </name>
                      <name>
                        <surname>Sudar</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Rutovitz</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Gray</surname>
                        <given-names>JW</given-names>
                      </name>
                      <name>
                        <surname>Waldman</surname>
                        <given-names>F</given-names>
                      </name>
                      <name>
                        <surname>Pinkel</surname>
                        <given-names>D</given-names>
                      </name>
                    </person-group>
                    <article-title>Comparative genomic hybridization for molecular cytogenetic analysis of solid tumors</article-title>
                    <source>Science</source>
                    <volume>258</volume>
                    <fpage>818</fpage>
                    <lpage>821</lpage>
                    <year>1992</year>
                    <pub-id pub-id-type="pmid">1359641</pub-id>
                  </citation>
                </ref>
                <ref id="bid.532">
                  <label>9</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Heselmeyer</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Schr&ouml;ck</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>du Manoir</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Blegen</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Shah</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Steinbeck</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Auer</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>Gain of chromosome 3q defines the transition from severe dysplasia to invasive carcinoma of the uterine cervix</article-title>
                    <source>Proc Natl Acad Sci U S A</source>
                    <volume>93</volume>
                    <fpage>479</fpage>
                    <lpage>484</lpage>
                    <year>1996</year>
                    <pub-id pub-id-type="pmid">8552665</pub-id>
                    <pub-id pub-id-type="other">40262</pub-id>
                  </citation>
                </ref>
                <ref id="bid.533">
                  <label>10</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Liyanage</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>du Manoir</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Heselmeyer</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Auer</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Macville</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Schr&ouml;ck</surname>
                        <given-names>E</given-names>
                      </name>
                    </person-group>
                    <article-title>Tumor cytogenetics revisited: comparative genomic hybridization and spectral karyotyping</article-title>
                    <source>J Mol Med</source>
                    <volume>75</volume>
                    <fpage>801</fpage>
                    <lpage>814</lpage>
                    <year>1997</year>
                    <pub-id pub-id-type="pmid">9428610</pub-id>
                  </citation>
                </ref>
                <ref id="bid.534">
                  <label>11</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Forozan</surname>
                        <given-names>F</given-names>
                      </name>
                      <name>
                        <surname>Karhu</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Kononen</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Kallioniemi</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Kallioniemi</surname>
                        <given-names>OP</given-names>
                      </name>
                    </person-group>
                    <article-title>Genome screening by comparative genomic hybridization</article-title>
                    <source>Trends Genet</source>
                    <volume>13</volume>
                    <fpage>405</fpage>
                    <lpage>409</lpage>
                    <year>1997</year>
                    <pub-id pub-id-type="pmid">9351342</pub-id>
                  </citation>
                </ref>
                <ref id="bid.535">
                  <label>12</label>
                  <citation>ISCN: An International System for Human Cytogenetic Nomenclature. In: Mitelman F, editor. Basel: S. Karger; 1995.

</citation>
                </ref>
                <ref id="bid.536">
                  <label>13</label>
                  <citation>Fritz A, Percy C, Jack Andrew, Shanmugaratnam K, Sobin L, Parkin DM, Whelan S, editors. International Classification of Diseases for Oncology, 3rd ed. Geneva: World Health Organization; 2000.

</citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.1776" book-part-type="chapter" book-part-number="11">
          <book-part-meta>
            <title-group>
              <title>The Major Histocompatibility Complex Database, dbMHC</title>
            </title-group>
            <contrib-group>
              <contrib>
                <name>
                  <surname>Kitts</surname>
                  <given-names>Adrienne</given-names>
                </name>
              </contrib>
              <contrib>
                <name>
                  <surname>Feolo</surname>
                  <given-names>Michael</given-names>
                </name>
              </contrib>
              <contrib>
                <name>
                  <surname>Helmberg</surname>
                  <given-names>Wolfgang</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>27</day>
                <month>5</month>
                <year>2003</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch11d1"/>
            <abstract>
              <title>Summary</title>
              <p>One of the most intensely studied regions of the human genome is the Major Histocompatibility Complex (MHC), a group of genes that occupies approximately 4&ndash;6 megabases on the short arm of chromosome 6. The MHC genes, known in humans as Human Leukocyte Antigen (HLA) genes, are highly polymorphic and encode molecules involved in the immune response. The MHC database, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mhc">dbMHC</ext-link>, was designed to provide a neutral platform where the HLA community can submit, edit, view, and exchange MHC data. It currently consists of an interactive Alignment Viewer for HLA and related genes, an MHC microsatellite database (dbMHCms), a sequence interpretation site for Sequencing Based Typing (SBT), and a Primer/Probe database. dbMHC staff are in the process of creating a new database that will house a wide variety of HLA data including:
                <list list-type="bullet">
                  <list-item id="bid.1777">
                    <p>Detailed single nucleotide polymorphism (SNP) mapping data of the HLA region</p>
                  </list-item>
                  <list-item id="bid.1778">
                    <p>KIR gene analysis data</p>
                  </list-item>
                  <list-item id="bid.1779">
                    <p>HLA diversity/anthropology data</p>
                  </list-item>
                  <list-item id="bid.1780">
                    <p>Multigene haplotype data</p>
                  </list-item>
                  <list-item id="bid.1781">
                    <p>HLA/disease association data</p>
                  </list-item>
                  <list-item id="bid.1782">
                    <p>Peptide binding prediction data</p>
                  </list-item>
                </list>
              </p>
              <p>The MHC database is fully integrated with other NCBI resources, as well as with the International Histocompatibility Working Group (IHWG) Web site, and provides links to the IMmunoGeneTics HLA (IMGT/HLA) database.</p>
              <p>This chapter provides a detailed description of current dbMHC resources and reviews dbMHC content and data computation protocols.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.1783">
              <title>Introduction</title>
              <p>The most important known function of HLA genes is the presentation of processed peptides (antigens) to T cells, although there are many genes in the HLA region yet to be characterized, and work continues to locate and analyze new genes in and around the MHC (<xref ref-type="bibr" rid="bid.1908">1</xref>&ndash;<xref ref-type="bibr" rid="bid.1911">4</xref>). The vast amount of work required to elucidate the MHC led to a series of International Histocompatibility Workshop/Conferences (IHWCs). These conferences were the genesis for the IHWG, a group of researchers actively involved in projects that lead to shared data resources for the HLA community (<xref ref-type="bibr" rid="bid.1912">5</xref>&ndash;<xref ref-type="bibr" rid="bid.1923">16</xref>). dbMHC is a permanent public archive that was conceived by the IHWG and is implemented and hosted by NCBI. This database is intended to be expandable; if you have ideas for additional resources, please send an email to <email>helmberg@ncbi.nlm.nih.gov</email>.</p>
              <p>The staff of the MHC database are currently accepting online submissions for the Primer/Probe/Mix component of dbMHC. Submissions can include typing data from any of the following: Sequence Specific Oligonucleotides (SSOs), Sequence Specific Primers (SSPs), SSO and/or SSP mixes, HLA typing kits, and Sequencing Based Typing (SBT). dbMHC allows submitters to edit their submissions online at any time.</p>
            </sec>
            <sec id="bid.1784">
              <title>dbMHC Resources</title>
              <boxed-text id="bid.1785" position="margin">
                <label>1</label>
                <title>dbMHC guests and dbMHC accounts.</title>
                <p>A user can access dbMHC resources as a guest, that is, without having or being a member of an account. However, dbMHC guests will be unable to submit data to dbMHC or edit existing data. Data from a guest session will not be saved from session to session or even from frame to frame. Guests do have the option to download data from a particular session.</p>
                <p>To create a dbMHC account, select the &ldquo;Create an Account&rdquo; from the left sidebar of the dbMHC homepage. Here, you will provide institutional information and specify an account administrator. Only the account administrator is allowed to do the following:
                  <list list-type="bullet">
                    <list-item id="bid.1786">
                      <p>Enter new users.</p>
                    </list-item>
                    <list-item id="bid.1787">
                      <p>View or edit existing users.</p>
                    </list-item>
                    <list-item id="bid.1788">
                      <p>Change user permissions. This includes permission to modify allele reactivities and to enter new primers/probes/mixes or typing kits and to modify existing ones.</p>
                    </list-item>
                    <list-item id="bid.1789">
                      <p>Edit institutional information.</p>
                    </list-item>
                  </list>
                </p>
              </boxed-text>
              <p><xref ref-type="boxed-text" rid="bid.1785">Box 1</xref> contains information on setting up accounts for using dbMHC.</p>
              <sec id="bid.1790">
                <title>Alignment Viewer</title>
                <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://yar.ncbi.nlm.nih.gov:6224/IEB/projects/MHC/luup/MHC.cgi?cmd=aligndisplay&amp;banner=1">Alignment Viewer</ext-link> is designed to display pre-compiled allele sequence alignments in an aligned or FASTA format for selected loci. It offers an interactive display of the alignment, where users can select alleles, highlight SNPs within the alignment, and change between a Codon display and a Decade display (blocks of 10 nucleotides). Users can also switch the type of display, from one where the sequence is completely written out to one where only the differences between selected sequences and a pre-selected reference sequence are seen. All sequences displayed in the Alignment Viewer can be downloaded and formatted as alignment, FASTA, or XML. If the <named-content content-type="book-spec-emph">Alignment</named-content> option is selected, it is up to the user to define how the alignment should be organized. The <named-content content-type="book-spec-emph">Alignment</named-content> option allows the user to download the entire alignment, or just a section of it, and also allows the user to specify how the sequence will be displayed (e.g., in groups of 10 nucleotides or amino acids).</p>
                <p>dbMHC can be used as a tool to help design primers/probes, to evaluate the reactivity patterns of potential probes, and to evaluate the polymorphism content of a particular gene.</p>
              </sec>
              <sec id="bid.1791">
                <title>MHC Microsatellite Database (dbMHCms)</title>
                <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mhc/html/STRsearch2.html">dbMHCms</ext-link> contains many of the known microsatellites across the HLA region and is designed to search for descriptive information, when available, about these highly heterozygotic short tandem repeats (STRs). The MHC microsatellite information that users will be able to extract using dbMHCms includes:
                  <list list-type="bullet">
                    <list-item id="bid.1792">
                      <p>Physical location</p>
                    </list-item>
                    <list-item id="bid.1793">
                      <p>Number and length of known alleles</p>
                    </list-item>
                    <list-item id="bid.1794">
                      <p>Allelic motifs</p>
                    </list-item>
                    <list-item id="bid.1795">
                      <p>Informativity</p>
                    </list-item>
                    <list-item id="bid.1796">
                      <p>Heterozygosity</p>
                    </list-item>
                    <list-item id="bid.1797">
                      <p>Primer sequences</p>
                    </list-item>
                  </list>
                </p>
                <p>Users will find the markers shown on the dbMHCms Web page useful because they provide evidence for genetic linkage and/or association of a genomic region with disease susceptibility. This resource was developed by NCBI in collaboration with A. Foissac, M. Salhi, and Anne Cambon-Thomsen, who provided the original data on MHC microsatellites in a series of updates (<xref ref-type="bibr" rid="bid.1924">17</xref>&ndash;<xref ref-type="bibr" rid="bid.1926">19</xref>). dbMHCms will be expanded to include the STR markers used in the 13th IHWG workshop.</p>
              </sec>
              <sec id="bid.1798">
                <title>Sequencing Based Typing (SBT) Interface</title>
                <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mhc/sbt.cgi?cmd=main">SBT interface</ext-link> is designed to analyze sequence information provided by users and provides an interpretation of the sequence that includes potential allele assignments and division of the sequence into exons and introns. Sequence can be submitted to the SBT interface by copying and pasting or by uploading from a file, either as haploid strands or heterozygote sequence. The current format restricts input to text or FASTA-formatted sequence.</p>
                <p>Submitted sequences are aligned to reference locus sequences located in the dbMHC allele database based on a user-defined degree of nucleotide mismatch. After a brief analysis, the SBT interface will display exons, introns, and untranslated regions for each sequence. Allele assignments are listed according to matching order. Mismatched nucleotide positions are listed separately (in the lower frame of the SBT interface) after selecting <named-content content-type="book-spec-emph">Alignment</named-content>. Note that the SBT interface analyzes one or more sequences for a single locus. If sequences from multiple loci that have identical group-specific amplifications are submitted, the SBT interface must be used to analyze sequences from only one locus at a time.</p>
              </sec>
              <sec id="bid.1799">
                <title>Primer/Probe Database</title>
                <p>This <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mhc/probe.cgi?cmd=probe">database</ext-link> is designed to provide a comprehensive and standardized characterization of individual typing primers and probes, as well as primer and/or probe mixes, and encompasses the following technologies:
                  <list list-type="arabic">
                    <list-item id="bid.1800">
                      <label>1</label>
                      <p>Probe hybridization using Sequence Specific Oligonucleotides (SSOs).</p>
                    </list-item>
                    <list-item id="bid.1801">
                      <label>2</label>
                      <p>DNA amplification using Sequence Specific Primers (SSPs).</p>
                    </list-item>
                    <list-item id="bid.1802">
                      <label>3</label>
                      <p>Sequencing Based Typing (SBT) protocols. The Primer/Probe database can be used as a reference resource for probe design and is intended to provide necessary tools for the exchange of DNA typing data based on the above technologies. This database is composed of a Primer/ Probe Interface and a Typing Kit Interface. Each interface has a series of interactive frames that allow the user to access different database functions.</p>
                    </list-item>
                  </list>
                </p>
                <sec id="bid.1803">
                  <title>Primer/Probe Interface</title>
                  <p>The Primer/Probe Interface provides information on individual reagents used for the typing of MHC or MHC-related loci. Users can access this information through the multiple functional frames within the interface. The words &ldquo;primers&rdquo; and &ldquo;probes&rdquo; represent SSOs, SSPs, SSP mixes, SSO mixes, and their combinations, which include nested PCR or SSO hybridization on group-specific amplifications. Nested PCR and group-specific amplification are both based on a primary amplification with an SSP mix.</p>
                </sec>
                <sec id="bid.1804">
                  <title>Selection of Primers/Probes</title>
                  <p>The process of selecting a primer/probe begins by entering information in the upper frame of the Primer/Probe Interface page. Users enter search options for primers, probes, and mixes, which are grouped by type (SSO, SSP, SSO mix, and SSP mix), locus, and source (submitting institution). Primers/probes can also be selected by entering either the global or local name into the search field. Once this information is entered, the lower frame fills in with the Primer/Probe Listing. Primers/probes that are then selected by the user within the Primer/Probe Listing will be displayed in the large scroll box (in the upper frame) and can be downloaded in XML format.</p>
                </sec>
                <sec id="bid.1805">
                  <title>Primer/Probe Listing</title>
                  <boxed-text id="bid.1806" position="margin">
                    <label>2</label>
                    <title>Global IDs.</title>
                    <p>Upon submission, dbMHC creates a unique global ID number for each primer/probe/mix/kit submitted. This global ID consists of three letters for the submitting institution and seven digits.</p>
                    <p>dbMHC uses the global ID number to store the entire history of a reagent. The global ID number will never change, but every time a user edits the sequence of a submitted primer/probe/mix/kit, that edit is given an incremental version number. Thus, all previous versions of a particular primer/probe/mix/kit are consequently accessible, and submitted primers/probes/mixes/kits can never be deleted.</p>
                    <p>A &ldquo;local name&rdquo; is the identifier given by submitters to their primers and probes.</p>
                  </boxed-text>
                  <p>This function is found in the lower frame of the Primer/Probe Interface page. Results of a Primer/Probe selection are displayed here and will include a unique global ID, local name, source, locus, and type for each primer/probe result. The Global Name (global ID) of each result provides a link to the Primer/Probe View or Mix View functions, and the check box associated with each primer/probe/mix can be used to generate a list of items in the Primer/Probe Interface for subsequent download in XML format.</p>
                  <p><xref ref-type="boxed-text" rid="bid.1806">Box 2</xref> contains information on global IDs.</p>
                </sec>
                <sec id="bid.1807">
                  <title>Primer/Probe View</title>
                  <p>The Primer/Probe View is located in the upper frame of the Primer/Probe Interface and appears as a result of selecting a global ID. This view will display the entire set of data of the SSP and/or SSO primers and/or probes that the user had selected in the Primer/Probe Interface. The data displayed include:
                    <list list-type="bullet">
                      <list-item id="bid.1808">
                        <p>Local name</p>
                      </list-item>
                      <list-item id="bid.1809">
                        <p>Locus as specified by the submitter</p>
                      </list-item>
                      <list-item id="bid.1810">
                        <p>Global ID</p>
                      </list-item>
                      <list-item id="bid.1811">
                        <p>Date of last change (&ldquo;last modified&rdquo;)</p>
                      </list-item>
                      <list-item id="bid.1812">
                        <p>Probe orientation as specified by the submitter</p>
                      </list-item>
                      <list-item id="bid.1813">
                        <p>Corresponding allele sequence (Allele Seq:)</p>
                      </list-item>
                      <list-item id="bid.1814">
                        <p>Reagent type</p>
                      </list-item>
                      <list-item id="bid.1815">
                        <p>Optional filter, i.e., pre-amplification</p>
                      </list-item>
                      <list-item id="bid.1816">
                        <p>Version number</p>
                      </list-item>
                      <list-item id="bid.1817">
                        <p>Probe sequence (Probe Seq:)</p>
                      </list-item>
                      <list-item id="bid.1818">
                        <p>The annealing position of the 3' end of the probe as specified by the submitter</p>
                      </list-item>
                      <list-item id="bid.1819">
                        <p>Probe stringency</p>
                      </list-item>
                    </list>
                  </p>
                  <p>Primer/Probe View offers links that will enable a user to edit primer/probe information, change primer/probe stringency, and allow users to list alleles detected by a probe with a sequence alignment. If a list of alleles detected by a primer/probe without sequence alignment is wanted, Primer/Probe View can do this as well. Select <named-content content-type="book-spec-emph">Listed</named-content>, and the results will be displayed in the Primer/Probe Allele Reactivities Listing.</p>
                </sec>
                <sec id="bid.1820">
                  <title>Primer/Probe Edit</title>
                  <p>The primer/probe edit function allows users to enter new primers/probes into dbMHC or allows users to edit existing primer/probe data. All but the following three fields can be edited (these fields are set by dbMHC):
                    <list list-type="bullet">
                      <list-item id="bid.1821">
                        <p>Global ID</p>
                      </list-item>
                      <list-item id="bid.1822">
                        <p>Version number</p>
                      </list-item>
                      <list-item id="bid.1823">
                        <p>Date of last change</p>
                      </list-item>
                    </list>
                  </p>
                  <p>Within the primer/probe edit frame, there is an alternative input field for primer/probe sequence called Allele Sequence. If primer or probe orientation is set to &ldquo;reverse&rdquo;, the Allele Sequence field will reverse complement the allele sequence as displayed in the alignment to generate the appropriate primer/probe sequence.</p>
                  <p>Before submitting primer/probe data, a user should define the matching stringency of the primer or probe in the Annealing Stringency field. Once the user has selected the matching stringency, the probe can be submitted. Submission of a primer/probe triggers dbMHC to begin an allele reactivity calculation that is based on the primer/probe sequence and the selected matching stringency. The result of this calculation is a list of alleles that might be detected by the submitted primer/probe. The user will find this detectable allele list in the Reactive Allele Listing (lower frame).</p>
                  <p><named-content content-type="book-spec-emph">Alignment</named-content>, on the Primer/Probe View, opens a sequence alignment in the lower frame, which is the reactive allele alignment.</p>
                </sec>
                <sec id="bid.1824">
                  <title>Reactive Allele Listing</title>
                  <boxed-text id="bid.1825" position="margin">
                    <label>3</label>
                    <title>Primer/probe reactivity.</title>
                    <p>Reactivity scores characterize the reactivity between a primer/probe and an allele for primers and probes and are listed below:
                      <list list-type="bullet">
                        <list-item id="bid.1826">
                          <p>Positive: probe anneals with allele.</p>
                        </list-item>
                        <list-item id="bid.1827">
                          <p>Weak: probe anneals sometimes with allele.</p>
                        </list-item>
                        <list-item id="bid.1828">
                          <p>Unclear: annealing cannot be predicted, no empirical information, allele has not been sequenced at the annealing position.</p>
                        </list-item>
                        <list-item id="bid.1829">
                          <p>Negative: probe does not anneal.</p>
                          <p>dbMHC calculates primer/probe reactivity based on sequence similarity only and does not take into account laboratory experimental conditions such as magnesium concentration, temperature, etc.</p>
                        </list-item>
                      </list>
                    </p>
                  </boxed-text>
                  <boxed-text id="bid.1830" position="margin">
                    <label>4</label>
                    <title>Submitter's reactivity scores vs. dbMHC reactivity scores.</title>
                    <p>A submitter's reactivity score can differ from the dbMHC's calculated reactivity score. If there is disparity between the submitter's score and the dbMHC score, users should regard the submitter's score as reliable. In cases of unexplainable disagreement, however, users should contact the submitting institution for further information.</p>
                  </boxed-text>
                  <p>The Reactive Allele Listing is located in the lower frame of the Primer/Probe Interface. It lists the alleles that might be detected by a selected primer/probe or mix (<xref ref-type="boxed-text" rid="bid.1825">Boxes 3</xref> and <xref ref-type="boxed-text" rid="bid.1830">4</xref>).</p>
                </sec>
                <sec id="bid.1831">
                  <title>Reactive Allele Alignment</title>
                  <boxed-text id="bid.1832" position="margin">
                    <label>5</label>
                    <title>Setting allele reactivity scores.</title>
                    <p>A submitter can set individual allele reactivity scores either one-by-one or in a batch. The allele reactivity score for all alleles or for individual alleles that are unedited can be set the same as dbMHC's proposed allele reactivity score.</p>
                    <p>If the user chooses to set the reactivity score as a batch, be aware that alleles within the batch that are not currently displayed will be scored as well. If the user makes a mistake in the scoring, <named-content content-type="book-spec-emph">Reset</named-content> will reset the reactivity scores to the values present at the beginning of the session.</p>
                    <p><named-content content-type="book-spec-emph">Submit</named-content> stores the edited scores in the database. Once submitted, allele reactivity scores cannot be automatically reset to their prior value.</p>
                    <p>Submitting allele reactivity scores will trigger a new dbMHC allele reactivity calculation. If the submitted primer/probe sequence is shorter than 10 nucleotides, dbMHC will use the score information of the alleles to extend the probe sequence.</p>
                  </boxed-text>
                  <p>Selecting <named-content content-type="book-spec-emph">Aligned</named-content> from the Primer/Probe View results in a display of the reactive allele alignment in the lower frame of the Primer/Probe Interface. It displays the alleles that might be detected by a certain primer/probe or mix in alignment with that primer/probe sequence. By default, the first 20 alleles detected by the primer/probe or mix will be displayed. The frame for the reactive allele alignment also allows a submitter to set or edit the Allele Reactivity Score of a primer/probe in this frame. If a submitter chooses to not set a reactivity score, then dbMHC sets the submitter reactivity score to &ldquo;not edited&rdquo;. dbMHC will also calculate a system reactivity score for the primer/probe based on the primer/probe sequence and matching stringency. Colored option boxes indicate dbMHC's reactivity score for the primer/probe, whereas dots within the option boxes indicate the submitter's reactivity score. If the submitter's score (see <xref ref-type="boxed-text" rid="bid.1832">Box 5</xref>) is set and differs from dbMHC's score, the dbMHC score box changes to a warning color.</p>
                  <p>The alignment position of the 3' end of the probe is recorded for each allele. Both the sense and the reverse alignment positions are displayed if the alignment represents an SSP mix.</p>
                  <p>Only users with permission from an account administrator will be able to edit or add additional alleles to a Reactive Allele Listing for a particular primer/probe or mix. To edit the allele reactivity list, mark the &ldquo;edit reactivity list&rdquo; check box in the reactive allele alignment frame.</p>
                </sec>
                <sec id="bid.1833">
                  <title>Mix View</title>
                  <p>The Mix View function is located in the upper frame of the Primer/Probe Interface and can be accessed via the link on the global ID of the Primer/Probe Listing page.</p>
                  <p>It displays an entire set of mix data for selected SSO and SSP mixes:
                    <list list-type="bullet">
                      <list-item id="bid.1834">
                        <p>Local name</p>
                      </list-item>
                      <list-item id="bid.1835">
                        <p>Mix type</p>
                      </list-item>
                      <list-item id="bid.1836">
                        <p>Locus as specified by the submitter</p>
                      </list-item>
                      <list-item id="bid.1837">
                        <p>Optional filter, i.e., pre-amplification</p>
                      </list-item>
                      <list-item id="bid.1838">
                        <p>Global ID</p>
                      </list-item>
                      <list-item id="bid.1839">
                        <p>Version number</p>
                      </list-item>
                      <list-item id="bid.1840">
                        <p>Date of last change</p>
                      </list-item>
                      <list-item id="bid.1841">
                        <p>List of probes as mix elements</p>
                      </list-item>
                      <list-item id="bid.1842">
                        <p>Mix stringency </p>
                      </list-item>
                    </list>
                  </p>
                  <p>The Mix View function contains links to the Mix Element function, where users can change or add elements of a mix. Users will also find links to the Reactive Allele Listing, where they may list alleles detected by a selected mix or selected individual elements of a mix that do not have a sequence alignment. Finally, users will find links to the reactive allele alignment function, where they can view alleles with sequence alignments that are detected by a selected mix or selected individual elements of a mix.</p>
                </sec>
                <sec id="bid.1843">
                  <title>Mix Edit</title>
                  <p>The Mix Edit function is located in the upper frame of the Primer/Probe Interface and allows users to enter new mixes or to edit existing mix data. Users can edit all but the following three fields (these fields are set by dbMHC):
                    <list list-type="bullet">
                      <list-item id="bid.1844">
                        <p>Global ID</p>
                      </list-item>
                      <list-item id="bid.1845">
                        <p>Version number</p>
                      </list-item>
                      <list-item id="bid.1846">
                        <p>Date of last change</p>
                      </list-item>
                    </list>
                  </p>
                  <p>Annealing stringency defines the cumulative matching stringency of the elements of the mix. For SSP mixes, both the sense and the reverse primer must react with the allele with at least the defined stringency.</p>
                </sec>
                <sec id="bid.1847">
                  <title>Mix Element</title>
                  <p>The Mix Element function is located in the lower frame of the Primer/Probe Interface and allows users to add elements to a mix or edit existing elements of a mix. To use the Mix Element function, the user must first specify the mix to be altered in the Mix Edit function and define a source for the primers/probes that the user wants listed. The mix element function displays only SSO probes for SSO mixes and displays SSP probes from SSP mixes in a sense column and a reverse orientation column. For mixes that contain probes from different sources, users must enter the mix elements separately.</p>
                </sec>
              </sec>
              <sec id="bid.1848">
                <title>Typing Kit Interface</title>
                <boxed-text id="bid.1849" position="margin">
                  <label>6</label>
                  <title>Versions of a typing kit.</title>
                  <p>All typing kits are identified by a global ID that is created by dbMHC, which will store the entire history of changes made to a typing kit. An incremental version number is given to every editing session of a typing kit; thus, even if some or all elements of a kit were deleted or altered by a user during an editing session, previous versions of the kit and kit elements will still be available.</p>
                </boxed-text>
                <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mhc/kit.cgi?cmd=main">Typing Kit Interface</ext-link> is the gateway to information on individual typing kits used for typing MHC or MHC-related loci. Users can access this information through multiple functional frames within the interface. Typing kits contained within the Typing Kit Interface consist of SSOs, SSO mixes, or SSP mixes. Elements of typing kits may interact with unamplified DNA, pre-amplified DNA, or several distinct groups of pre-amplified DNA within one locus (e.g., two different amplifications of a certain exon using a distinct variation). The elements of each typing kit will react in characteristic patterns with individual alleles. These patterns can be used to determine allelic variants or a group of allelic variants within a locus (see <xref ref-type="boxed-text" rid="bid.1849">Box 6</xref>).</p>
                <sec id="bid.1850">
                  <title>Selection of the Typing Kit</title>
                  <p>Use the Typing Kit Interface to enter search parameters for typing kits. Typing kits are grouped by:
                    <list list-type="bullet">
                      <list-item id="bid.1851">
                        <p>Type: SSO, SSP</p>
                      </list-item>
                      <list-item id="bid.1852">
                        <p>Locus</p>
                      </list-item>
                      <list-item id="bid.1853">
                        <p>Source: the submitting institution </p>
                      </list-item>
                    </list>
                  </p>
                  <p>Typing kits can also be selected by typing either the global or local name into the search field.</p>
                  <p>Search results, based on user-selected parameters, will be displayed in the Typing Kit Listing, which appears in the lower frame of the Typing Kit Interface. Typing kits selected from the Typing Kit Listing will be displayed in the large scroll box (of the upper frame). They either can be used for combined pattern interpretation of multiple kits or downloaded in XML format.</p>
                </sec>
                <sec id="bid.1854">
                  <title>Typing Kit Listing</title>
                  <p>The Typing Kit Listing is displayed in the lower frame of the Typing Kit Interface. On the basis of the criteria selected, kits are listed with their unique global ID, local name, source, locus, and type. The global ID for each kit provides a link to the Typing Kit View (upper frame of the Typing Kit Interface). The check boxes associated with each typing kit are used to generate a list of items in the Kit Select view for subsequent interpretation of a pattern or for download in XML format.</p>
                </sec>
                <sec id="bid.1855">
                  <title>Typing Kit View</title>
                  <boxed-text id="bid.1862" position="margin">
                    <label>7</label>
                    <title>Creating a virtual kit.</title>
                    <p>If a primer/probe within an existing typing kit malfunctions and a user created a modification of the primer/probe, that modified primer/probe is considered by dbMHC to be an entirely new probe.</p>
                    <p>The kit that contains this modified primer/probe is considered by dbMHC as an entirely new kit. Therefore, when entering a sequence modification of primer/probe within a kit, a submitter must create a new kit by using <named-content content-type="book-spec-emph">Save kit as</named-content> located within New Typing Kit.</p>
                    <p><named-content content-type="book-spec-emph">Save kit as</named-content> will create a copy of the existing kit. The user can then rename the copy of the existing kit, virtually creating a new kit. The user can then go to the kit element frame, Edit Kit Locus Groups and Elements, to remove the old primer/probe from the new kit and replace it with the modified primer/probe.</p>
                  </boxed-text>
                  <p>The Typing Kit View (upper frame of the Typing Kit Interface) lists the entire set of primer/probe data for selected SSO and SSP kits. Typing kit probe data include:
                    <list list-type="bullet">
                      <list-item id="bid.1856">
                        <p>Local name</p>
                      </list-item>
                      <list-item id="bid.1857">
                        <p>Kit type</p>
                      </list-item>
                      <list-item id="bid.1858">
                        <p>Locus (with or without group-specific pre-amplifications)</p>
                      </list-item>
                      <list-item id="bid.1859">
                        <p>Global ID</p>
                      </list-item>
                      <list-item id="bid.1860">
                        <p>Version number</p>
                      </list-item>
                      <list-item id="bid.1861">
                        <p>Batch</p>
                      </list-item>
                    </list>
                  </p>
                  <p>If <named-content content-type="book-spec-emph">List Elements</named-content> from the Typing Kit View is selected, the Typing Kit Elements list provides links to the kit components. Users will also find links to the Edit Kit Locus Groups and Elements page, where they can edit kit information, or they can use <named-content content-type="book-spec-emph">Save kit as...</named-content> to access the New Typing Kit function, where they can create a new kit based on a currently displayed one (see <xref ref-type="boxed-text" rid="bid.1862">Box 7</xref>).</p>
                </sec>
                <sec id="bid.1863">
                  <title>Typing Kit Elements</title>
                  <p>Typing Kit Elements is displayed in the lower frame of the Typing Kit Interface and is used to display all of the elements (components) within a particular typing kit. This function will group the kit elements according to the kit locus groups and will display them as an ordered list. The global ID of each kit element serves as a link to either the Mix View or the Primer/Probe View of this element.</p>
                </sec>
                <sec id="bid.1864">
                  <title>Edit Kit Locus: Groups and Elements</title>
                  <boxed-text id="bid.1865" position="margin">
                    <label>8</label>
                    <title>Typing kit locus &ldquo;groups&rdquo;.</title>
                    <p>HLA typing kits are usually used to detect alleles in one or more loci. Within a typing kit, a particular locus and an optional pre-amplification define what is termed a &ldquo;group&rdquo;. Thus, one typing kit can contain several groups, with each group either consisting of the same locus and a different pre-amplification (or no pre-amplification) or consisting of different loci (with or without pre-amplification).</p>
                  </boxed-text>
                  <p>The Typing Kit Interface <named-content content-type="book-spec-emph">Locus:</named-content> function (upper frame) can be considered a switchboard for assembling a typing kit for dbMHC, although kit name, batch, and type must be edited using the New Typing Kit function, accessed from <named-content content-type="book-spec-emph">New Kit</named-content> on the Typing Kit Interface. The kit's global ID number, version number, and the date of last change are set by dbMHC. Users wanting to edit elements of a particular kit must select a particular locus first and then select <named-content content-type="book-spec-emph">Add locus group</named-content> or <named-content content-type="book-spec-emph">Edit locus group</named-content>, which will open the Edit Locus Group Elements page in the lower frame of the browser (see <xref ref-type="boxed-text" rid="bid.1865">Box 8</xref>).</p>
                </sec>
                <sec id="bid.1866">
                  <title>Edit Locus Group Elements</title>
                  <p>The kit group function is located in the lower frame of the Typing Kit Interface and allows a user to add single elements to, or remove single elements from, a kit group. The kit group frame displays the elements for a particular locus selected by the user in the <named-content content-type="book-spec-emph">kit locus</named-content> function. If the typing kit of interest is an SSO kit, only SSO or SSO mixes will be displayed. If the typing kit is an SSP kit, only SSP mixes will be displayed. Users can then add a primer/probe to the displayed group elements by clicking on a particular primer/probe in the left column and can remove a primer/probe from the group by selecting that element and clicking <named-content content-type="book-spec-emph">Remove</named-content>.</p>
                </sec>
                <sec id="bid.1867">
                  <title>New Typing Kit</title>
                  <p>Access to the kit edit function, <named-content content-type="book-spec-emph">New Kit</named-content>, is found in the upper frame of the Typing Kit Interface and allows a user to enter or modify the name, batch, and type of typing kit. Users may change the kit type so long as the kit does not yet contain any element or group.</p>
                </sec>
                <sec id="bid.1868">
                  <title>Typing Kit Interpretation</title>
                  <p>Access to the typing kit interpretation function, <named-content content-type="book-spec-emph">Interpret</named-content>, is found on the Typing Kit Interface. It provides an online typing pattern interpretation tool. Typing kit reactivity patterns can be analyzed one at a time. Several typing kits that are used to type one sample can be analyzed in combination. An order number represents each element of a typing kit. This number provides a link to the probe reactivity alignment view of each element. The display also indicates whether, and for which elements, an allele-specific amplification has been used. Reactivity patterns can be entered either as a string or via the graphical display. Both the graphical display and the text field will update each other. The string entry accepts &ldquo;1&rdquo;, &ldquo;+&rdquo; for positive reactivity, &ldquo;0&rdquo;,&ldquo;-&rdquo; for negative reactivity, &ldquo;n&rdquo; for not tested, &ldquo;w&rdquo; for weak reactivity, and &ldquo;?&rdquo; for undetermined. The graphic entry symbols are green for positive reactivity, orange for weak reactivity, yellow for undetermined, white for negative reactivity, and gray for &ldquo;not tested&rdquo;. Users can preset the main reactivity by selecting one reactivity option and clicking on <named-content content-type="book-spec-emph">Set All</named-content>. If the option <named-content content-type="book-spec-emph">cycle</named-content> has been selected, repeated clicking on one reactivity field of a kit will cycle through all possible reactivities.</p>
                  <p>The reactivity string or pattern entered will be interpreted as a heterozygote allele combination. Multiple kits can be combined. Users can set the degree of tolerance, which limits the number of false-positive or false-negative reactivities per locus. Each locus is analyzed separately. Allele assignments for each locus are listed according to the number of false-positive/false-negative reactivities.</p>
                </sec>
                <sec id="bid.1869">
                  <title>Typing Kit Pattern</title>
                  <boxed-text id="bid.1870" position="margin">
                    <label>9</label>
                    <title>Some dbMHC/browser limitations.</title>
                   
                      <list list-type="bullet">
                        <list-item id="bid.1871">
                          <p>Because of operating system limitations, the Alignment Viewer can only display up to 300 nucleotides in one line if the browser is Internet Explorer, whereas Netscape is not restricted by this limitation.</p>
                        </list-item>
                        <list-item id="bid.1872">
                          <p>Several essential parts of dbMHC are based on Javascript interaction and dynamic text generation within a page. Users must be aware that many browsers are unable to properly interpret and display Javascript-generated text.</p>
                        </list-item>
                        <list-item id="bid.1873">
                          <p>Netscape version 4.76 does not check for browser content size changes; therefore, users must manually resize to trigger the correct size recognition. Users may resize contents by using a &ldquo;post&rdquo; command, which will lead to a new download of the initial request instead of simply resizing the window.</p>
                        </list-item>
                        <list-item id="bid.1874">
                          <p>Netscape 4.7 may sometimes cause fatal errors and does not allow users to copy and paste sequences from the alignment to the probe sequence field in the probe edit function.</p>
                        </list-item>
                        <list-item id="bid.1875">
                          <p>Internet Explorer version 5.5 and Netscape version 6.2 will correctly interface with dbMHC.</p>
                        </list-item>
                      </list>
                
                  </boxed-text>
                  <p>Selecting <named-content content-type="book-spec-emph">Pattern</named-content> from the Typing Kit Interface results in a display (in the lower frame) of a cross-tab view of allele reactivities of an individual typing kit. Alleles are listed in rows, and kit elements are listed in columns. Each element of a typing kit is represented by the respective order number. This number provides a link to the probe reactivity alignment view of each element. The display also indicates whether and for which elements an allele-specific amplification has been used.</p>
                  <p>A &ldquo;+&rdquo; with a green background signals a detection of an allele by a kit element, a &ldquo;w&rdquo; with an orange background signals a weak detection, a &ldquo;?&rdquo; with a yellow background signals a lack of information, and an &ldquo;r&rdquo; with a red background signals a rejected interaction, although originally suggested by the prediction algorithm. If a kit is designed to detect only a subset of alleles, the display will be limited to this subset (see <xref ref-type="boxed-text" rid="bid.1870">Box 9</xref>).</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.1876">
              <title>Database Content</title>
              <table-wrap id="bid.1877">
                <label>1</label>
                <caption>
                  <title>Data dictionary.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Table name</th>
                    <th align="left" valign="top">Column name</th>
                    <th align="left" valign="top">Data type</th>
                    <th align="left" valign="top">Column comment</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">allele</td>
                    <td align="left" valign="top">allele_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier allele instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">allele_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier for each allele.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">allele_name</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Full allele name defined by IHWG.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">allele_short</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Common use allele name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">allele_group</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Group associated with allele.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">compound_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Number identifying to which compound allele is a member.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Number identifying to which locus allele is a member.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active allele for this allele_id.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">db_version_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">IHWG database version to which this allele belongs.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier or user who created/updated allele.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date when record was created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">block</td>
                    <td align="left" valign="top">block_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for block instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">block_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier for each block.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">block_ord</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Sequence order for each block.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">block_type</td>
                    <td align="left" valign="top">varchar(1)</td>
                    <td align="left" valign="top">Defines block type.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier for locus for which block is a member.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">ref_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Start reference position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">ref_length</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reference length.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">working_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Start working position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">working_length</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Working length.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">block_name</td>
                    <td align="left" valign="top">varchar(20)</td>
                    <td align="left" valign="top">Common name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status of row.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">db_version_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">IHWG database version to which this block belongs.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="top">compound</td>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Locus identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">compound_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Compound number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">job_trans</td>
                    <td align="left" valign="top">trans_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for transaction.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">app_name</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Application name that processes transaction.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">app_arg</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Arguments to processing application.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">create_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Time created.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">start_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Time application started processing transaction.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">end_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Time application completed processing transaction.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">status</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Current status of transaction.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">priority</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Defines priority level.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top">message</td>
                    <td align="left" valign="top">varchar(255)</td>
                    <td align="left" valign="top">Any message that the processing application wants to record.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">kit</td>
                    <td align="left" valign="top">kit_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for kit instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">kit_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Kit identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">kit_local_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Submitter identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">version_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Version number of kit.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">kit_name</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Submitter name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">kit_global_id</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">NCBI-defined name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">kit_batch</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Submitter batch.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">type</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Type.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active kit for this kit_id.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Source identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">kit_locus</td>
                    <td align="left" valign="top">kit_locus_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for kit locus instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">kit_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">kit_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Kit identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Locus identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_order_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Order for loci to be used.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">filter_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Filter identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Tube identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_order_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Order for tube to be used.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active kit locus for this kit id.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">locus</td>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier for locus.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_NCBI_id</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">NCBI locus identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_name</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Common name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">display_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier of allele that should be used as the display reference.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier of reference allele.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_MIM_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">MIM locus identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_pub_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">PubMed locus identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">display_order</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Order number for display.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">mhc_snp</td>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Identifier for locus.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reference position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_offset</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reference offset.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">subsnp_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">dbSNP's ss number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">snp_length</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">SNP's length.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">pos_conv</td>
                    <td align="left" valign="top">compound_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Compound number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">working_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Working position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reference position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_offset</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reference offset.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">blast_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">BLAST position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="middle"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">probe_working</td>
                    <td align="left" valign="top">probe_working_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for each row.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for tube instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">probe_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Tube (probe) identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Active status.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">sequence</td>
                    <td align="left" valign="top">varchar(255)</td>
                    <td align="left" valign="top">Working sequence.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">sequence</td>
                    <td align="left" valign="top">seq_type</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Sequence type.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">seq_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Sequence number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">seq_ord</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Sequence order.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">sequence</td>
                    <td align="left" valign="top">varchar(255)</td>
                    <td align="left" valign="top">Character sequence.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">session</td>
                    <td align="left" valign="top">cur_session_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Current session number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">next_session_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Next avaliable session number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">create_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">source</td>
                    <td align="left" valign="top">source_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier source instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Source identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_code</td>
                    <td align="left" valign="top">varchar(3)</td>
                    <td align="left" valign="top">Source code for use in naming tubes and kits.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">institution</td>
                    <td align="left" valign="top">varchar(50)</td>
                    <td align="left" valign="top">Source name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">admin_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier for administrator for source.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">address</td>
                    <td align="left" valign="top">varchar(50)</td>
                    <td align="left" valign="top">Address.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">city</td>
                    <td align="left" valign="top">varchar(50)</td>
                    <td align="left" valign="top">City.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">state</td>
                    <td align="left" valign="top">varchar(50)</td>
                    <td align="left" valign="top">State.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">postal_code</td>
                    <td align="left" valign="top">varchar(20)</td>
                    <td align="left" valign="top">Postal code.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">country</td>
                    <td align="left" valign="top">varchar(50)</td>
                    <td align="left" valign="top">Country.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">phone</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Phone number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">fax</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Fax number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">email</td>
                    <td align="left" valign="top">varchar(50)</td>
                    <td align="left" valign="top">Email address.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active source for this source id.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">options</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Not used.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">info</td>
                    <td align="left" valign="top">varchar(255)</td>
                    <td align="left" valign="top">Comments/general information.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">tube</td>
                    <td align="left" valign="top">tube_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for tube instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Tube identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_name</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Submitter name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_global_id</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">NCBI-defined name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">filter_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Filter identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Source identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Not used.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">type</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Type.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">sense</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Orientation.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Locus identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Submitted reference position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_offset</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Submitted reference offset.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">stringency</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Stringency set for annealing results.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active tube for this tube_id.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">recalc_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Not used.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">db_version</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">IHWG database version to which this tube belongs.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">info</td>
                    <td align="left" valign="top">varchar(255)</td>
                    <td align="left" valign="top">Comments/general information.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">sequence</td>
                    <td align="left" valign="top">varchar(255)</td>
                    <td align="left" valign="top">Submitted sequence.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">tube_allele</td>
                    <td align="left" valign="top">tube_allele_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for tube allele instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Tube identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">locus_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Locus identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">allele_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Allele identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_status</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User status.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">system_status</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">System status.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reference position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">ref_offset</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reference offset.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">block_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Not used.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">score</td>
                    <td align="left" valign="top">real</td>
                    <td align="left" valign="top">Annealing score.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">working_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Working position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">for_tube_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Forward tube identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">for_working_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Forward working position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">rev_tube_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reverse tube identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">rev_working_pos</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Reverse working position.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">sense</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Orientation.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">seq_known</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Sequence known (sequence fully defined).</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active tube allele for this tube_id allele_id combination.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top"/>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="middle"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">tube_probe</td>
                    <td align="left" valign="top">tube_probe_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for tube probe instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for tube instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">probe_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Tube identifier for sub tube in mix.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">tube_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Tube identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active tube probe for this tube_id probe_id combination.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">users</td>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_nr</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Unique identifier for user instance.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">user_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">User identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">first_name</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">First name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">last_name</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Last name.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">login</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Login.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">pwd</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Password.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Source identifier.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">phone</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Phone number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">fax</td>
                    <td align="left" valign="top">varchar(30)</td>
                    <td align="left" valign="top">Fax number.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">email</td>
                    <td align="left" valign="top">varchar(50)</td>
                    <td align="left" valign="top">Email address.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">create_probe</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Permission to create probe.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">create_kit</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Permission to create kit.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">modify_probe</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Permission to modify probe.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">modify_react</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Permission to modify reactivity.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">modify_kit</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Permission to modify kit.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">info</td>
                    <td align="left" valign="top">varchar(255)</td>
                    <td align="left" valign="top">Comments/general information.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">active</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Status labeling current active user for this user_id.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">createdby_id</td>
                    <td align="left" valign="top">int</td>
                    <td align="left" valign="top">Created by identifier (another user_id).</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">created_date</td>
                    <td align="left" valign="top">datetime</td>
                    <td align="left" valign="top">Date created/updated.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle"/>
                    <td align="left" valign="top">source_admin</td>
                    <td align="left" valign="top">tinyint</td>
                    <td align="left" valign="top">Is user_id an administrator for source?</td>
                  </tr>
                </table>
              </table-wrap>
              <p>The data submitted to dbMHC are stored in a Microsoft SQL (MSSQL) relational database. <xref ref-type="table" rid="bid.1877">Table 1</xref> is a data dictionary that defines dbMHC's database tables or record sets. For each dbMHC table, the data dictionary provides the table name, a column name, the data type in a particular table, and a summary comment.</p>
              <p>The relationships between dbMHC database tables are depicted in <xref ref-type="fig" rid="bid.1878">Figure 1</xref>.<fig id="bid.1878"><label>1</label><caption><title>Physical model of the relationships among dbMHC database tables.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch11f1" mime-subtype="gif"/></fig></p>
              <sec id="bid.1879">
                <title>Primer/Probe Database</title>
                <p>dbMHC's primer/probe database or interface is not curated. As such, the accuracy of the information presented in this database is dependent entirely upon the accuracy of the data submitted to dbMHC.</p>
                <p>We suggest that each primer or probe submitted should be characterized by its complete sequence. In cases where submission of the complete sequence is impossible because the sequence information is considered proprietary, dbMHC offers the option to submit partial sequences. Submitted primer/probe specifications should comply with American Society of Histocompatibility and Immunogenetics (ASHI) (<xref ref-type="bibr" rid="bid.1927">20</xref>) and European Federation of Immunogenetics (EFI) (<xref ref-type="bibr" rid="bid.1928">21</xref>) standards for primers/probes used in histocompatibility DNA testing. If the submitted sequence contains fewer than 10 nucleotides, the submitter must also provide the position of the 3' end of the probe within the alignment of a locus. This additional information is necessary because the likelihood of a primer/probe reactivity calculation producing multiple erroneous alignment positions becomes too great if dbMHC calculates it with fewer than 10 nucleotides.</p>
              </sec>
              <sec id="bid.1880">
                <title>Reagent Allele Reactivity Prediction</title>
                <p>The allele reactivity prediction tool for primers and probes available through dbMHC presents calculations based solely on sequence similarity and can therefore offer only suggestions for the possible allele reactivities of individual primers and probes. Users should not consider this tool to be 100% accurate. A reliable prediction of primer/probe allele reactivities requires complete sequence information, as well as reaction data that include annealing temperature and magnesium concentration. Because these data are not consistently available to dbMHC, we are unable to offer more than a suggestion for primer/probe allele reactivity.</p>
              </sec>
              <sec id="bid.1881">
                <title>Computation of Primer/Probe&ndash;Allele Interaction Stringency</title>
                <p>dbMHC uses sequence comparisons to create a sequence interaction match or stringency grade that is compiled within a penalizing system. dbMHC's starting point for grading the interaction between a primer/probe and an allele is 100%, with each difference in sequence between the primer/probe and the allele causing a reduction in the remaining match grade by a certain percentage. The primary factors dbMHC uses to compute stringency grading include:
                <list list-type="bullet">
                  <list-item id="bid.1882">
                    <p>Nucleotide differences</p>
                  </list-item>
                  <list-item id="bid.1883">
                    <p>Nucleotide position and primer/probe type</p>
                  </list-item>
                </list>
                </p>
                <sec id="bid.1884">
                  <title>Nucleotide Differences</title>
                  <table-wrap id="bid.1885">
                    <label>2</label>
                    <caption>
                      <title>Penalties for nucleotide mismatches.</title>
                    </caption>
                    <table>
                      <tr>
                        <th align="left" valign="top">Probe</th>
                        <th align="center" valign="top" colspan="3">DNA template</th>
                      </tr>
                      <tr>
                        <td align="left" valign="top"/>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">C</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.56</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-green">0</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.42</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.69</named-content>
                        </td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-green">0</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.54</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.42</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.65</named-content>
                        </td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.67</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.63</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.56</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-green">0</named-content>
                        </td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.9</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.94</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-green">0</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">1</named-content>
                        </td>
                      </tr>
                    </table>
                    <table-wrap-foot>
                        <p>Penalties calculated according to Peyret et al. (<xref ref-type="bibr" rid="bid.2012">22</xref>). </p>
                    </table-wrap-foot>
                  </table-wrap>
                  <p>dbMHC divides nucleotide interactions into five categories: perfect match, high match, medium match, poor match, and no match. dbMHC defines these categories by using the purine and pyrimidine interaction between the allele and the primer/probe, as well as the number of hydrogen bonds affected during the virtual pairing of the allele with the primer/probe. See <xref ref-type="table" rid="bid.1885">Table 2</xref>.</p>
                </sec>
                <sec id="bid.1886">
                  <title>Nucleotide Position and Primer/Probe Type</title>
                  <table-wrap id="bid.1887">
                    <label>3</label>
                    <caption>
                      <title>Position penalties for SSP nucleotide mismatches.</title>
                    </caption>
                    <table>
                      <tr>
                        <th align="left" valign="top">5'</th>
                        <th align="left" valign="top">18</th>
                        <th align="left" valign="top">17</th>
                        <th align="left" valign="top">16</th>
                        <th align="left" valign="top">15</th>
                        <th align="left" valign="top">14</th>
                        <th align="left" valign="top">13</th>
                        <th align="left" valign="top">12</th>
                        <th align="left" valign="top">11</th>
                        <th align="left" valign="top">10</th>
                        <th align="left" valign="top">9</th>
                        <th align="left" valign="top">8</th>
                        <th align="left" valign="top">7</th>
                        <th align="left" valign="top">6</th>
                        <th align="left" valign="top">5</th>
                        <th align="left" valign="top">4</th>
                        <th align="left" valign="top">3</th>
                        <th align="left" valign="top">2</th>
                        <th align="left" valign="top">1</th>
                        <th align="left" valign="top">3'</th>
                      </tr>
                      <tr>
                        <td align="left" valign="top"/>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.03</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.03</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.03</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.03</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.03</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.04</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.04</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.05</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.05</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.06</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.07</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.12</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.17</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.21</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.25</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.30</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.50</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.90</named-content>
                        </td>
                        <td align="left" valign="top"/>
                      </tr>
                      <tr>
                        <td align="left" valign="top"/>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Primer</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">C</named-content>
                        </td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top"/>
                      </tr>
                      <tr>
                        <td align="left" valign="top">DNA template</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">A</named-content>
                        </td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top"/>
                      </tr>
                    </table>
                    <table-wrap-foot>
                      <title>Match score calculation for SSPs.</title>
                        <p>An 18-mer primer anneals to a mismatched DNA. Primer and template are shown in the sense orientation. The substitution of C with A leads to the mismatch C-T with a penalty of 0.94 (refer to <xref ref-type="table" rid="bid.1885">Table 2</xref>).</p>
                        <p>This position has a 6% influence on the overall probe reactivity.</p>
                        <p>The overall probe score is 1 &minus; (0.06 &times; 0.94) = 0.94.</p>
                    </table-wrap-foot>
                  </table-wrap>
                  <table-wrap id="bid.1888">
                    <label>4</label>
                    <caption>
                      <title>Position penalties for SSO nucleotide mismatches.</title>
                    </caption>
                    <table>
                      <tr>
                        <th align="left" valign="top">
                          <bold>5'</bold>
                        </th>
                        <th align="left" valign="top">10</th>
                        <th align="left" valign="top">9</th>
                        <th align="left" valign="top">8</th>
                        <th align="left" valign="top">7</th>
                        <th align="left" valign="top">6</th>
                        <th align="left" valign="top">5</th>
                        <th align="left" valign="top">4</th>
                        <th align="left" valign="top">3</th>
                        <th align="left" valign="top">2</th>
                        <th align="left" valign="top">1</th>
                        <th align="left" valign="top">1</th>
                        <th align="left" valign="top">2</th>
                        <th align="left" valign="top">3</th>
                        <th align="left" valign="top">4</th>
                        <th align="left" valign="top">5</th>
                        <th align="left" valign="top">6</th>
                        <th align="left" valign="top">7</th>
                        <th align="left" valign="top">8</th>
                        <th align="left" valign="top">9</th>
                        <th align="left" valign="top">10</th>
                        <th align="left" valign="top">
                          <bold>3'</bold>
                        </th>
                      </tr>
                      <tr>
                        <td align="left" valign="top"/>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.22</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.22</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.34</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.34</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.70</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.70</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-pink">0.80</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.70</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">0.70</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.34</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-yellow">0.34</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.22</named-content>
                        </td>
                        <td align="left" valign="top">
                          <named-content content-type="light-blue">0.22</named-content>
                        </td>
                        <td align="left" valign="top"/>
                      </tr>
                      <tr>
                        <td align="left" valign="top"/>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Probe</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">T</named-content>
                        </td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top"/>
                      </tr>
                      <tr>
                        <td align="left" valign="top">DNA template</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">
                          <named-content content-type="light-orange">C</named-content>
                        </td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">A</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">G</td>
                        <td align="left" valign="top">T</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top">C</td>
                        <td align="left" valign="top"/>
                      </tr>
                    </table>
                    <table-wrap-foot>
                      <title>Match score calculation for SSO probes.</title>
                        <p>A 20-mer SSO anneals to a mismatched DNA. Both probe and template are shown in the sense orientation. The substitution of T with C leads to the probe-template mismatch T-G with a penalty of 0.42 (refer to <xref ref-type="table" rid="bid.1885">Table 2</xref>).</p>
                        <p>This position has a 70% influence on the overall probe reactivity.</p>
                        <p>The overall probe score is 1 &minus; (0.7 &times; 0.42) = 0.71.</p>
                    </table-wrap-foot>
                  </table-wrap>
                  <p>dbMHC uses a penalty system based on nucleotide position in its calculation of SSO/SSP interactions with different alleles. The system dbMHC uses indicates the extent to which an individual nucleotide position will affect the interaction stringency grade in a worst-case scenario, where guanines are in opposing sequence positions or cytosines are in opposing sequence positions. The maximum penalty given in the system is in the middle of an SSO and at the 3' end of an SSP, as shown in <xref ref-type="table" rid="bid.1887">Tables 3</xref> and <xref ref-type="table" rid="bid.1888">4</xref>.</p>
                  <p>If the initial interaction stringency grade for a particular primer/probe&ndash;allele interaction is 100% and a mismatch occurs in a position that carries a penalty of 100%, then dbMHC will reduce the interaction stringency grade to 0%. If the mismatch occurs in a position that carries a penalty of 95%, dbMHC will reduce the interaction stringency grade to 5%. A more detailed <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="ncbi.nlm.nih.gov/mhc/html/annealinghelp.html">example</ext-link> of how dbMHC computes primer/probe interaction stringency is available online.</p>
                  <p>dbMHC's reactive allele alignment algorithm searches for sequences with a match grade above the stringency level set by the submitter of each probe. Currently, the search algorithm searches within all loci that are part of dbMHC. Within each locus, the algorithm constructs compound sequences that represent the combined polymorphic positions of alleles of the same length. Alleles with insertions or deletions are handled separately. If a primer or probe matches with a certain position within the compound, all contributing alleles are checked for that position, whether or not they match.</p>
                  <p>The reactive allele alignment algorithm will check within a specified locus for sequences that have a match grade in accordance with primer/probe specifications at a user-indicated position. The algorithm then extends SSPs toward the 5' end of the probe, to a maximum of 10 nucleotides, and extends both sides of SSOs a maximum of 15 nucleotides, observing all polymorphic positions in matching alleles. The resulting probe extension is then used by the alignment algorithm to check for cross-reactivities within other loci.</p>
                  <p>If the submitted primer/probe sequence is shorter than 10 nucleotides and the user submits a score that accepts certain alleles and rejects others, the reactive allele alignment algorithm will use this information to refine the probe extension. If the probe submitter rejects all alleles containing a unique sequence motif in the vicinity of the probe sequence, the algorithm will generate the extended probe sequence such that it does not match the unique sequence motif.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.1889">
              <title>Integration with Other Resources</title>
              <sec id="bid.1890">
                <title>NCBI Links</title>
                <p>The dbMHC provides a set of internal links to other NCBI resource homepages through either the black &ldquo;quick link&rdquo; bar located at the top of the dbMHC homepage or through the rotating HLA molecule located at the top left-hand corner of the dbMHC homepage. These links will take users to the following NCBI resources:
                  <list list-type="bullet">
                    <list-item id="bid.1891">
                      <p>PubMed</p>
                    </list-item>
                    <list-item id="bid.1892">
                      <p>Nucleotide</p>
                    </list-item>
                    <list-item id="bid.1893">
                      <p>Protein</p>
                    </list-item>
                    <list-item id="bid.1894">
                      <p>OMIM</p>
                    </list-item>
                    <list-item id="bid.1895">
                      <p>BLAST</p>
                    </list-item>
                    <list-item id="bid.1896">
                      <p>Molecular Modeling database (MMdb)</p>
                    </list-item>
                  </list>
                </p>
                <p>The dbMHC Alignment Viewer provides locus-level linkage to NCBI's LocusLink and to dbSNP at the individual SNP level.</p>
              </sec>
              <sec id="bid.1897">
                <title>Graphic View</title>
                <p>The Graphic View page is a hyperlink-enabled representation of the MHC region on chromosome 6. Locus-level links from the Graphic View page include:
                  <list list-type="bullet">
                    <list-item id="bid.1898">
                      <p>dbSNP (at the SNP haplotype level)</p>
                    </list-item>
                    <list-item id="bid.1899">
                      <p>Map View</p>
                    </list-item>
                    <list-item id="bid.1900">
                      <p>LocusLink</p>
                    </list-item>
                    <list-item id="bid.1901">
                      <p>OMIM</p>
                    </list-item>
                    <list-item id="bid.1902">
                      <p>Nucleotide</p>
                    </list-item>
                    <list-item id="bid.1903">
                      <p>Protein</p>
                    </list-item>
                    <list-item id="bid.1904">
                      <p>PubMed</p>
                    </list-item>
                    <list-item id="bid.1905">
                      <p>Structure</p>
                    </list-item>
                    <list-item id="bid.2011">
                      <p>Books</p>
                    </list-item>
                  </list>
                </p>
                <p>The header and link column sections of dbMHC's Graphic View are the same as those on the main page. The content section of this page contains a selection box, called Choose Linked Resource, and three horizontally arranged sections, called Chromosome 6, HLA Class I, and HLA Class II. Genes listed in the HLA Class II section act as hyperlinks to the Web resource selected in Choose Linked Resource.</p>
              </sec>
              <sec id="bid.1906">
                <title>External Links</title>
                <p>Currently, dbMHC provides links (from the left sidebar) to the following external sites:</p>
                <p>
                  <bold>The MHC haplotype project.</bold>
                </p>
                <p><bold>The International Histocompatibility Working Group (IHWG).</bold> Links to individual IHWG projects are listed in the text portion of the homepage.</p>
                <p><bold>The IMmunoGeneTics (IMGT)/HLA database.</bold> Links to the IMGT/HLA database are also located on the allele selection page. Users can access the allele selection page by selecting <named-content content-type="book-spec-emph">Alleles</named-content> on the dbMHC Alignment Viewer. When <named-content content-type="book-spec-emph">Alleles</named-content> is selected, each of the listed allele names on the allele selection page links directly to the IMGT/HLA allele-specific information page at the gene&ndash;allele level.</p>
              </sec>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.1908">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Bodmer</surname>
                        <given-names>J G</given-names>
                      </name>
                      <name>
                        <surname>Marsh</surname>
                        <given-names>S G E</given-names>
                      </name>
                      <name>
                        <surname>Albert</surname>
                        <given-names>E D</given-names>
                      </name>
                      <name>
                        <surname>Bodmer</surname>
                        <given-names>W F</given-names>
                      </name>
                      <name>
                        <surname>Bontrop</surname>
                        <given-names>R E</given-names>
                      </name>
                      <name>
                        <surname>Dupont</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Erlich</surname>
                        <given-names>H A</given-names>
                      </name>
                      <name>
                        <surname>Hansen</surname>
                        <given-names>J A</given-names>
                      </name>
                      <name>
                        <surname>Mach</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Mayr</surname>
                        <given-names>W R</given-names>
                      </name>
                      <name>
                        <surname>Parham</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Petersdorf</surname>
                        <given-names>E W</given-names>
                      </name>
                      <name>
                        <surname>Sasazuki</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Schreuder</surname>
                        <given-names>G M T</given-names>
                      </name>
                      <name>
                        <surname>Strominger</surname>
                        <given-names>J L</given-names>
                      </name>
                      <name>
                        <surname>Svejgaard</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Terasaki</surname>
                        <given-names>P I</given-names>
                      </name>
                    </person-group>
                    <article-title>Nomenclature for factors of the HLA system, 1998</article-title>
                    <source>Tissue Antigens</source>
                    <year>1999</year>
                    <volume>53</volume>
                    <issue>4</issue>
                    <fpage>407</fpage>
                    <lpage>446</lpage>
                    <pub-id pub-id-type="pmid">10321590</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1909">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Robinson</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Bodmer</surname>
                        <given-names>J G</given-names>
                      </name>
                      <name>
                        <surname>Malik</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Marsh</surname>
                        <given-names>S G E</given-names>
                      </name>
                    </person-group>
                    <article-title>Development of the International Immunogenetics HLA Database.</article-title>
                    <source>Hum Immunol</source>
                    <year>1998</year>
                    <volume>59</volume>
                    <issue>Suppl</issue>
                    <fpage>1</fpage>
                    <lpage>17</lpage>
                  </citation>
                </ref>
                <ref id="bid.1910">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Robinson</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Marsh</surname>
                        <given-names>S G E</given-names>
                      </name>
                      <name>
                        <surname>Bodmer</surname>
                        <given-names>J G</given-names>
                      </name>
                    </person-group>
                    <article-title/>
                    <source>Eur J Immunogenet</source>
                    <year>1999</year>
                    <volume>26</volume>
                    <fpage>75</fpage>
                  </citation>
                </ref>
                <ref id="bid.1911">
                  <label>4</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Robinson</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Bodmer</surname>
                        <given-names>J G</given-names>
                      </name>
                      <name>
                        <surname>Marsh</surname>
                        <given-names>S G E</given-names>
                      </name>
                    </person-group>
                    <article-title>The American Society for Histocompatibility and Immunogenetics 25th annual meeting. New Orleans, Louisiana, USA. October 20-24, 1999</article-title>
                    <source>Hum Immunol</source>
                    <year>1999</year>
                    <volume>60</volume>
                    <fpage>S1</fpage>
                    <pub-id pub-id-type="pmid">10549324</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1912">
                  <label>5</label>
                  <citation>Histocompatibility testing: report of a conference and workshop. Washington, DC: National Academy of Sciences - National Research Council, 1965. </citation>
                </ref>
                <ref id="bid.1913">
                  <label>6</label>
                  <citation>Histocompatibility testing 1965. Copenhagen: Munksgaard, 1965.</citation>
                </ref>
                <ref id="bid.1914">
                  <label>7</label>
                  <citation>Curtoni ES, Mattiuz PL, Tosi RM, eds. Histocompatibility testing 1967. Copenhagen: Munksgaard, 1967.</citation>
                </ref>
                <ref id="bid.1915">
                  <label>8</label>
                  <citation>Terasaki PI, ed. Histocompatibility testing 1970. Copenhagen: Munksgaard, 1970.</citation>
                </ref>
                <ref id="bid.1916">
                  <label>9</label>
                  <citation>Dausset J, Colombani J, eds. Histocompatibility testing 1972. Copenhagen: Munksgaard, 1973.</citation>
                </ref>
                <ref id="bid.1917">
                  <label>10</label>
                  <citation>Kissmeyer-Nielsen F, ed. Histocompatibility      testing 1975. Copenhagen: Munksgaard, 1975.</citation>
                </ref>
                <ref id="bid.1918">
                  <label>11</label>
                  <citation>Bodmer WF, Batchelor JR, Bodmer JG, Festenstein H, Morris PJ, eds. Histocompatibility testing 1977. Copenhagen: Munksgaard, 1978.</citation>
                </ref>
                <ref id="bid.1919">
                  <label>12</label>
                  <citation>Terasaki PI, ed. Histocompatibility testing 1980. Los Angeles: UCLA Tissue Typing Laboratory, 1980.</citation>
                </ref>
                <ref id="bid.1920">
                  <label>13</label>
                  <citation>Albert ED, Baur MP, Mayr WR, eds. Histocompatibility testing 1984. Heidelberg: Springer-Verlag, 1984.</citation>
                </ref>
                <ref id="bid.1921">
                  <label>14</label>
                  <citation>Dupont B, ed. Immunobiology of HLA. Vol. I. Histocompatibility testing 1987, and Vol. II. Immunogenetics and histocompatibility. New York: Springer-Verlag, 1988.</citation>
                </ref>
                <ref id="bid.1922">
                  <label>15</label>
                  <citation>Tsuji K, Aizawa M, Sasazuki T, eds. HLA 1991. New York: Oxford University Press, 1992.</citation>
                </ref>
                <ref id="bid.1923">
                  <label>16</label>
                  <citation>Charron, D, ed. Genetic diversity of HLA. Functional and medical implications. Paris: EDK Publishers, 1996.</citation>
                </ref>
                <ref id="bid.1924">
                  <label>17</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Foissac</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Salhi</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Cambon-Thomsen</surname>
                        <given-names>A</given-names>
                      </name>
                    </person-group>
                    <article-title>Microsatellites in the HLA region: 1999 update</article-title>
                    <source>Tissue Antigens</source>
                    <year>2000</year>
                    <volume>55</volume>
                    <issue>6</issue>
                    <fpage>477</fpage>
                    <lpage>509</lpage>
                    <pub-id pub-id-type="pmid">10902606</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1925">
                  <label>18</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Foissac</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Cambon-Thomsen</surname>
                        <given-names>A</given-names>
                      </name>
                    </person-group>
                    <article-title>Microsatellites in the HLA region: 1998 update</article-title>
                    <source>Tissue Antigens</source>
                    <year>1998</year>
                    <volume>52</volume>
                    <issue>4</issue>
                    <fpage>318</fpage>
                    <lpage>352</lpage>
                    <pub-id pub-id-type="pmid">9820597</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1926">
                  <label>19</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Foissac</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Crouau-Roy</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Faure</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Thomsen</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Cambon-Thomsen</surname>
                        <given-names>A</given-names>
                      </name>
                    </person-group>
                    <article-title>Microsatellites in the HLA region: an overview.</article-title>
                    <source>Tissue Antigens</source>
                    <year>1997</year>
                    <volume>49</volume>
                    <issue>3 Pt 1</issue>
                    <fpage>197</fpage>
                    <lpage>214</lpage>
                    <pub-id pub-id-type="pmid">9098926</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1927">
                  <label>20</label>
                  <citation>ASHI Governance. Standards for histocompatibility testing, 1998, P2.151.</citation>
                </ref>
                <ref id="bid.1928">
                  <label>21</label>
                  <citation>European Federation for Immunogenetics. Modifications of the EFI standards for histocompatibility testing, 1996, P2.3100.</citation>
                </ref>
                <ref id="bid.2012">
                  <label>22</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Peyret</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Seneviratne</surname>
                        <given-names>P A</given-names>
                      </name>
                      <name>
                        <surname>Allawi</surname>
                        <given-names>H T</given-names>
                      </name>
                      <name>
                        <surname>SantaLucia</surname>
                        <given-names>J</given-names>
                        <suffix>Jr.</suffix>
                      </name>
                    </person-group>
                    <article-title>Nearest-neighbor thermodynamics and NMR of DNA sequences with internal A.A. C.C, G.G, and T.T mismatches.</article-title>
                    <source>Biochemistry</source>
                    <year>1999</year>
                    <volume>38</volume>
                    <fpage>3468</fpage>
                    <lpage>3477</lpage>
                    <pub-id pub-id-type="pmid">10090733</pub-id>
                  </citation>
                </ref>
              </ref-list>
           </back>
        </book-part>
      </body><?mtl hellip?>
</book-part>
<?mtl hellip?><book-part id="bid.537" book-part-type="part" book-part-number="Part 2">
      <book-part-meta>
        <title-group>
          <title>Data Flow and Processing</title>
        </title-group>
      </book-part-meta>
      <body>
        <book-part id="bid.538" book-part-type="chapter" book-part-number="12">
          <book-part-meta>
            <title-group>
              <title>Sequin: A Sequence Submission and Editing Tool</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Kans</surname>
                  <given-names>Jonathan</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch12d1"/>
            <abstract>
              <title>Summary</title>
              <p>Sequin is a stand-alone sequence record editor, designed for preparing new sequences for submission to GenBank and for editing existing records. Sequin runs on the most popular computer platforms found in biology laboratories, including PC, Macintosh, UNIX, and Linux. It can handle a wide range of sequence lengths and complexities, including entire chromosomes and large datasets from population or phylogenetic studies. Sequin is also used within NCBI by the GenBank and Reference Sequence indexers for routine processing of records before their release.</p>
              <p>Sequin has a modular construction, which simplifies its use, design, and implementation. Sequin relies on many components of the NCBI Toolkit and thus acts as a quality assurance that these functions are working properly.</p>
              <p>Detailed information on how to use Sequin to submit records to GenBank or edit sequence records can be found in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/QuickGuide/sequin.htm">Sequin Quick Guide</ext-link>. Although this chapter will make frequent reference to that help document, the focus will be mostly on the underlying concepts and software components upon which Sequin is built.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.539">
              <title>Sequin: A Brief Overview</title>
              <p>As input, Sequin takes a biological sequence(s) from a scientist wanting to submit or edit sequence data. The sequence (or set of sequences) can be new information that has not yet been assigned a GenBank Accession number, or it can be an existing GenBank sequence record. If Sequin is being used to submit a sequence(s) to GenBank, then the scientist is prompted to include his/her contact information, information about other authors, and the sequence, at the start of the submission process. Once all the necessary information has been entered, it is then possible to view the sequence in a variety of displays and edit it using Sequin's suite of editing tools.</p>
              <p>Sequin is designed for use by people with different levels of expertise. Thus, it has several built-in functions that can, for example, ensure that a new user submits a valid sequence record to GenBank, or it can be prompted to automatically generate a sequence <xref ref-type="other" rid="bid.1273">definition line</xref>. At the other end of the scale, for computer-literate users, Sequin can be customized by the addition of more (perhaps research-specific) analysis functions. Furthermore, there are some extremely powerful functions built into Sequin that are only available to NCBI Indexing staff. These are switched off by default in the public download version of Sequin because they include the ability to make the kinds of changes to a sequence record that can also completely destroy it, if handled incorrectly. These various built-in Sequin functions are discussed further below.</p>
              <p>Sequin's versatility is based on its design: (<italic>a</italic>) Sequin holds the sequence(s) being manipulated in memory, in a structured format that allows a rapid response to the commands initiated by the person who is using Sequin; and (<italic>b</italic>) it makes use of many standard functions found in the <xref ref-type="other" rid="bid.1354">NCBI Toolkit</xref> for both basic data manipulations and as components of Sequin-specific tasks. In particular, Sequin makes heavy use of the Toolkit's object manager, a &ldquo;behind the scenes&rdquo; support system that keeps track of Sequin's internal data structures and the relationships of each piece of information to others. This allows many of Sequin's functions to operate independently of each other, making data manipulation much faster and making the program easier to maintain.</p>
            </sec>
            <sec id="bid.540">
              <title>Sequence Submission</title>
              <p>Sequin is used to edit and submit sequences to GenBank and handles a wide range of sequence lengths and complexities. After <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/download.html">downloading</ext-link> and installing Sequin, a scientist wanting to submit or edit a sequence(s) is led through a series of forms to input information about the sequence to be submitted. The forms are &ldquo;smart&rdquo;, and different forms will appear, customized to the type of submission. Detailed information on how to fill in these forms can be found within the <named-content content-type="book-spec-emph">Help</named-content> feature of the Sequin application itself and in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/QuickGuide/sequin.htm">online Help</ext-link> documents.</p>
              <p>Sequin expects sequence data in <xref ref-type="other" rid="bid.1290">FASTA</xref> formatted files, which should be prepared as plain text before uploading them into Sequin. Population, phylogenetic, and mutation studies can also be entered in <xref ref-type="other" rid="bid.1373">PHYLIP</xref>, <xref ref-type="other" rid="bid.1356">NEXUS</xref>, <xref ref-type="other" rid="bid.1335">MACAW</xref>, or <!--<glossref>-->FASTA+GAP<!--</glossref>--> formats.</p>
              <p>A sequence in FASTA format consists of a definition line, which starts with a &ldquo;&gt;&rdquo;, and the sequence itself, which starts on a new line directly below the definition line. The definition line should contain the name or identifier of the sequence but may also include other useful information. In the case of nucleotides, the name of the source organism and strain should be included; for proteins, it is useful to include the gene and protein names. Given all this information, Sequin can automatically assemble a record suitable for inclusion in GenBank (see below). Detailed information on how to prepare FASTA files for Sequin can be found in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/QuickGuide/sequin.htm#before">Quick Guide</ext-link>.</p>
              <sec id="bid.541">
                <title>Single Sequences</title>
                <p>For single nucleotide sequence submissions to GenBank, the submitter supplies Sequin with the nucleotide sequence and any translated protein sequence(s). For example, a submission consisting of a nucleotide from mouse strain BALB/c that contains the &beta;-hemoglobin gene, encoding the adult major chain &beta;-hemoglobin protein, would have two sequences with the following definition lines, where &ldquo;BALB23g&rdquo; and &ldquo;BALB23p&rdquo; are nucleotide and protein IDs provided by the submitter:
                  <preformat>
    &gt; BALB23g [organism=Mus musculus] [strain=BALB/c]
    &gt; BALB23p [gene=Hbb-b1] [protein=hemoglobin, beta adult major chain]</preformat>
                </p>
                <p>The organism name is essential to make a legal GenBank <xref ref-type="other" rid="bid.1294">flatfile</xref>. It can be included in the definition line as shown above, for the convenience of the submitter, or one of the Sequin submission forms will prompt for its clarity.</p>
                <p>Although it is not necessary to include a protein translation with the nucleotide submission, scientists are strongly encouraged to do so because this, along with the source organism information, enables Sequin to automatically calculate the coding region (<xref ref-type="other" rid="bid.1259">CDS</xref>) on the nucleotide being submitted. Furthermore, with gene and protein names properly annotated, the record becomes informative to other scientists who may retrieve it through a <xref ref-type="other" rid="bid.1246">BLAST</xref> or <xref ref-type="other" rid="bid.1282">Entrez</xref> search (see also <xref ref-type="other" rid="bid.588">Chapter 15</xref> and <xref ref-type="other" rid="bid.610">Chapter 16</xref>).</p>
              </sec>
              <sec id="bid.542">
                <title>Segmented Nucleotide Sets</title>
                <p>A segmented nucleotide entry is a set of non-contiguous sequences that has a defined order and orientation. For example, a genomic DNA segmented set could include encoding exons along with fragments of their flanking introns. An example of an mRNA segmented pair of records would be the 5&prime; and 3&prime; ends of an mRNA where the middle region has not been sequenced. To import nucleotides in a segmented set, each individual sequence must be in FASTA format with an appropriate definition line, and all sequences should be in the same file.</p>
              </sec>
              <sec id="bid.543">
                <title>High-Throughput Genomic Sequences</title>
                <p>Genome Sequencing Centers use automated sequencing machines to rapidly produce large quantities of &ldquo;unfinished&rdquo; DNA sequence, called high-throughput genomic sequence (<xref ref-type="other" rid="bid.1311">HTGS</xref>). These sequences are not usually annotated with any features, such as coding regions, at all, and in the initial phases are not of high (&ldquo;finished&rdquo;) quality.</p>
                <p>The sequencing machines produce intensity traces for the four fluorescent dyes that correspond to the four bases adenine, cytosine, guanine, and thymine. Software such as <xref ref-type="other" rid="bid.1371">PHRED</xref> and <xref ref-type="other" rid="bid.1370">PHRAP</xref> convert these raw traces into the sequence letters A, C, G, or T. PHRED is a base-calling program that &ldquo;reads&rdquo; the sequences of the DNA fragments and produces a quality score. With multiple overlapping reads to work on, PHRAP assembles the DNA fragments using the quality scores of PHRED, itself producing a quality score for each base. The resulting file, which PHRAP outputs in &ldquo;.ace&rdquo; format, consists of the sequence itself plus the associated quality scores. Sequin can use these files as input and assemble valid GenBank records from them. Further information on using Sequin to prepare a HTGS record can be found <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/HTGS/sequininfo.html">here</ext-link>.</p>
              </sec>
              <sec id="bid.544">
                <title>Feature Tables</title>
                <p>Some Genome Centers now analyze their sequences and record the base positions of a number of sequence features such as the gene, mRNA, or coding regions. Sequin can capture this information and include it in a GenBank submission as long as it is formatted correctly in a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sequin/table.html">feature table</ext-link>. Sequin can read a simple, five-column, tab-delimited file in which the first and second columns are the start and stop locations of the feature, respectively, the third column is the type of feature (the feature key&mdash;gene, mRNA, CDS, etc.), the fourth column is the qualifier name (e.g., &ldquo;product&rdquo;), and the fifth the qualifier value (e.g., the name of the protein or gene). The features for an entire bacterial genome can be read in seconds using 
this format. A sample features table is shown below.
                  <preformat>
    &gt;Feature sde3g
    240     4048    gene
                            gene       SDE3
    240     1361    mRNA
    1450    1614
    1730    3184
    3275    4048
                            product    RNA helicase SDE3
    579     1361    CDS
    1450    1614
    1730    3184
    3275    3880
                            product    RNA helicase SDE3
				</preformat>
                </p>
              </sec>
              <sec id="bid.545">
                <title>Alignments</title>
                <p>Population, phylogenetic, and mutation studies all involve the alignment of a number of sequences with each other so that regions of sequence similarity are emphasized. Sometimes it is necessary to introduce gaps into the sequences to give the best alignment. Sequin reads several output formats from sequence and phylogenetic analysis programs, including PHYLIP, NEXUS, PAUP, or FASTA+GAP.</p>
                <p>The submitted sequence alignment represents the relationship between sequences. This inferred relationship allows Sequin to propagate features annotated on one sequence to the equivalent positions on the remaining sequences in the alignment. Feature propagation is one of the many editing functions possible in Sequin. Using this tool significantly reduces the time required to annotate an alignment submission.</p>
              </sec>
              <sec id="bid.1767">
                <title>Automated Submission</title>
                <p>tbl2asn is a program that automates the submission of sequence records to GenBank. It uses many of the same functions as Sequin but is driven entirely by data files, and records need no additional manual editing before submission. Entire genomes, consisting of many chromosomes with feature annotation, can be processed in seconds using this method.</p>
              </sec>
            </sec>
            <sec id="bid.546">
              <title>Packaging the Submissions</title>
              <p>Sequences given to Sequin in the input data formats described in this chapter are retained within Sequin memory, allowing them to be manipulated in real time. For example, for submission to GenBank, the sequence is transformed from the Sequin internal structure to Abstract Syntax Notation 1 (<xref ref-type="other" rid="bid.1242">ASN.1</xref>), the data description language in which GenBank records are stored. This is the format transmitted over the Internet when submitting to GenBank. Sequin can also output information in other formats, such as GenBank flatfile or <xref ref-type="other" rid="bid.1435">XML</xref>, for saving to a local file.</p>
              <p>Most sequence submissions are packaged into a BioseqSet, which contains one or more sequences (Bioseqs), along with supporting information that has been included by the submitter, such as source organism, type of molecule, sequence length, and so on (<xref ref-type="fig" rid="bid.547">Figure 1</xref>). There are different classes of BioseqSets; thus, a simple single nucleotide submission is called a nuc-prot set (a BioseqSet of class nuc-prot) containing the nucleotide and protein Bioseqs. Similarly, population, phylogenetic, and mutation sequences, along with alignments, are packaged into BioseqSets of classes pop-set, phy-set, and mut-set, respectively. The alignment information is extracted into a Seq-align, which is packaged as annotation (Seq-annot) associated with the BioseqSet. In the case of PHRAP quality scores, these are converted into a Seq-graph, which, similar to alignment information, is packaged in a Seq-annot; however, in this case it is associated with the nucleotide sequence and not the higher-level BioseqSet. The Seq-graph of PHRAP scores can be displayed in Sequin's Graphical view.<fig id="bid.547"><label>1</label><caption><title>The internal structure of a sequence record in Sequin, as seen in the Desktop window.</title><p>The display can be understood as a Venn diagram. Selecting the up or down arrow expands or contracts, respectively, the level of detail shown. In a typical submission of a protein-coding gene, a BioseqSet (of class &ldquo;nuc-prot&rdquo;) contains two Bioseqs, one for the nucleotide and one for the protein. Descriptors, such as BioSrc, can be packaged on the set and thus apply to all Bioseqs within the set. Features allow annotation on specific regions of a sequence. For example, the CDS location provides instructions to translate the DNA sequence into the protein product.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch12f1" mime-subtype="gif"/></fig></p>
              <p>Features are usually packaged on the sequence indicated by their location. For example, the gene feature is packaged on the nucleotide Bioseq, and a protein feature is packaged on the protein Bioseq. Proteins are real sequences, and features such as mature peptides are annotated on the proteins in protein coordinates (although they can be mapped to nucleotide coordinates for display in a GenBank flatfile). A CDS (coding region) feature location points to the nucleotide, but the feature product points to the protein. For historical reasons, the CDS is usually packaged on the nuc-prot set instead of on the nucleotide sequence.</p>
            </sec>
            <sec id="bid.548">
              <title>Viewing and Editing the Sequences</title>
              <table-wrap id="bid.549">
                <label>1</label>
                <caption>
                  <title>The display formats available in Sequin.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Format</th>
                    <th align="left" valign="top">Notes</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Alignment</td>
                    <td align="left" valign="top">For sets of aligned sequences</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">ASN.1</td>
                    <td align="left" valign="top">Abstract Syntax Notation 1 format</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Desktop</td>
                    <td align="left" valign="top">The internal structure of a record</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">EMBL</td>
                    <td align="left" valign="top">As the record would appear in the EMBL   database</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">FASTA</td>
                    <td align="left" valign="top">FASTA format</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GBSeq</td>
                    <td align="left" valign="top">XML structured representation of GenBank format</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GenBank</td>
                    <td align="left" valign="top">As the record would appear in GenBank or DDBJ</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GenPept</td>
                    <td align="left" valign="top">Flatfile view of a protein</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Graphic</td>
                    <td align="left" valign="top">Graphical representation of the sequence (several styles are available)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Quality</td>
                    <td align="left" valign="top">Displays the quality scores for each base in biological order</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Sequence</td>
                    <td align="left" valign="top">Nucleotide sequence as letters plus any annotated features</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Summary</td>
                    <td align="left" valign="top">Similar to graphical view but with no labels</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Table</td>
                    <td align="left" valign="top">Sequin's 5-column feature table format</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">XML</td>
                    <td align="left" valign="top">eXtensible Markup Language, representation of ASN.1 data</td>
                  </tr>
                </table>
              </table-wrap>
              <p>After the record has been constructed, the features can be viewed in a variety of display formats (<xref ref-type="table" rid="bid.549">Table 1</xref>). These include the traditional GenBank or GenPept flatfiles, a graphical overview, the feature spans displayed over the actual sequence letters, and ASN.1. These formats are generated by components of the NCBI Toolkit.</p>
              <p>The different format generators all work independently from one another. When Sequin starts up, it registers a set of function procedures used to generate each display format. While issuing Sequin commands during manipulation of the sequence, appropriate messages (for example, &ldquo;generate the view from the internal sequence record&rdquo;, &ldquo;highlight this feature&rdquo;, &ldquo;export the view to the clipboard&rdquo;, etc.) are sent to the viewer by calling one of these procedures. Separate lists of registered formats are maintained for nucleotide, protein, and genome record types.</p>
              <p>Just as the different format generators do not need to know about each other, Sequin's viewer windows do not need to know about other Sequin viewer or editor windows that are active at the same time. When editing a sequence, the user may have several different views of the same sequence open at the same time (for example, a GenBank flatfile and a graphical view). Clicking on a feature in the graphical view will select the same feature in the GenBank flatfile, and double-clicking on a feature launches the specific editor for that feature. This type of communication between different windows is orchestrated by the NCBI Toolkit's object manager.</p>
              <sec id="bid.550">
                <title>The Sequence Editor</title>
                <p>The sequence editor is used like a text editor, with new sequence added at the position of the cursor. Furthermore, the sequence editor automatically adjusts the biological feature intervals as editing proceeds. For example, if 60 bases are pasted or typed onto the 5&prime; end of a sequence record, the sequence editor will shift all the features by 60 bases. This means that interval correction does not need to be done by hand. Prior to Sequin, it was usually easier to resubmit from scratch than to edit all of the feature intervals manually.</p>
              </sec>
              <sec id="bid.551">
                <title>The Feature Editors</title>
                <p>Feature editor windows have a common structure, organized by tabs (Figures 2&ndash;4). The first tab is for elements specific to the given feature (<xref ref-type="fig" rid="bid.552">Figure 2</xref>). For example, in a coding region editor, the first tab has controls for entering the coding region and reading frame. The second tab is for elements common to all sequence features (<xref ref-type="fig" rid="bid.1768">Figure 3</xref>). These include an exception flag, which allows explanations to be given for unusual events (e.g., RNA editing or non-consensus splice sites) and a comment (a free-text statement shown as &ldquo;/note&rdquo; in the flatfile). The last tab is a location spreadsheet allowing multiple feature intervals to be entered (<xref ref-type="fig" rid="bid.1769">Figure 4</xref>). For a coding region, these would reflect the boundaries of the exons used to encode the protein.<fig id="bid.552"><label>2</label><caption><title>The first feature editor tab, <named-content content-type="book-spec-emph">Coding Region</named-content>, allows entry of information specific to the given type of feature.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch12f2" mime-subtype="gif"/></fig><fig id="bid.1768"><label>3</label><caption><title>The <named-content content-type="book-spec-emph">Properties</named-content> tab is for entry of information common to all types of features.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch12f3" mime-subtype="gif"/></fig><fig id="bid.1769"><label>4</label><caption><title>The <named-content content-type="book-spec-emph">Location</named-content> tab has a spreadsheet for entry of feature locations.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch12f4" mime-subtype="gif"/></fig></p>
              </sec>
            </sec>
            <sec id="bid.553">
              <title>Computational Functions of Sequin</title>
              <p>Sequin combines several NCBI Toolkit functions to perform many useful computations on the data in Sequin's memory.</p>
              <sec id="bid.554">
                <title>Automatic Annotation of Coding Regions</title>
                <p>When a nucleotide is submitted to GenBank using Sequin, it is essential to give the name of the source organism.  The submitter is also strongly encouraged to supply the translated protein sequence(s) for the nucleotide.</p>
                <p>Supplying the organism name allows Sequin to automatically find and use the correct genetic code for translating the nucleotide sequence to protein for the most frequently sequenced organisms. On the basis of only the genetic code and sequences of the nucleotide and protein products, Sequin will then calculate the location of the protein-coding region(s) on the nucleotide sequence. This is an extremely powerful function of Sequin. The ability to do this automatically, instead of by hand, has made sequence submission much faster and less error-prone.</p>
                <p>Sequin uses a reverse translating alignment algorithm, called Suggest Intervals, to locate the protein-coding region(s) on the nucleotide sequence. The algorithm builds a table of the positions of all possible stretches of three amino acids in the protein. It then translates the nucleotide in all six reading frames and searches for a match to one of these triplets. When it finds one, it attempts to extend the match on each side of the initial hit. If the extension hits a mismatch or an intron, it stops. Given these candidate regions of matching, Sequin then tries to find the best set of other identical regions that will generate a complete protein. While doing this, the algorithm takes <xref ref-type="other" rid="bid.1407">splice sites</xref> into account when deciding where to start an intron in eukaryotic sequences and fuses regions split by a single amino acid mismatch.</p>
              </sec>
              <sec id="bid.555">
                <title>Updating Sequence Records</title>
                <p>The ability to propagate features through an alignment and the way the sequence editor can adjust feature positions as the sequence is edited are combined in Sequin to provide a simple and automatic method for updating an existing sequence.</p>
                <p>The <named-content content-type="book-spec-emph">Update Sequence</named-content> function allows overlapping sequence or a replacement sequence to be included in an existing sequence record. Sequin makes an alignment (using the BLAST functions in the NCBI Toolkit), merges the sequence if necessary, and propagates features onto the new sequence in the new positions. This effectively replaces the old sequence and features. The <named-content content-type="book-spec-emph">Update Sequence</named-content> function is based on the NCBI Toolkit's alignment indexing, which allows Sequin to produce several displays that help the user to confirm that the correct sequence is in fact being processed (<xref ref-type="fig" rid="bid.556">Figure 5</xref>).
<fig id="bid.556"><label>5</label><caption><title>The Update Sequence window.</title><p>The <italic>top panel</italic> shows the overall relationship between the two sequences, including the parts that align and any parts that do not align. From these data, Sequin determines whether the user is updating with a 5&prime; overlap, 3&prime; overlap, or full replacement sequence, and it presets radio buttons to indicate the relationship. The <italic>second panel</italic> shows a simple graphical view of the positions of gaps and base mismatches in the old and new sequences. The <italic>third panel</italic> shows the same information but with the actual sequence letters. Clicking on the second panel scrolls to the same place in the third panel.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch12f5" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.557">
                <title>The Validator</title>
                <p>The final version of the sequence, complete with all the annotated features, can be checked using the validator. This function checks for consistency and for the presence of required information for submission to GenBank. The validator searches for missing organism information, incorrect coding-region lengths (compared to the submitted protein sequence), internal stop codons in coding regions, mismatched amino acids, and non-consensus splice sites. Double-clicking on an item in the error report launches an editor on the &ldquo;offending&rdquo; feature. The NCBI Toolkit has a program (testval), which is a stand-alone version of the validator.</p>
                <p>The validator also checks for inconsistency between nucleotide and protein sequences, especially in coding regions, the protein product, and the protein feature on the product. For example, if the coding region is marked as incomplete at the 5&prime; end, the protein product and protein feature should be marked as incomplete at the amino end. (Unless told otherwise, the CDS editor will automatically synchronize these anomalies, facilitating the correction of this kind of inconsistency.)</p>
                <p>Additional checks include ensuring that all features are annotated within the range of the sequence, all feature location intervals are noted on the same DNA strand, tRNA codons conform to the given genetic code, and that there are no duplicate features or different genes with the same names. The validator even checks that the sequence letters are valid for the indicated alphabet (e.g., the letter &ldquo;E&rdquo; may appear in proteins but not in nucleotides).</p>
                <p>In cases where an exception has been flagged in a feature editor, specific validator tests can be disabled. For example, if the reason given for an exception is &ldquo;RNA editing&rdquo;, this turns off CDS translation checking in the validator. Likewise, &ldquo;ribosomal slippage&rdquo; disables exon splice checking, and &ldquo;trans splicing&rdquo; suppresses the error message that usually appears when feature intervals are indicated on different DNA strands.</p>
              </sec>
              <sec id="bid.559">
                <title>Automatic Definition Line Generator</title>
                <p>NCBI has a preferred format for the definition line in the GenBank format. It starts with the organism name, then the names of the protein products of coding regions (with the gene name in parentheses), with &ldquo;complete&rdquo; or &ldquo;partial&rdquo; at the end:
                  <preformat>
    DEFINITION Human T-cell lymphotropic virus type 1 isolate ES-TMD envelope 
               glycoprotein (env) gene, partial cds.</preformat>
                </p>
                <p>There is also a standard style for explaining alternative splice products. Sequin's Automatic Definition Line Generator collects CDS, RNA, and exon features in the order that they appear on the nucleotide sequence, finds the relevant genes (usually by location overlap), and prepares a definition line that conforms to GenBank policy.</p>
              </sec>
              <sec id="bid.560">
                <title>Recalculation of Multiple CDS Features</title>
                <p>In spite of all the safeguards built into Sequin, a submitter sometimes uses an incorrect genetic code for an organism. This means that the protein products of CDS translations may be incorrect. Sequin can retranslate all CDS features with a single command. Even so, if the sequence being edited is large, for example, a whole chromosome, this can be a time-consuming operation. To speed it up, the NCBI Toolkit uses a finite state machine (an efficient pattern search algorithm) for rapid translation. The machine is primed with a given genetic code, and then nucleotide sequence letters are fed into the algorithm one at a time, in the order they appear in the sequence. This allows all six frames (three frames on each strand) to be translated in the least possible time. The <named-content content-type="book-spec-emph">Open Reading Frame</named-content> search in the NCBI Toolkit's tbl2asn program also uses this function.</p>
              </sec>
            </sec>
            <sec id="bid.561">
              <title>Advanced Topics</title>
              <sec id="bid.1653">
                <title>The Special Menu</title>
                <p>The Special Menu of Sequin encompasses a powerful set of tools that are available to GenBank and Reference Sequence indexers only. The Special Menu is not available to the public in the standard release because without a thorough understanding of the NCBI Data Model, use of the functions can cause irreparable damage to a record. It allows indexers to globally edit features, qualifiers, or descriptors in all sequences in a record, so that the same correction does not have to be made at each occurrence of the error. For example, all CDS features with internal stop codons can be converted to pseudogenes. Another common error made by submitters is to enter a repeat unit (a rpt_unit; e.g., ATTGG) in the repeat-type field (a rpt_type; e.g., tandem repeat). The Special Menu allows indexing staff to convert rpt_type to rpt_unit throughout the record.</p>
              </sec>
              <sec id="bid.1654">
                <title>NCBI Desktop Window</title>
                <p>Although Sequin has editors for changing specific fields and Special Menu functions for doing bulk changes on several features, it is not possible to anticipate all of the manipulations NCBI indexers might need to do to clean up a problem record. The NCBI Desktop window shows the internal structure of a record (<xref ref-type="fig" rid="bid.547">Figure 1</xref>), i.e., how Bioseqs are packaged within Bioseq-sets, and where features, alignments, graphs, and descriptors are packaged on the sequences or sets. Objects (sets, sequences, features, etc.) can be dragged out of the record or moved to a different place in the record. Such manipulations could break the validity of the sequence record; therefore, great care must be taken when using it.</p>
                <p>For technically adept Sequin users, the Desktop is where additional analysis functions can be added to Sequin without building a complicated user interface. With a feature or sequence selected, items in the Filter menu perform specific analyses on the selected objects. The standard filters include reverse, complement, and reverse complement of a sequence, and reverse complement of a sequence and all of its features. These are needed to repair the occasional record that came in on the wrong strand or in 3&prime;&rarr;5&prime; direction. Adding new filter functions requires adding code to one of the Sequin source code files (one is provided with no other code in it for this purpose) and recompiling the program.</p>
              </sec>
              <sec id="bid.1655">
                <title>Network Analysis Functions</title>
                <p>The functions of Sequin can be expanded by the addition of a configuration file that specifies the URLs for other programs (CGI scripts) available from the Internet. For example, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.genetics.wustl.edu/eddy/tRNAscan-SE/">tRNAscan-SE</ext-link>, a program by Sean Eddy and colleagues at Washington University (St. Louis, MO), can be used on sequences in Sequin in this way.</p>
                <p>At a minimum, the CGI should be able to read FASTA format and should return either sequence data or the five-column feature table (discussed above) as its result. The external programs that Sequin knows about appear in the <named-content content-type="book-spec-emph">Analysis</named-content> menu. When one of these analyses is triggered, Sequin sends a message to the URL and checks for the result with a timer so that the user can continue to work while the (Web) server is processing the request. The code for the CGI that chaperones data from Sequin to tRNAscan-SE and converts the tRNAscan results to a five-column feature table is in the demo directory of the public NCBI Toolkit.</p>
              </sec>
              <sec id="bid.1656">
                <title>Network Fetch Functions</title>
                <p>As sequencing of complete bacterial genomes and eukaryotic chromosomes becomes more commonplace, the demand to break up long sequences into more manageable bites has increased (although Sequin is perfectly capable of editing these large records). Some genomes at NCBI are thus represented as segmented (or delta) Bioseqs, which are composed of pointers to other raw sequences in GenBank. Obtaining the entire sequence and all of the features requires fetching the individual components from a network service.</p>
                <p>The object manager allows Sequin to know about different fetch functions that can be used. When a sequence is needed, these functions will be called until one of them satisfies the request. For example, the lsqfetch configuration file can be edited to point to a directory containing sequence files on a user's disk. The SeqFetch function calls a network service at NCBI to obtain sequences and to look up Accession numbers given gi numbers or gi numbers given Accession numbers.</p>
                <p>When used internally by NCBI indexers, Sequin can also fetch records from the DirSub and TMSmart databases. To ensure the confidentiality of pre-released records, this access requires the indexer to have a database password and to be working from a computer within NCBI. For additional protection, the paths to the database scripts are stored in a configuration file and are not encoded in the public Sequin source code.</p>
              </sec>
            </sec>
            <sec id="bid.1770">
              <title>Conclusion</title>
              <p>Sequin has an important dual role as a primary submission tool and as a full-featured sequence record editor. It is designed to be modular on several levels, which simplifies the design and implementation of its components. Sequin sits at the top of the NCBI software Toolkit, relying on many of the underlying components, and thus acts as a quality assurance that these functions are working properly.</p>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.566" book-part-type="chapter" book-part-number="13">
          <book-part-meta>
            <title-group>
              <title>The Processing of Biological Sequence Data at NCBI</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Sirotkin</surname>
                  <given-names>Karl</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Tatusova</surname>
                  <given-names>Tatiana</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Yaschenko</surname>
                  <given-names>Eugene</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Cavanaugh</surname>
                  <given-names>Mark</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch13d1"/>
            <abstract>
              <title>Summary</title>
              <p>The biological sequence information that makes up the foundation of NCBI's databases and curated resources comes from many sources. How are these data managed and processed once they reach NCBI? This chapter will discuss the flow of sequence data, from the management of data submission to the making of data available for use.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.567">
              <title>Overview</title>
              <table-wrap id="bid.568">
                <label>1</label>
                <caption>
                  <title>Access to genetic sequences at NCBI.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="middle">Access</th>
                    <th align="left" valign="middle">URL</th>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">FTP</td>
                    <td align="left" valign="middle">
                      <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sitemap/index.html#FTPSite">http://www.ncbi.nlm.nih.gov/Sitemap/index.html#FTPSite</ext-link>
                    </td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">BLAST</td>
                    <td align="left" valign="middle">
                      <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/">http://www.ncbi.nlm.nih.gov/BLAST/</ext-link>
                    </td>
                  </tr>
                  <tr>
                    <td align="left" valign="middle">Entrez</td>
                    <td align="left" valign="middle">
                      <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sitemap/index.html#Entrez">http://www.ncbi.nlm.nih.gov/Sitemap/index.html#Entrez</ext-link>
                    </td>
                  </tr>
                </table>
              </table-wrap>
              <p>The sequence database GenBank (<xref ref-type="other" rid="bid.2">Chapter 1</xref>) is made up of nucleotide sequences
		submitted by individual scientists and sequencing centers from around the world. These sequences have been submitted
directly to GenBank or are replicated from one of the collaborating databases,the European Molecular Biology Laboratory (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.embl-heidelberg.de/">EMBL</ext-link>) Data Library or the DNA Data Bank of Japan (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ddbj.nig.ac.jp">DDBJ</ext-link>). Once processed, the data are made available for public use in several ways (<xref ref-type="table" rid="bid.568">Table 1</xref>). The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Nucleotide">Entrez</ext-link>
search and retrieval system (<xref ref-type="other" rid="bid.588">Chapter
14</xref>) can be used to find one or many sequences in GenBank by using text queries into fields that are linked to the raw sequence data, such as author, definition line, organism, and so on. The sequences themselves are searched directly by using <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/">BLAST</ext-link> (<xref ref-type="other" rid="bid.610">Chapter 16</xref>), which takes a sequence as a query to find similar sequences. Large subsets or the latest, complete version of GenBank can also be downloaded by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/">FTP</ext-link>. Underlying the submission, storage, and access processes seen by users of GenBank, BLAST, and other curated data resources [such as the Reference Sequences (<xref ref-type="other" rid="bid.681">Chapter 18</xref>), the Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>), or LocusLink (<xref ref-type="other" rid="bid.788">Chapter 19</xref>)] is an information management system that consists of two major components, the ID database and the IQ database. Whereas ID handles incoming sequences and feeds other databases with subsets to suit different needs, IQ holds links between sequences from ID and links from these sequences to other resources. This chapter discusses these resources in more detail.</p>
              <sec id="bid.1657">
                <title>Abstract Syntax Notation 1 (ASN.1) Is the Data Format Used by the ID System</title>
                <p><xref ref-type="other" rid="bid.1242">ASN.1</xref> is the data description language in which all sequence data at NCBI are structured. ASN.1 allows a detailed description of both the sequences and the information associated with them, e.g., author names, source organism, and sequence features. Maintaining all of the data in the same structured format simplifies data parsing, manipulation, and quality assurance, as well as making data integration and the development of software for sequence analysis easier. All of the various divisions of GenBank can be downloaded in ASN.1 from the NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/ncbi-asn1">FTP site</ext-link>. Similar to an XML <xref ref-type="other" rid="bid.1277">DTD</xref>, ASN.1 has associated with it a description of the legal data structures; this file is called asn.all and is available by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools/">FTP</ext-link> in the ASN directory from the file, ncbi.tar.Z. The DEMO directory of this same file contains the tool, testval, which validates the data against asn.all. There is also a set of utilities for producing ASN.1 while programming in C; these are largely in the subutil.c file of the API directory. In the ID data management system, data are stored as ASN.1 blob, minimizing the amount of biological information that needs be captured and updated in the relational database schema.</p>
              </sec>
              <sec id="bid.1658">
                <title>Sources of Sequence Data</title>
                <p>The sequence data available at NCBI comes from many different sources (<xref ref-type="fig" rid="bid.570">Figure 1</xref>). In summary, the data consist of:<fig id="bid.570"><label>1</label><caption><title>Sources of sequence data available at NCBI.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch13f1" mime-subtype="gif"/></fig></p>
                  <list list-type="arabic">
                    <list-item id="bid.571">
                      <label>1</label>
                      <p>GenBank sequences (<xref ref-type="other" rid="bid.2">Chapter 1</xref>)</p>
                    </list-item>
                    <list-item id="bid.572">
                      <label>2</label>
                      <p>Reference sequences (<xref ref-type="other" rid="bid.681">Chapter 18</xref>)</p>
                    </list-item>
                    <list-item id="bid.573">
                      <label>3</label>
                      <p>Sequences from other databases, such as SWISS-PROT, PIR, PRF, and PDB</p>
                    </list-item>
                    <list-item id="bid.574">
                      <label>4</label>
                      <p>Sequences from United States patents</p>
                    </list-item>
                  </list>
           
                <p>The submission pathway depends on the data source (see <xref ref-type="fig" rid="bid.570">Figure 1</xref>) and volume. Generally, large-volume submitters, such as HTGS, use FTP, often after using tools such as fa2htgs to convert their data to ASN.1. Small-volume submitters typically use either BankIt (<xref ref-type="other" rid="bid.2">Chapter 1</xref>) or Sequin (<xref ref-type="other" rid="bid.538">Chapter 12</xref>) to prepare the ASN.1 for submission.</p>
                <p>The data received are subject to some quality control. The submission tools BankIt, Sequin, and fa2htgs have built-in validation mechanisms that ensure that the data submitted have the correct structure and contain the essential information. The work of the GenBank indexing staff, who also use Sequin, adds an additional layer of quality control and provides assistance with problems or complex submissions and updates.</p>
              </sec>
            </sec>
            <sec id="bid.1659">
              <title>Data Flow Components</title>
              <sec id="bid.580">
                <title>The ID Database</title>
                <p>The ID database is a group of standard relational databases that holds both ASN.1 objects and sequence identifier-related information. ASN.1 objects follow the specifications in the asn.all file for NCBI sequence data objects. ID holds data for GenBank and other data, all of which are included in Entrez. All of the sequences for the International Nucleic Acid Database Collaboration are in GenBank and must have Accession numbers assigned to them. These Accession numbers define an entity, the sequence of which is described. When the understanding of that sequence changes, the sequence can have a new version. This leads to two parallel ways of tracking sequence versions for an object. In the GenBank flatfile format, there is an Accession.Version, where the version gives the ordinal instance (version) of the sequence. Within ID, each unique sequence is assigned a gi number; therefore, the chain of gi numbers for an Accession also gives all of the instances of the sequence for that Accession. The chain identifier within ID can be used to link these gi numbers. (Not all sequences within ID are in GenBank and not all have sequence versions, but all sequences have a chain of gi numbers. For this reason, internally, the gi number is the universal pointer to a particular sequence, as opposed to the Accession.Version, which would work only for versioned sequences.) The ID database is also the controller for allowed &ldquo;takeovers&rdquo; of one Accession by another. A takeover would occur, for example, when two clones' sequences were now to be merged into a single clone. One or both of the two older clones' Accessions could be taken over by the new clone. The actual details of the architecture of this group of databases and software associated with it are described later in this chapter.</p>
              </sec>
              <sec id="bid.1660">
                <title>Output of Data from the ID System</title>
                <p>Once all incoming data have been converted to ASN.1 format and have been entered into ID, the data are then replicated to several different servers and also transformed into several different formats (<xref ref-type="fig" rid="bid.579">Figure 2</xref>). The replication is necessary for a number of reasons: (1) it separates the &ldquo;incoming&rdquo; data system (ID) from the &ldquo;outgoing&rdquo; data, i.e., the data served in response to scientific queries by users; (2) replicating the data to different servers helps balance the load of queries, thus providing quicker response times, and allows different servers to specialize in different functions; and (3) it protects against data loss; having copies kept at different locations means that the data are safe, should one server fail.<fig id="bid.579"><label>2</label><caption><title>Products of the ID system.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch13f2" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.1661">
                <title>The IQ Database</title>
                <p>The IQ database is a <xref ref-type="other" rid="bid.1413">Sybase</xref> data-warehousing product that preserves its SQL language interface but inverts the data it stores so that the data are stored by column, not by row. Its strength is that it can speed up results from queries based on the anticipated indexing. This non-relational database holds links between many different objects.</p>
                <p>For example, as part of the processing of incoming sequences, each protein and nucleotide sequence is searched (using BLAST, <xref ref-type="other" rid="bid.610">Chapter 16</xref>) against the rest of the database for similar sequences. The resulting set of similar sequences (sometimes known as &ldquo;neighbors&rdquo;) are made available to users by selecting the <named-content content-type="book-spec-emph">related sequences</named-content> link next to a record in Entrez (<xref ref-type="other" rid="bid.588">Chapter 15</xref>). The IQ database keeps track of the neighbors for any given sequence; these relationships are all precomputed to save time for the people using this feature.</p>
                <p>As well as storing the relationships between nucleotide or between proteins, IQ also holds information on the links between entries in different Entrez databases. This might include, for example, information on the publications cited within sequence records (which link to PubMed), or to an organism in the Taxonomy database. Some of this information comes from the analysis of the ASN.1 in ID by e2index.</p>
              </sec>
              <sec id="bid.581">
                <title>The BLAST Control Database</title>
                <p>The BLAST Control database receives information from ID and uses it to generate BLAST databases (<xref ref-type="other" rid="bid.610">Chapter 16</xref>) for the BLAST query service and for stand-alone BLAST users, as well as for internal use, to generate the sequence neighbors detailed in IQ.</p>
              </sec>
              <sec id="bid.582">
                <title>The GenBank Flatfile and Error Capture Databases</title>
                <p>Many people who use NCBI services think of the GenBank flatfile as the archetypal sequence data format. However, in terms of the ID internal data flow system, ASN.1 is considered the original format from which reports such as the GenBank flatfile can be generated. Generally, the GenBank flatfile is generated on demand from the ASN.1. However, for certain products, such as complete GenBank releases, a GenBank flatfile image is made for each active sequence and is stored in a database called FF4Release, which consists of the latest transformation of ASN.1 to the GenBank flatfile format.</p>
                <p>This database also doubles as the place where internal error reports are captured. The reports can be analyzed and displayed for different time points in the data processing pathway: (<italic>a</italic>) ASN.1 itself can be validated using the testval tool. (Syntax checking is not necessary, because the underlying ASN.1 libraries enforce proper syntax according to the definition file.); (<italic>b</italic>) errors can be discovered during conversion to GenBank flatfile format; and (<italic>c</italic>) through a reparse from GenBank flatfile format to ASN.1. This is done as a further check for legality of the ASN.1, and our current software for producing GenBank format reports from it.</p>
              </sec>
              <sec id="bid.583">
                <title>Entrez Postings Files</title>
                <table-wrap id="bid.584">
                  <label>2</label>
                  <caption>
                    <title>Products of the ID system.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="middle">Type</th>
                      <th align="left" valign="middle">Source</th>
                      <th align="center" valign="middle">ASN.1</th>
                      <th align="center" valign="middle">GBFF<xref ref-type="fn" rid="mul12"><sup><italic>a</italic></sup></xref></th>
                      <th align="center" valign="middle">Qscore</th>
                      <th align="center" valign="middle">GenPept</th>
                      <th align="center" valign="middle">Protein FASTA</th>
                    </tr>
                    <tr>
                      <td align="left" valign="middle">Cumulative</td>
                      <td align="left" valign="middle">GenBank</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle"/>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="middle">Incremental</td>
                      <td align="left" valign="middle">GenBank</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle"/>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="middle">Incremental</td>
                      <td align="left" valign="middle">GenBank<xref ref-type="fn" rid="mul13"><sup><italic>b</italic></sup></xref></td>
                      <td align="center" valign="middle"/>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle"/>
                      <td align="center" valign="middle"/>
                    </tr>
                    <tr>
                      <td align="left" valign="middle">Cumulative</td>
                      <td align="left" valign="middle">RefSeq</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle"/>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="middle">Incremental</td>
                      <td align="left" valign="middle">RefSeq</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle"/>
                      <td align="center" valign="middle">X</td>
                      <td align="center" valign="middle">X</td>
                    </tr>
                  </table>
                  <table-wrap-foot>
                    <fn symbol="a" id="mul12">
                      <p>GBFF, GenBank flatfile.</p>
                    </fn>
                    <fn symbol="b" id="mul13">
                      <p>NCBI records only.</p>
                    </fn>
                  </table-wrap-foot>
                </table-wrap>
                <p>When sequences are submitted to GenBank or one of the collaborating databases, useful information about the sequence is often also included. This might include a brief description of a gene in the definition line, along with annotated sequence features, e.g., the source organism name. To make this information searchable via Entrez, these words have to be indexed. They are therefore extracted from the ASN.1 (using a tool called e2index) and are then stored in the Entrez posting files, which are optimized for Boolean queries by the Entrez system (see <xref ref-type="other" rid="bid.588">Chapter 15</xref>).</p>
                <p>All of these products from the ID system are listed in <xref ref-type="table" rid="bid.584">Table 2</xref>. NCBI also generates weekly &ldquo;LiveLists&rdquo; for public, collaborator, and in-house use. LiveLists show currently known Accession numbers that are in use (i.e., they have not been replaced or otherwise removed from circulation for one reason or another).</p>
              </sec>
            </sec>
            <sec id="bid.585">
              <title>Data Flow Architecture</title>
              <p>ID has a distributed architecture, the features of which are shown in <xref ref-type="fig" rid="bid.586">Figure 3</xref>. The system is activated when a client (internal to NCBI) is ready to load data into the ID system. A stand-alone program takes ASN.1-containing files, while a client API can submit data to ID through the IDProdOS. The rationale behind the various architecture components is discussed below.</p>
              <p>IDProdOS is an open server (commonly called &ldquo;middleware&rdquo;), which sits between the clients and the database system. It hides details of the underlying complexity from the client API. For example, the previous version of the ID system consisted of a single database. &ldquo;Using the Open Server&rdquo; meant that when the conversion to the current system took place, only the open server had to be changed, and none of the clients. IDProdOS begins the complex process of checking that the actions implied by the load are allowed and changes the SeqId pointers to gi numbers to be used as sequence-specific pointers. For example, in a record that has DNA and protein sequences, including annotation and sequence identifiers, the identifier on the protein has to be unique. If the identifier has been used previously on another DNA sequence and the current sequence is not replacing the previous sequence, then this is not allowed because proteins are not allowed to move between GenBank records. As an exception, this rule is sometimes relaxed for proteins moving between segments of a complete genome submission.</p>
              <p>The IdMain database contains the sequence identifiers for each of the sequence records, including all those for ASN.1 blobs that contain multiple sequences. It enforces sequence version rules, among others.</p>
              <p>Relational satellite databases are fully normalized databases that hold records for which there is only one sequence per intended ASN.1 blob. Few, if any, features are allowed on records intended for relational satellite databases. (The PubSeqOS produces the ASN.1 by querying the relational tables.) This is in contrast to the Blob satellite databases, from which ASN.1 is retrieved pre-made.</p>
              <p>Blob satellite databases, although relational databases, contain ASN.1 objects as unnormalized data objects.</p>
              <p>The SnpAnnot database contains only feature information. It has simple mutation information from dbSNP (<xref ref-type="other" rid="bid.1143">Chapter 5</xref>), which is added to NCBI-curated records by the PubSeqOS when the records are requested.</p>
              <p>To visualize the role of replication, the rectangle in the middle of <xref ref-type="fig" rid="bid.586">Figure 3</xref> represents the use of the Sybase Replication Server to copy information from the loading side of the system to the query side.</p>
              <p>PubSeqOS is an open server (commonly called &ldquo;middleware&rdquo;) that sits between the clients and the database system. It hides details of the underlying complexity from the client API. It actually has an almost identical code base to IDProdOS because it serves similar functions. When the ASN.1 that has been requested is to be presented in a format other than ASN.1, psansconvert is called to do the conversion. This separate process allows both insulation from any possible instability and allows for use of multiple central processing units in a natural way.</p>
              <p>The EntrezControl and The Graveyards databases are uniquely on the query side. The EntrezControl database controls information on records that await indexing and have been indexed by e2index.The Graveyards are blob databases and contain those blobs killed by replacement or that were taken over by other blobs. Once taken over, blobs are not changed; therefore, they need not be on the loading side.<fig id="bid.586"><label>3</label><caption><title>The ID system architecture.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch13f3" mime-subtype="gif"/></fig></p>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.1440" book-part-type="chapter" book-part-number="14">
          <book-part-meta>
            <title-group>
              <title>Genome Assembly and Annotation Process</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Kitts</surname>
                  <given-names>Paul</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch14d1"/>
            <abstract>
              <title>Summary</title>
              <boxed-text id="bid.1742" position="margin">
                <label>1</label>
                <title>Annotation of other genomes.</title>
                <p>NCBI may assemble a genome prior to annotation, add annotations to a genome assembled elsewhere, or simply process an annotated genome to produce RefSeqs and maps for display in Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>).</p>
                <p>The basic procedures used to annotate other eukaryotic genomes are essentially the same as those used to annotate the human genome. However, the overall process is adjusted to accommodate the different types of input data that are available for each organism. Genes can be annotated on any genome for which a significant number of mRNA, EST, or protein sequences are available. Other features, such as clones, STS markers, and SNPs, can also be annotated whenever the relevant data are available for an organism.</p>
                <p>For example, genes and other features are placed on the mouse Whole Genome Shotgun (WGS) assembly from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ensembl.org/Mus_musculus/credits.html">Mouse Genome Sequencing Consortium</ext-link> (MGSC) by skipping the assembly steps used in the human process but following the annotation steps with relatively minor adjustments. A variation of the human process is also used to assemble and annotate genomic contigs from finished mouse clone sequences (see the Map Viewer <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?ORG=mouse&amp;chr=1">display</ext-link> of the mouse genome).</p>
              </boxed-text>
              <p>The primary data produced by genome sequencing projects are often highly fragmented and sparsely annotated. This is especially true for the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.genome.gov/page.cfm?pageID=10001772">Human Genome Project</ext-link> as a result of its policy of releasing sequence data to the public sequence databases every day (<xref ref-type="bibr" rid="bid.1536">1</xref>, <xref ref-type="bibr" rid="bid.1537">2</xref>). So that individual researchers do not have to piece together extended segments of a genome and then relate the sequence to genetic maps and known genes, NCBI provides annotated assemblies of public genome sequence data. NCBI assimilates data of various types, from numerous sources, to provide an integrated view of a genome, making it easier for researchers to spot informative relationships that might not have been apparent from looking at the primary data. The annotated genomes can be explored using Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>) to display different types of data side-by-side and to follow links between related pieces of data.</p>
              <p>This chapter describes the series of steps, the &ldquo;pipeline&rdquo;, that produces NCBI's annotated genome assembly from data deposited in the public sequence databases. A variant of the annotation process developed for the human genome is used to annotate the mouse genome, and similar procedures will be applied to other genomes (<xref ref-type="boxed-text" rid="bid.1742">Box 1</xref>).</p>
              <p>NCBI constantly strives to improve the accuracy of its human genome assembly and annotation, to make the data displays more informative, and to enhance the utility of our access tools. Each run through the assembly and annotation procedure, together with feedback from outside groups and individual users, is used to improve the process, refine the parameters for individual steps, and add new features. Consequently, the details of the assembly and annotation process change from one run to the next. This chapter, therefore, describes the overall human genome assembly and annotation process and provides short descriptions of the key steps, but it does not detail specific procedures or parameters. However, sufficient detail is provided to enable users of our assembly and annotations to become familiar with the complexities and possible limitations of the data we provide.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.1441">
              <title>Overview of the Genome Assembly and Annotation Process</title>
              <p><xref ref-type="fig" rid="bid.1442">Figure 1</xref> shows how the main steps in the human genome assembly and annotation process are organized and also shows the most significant interdependencies between the steps. The pipeline is not linear, because whenever possible, steps are performed in parallel to reduce the overall time taken to produce an annotated assembly from a new set of data. Some of the steps are run incrementally on a timetable that is independent from that of the main pipeline to produce a new assembly more quickly.<fig id="bid.1442"><label>1</label><caption><title>The human genome assembly and annotation process.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch14f1" mime-subtype="gif"/></fig></p>
              <sec id="bid.1443">
                <title>Data Freeze</title>
                <p>New sequence data that could be used to improve the genome assembly and annotation become available on a daily basis. Since the assembly and annotation process takes several weeks to complete, the data are &ldquo;frozen&rdquo; at the start of the build process by making a copy of all of the data available for use at that time. Freezing the data provides a stable set of inputs for the remainder of the build process. Additional or revised data that become available during the period taken to complete the process are not used until the next build.</p>
              </sec>
              <sec id="bid.1444">
                <title>The Build Cycle</title>
                <p>A build begins with a freeze of the input data and ends with the public release of an annotated assembly of genomic sequences (<xref ref-type="fig" rid="bid.1442">Figure 1</xref>). There are usually a few months between builds so that the latest build can be evaluated and improvements can be made.</p>
              </sec>
              <sec id="bid.1445">
                <title>Steps Run Incrementally</title>
                <p>The early steps of the genome assembly and annotation pipeline involve many computationally intensive processes, including masking the repetitive sequences and aligning each genomic sequence to the other genomic sequences, mRNAs, and Expressed Sequence Tags (<xref ref-type="other" rid="bid.1283">EST</xref>s). Running these steps incrementally minimizes the time between starting a new build and being ready to start the assembly and annotation steps. Approximately once a week, independent from the build cycle, new or updated sequences that have been deposited in GenBank are retrieved for processing. Periodically, old versions of the sequences are purged from the set of accumulated data files.</p>
                <p>The manual refinement of the set of assembled genomic <xref ref-type="other" rid="bid.1267">contig</xref>s produced entirely from <xref ref-type="other" rid="bid.1292">finished</xref> sequence is another time-consuming step that is carried out incrementally, approximately once a week.</p>
              </sec>
              <sec id="bid.1446">
                <title>Steps Run Irregularly</title>
                <p>Because some data change infrequently, some relatively quick steps are executed on time frames that are not tied to the build cycle. For example, the list of special cases used to override the automatic process is updated whenever the need becomes apparent.</p>
              </sec>
            </sec>
            <sec id="bid.1447">
              <title>The Input Data</title>
              <p>The main inputs for the genome assembly and annotation process are genomic sequences, transcript sequences, and Sequence Tagged Site (STS) maps.</p>
              <sec id="bid.1448">
                <title>Human Genomic Sequences</title>
                <sec id="bid.1449">
                  <title>
                    <bold>Genomic Sequences Used for Assembly</bold>
                  </title>
                  <p>Genomic sequences from the following five data sets are processed for use in the assembly:</p>
                  <p><bold>High-Throughput Genomic Sequences.</bold> Human high-throughput genomic sequences (<xref ref-type="other" rid="bid.1311">HTGS</xref>s) are retrieved using the <xref ref-type="other" rid="bid.1282">Entrez</xref> query system (<xref ref-type="other" rid="bid.588">Chapter 15</xref>). The query used returns sequences for all entries that contain the HTG keyword, regardless of whether the sequence is <xref ref-type="other" rid="bid.1292">finished</xref> or is in any of the unfinished <xref ref-type="other" rid="bid.1276">draft</xref> phases.</p>
                  <p><bold>Finished Chromosome Sequences.</bold> The center that coordinates the sequencing of a finished chromosome submits a specification regarding how to build the sequence of that chromosome from its component clone sequences to a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://cvsweb.sanger.ac.uk/cgi-bin/cvsweb.cgi/genome-map/data/?cvsroot=Ensembl">data repository</ext-link> at the European Bioinformatics Institute (EBI). The sequences specified for any finished chromosomes are retrieved from GenBank and included in the set of genomic sequences to be processed for an assembly.</p>
                  <p><bold>Genomic Sequences from the Tiling Paths of Individual Chromosomes.</bold> The human genome sequencing centers use a variety of experimental evidence to compile an <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://genome.cse.ucsc.edu/goldenPath/chromReports/">ordered list</ext-link> of clones they believe provides the best coverage for each chromosome. At least once every 2 months, the sequencing centers submit an updated minimal tiling path for each chromosome to a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://cvs.sanger.ac.uk/cgi-bin/cvsweb.cgi/genome-map/data/?cvsroot=Ensembl">data repository</ext-link> at the EBI. These tiling path files (TPFs) include Accession numbers if sequence for a clone is available. The tiling path repository is checked each day for the Accession numbers from any tiling path that has been updated. Any secondary Accession numbers are replaced with the corresponding primary Accession numbers, and any invalid Accession numbers are flagged to prevent those sequences from being used for assembly. The latest version of the sequence for each Accession number in the most recent clone tiling paths is retrieved from GenBank and included in the set of genomic sequences to be processed for assembly.</p>
                  <p><bold>Assembled Blocks of Contiguous Finished Genomic Sequence.</bold> As the sequences for individual clones are finished, they are merged with overlapping finished sequences to form contigs (<xref ref-type="bibr" rid="bid.1538">3</xref>). The primary source for identifying neighboring clones is the clone tiling path for each chromosome. Additional information is obtained from some GenBank entries that contain annotation specifying the neighboring clones. <xref ref-type="other" rid="bid.1246">BLAST</xref> (<xref ref-type="bibr" rid="bid.1539">4</xref>) is used to align the sequences of candidate pairs of clones, and a merged sequence is produced automatically if the expected overlap is confirmed by the sequence alignment. When the automatic processes do not find an expected overlap, there is a manual review to find the correct overlap, refining the clone order if necessary. The most recent set of finished contigs are processed for assembly.</p>
                  <p><bold>Additional Genomic Sequences.</bold> A few other specific human genomic sequences are added to the assembly set because they contain genes that may not be represented in the genomic sequences from the other sources.</p>
                </sec>
                <sec id="bid.1450">
                  <title>
                    <bold>Genomic Sequences Used for Ordering and Orienting</bold>
                  </title>
                  <p>Sequences known to come from both ends of the same cloned genomic fragment provide valuable linking information that helps to order and orient sequence contigs in the assembly step (<xref ref-type="bibr" rid="bid.1540">5</xref>, <xref ref-type="bibr" rid="bid.1541">6</xref>). The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://snp.cshl.org">SNP Consortium</ext-link> sequenced the ends of the inserts in several million plasmid clones containing small (0.8&ndash;6 Kbp) fragments of human genomic DNA. In many cases, both ends of the same insert were sequenced (Table 4 in Ref. <xref ref-type="bibr" rid="bid.1542">7</xref>), thereby providing a set of plasmid paired-end sequences.</p>
                </sec>
                <sec id="bid.1451">
                  <title>
                    <bold>Genomic Sequences Used for Annotation</bold>
                  </title>
                  <p/>
                  <p><bold>Curated Genomic Regions.</bold> The Reference Sequence project (<xref ref-type="other" rid="bid.681">Chapter 18</xref>; Refs. <xref ref-type="bibr" rid="bid.1543">8</xref>, <xref ref-type="bibr" rid="bid.1544">9</xref>) provides reviewed annotated sequences for a number of genomic regions that are difficult to annotate correctly by automated processes (e.g., immunoglobulin gene regions). These Reference Sequences (RefSeqs) are aligned to the assembled genome so that the curated annotation can be transferred to the assembled genomic sequence. RefSeqs for known pseudogenes are also aligned, not only to enable transfer of the correct annotation but, more importantly, to prevent prediction of erroneous model transcripts and proteins.</p>
                  <p><bold>BAC End Sequences</bold>. Sequences from the ends of human genomic inserts in Bacterial Artificial Chromosome (<xref ref-type="other" rid="bid.1243">BAC</xref>) clones are used to help map the location of specific clones onto the assembled genome sequence. The BAC end sequences are obtained from <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/dbGSS/">dbGSS</ext-link> (see         <xref ref-type="other" rid="bid.2">Chapter 1</xref>). The clone names are extracted and converted to a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/clone/nomenclature.shtml">standardized format</ext-link> to facilitate linking of the BAC end sequences with mapping data and additional sequences for the same clone, when these are available.</p>
                </sec>
                <sec id="bid.1452">
                  <title>
                    <bold>Genomic Sequences Used for Alignment</bold>
                  </title>
                  <p>Human genomic sequences that are not used for either assembly or annotation are processed so that their relationships to the assembled genome can be displayed in Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>). Most genomic sequences deposited in GenBank by individual scientists will not be HTG and therefore will not be used for assembly; however, they are used for alignment. The exceptions are a few non-HTG sequences that are used for assembly, because they are included in the clone tiling path of an individual chromosome or in an assembled block of finished sequence. Any sequences intended for assembly but not used, either because they are redundant or are rejected by one of the quality screens in the assembly process, are also used for alignment.</p>
                </sec>
              </sec>
              <sec id="bid.1453">
                <title>Human Transcript Sequences</title>
                <p>Human transcript sequences are used to help order and orient genomic fragments in the assembly step, for feature annotation and also to produce maps that show the locations of the transcripts on the assembled genome. Transcripts used include: (<italic>a</italic>) human mRNA RefSeqs (<xref ref-type="bibr" rid="bid.1543">8</xref>, <xref ref-type="bibr" rid="bid.1544">9</xref>), except model transcripts produced from previous rounds of genome annotation; (<italic>b</italic>) human mRNA sequences deposited in GenBank by individual scientists, except those mRNAs produced after a translocation or other rearrangement of the genome; and (<italic>c</italic>) a nonredundant set of EST sequences from the BLAST <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/blast/db">FTP</ext-link> site. Additional information relating to these EST sequences is obtained from <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/index.html">UniGene</ext-link> (<xref ref-type="other" rid="bid.857">Chapter 21</xref>).</p>
              </sec>
              <sec id="bid.1454">
                <title>Transcripts from Other Organisms</title>
                <p>Transcripts from other organisms may be aligned to the genome being processed. These data may reveal the location of potential genes not identified by other means. RefSeq mRNAs, GenBank mRNAs, and ESTs, obtained from the same sources that provide the human transcripts, are used. Their alignments are processed for display in Map Viewer but are not used in the assembly step or for feature annotation.</p>
              </sec>
              <sec id="bid.1455">
                <title>Sequence Tagged Site (STS) Maps</title>
                <table-wrap id="bid.1456">
                  <label>1</label>
                  <caption>
                    <title>STS maps used for assembly or display.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="bottom">Map type</th>
                      <th align="center" valign="bottom">Map</th>
                      <th align="center" valign="bottom">Contig assembly</th>
                      <th align="center" valign="bottom">Contig placement</th>
                      <th align="center" valign="bottom">Display</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Genetic linkage</td>
                      <td align="left" valign="top">Genethon (<xref ref-type="bibr" rid="bid.1555">20</xref>)</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Genetic linkage</td>
                      <td align="left" valign="top">Marshfield (<xref ref-type="bibr" rid="bid.1556">21</xref>)</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Genetic linkage</td>
                      <td align="left" valign="top">Decode (<xref ref-type="bibr" rid="bid.1557">22</xref>)</td>
                      <td align="center" valign="top"/>
                      <td align="center" valign="top"/>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Radiation hybrid</td>
                      <td align="left" valign="top">GeneMap99-G3 (<xref ref-type="bibr" rid="bid.1558">23</xref>, <xref ref-type="bibr" rid="bid.1559">24</xref>)</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Radiation hybrid</td>
                      <td align="left" valign="top">GeneMap99-GB4 (<xref ref-type="bibr" rid="bid.1558">23</xref>, <xref ref-type="bibr" rid="bid.1559">24</xref>)</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Radiation hybrid</td>
                      <td align="left" valign="top">NCBI RH (<xref ref-type="bibr" rid="bid.1560">25</xref>)</td>
                      <td align="center" valign="top"/>
                      <td align="center" valign="top"/>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Radiation hybrid</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www-shgc.stanford.edu/Mapping/rh/index.html">Stanford G3</ext-link>
                      </td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Radiation hybrid</td>
                      <td align="left" valign="top">Stanford TNG (<xref ref-type="bibr" rid="bid.1561">26</xref>)</td>
                      <td align="center" valign="top"/>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Radiation hybrid</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www-genome.wi.mit.edu/cgi-bin/contig/phys_map">Whitehead-RH</ext-link>
                      </td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">YAC</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www-genome.wi.mit.edu/cgi-bin/contig/phys_map">Whitehead-YAC</ext-link>
                      </td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                      <td align="center" valign="top">X</td>
                    </tr>
                  </table>
                </table-wrap>
                <p>Genetic linkage maps, radiation hybrid (<xref ref-type="other" rid="bid.1394">RH</xref>) maps, and a <xref ref-type="other" rid="bid.1438">YAC</xref> map are used to help avoid assembling genomic contigs incorrectly and to help place the contigs along the chromosomes. The positions of <xref ref-type="other" rid="bid.1410">STS</xref> markers on various maps (<xref ref-type="table" rid="bid.1456">Table 1</xref>) are transformed into a common format that allows us to compare the maps to each other during the assembly process. Additional maps are processed so that they can be displayed in <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/humansearch.html">Map Viewer</ext-link>.</p>
                <p>The maps listed in <xref ref-type="table" rid="bid.1456">Table 1</xref> are static and are not updated with additional markers. Any new STS maps are added to our data set soon after they are released.</p>
              </sec>
              <sec id="bid.1457">
                <title>Special Cases</title>
                <p>Our own review of previous genome assemblies or feedback from users sometimes identify particular cases in which bad data or overlooked data prevent the automated processes from producing the best possible assembly of a particular segment of the genome. To help guide the assembly process, a list of such special cases is maintained. The list is used to provide supplemental data that override the automatic processes that assign a particular input genomic sequence to a chromosome or determine whether it is used for assembly.</p>
              </sec>
            </sec>
            <sec id="bid.1458">
              <title>Preparation of the Input Sequences</title>
              <p>The raw input genomic sequences are screened for contaminants, the repetitive sequences are masked, and the draft genomic sequences are split into fragments in preparation for alignment to other sequences. The input transcript sequences are also screened for contaminants before they are aligned to the genomic sequences. The STS content of the input genomic sequences is determined.</p>
              <sec id="bid.1459">
                <title>Preparation of Genomic Sequences</title>
                <sec id="bid.1460">
                  <title>
                    <bold>Removing Contaminants</bold>
                  </title>
                  <p>Draft-quality HTGSs sometimes contain segments of sequence derived from foreign sources, most commonly the cloning vector or bacterial host. Finished sequences are usually, but not always, free of such contaminants. Common contaminants introduce artificial blocks of homologous sequence that can give rise to misleading alignments between two unrelated genomic sequences. <xref ref-type="other" rid="bid.1340">MegaBLAST</xref> (<xref ref-type="bibr" rid="bid.1545">10</xref>) is used to compare the raw genomic sequences to a database of contaminant sequences (including the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/VecScreen/UniVec.html">UniVec</ext-link> database of vector sequences, the <italic>Escherichia coli</italic> genome, bacterial insertion sequences, and bacteriophage). Any foreign segments are removed from draft-quality sequence or masked in finished sequence to prevent them from participating in alignments.</p>
                </sec>
                <sec id="bid.1461">
                  <title>
                    <bold>Masking of Repetitive Sequences</bold>
                  </title>
                  <p>Sequences that occur in many copies in the genome will align to many different clones. Such repetitive sequences include interspersed repeats (SINEs, LINEs, LTR elements, and DNA transposons), satellite sequences, and low-complexity sequences (<xref ref-type="bibr" rid="bid.1542">7</xref>, <xref ref-type="bibr" rid="bid.1546">11</xref>, <xref ref-type="bibr" rid="bid.1547">12</xref>). Matches between repetitive sequences on unrelated clones make it difficult to identify alignments that indicate a genuine overlap between clones. To eliminate the confounding matches that are based only on repetitive sequences, the genomic sequences are run through <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ftp.genome.washington.edu/cgi-bin/RepeatMasker">RepeatMasker</ext-link> to identify known repeats. Repeats are masked by converting the sequence to lowercase letters so that they do not initiate alignments.</p>
                </sec>
                <sec id="bid.1462">
                  <title>
                    <bold>Fragmentation of Draft Sequences</bold>
                  </title>
                  <p>Draft HTGSs consist of a set of sequence contigs derived from a particular clone artificially linked together to form a single sequence. The masked, draft genomic sequences are split at the gaps between their constituent contigs to create separate sequence fragments that can be aligned independently. Vector sequences and other contaminants are also removed at this stage by trimming or further splitting the sequence fragments.</p>
                </sec>
                <sec id="bid.1463">
                  <title>
                    <bold>Determination of STS Content</bold>
                  </title>
                  <p>Any STS markers contained within the input genomic sequences are identified by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/sts/epcr.cgi">e-PCR</ext-link> (<xref ref-type="bibr" rid="bid.1548">13</xref>) using the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=unists">UniSTS</ext-link> database. The resulting data are used primarily to relate the genomic sequences to independently derived STS maps (genetic, radiation hybrid) but are also used to identify some foreign sequences.</p>
                </sec>
                <sec id="bid.1464">
                  <title>
                    <bold>Filtering</bold>
                  </title>
                  <p>Sequences from other clones being sequenced at the same institution can occasionally cross-contaminate draft HTGSs. The contaminating sequences may come from another clone from the same organism or from another organism. The raw genomic sequences are screened in several ways to detect cross-contamination: (<italic>a</italic>) they are compared with the genome sequences from completely sequenced organisms using MegaBLAST (<xref ref-type="bibr" rid="bid.1545">10</xref>); (<italic>b</italic>) they are screened for the presence of organism-specific interspersed repeats using <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ftp.genome.washington.edu/cgi-bin/RepeatMasker">RepeatMasker</ext-link>; and (<italic>3</italic>) they are screened for the presence of mapped STS markers from other organisms using e-PCR (<xref ref-type="bibr" rid="bid.1548">13</xref>). Any input sequence that contains foreign sequences, repeats, or markers is flagged for removal from the data set used for assembly. Draft sequences longer than the maximum insert length expected for a genomic clone are also rejected because it is likely they are contaminated with sequences from at least one other clone.</p>
                  <p>At this stage, draft sequences composed of fragments that are too small to contribute significantly to the assembly or that are tagged with the HTGS_CANCELLED keyword are also flagged for removal. Another filter rejects sequences annotated as being from another organism or as being RNA, erroneously included in the input sequences.</p>
                </sec>
                <sec id="bid.1465">
                  <title>
                    <bold>Chromosome Assignment</bold>
                  </title>
                  <p>To improve assembly of the genomic sequences, the input genomic sequences are assigned to a specific chromosome before attempting to merge the sequences. Genomic sequences that appear on any of the chromosome tiling paths are automatically assigned to the designated chromosome. Other genomic sequences are assigned to a chromosome based on: (<italic>a</italic>) annotation on the submitted GenBank record; (<italic>b</italic>) the presence of multiple STS markers that have been mapped to the same chromosome; (<italic>c</italic>) fluorescence <italic>in situ</italic> hybridization (<xref ref-type="other" rid="bid.1293">FISH</xref>) mapping (<xref ref-type="bibr" rid="bid.1549">14</xref>, <xref ref-type="bibr" rid="bid.1550">15</xref>); or (<italic>d</italic>) personal communication from a scientist with specialized knowledge. If there is no assignment, or the assignments are conflicted, the sequences are treated as unassigned and assembled without constraint by chromosome.</p>
                </sec>
              </sec>
              <sec id="bid.1466">
                <title>Filtering of Transcript Sequences</title>
                <p>Transcript sequences that contain sequences derived from vectors or other common contaminants can produce artificial alignments to the assembled genomic sequence. The input transcript sequences are therefore compared with a database of contaminants using MegaBLAST (<xref ref-type="bibr" rid="bid.1545">10</xref>), as described for genomic sequences. Any transcripts with significant matches to the sequence of a contaminant are excluded from the set of transcript sequences used for genome assembly or annotation.</p>
                <p>mRNA sequences shorter than 300 bases are excluded from the set of sequences that are aligned to the genomic sequences because they are too small to contribute significantly to genome assembly or annotation. Also excluded are any mRNA sequences flagged because they do not represent the true sequence of a transcript, e.g., those that are chimeric or contain genomic sequences.</p>
              </sec>
            </sec>
            <sec id="bid.1467">
              <title>Alignment of Sequences to the Input Genomic Sequences</title>
              <p>Alignment of the input genomic sequences to each other and to various other sequences is essential for both genome assembly and genome annotation. All relevant sequences are initially aligned to the unassembled genomic sequences because this means that the computationally intensive alignment processes can be run incrementally at an early stage in the pipeline. If necessary, these alignments are remapped to the sequence of the assembled genome at a later stage by a process that requires relatively little computation.</p>
              <sec id="bid.1468">
                <title>Alignment of Genomic Sequences to Each Other</title>
                <p>Assembly of the genomic sequences from individual clones into longer contiguous sequences (contigs) requires knowledge of which sequences overlap. The overlaps between genomic sequences are evaluated by aligning the sequences from individual genomic clones to each other. After masking of repeats, decontamination and fragmentation, each fragment of genomic sequence is aligned pairwise to all of the other fragments using MegaBLAST (<xref ref-type="bibr" rid="bid.1545">10</xref>). Alignments that are sufficiently long and of sufficiently high percentage identity are saved for consideration in the assembly step.</p>
              </sec>
              <sec id="bid.1469">
                <title>Alignment of Clone End Sequences to the Genomic Sequences</title>
                <p>
				The pairs of short genomic sequences derived from the ends of plasmid clones help to order and orient sequence fragments in the assembly step. These clone end sequences are aligned to the processed genomic sequences, as described for <italic>Alignment of Genomic Sequences to Each Other</italic>.</p>
              </sec>
              <sec id="bid.1470">
                <title>Alignment of Transcripts to the Genomic Sequences</title>
                <p>Annotation of genes requires knowledge of where the sequences for known transcripts align to the assembled genomic sequences. RefSeq RNA sequences, mRNA sequences from GenBank, and EST sequences from dbEST are aligned to the processed genomic sequences, as described for <italic>Alignment of Genomic Sequences to Each Other</italic>. Later, the alignments are remapped to the assembled genomic sequence.</p>
              </sec>
              <sec id="bid.1471">
                <title>Alignment of Curated Genomic Regions to the Genomic Sequences</title>
                <p>Curated genomic regions provide accurate annotation for regions of the genome that are difficult to annotate correctly by automated processes. Sequences from curated genomic regions are initially aligned to the unassembled genomic sequences and later remapped to the sequence of the assembled genome, as described for <italic>Alignment of Transcripts to the Genomic Sequences</italic>.</p>
              </sec>
              <sec id="bid.1472">
                <title>Alignment of Translated Genomic Sequences to Proteins</title>
                <p>Homologies between the polypeptides encoded by the genomic sequences and known proteins/conserved protein domains provide hints for the gene prediction process. The repeat-masked genomic sequences are compared with a non-redundant database of vertebrate proteins and to the NCBI Conserved Domain Database (CDD; Ref. <xref ref-type="bibr" rid="bid.1551">16</xref>) using different versions of BLAST (<xref ref-type="bibr" rid="bid.1539">4</xref>) (BLASTX and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi">RPS-BLAST</ext-link>, respectively). Significant alignments are saved for use in the gene prediction step.</p>
              </sec>
            </sec>
            <sec id="bid.1473">
              <title>Genome Assembly</title>
              <p>The input genomic sequences are assembled into a series of genomic sequence contigs. These are then ordered, oriented with respect to each other, and placed along each chromosome with appropriately sized gaps inserted between adjacent contigs. The resulting genome assembly thus consists of a set of genomic sequence contigs and a specification for how to arrange the sequence contigs along each chromosome.</p>
              <sec id="bid.1474">
                <title>Finished Chromosomes</title>
                <p>A chromosome sequence is considered <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.genome.gov/page.cfm?pageID=10000923">finished</ext-link> when any gaps that remain cannot be closed using current cloning and sequencing technology. In practice, therefore, the sequence for a finished chromosome usually consists of a small number of genomic sequence contigs. These are assembled from their component clone sequences according to the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://cvsweb.sanger.ac.uk/cgi-bin/cvsweb.cgi/genome-map/data/?cvsroot=Ensembl">specification</ext-link> provided by the center responsible for sequencing that chromosome. This specification also prescribes the order, orientation, and estimated sizes for the gaps between contigs.</p>
              </sec>
              <sec id="bid.1475">
                <title>Unfinished Chromosomes</title>
                <p>Genomic sequence contigs for unfinished chromosomes are assembled and laid out based largely on the clone <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://genome.cse.ucsc.edu/goldenPath/chromReports/">tiling path</ext-link>. However, the tiling paths do not specify the orientation of the clone sequences or how they should be joined; therefore, data on the alignment of the input genomic sequences to each other and to other sequences are also used to guide the assembly. Genomic sequences that augment the initial set of genomic contigs based on the tiling path clones are also incorporated.</p>
                <sec id="bid.1476">
                  <title>
                    <bold>Resolution of Conflicts in the Chromosome Tiling Paths</bold>
                  </title>
                  <p>Before the tiling paths are used in the genome assembly, the order of the finished clone sequences included in the tiling paths is compared with the specifications used to assemble the curated contigs of finished sequence. Discrepancies are resolved before proceeding with the assembly. Sequence from any clone should appear at just one place in the assembled genome; therefore, if a clone is listed more than once in the tiling paths, only the location with the best evidence is used in the assembly step.</p>
                </sec>
                <sec id="bid.1477">
                  <title>
                    <bold>Genomic Sequences Excluded from the Assembly Step</bold>
                  </title>
                  <p>Clone sequences that consist only of unassembled reads (HTGS_PHASE0) or that were flagged because of suspected cross-contamination or other problem detected in the pre-assembly screens are not used in the assembly step.</p>
                </sec>
                <sec id="bid.1478">
                  <title>
                    <bold>Assembly of the Genomic Sequence Contigs</bold>
                  </title>
                  <p>Adjacent, finished clone sequences from the chromosome tiling path that have good sequence overlap are merged. Tiling path draft sequences that are adjacent to and overlap the finished clone sequences or other draft clone sequences are added to extend the initial genomic sequence contigs. After that, genomic sequences from clones not on any chromosome tiling path are added, provided they have good overlaps with the assembled tiling path clones. Genomic sequences from additional clones may be added if they provide the sequence for a known gene that is missing from the existing genomic sequence contigs. Finally, the individual fragments of draft sequences are ordered and oriented.</p>
                  <p><bold>Assembly of Finished Sequences from Tiling Path Clones.</bold> The quality of any overlaps between finished clone sequences that are adjacent in the clone tiling path are assessed using the alignments between pairs of genomic sequences that were produced in advance. Sequences that have high-quality overlaps, or that are known from annotation or other data to abut, are merged to form a genomic sequence contig. Clone sequences that have no good overlaps are retained as separate contigs.</p>
                  <p><bold>Addition of Draft Sequences from Tiling Path Clones.</bold> The procedure used for merging draft sequence from tiling path clones is similar to that described for merging finished sequences, except that the minimum-overlap quality required for merging is different. An overlap involving a draft sequence can contain more mismatches, but must be longer, than an overlap between two finished sequences. Preference is given to finished sequences so that a contig made by merging finished and draft sequences will contain the finished sequence for the overlapping portion. Draft clone sequences that have no good overlaps are retained as separate contigs.</p>
                  <p><bold>Addition of Sequences from Other Genomic Clones.</bold> Genomic sequences from clones that are not on any chromosome tiling path are used to close, or extend into, gaps in the backbone of genomic contigs assembled based on the tiling paths. Genomic sequences that are fully contained within the existing genomic contigs are not used. Any remaining genomic sequences that were either assigned to the relevant chromosome or could not be assigned to any chromosome are evaluated, and sequences that have good-quality overlaps with genomic contigs are merged in to extend a contig or to join two adjacent contigs, if two additional conditions are met: (1) the gap must be sufficiently large to accommodate the additional sequence; and (2) the Sequence Tagged Site (STS) marker content of the additional clone sequence must be compatible with that of the flanking clone sequences when compared with various STS maps.</p>
                  <p>After all of the chromosomes have been assembled, any remaining genomic clone sequence that contains a known gene not present in the other contigs is added to the assembly as a separate contig.</p>
                  <p><bold>Ordering and Orienting Draft Sequence Fragments.</bold> The order and orientation of the fragments of HTGS_PHASE1 draft sequence need to be defined before sequence made from contigs that include this category of draft sequence can be completed. Some fragments may be ordered and oriented by overlaps with sequences from adjacent clones. Many more can be defined by aligning them with mRNAs, ESTs, or plasmid paired-end sequences. Any fragments whose order and orientation remain undefined are placed in the nearest open gap and given an arbitrary orientation. Fragments of draft sequence are connected to flanking sequences and to each other by runs of 100 unknown bases (Ns), which represent an arbitrarily sized gap in the sequence.</p>
                </sec>
                <sec id="bid.1483">
                  <title>
                    <bold>Placement of the Genomic Contigs</bold>
                  </title>
                  <p>After the genomic sequence contigs are assembled, they are oriented and placed in order along each chromosome with appropriately sized gaps inserted between adjacent contigs. The chromosome <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://genome.cse.ucsc.edu/goldenPath/chromReports/">tiling paths</ext-link> specify the order of the clone sequences and the sizes for some gaps. Therefore, the order and orientation of most of the genomic sequence contigs are derived from the tiling paths. Many of the remaining contigs are placed by comparison of the STS marker content of the contigs, as determined by e-PCR (<xref ref-type="bibr" rid="bid.1548">13</xref>) using the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=unists">UniSTS</ext-link> database, to various STS maps. There are some contigs that can be assigned to a specific chromosome but cannot be placed along that chromosome. Others cannot even be assigned to a specific chromosome and therefore remain unplaced within the genome assembly.</p>
                  <p>Gaps between the clone contigs laid out in the chromosome tiling paths are arbitrarily set at 50 Kbp, and 3 Mbp for the centromere, unless another gap size is specified in the tiling path. Any remaining gaps between genomic sequence contigs are arbitrarily set to 10 Kbp.</p>
                </sec>
              </sec>
              <sec id="bid.1484">
                <title>Preparing a Provisional Genome Assembly</title>
                <p>A set of sequences and data files is produced to represent the provisional assembly. This set includes: (<italic>a</italic>) sequences for each genomic contig in FASTA format; (<italic>b</italic>) specifications describing how to assemble each genomic contig from its components; and (<italic>c</italic>) a specification for how to arrange the contigs along each chromosome. A raw RefSeq entry is also made for each genomic sequence contig.</p>
              </sec>
              <sec id="bid.1485">
                <title>Quality Control</title>
                <p>The provisional assembly is checked for consistency with the chromosome tiling paths and various STS maps. The order in which the component clone sequences appear in the assembled chromosomes is compared with their order in the tiling paths on which the assembly was based. The STS marker order along each chromosome in the provisional assembly, as determined by e-PCR (<xref ref-type="bibr" rid="bid.1548">13</xref>) using the UniSTS database, is checked for consistency with a set of STS maps. Haussler et al. at the University of California at Santa Cruz (UCSC) also perform a set of independent quality checks on the provisional assembly. In addition to comparing the assembly to the chromosome tiling paths and to various STS maps, they also look for potentially misassembled contigs using alignments of BAC end sequences. Any serious errors in the assembly may be corrected by repeating the assembly steps using different parameters or by manually editing the assembly.</p>
              </sec>
            </sec>
            <sec id="bid.1486">
              <title>Annotation of Genes</title>
              <p>Identification of genes within the genome assembly reveals the functional significance of particular stretches of genomic sequence. Genes are found using three complementary approaches: (<italic>a</italic>) known genes are placed primarily by aligning mRNAs to the assembled genomic contigs; (<italic>b</italic>) additional genes are located based on alignment of ESTs to the assembled genomic contigs; and (<italic>c</italic>) previously unknown genes are predicted using hints provided by protein homologies. Whenever possible, predicted genes are identified by homology between the protein they encode and known protein sequences.</p>
              <sec id="bid.1487">
                <title>Generation of Transcript-based Gene Models</title>
                <p>Alignments between known transcripts and the assembled genomic sequences are processed to produce gene models. Each gene model consists of an ordered series of exons. The transcripts defining each gene model are used as evidence to support that model.</p>
                <sec id="bid.1488">
                  <title>
                    <bold>Alignment of Transcripts to the Assembled Genome</bold>
                  </title>
                  <p>The alignments between RefSeq RNA sequences, mRNA and EST sequences from GenBank and the component genomic sequences are remapped to produce alignments of these transcripts to the assembled genomic contigs.</p>
                </sec>
                <sec id="bid.1489">
                  <title>
                    <bold>Production of Candidate Gene Models</bold>
                  </title>
                  <p>A candidate gene model is produced from each set of alignments between a particular transcript and one strand of a particular genomic contig as follows: (1) putative exons are identified by looking for mRNA splice sites near the ends of those alignments that satisfy minimum length and percentage identity criteria; (2) a mutually compatible set of exons for the model is selected by applying rules, such as restrictions on the size of an intron, that define plausible exon&ndash;intron structures; and (3) BLAST (<xref ref-type="bibr" rid="bid.1539">4</xref>) may be used to produce additional alignments to try to identify exons that were missed because they were too short to be represented in the initial set of transcript alignments. Candidate gene models are only retained if good-quality alignments between their exons and the defining transcript cover either more than half the length of the transcript or more than 1 Kbp.</p>
                </sec>
                <sec id="bid.1490">
                  <title>
                    <bold>Selection of the Best RefSeq RNA-based Gene Models</bold>
                  </title>
                  <p>Each RefSeq RNA represents a distinct transcript produced from a particular gene (<xref ref-type="other" rid="bid.681">Chapter 18</xref>; Refs. <xref ref-type="bibr" rid="bid.1543">8</xref>, <xref ref-type="bibr" rid="bid.1544">9</xref>). Hence, there should not be more than one gene model corresponding to any given RefSeq RNA. Therefore, all gene models based on a particular RefSeq RNA are compared, and the best one is selected. Because the RefSeq RNA is taken to be the best representation of a particular transcript, this gene model is preserved without any further modification. Any extra models may represent paralogs; therefore, they are included with the mRNA- and EST-based models for further processing. Between builds, RefSeq RNAs are refined based on a review of related gene models and transcript alignments produced during the genome annotation process.</p>
                </sec>
                <sec id="bid.1491">
                  <title>
                    <bold>Exon Refinement</bold>
                  </title>
                  <p>Many gene models may be produced for the same gene because the input data set frequently contains multiple EST or mRNA sequences representing the same transcript. This redundancy is used to refine the splice sites defining a particular exon. Similar exons are clustered, and splice sites may be adjusted in some models to match those used by the majority of models containing the same exon. Inconsistent models may be discarded at this stage, unless they have sufficient support to be retained as likely splice variants.</p>
                </sec>
                <sec id="bid.1492">
                  <title>
                    <bold>Chaining of Transcript-based Gene Models</bold>
                  </title>
                  <p>Many of the mRNAs and most of the ESTs used to generate the initial gene models provide sequence for only part of the native transcript. Overlapping gene models that are compatible with each other are combined into an extended model. This chaining step produces models more likely to represent the full gene.</p>
                </sec>
              </sec>
              <sec id="bid.1493">
                <title><italic>Ab Initio</italic> Gene Prediction</title>
                <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://genes.mit.edu/genomescan/">GenomeScan</ext-link>, an <italic>ab initio</italic> gene prediction program, is used to provide models for genes inferred from the genomic sequence using hints provided by protein homologies (<xref ref-type="bibr" rid="bid.1552">17</xref>). The genes predicted by GenomeScan are combined with the transcript-based gene models, but they are also retained as a distinct set of models that can be viewed or searched separately.</p>
                <sec id="bid.1494">
                  <title>
                    <bold>Dividing the Genomic Sequences into Segments</bold>
                  </title>
                  <p>GenomeScan produces better results when long genomic sequences are broken into shorter segments at putative gene boundaries. The locations of gene models based on RefSeq RNA alignments are, therefore, used to divide the assembled genomic contigs into segments. Repetitive sequences are masked by remapping the repeats found in the component genomic sequences.</p>
                </sec>
                <sec id="bid.1495">
                  <title>
                    <bold>Producing Protein Hints for GenomeScan</bold>
                  </title>
                  <p>GenomeScan can use data on protein homologies to improve its gene predictions (<xref ref-type="bibr" rid="bid.1552">17</xref>). The locations of genomic sequences that potentially code for polypeptides with homology to other proteins are obtained from three sources. Significant alignments between translated genomic segments and vertebrate proteins are obtained by filtering and remapping the precomputed alignments. Significant alignments between translated genomic sequences and conserved protein domains are obtained in the same manner. A third set of alignments comes from running GenomeScan without any hints. The proteins predicted by this initial run are aligned to proteins from <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.expasy.ch/sprot/">SWISS-PROT</ext-link> (<xref ref-type="bibr" rid="bid.1553">18</xref>) and NCBI RefSeq proteins (<xref ref-type="bibr" rid="bid.1543">8</xref>, <xref ref-type="bibr" rid="bid.1544">9</xref>) using blastp (<xref ref-type="bibr" rid="bid.1539">4</xref>). The eukaryotic protein with the best match is then aligned to the genomic sequence segments using tblastn (<xref ref-type="bibr" rid="bid.1539">4</xref>). These three sets of data are converted into the format required by GenomeScan and merged to produce a single set of protein hints.</p>
                </sec>
                <sec id="bid.1496">
                  <title>
                    <bold>Predicting Genes Using GenomeScan</bold>
                  </title>
                  <p>Each segment of genomic sequence is processed by GenomeScan using the combined set of protein homology-based hints as an additional input. This produces one model containing all of the predicted exons for each putative gene. Models with coding sequences shorter than 90 amino acids are discarded. Each remaining model is aligned to proteins from SWISS-PROT and NCBI RefSeq proteins using blastp. The eukaryotic protein with the best match to any model is used as evidence for that model and to provide a clue as to the possible function of that model.</p>
                </sec>
              </sec>
              <sec id="bid.1497">
                <title>Consolidation of Gene Models</title>
                <p>Consolidation of the transcript-based gene models and the predicted gene models forms a single set of models. Models are clustered into genes if they share one or more exons or if LocusLink (<xref ref-type="other" rid="bid.788">Chapter 19</xref>; Refs. <xref ref-type="bibr" rid="bid.1543">8</xref>, <xref ref-type="bibr" rid="bid.1544">9</xref>) indicates that the transcripts used as evidence for the models come from the same gene. If a model is entirely contained within a longer model, it is redundant and, therefore, eliminated. Sets of identical models are reduced to a single representative model linked to all of the supporting evidence. For sets of very similar models, a single model is picked as a representative, giving preference to models based on RefSeq RNAs or on GenBank mRNAs. Predicted gene models that significantly overlap transcript-based models but that are not sufficiently similar to consolidate are discarded.</p>
              </sec>
              <sec id="bid.1498">
                <title>Pruning of Gene Models</title>
                <p>Some gene models are discarded because: superior gene annotation is available from a curated genomic region, they are likely to represent pseudogenes, or they are incompatible with other gene models.</p>
                <sec id="bid.1499">
                  <title>
                    <bold>Gene Models Superceded by Curated Genomic Regions</bold>
                  </title>
                  <p>The manually reviewed annotations from curated genomic region RefSeqs are used in preference to any corresponding gene models generated by automated processing. The curated genomic regions are aligned to the assembled genomic contigs by remapping the alignments between these RefSeqs and the component genomic sequences. Any gene model that significantly overlaps a segment of the assembled sequence that corresponds to a curated genomic region is discarded.</p>
                </sec>
                <sec id="bid.1500">
                  <title>
                    <bold>Gene Models Likely to Be Pseudogenes</bold>
                  </title>
                  <p>When transcripts from a particular gene are aligned to the genomic sequences, they will align not only to the active copy of the gene but also to any segment of the genome containing a pseudogene derived from the active gene. Because model transcripts or model proteins that represent nontranscribed pseudogenes are undesirable, an attempt is made to identify and remove such models.</p>
                  <p>Whenever possible, alignments of RefSeqs for pseudogenes, either curated genomic regions or RNAs, are used to annotate pseudogenes. Some additional models derived from pseudogenes that are not yet represented by RefSeqs are eliminated by the following mechanism. All models based on the same supporting mRNA are compared with respect to the percent identity of the alignments and the number of exons. Only the model with the strongest evidence is retained.</p>
                </sec>
                <sec id="bid.1501">
                  <title>
                    <bold>Conflicting Gene Models</bold>
                  </title>
                  <p>When two gene models are found to have an extensive overlap, then in general only the model with the stronger evidence is retained. However, models based on RefSeqs are always retained. Whereas any model not based on a RefSeq is discarded if it overlaps a model that is RefSeq based, two RefSeq-based models that overlap are both retained.</p>
                </sec>
              </sec>
              <sec id="bid.1502">
                <title>Location of Model Coding Regions</title>
                <p>Initially, the longest open reading frame from each gene model is annotated as the protein coding sequence. This annotation can be revised if evidence associated with that model provides support for an alternative coding region. The protein coding sequence from any transcript used as evidence for a gene model is compared with the longest open reading frame in that model using BLAST (<xref ref-type="bibr" rid="bid.1539">4</xref>). If the two do not match, the conflict is noted, and the annotation is revised if there is evidence to support an alternative coding region. For example, the coding sequence from the transcript evidence may indicate that an alternate translation start site is used, or that the model contains a premature termination codon. Models with coding regions less than 90 amino acids long are discarded, unless they are based on a RefSeq.</p>
              </sec>
              <sec id="bid.1503">
                <title>Relating Gene Models to Known Genes, Transcripts, and Proteins</title>
                <p>The set of gene models produced by the preceding steps is a mixture of models for predicted genes and for known genes. To help identify models representing known genes, the model transcripts are compared with known transcripts. To help name the predicted genes, the proteins encoded by the models are also compared with known proteins.</p>
                <sec id="bid.1504">
                  <title>
                    <bold>Relating the Model Transcripts to Known Transcripts</bold>
                  </title>
                  <p>To provide continuity from build to build and to identify genes based on their predicted transcripts, MegaBLAST (<xref ref-type="bibr" rid="bid.1545">10</xref>) is used to compare model RNAs to: (<italic>a</italic>) RefSeq RNAs; (<italic>b</italic>) mRNAs from GenBank; and (<italic>c</italic>) model RNAs from the previous build. These comparisons are reported as reciprocal best hits if: (<italic>a</italic>) they produce a significant hit; (<italic>b</italic>) no other model has a better hit to that particular RNA; and (<italic>c</italic>) no other RNA has a better hit to that particular model.</p>
                </sec>
                <sec id="bid.1505">
                  <title>
                    <bold>Relating the Model Proteins to Known Proteins</bold>
                  </title>
                  <p>The eukaryotic proteins with the best match to each protein predicted by the annotation process are used to identify the best model for a possible gene and to assign a name to gene models that are novel. The proteins encoded by the models are aligned to proteins from SWISS-PROT (<xref ref-type="bibr" rid="bid.1553">18</xref>), NCBI RefSeq proteins (<xref ref-type="bibr" rid="bid.1543">8</xref>, <xref ref-type="bibr" rid="bid.1544">9</xref>), and the NCBI non-redundant protein <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/blast/html/blastcgihelp.html#protein_databases">database</ext-link> using blastp (<xref ref-type="bibr" rid="bid.1539">4</xref>). The name of the eukaryotic protein with the best match, its sequence identifier, and match score are recorded for each predicted protein with a significant hit.</p>
                </sec>
              </sec>
              <sec id="bid.1506">
                <title>Assigning Gene Identifiers to Models</title>
                <p>Gene models are attributed to known genes whenever the correspondence is clear. If a model RNA has a reciprocal best hit with a known RNA, then the annotation of the known RNA is used to identify the gene. The first models to be assigned to genes are those that have reciprocal best hits with RefSeq RNAs. This is followed by assignment of those models that have reciprocal best hits to models from the previous build or to GenBank mRNAs. Gene data for models that match a mRNA not yet represented by a RefSeq are obtained from NCBI gene-specific databases (currently LocusLink, <xref ref-type="other" rid="bid.788">Chapter 19</xref>). If the mRNA is associated with an entry in one of these databases, then the information attached to that gene record (e.g., symbols, names, and database cross-references) is used in the annotation. If the correspondence with known genes is ambiguous, as may occur if there are undocumented paralogs, then an interim gene identifier is assigned.</p>
              </sec>
              <sec id="bid.1511">
                <title>Selection of Transcript Models to Represent Each Gene</title>
                <p>Multiple models based on alternative transcripts for some genes may be produced. In most of these cases, one transcript model is selected to represent the product of the gene for annotation purposes. Any homology between eukaryotic proteins and proteins encoded by the models guides the choice between alternative models. Multiple transcripts are annotated only if the models are based on RefSeq mRNAs representing alternative transcripts from the same gene.</p>
                <p>Although alternative transcript models are not annotated, the alignments between the transcripts that represent alternative splicing and genomic contigs are processed for display in Map Viewer, Evidence Viewer, and Model Maker (see <xref ref-type="other" rid="bid.1562">Chapter 20</xref>).</p>
              </sec>
              <sec id="bid.1512">
                <title>Naming of Gene Products</title>
                <p>The transcripts and protein products of any models that have been assigned to a known gene are given the product names that appear in the LocusLink entry for that gene. The gene products from other genes are named based on any significant homology to other eukaryotic proteins, provided that the matching protein has a meaningful name (i.e., names such as &ldquo;Hypothetical...&rdquo; are ignored).</p>
              </sec>
              <sec id="bid.1513">
                <title>Annotation of the Assembled Genomic Contigs</title>
                <p>The genomic contig RefSeqs are annotated with features that provide information about the location of genes, mRNAs, and coding regions. Features from curated genomic region RefSeqs are copied to the contigs based on the alignment between the curated sequence and the corresponding contig. Protein domains from the Conserved Domain Database (CDD; Ref. <xref ref-type="bibr" rid="bid.1551">16</xref>) are identified using reverse position-specific BLAST (RPS-BLAST; Ref. <xref ref-type="bibr" rid="bid.1539">4</xref>), and their locations are annotated. A description of the evidence supporting those RNAs and proteins that are not curated RefSeqs, i.e., those that are models, is also recorded.</p>
              </sec>
            </sec>
            <sec id="bid.1514">
              <title>Annotation of Other Features</title>
              <p>Reference sequences produced by the genome assembly process are annotated with features that provide landmarks valuable for making connections between maps based on different coordinate systems and for associating genes with diseases.</p>
              <sec id="bid.1515">
                <title>Annotation of STSs</title>
                <p>Placement of STSs on the genome assembly allows sequence-based data to be integrated with non-sequence-based maps that contain STS markers, such as genetic and radiation hybrid maps. STSs are identified by using e-PCR (<xref ref-type="bibr" rid="bid.1548">13</xref>) to find sequences that match the STS primer pairs from UniSTS, the spacing of which is consistent with the reported PCR product size. The number of times that each STS appears in the assembled genome is recorded so that only those STSs that appear at only one or two locations in the assembled genome are annotated.</p>
              </sec>
              <sec id="bid.1516">
                <title>Annotation of Clones</title>
                <p>Placement on the genome assembly of clones that have been mapped to cytogenetic bands by FISH provides the means to determine the correspondence between the sequence and cytogenetic coordinate systems (<xref ref-type="bibr" rid="bid.1549">14</xref>, <xref ref-type="bibr" rid="bid.1550">15</xref>). Knowing this correspondence allows the integration of sequence-based data with cytogenetic data. For human, only those clones mapped by fluorescence <italic>in situ</italic> hybridization (FISH) by the human BAC Resource Consortium (see the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/cyto/hbrc.shtml">Human BAC Resource</ext-link>) are annotated. Clones are placed using three types of sequence tags. Clones that have sequence for the genomic insert, either draft or finished, with a GenBank Accession number are localized by remapping the alignment between the clone sequence and other genomic clones to the assembled genomic contigs. Similarly, clones that have BAC end sequences are localized by remapping the alignment between the BAC end sequences and genomic clone sequences to the assembled genomic contigs. Clones that have STS markers confirmed by PCR or hybridization experiments are mapped using the locations in the assembled contigs of STS markers that were identified by e-PCR. The number of places that each clone appears in the assembled genome is recorded so that only those clones that either have a unique placement in the assembled genome or are placed twice on the same chromosome are annotated.</p>
              </sec>
              <sec id="bid.1517">
                <title>Annotation of Sequence Variation</title>
                <p>Placement of Single Nucleotide Polymorphisms (SNPs) and other variations on the genome provides numerous landmarks that are valuable for associating genes with diseases (<xref ref-type="other" rid="bid.1143">Chapter 5</xref>). Variations from <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/SNP">dbSNP</ext-link> (<xref ref-type="bibr" rid="bid.1554">19</xref>) are placed in their genomic contexts using the sequences that flank the variation. Flanking sequences are first run through <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ftp.genome.washington.edu/cgi-bin/RepeatMasker">RepeatMasker</ext-link> to mask repetitive sequences and then aligned to the assembled genomic sequence contigs using MegaBLAST (<xref ref-type="bibr" rid="bid.1545">10</xref>). The resulting matches are classified as either high or low confidence, depending on the quality of the alignment, and the number of matches for each SNP is recorded so that only those SNPs that map to one or two locations in the assembled genome are annotated.</p>
              </sec>
            </sec>
            <sec id="bid.1519">
              <title>Product Data Sets</title>
              <p>The products of our assembly and annotation process are made available to the public as RefSeqs of assembled chromosome sequences, genomic sequence contigs, model transcripts, and model proteins. RefSeqs are produced in alternative formats so that they can be retrieved by Entrez, BLAST, or FTP.</p>
              <sec id="bid.1520">
                <title>RefSeqs</title>
                <p>A fully annotated Refseq entry is made for each genomic sequence contig. Separate RefSeq model RNA and protein entries are also made for any of the transcripts and coding regions annotated on genomic contigs not identified as existing RefSeqs. Finally, a RefSeq entry is made for each chromosome by combining the annotated sequences of the genomic contigs in the appropriate order and with the appropriate spacing. The resulting RefSeqs can be retrieved through Entrez.</p>
              </sec>
              <sec id="bid.1521">
                <title>BLAST Databases</title>
                <p>The assembled genomic contig RefSeqs are formatted as a BLAST database (<xref ref-type="other" rid="bid.610">Chapter 16</xref>). Separate BLAST databases are also produced from the set of transcripts and the set of proteins annotated on the assembled genome. These databases include both known and model RefSeqs. In addition, separate BLAST databases are produced from the complete sets of transcripts and proteins predicted by GenomeScan.</p>
              </sec>
              <sec id="bid.1522">
                <title>Data Files for FTP</title>
                <p>The annotated genomic sequence contig, model transcript, and model protein RefSeqs are saved in GenBank flatfile and ASN.1 formats. The same sets of sequences that are used to make BLAST databases are also saved in FASTA format. All of these data files, together with files that specify the construction of the genomic contigs and their arrangement along the chromosomes, are made available for download by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/genomes/H_sapiens/">FTP</ext-link>.</p>
              </sec>
            </sec>
            <sec id="bid.1523">
              <title>Production of Maps That Display Genome Features</title>
              <p>We produce many maps showing the locations of various features annotated on our genome assembly. Maps containing whatever combination of features that interests the user can be selected and displayed side-by-side using <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/MapViewerHelp.html">Map Viewer</ext-link> (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>). Detailed descriptions of the maps available for each genome are available in the relevant Genome Map Viewer <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/humansearch.html">help document</ext-link>.</p>
              <sec id="bid.1524">
                <title>Preparation of Map Data</title>
                <p>Basic map data are prepared for each map to identify each feature, delineate its position on the chromosome, and specify how it is to be displayed. For many maps, supplemental data are prepared to provide more information about each feature. Map Viewer displays this map-specific supplemental information when users select a particular map as the Master Map (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>).</p>
                <sec id="bid.1525">
                  <title>
                    <bold>Maps Based on Sequence Coordinates</bold>
                  </title>
                  <p>Maps that display those features annotated on the genomic sequence contigs (genes, STSs, clones, and SNPs) are generated by translating the positions of the features on the contigs into chromosome coordinates. Contig coordinates are translated into chromosome coordinates using the positions of the contigs along each chromosome, as determined in the genome assembly step. Using this same method, alignments between various sequences and the genomic contigs are translated into chromosome coordinates to produce additional maps that show the locations of the aligned sequences on the chromosomes. Maps generated from sequence alignments include maps that show the genomic positions of mRNA plus EST sequences, or genomic sequences from GenBank. The specifications used to build each genomic contig are also translated into chromosome coordinates to produce one map that shows the component sequences used to assemble each contig and another that simply shows the finished and draft sections of the contigs.</p>
                </sec>
                <sec id="bid.1526">
                  <title>
                    <bold>Maps Based on Other Coordinate Systems</bold>
                  </title>
                  <p>Cytogenetic maps, genetic linkage maps, and radiation hybrid maps use different coordinate systems that are not based on sequence. To generate data for these types of maps, the locations of the map elements are listed in the coordinate system appropriate to each map. Map Viewer can scale maps defined in different coordinate systems so that they can de displayed side by side.</p>
                </sec>
              </sec>
              <sec id="bid.1527">
                <title>Making the Map Data Available for Use</title>
                <p>All of the map data for the new genome assembly are loaded into the Map Viewer database. Next, the objects in the new maps are indexed so that users can search for and then display specific features (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>). The data from the Map Viewer database are exported to produce a set of map data files that is made available via <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/genomes/H_sapiens/maps/">FTP</ext-link>.</p>
              </sec>
            </sec>
            <sec id="bid.1528">
              <title>Public Release of Assembly and Models</title>
              <p>To ensure that a consistent view of the annotated genome assembly is presented, the release of databases and FTP files is coordinated. When everything is ready for release, the Map Viewer <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/map_search.cgi">display</ext-link> is switched to the new build, the BLAST <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/seq/page.cgi?F=HsBlast.html&amp;org=Hs">databases</ext-link> are swapped, and the files on the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nlm.nih.gov/genomes/H_sapiens">FTP</ext-link> site are replaced. Several associated databases are then refreshed, including LocusLink, dbSNP, and UniSTS, so that the data they contain reflect the new build. Finally, the Web pages that provide <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/guide/human/HsStats.html">statistics</ext-link> for the build and record <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/guide/human/release_notes.html">changes</ext-link> to the genome assembly and annotation process are updated.</p>
            </sec>
            <sec id="bid.1529">
              <title>Integration with Other Resources</title>
              <p>The products of the genome assembly and annotation process are linked extensively to various NCBI resources. These links provide different views of the data and more information for researchers as they follow a particular line of investigation.</p>
              <sec id="bid.1530">
                <title>Links between Map Viewer and Other Resources</title>
                <table-wrap id="bid.1531">
                  <label>2</label>
                  <caption>
                    <title>Links from Map Viewer objects to other NCBI resources.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Map object</th>
                      <th align="left" valign="top">Linked NCBI resource</th>
                      <th align="left" valign="top">Resource description</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Accession number</td>
                      <td align="left" valign="top">Entrez</td>
                      <td align="left" valign="top">
                        <xref ref-type="other" rid="bid.588">Chapter 15</xref>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Clone</td>
                      <td align="left" valign="top">Clone Registry</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/clone/">http://www.ncbi.nlm.nih.gov/genome/clone/</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Disease gene</td>
                      <td align="left" valign="top">OMIM</td>
                      <td align="left" valign="top">
                        <xref ref-type="other" rid="bid.405">Chapter 7</xref>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">EST or mRNA</td>
                      <td align="left" valign="top">UniGene</td>
                      <td align="left" valign="top">
                        <xref ref-type="other" rid="bid.857">Chapter 21</xref>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Gene</td>
                      <td align="left" valign="top">LocusLink</td>
                      <td align="left" valign="top">
                        <xref ref-type="other" rid="bid.788">Chapter 19</xref>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Gene or transcript</td>
                      <td align="left" valign="top">Evidence Viewer</td>
                      <td align="left" valign="top">
                        <xref ref-type="other" rid="bid.1562">Chapter 20</xref>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Gene or transcript</td>
                      <td align="left" valign="top">Model Maker</td>
                      <td align="left" valign="top">
                        <xref ref-type="other" rid="bid.1562">Chapter 20</xref>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Gene</td>
                      <td align="left" valign="top">Human&ndash;mouse homology map</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Homology/">http://www.ncbi.nlm.nih.gov/Homology/</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">STS</td>
                      <td align="left" valign="top">UniSTS</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=unists">http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=unists</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Variation (SNP)</td>
                      <td align="left" valign="top">dbSNP</td>
                      <td align="left" valign="top">
                        <xref ref-type="other" rid="bid.1143">Chapter 5</xref>
                      </td>
                    </tr>
                  </table>
                </table-wrap>
                <p>The maps displayed by Map Viewer have embedded links between map objects and relevant NCBI resources (<xref ref-type="table" rid="bid.1531">Table 2</xref>). Many of these resources also have reciprocal links back to Map Viewer, allowing, for example, a gene in LocusLink to be displayed in its genomic context.</p>
              </sec>
              <sec id="bid.1532">
                <title>Links between Reference Sequences and Other Resources</title>
                <p>During the production of RefSeqs, links between the annotated features (clones, genes, SNPs, and STSs) and the relevant resources listed in <xref ref-type="table" rid="bid.1531">Table 2</xref> are created. Links are also made between the genomic contig RefSeqs and the RefSeqs for the model transcripts and proteins that they encode.</p>
              </sec>
              <sec id="bid.1533">
                <title>Integration with BLAST</title>
                <p>A customized BLAST <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/seq/page.cgi?F=HsBlast.html&amp;ORG=Hs">Web page</ext-link> allows the comparison of any sequence to a BLAST database of model transcript, model protein, or genomic contig RefSeqs. Users can choose to view any hits that result from such a search on a diagram showing the chromosomal location of the hits, with each hit linked to a Map Viewer display of the region encompassing the sequence alignment.</p>
              </sec>
            </sec>
            <sec id="bid.1534">
              <title>Contributors</title>
              <p>Richa Agarwala, Jonathan Baker, Hsiu-Chuan Chen, Vyacheslav Chetvernin, Deanna Church, Cliff Clausen, Dmitry Dernovoy, Olga Ermolaeva, Wratko Hlavina, Wonhee Jang, Philip Johnson, Jonathan Kans, Paul Kitts, Alex Lash, David Lipman, Donna Maglott, Jim Ostell, Keith Oxenrider, Kim Pruitt, Sergei Resenchuk, Victor Sapojnikov, Greg Schuler, Steve Sherry, Andrei Shkeda, Alexandre Souvorov, Tugba Suzek, Tatiana Tatusova, Lukas Wagner, and Sarah Wheelan</p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.1536">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Bently</surname>
                        <given-names>DR</given-names>
                      </name>
                    </person-group>
                    <article-title>Genomic sequence information should be released immediately and freely in the public domain</article-title>
                    <source>Science</source>
                    <year>1996</year>
                    <volume>274</volume>
                    <fpage>533</fpage>
                    <lpage>534</lpage>
                    <pub-id pub-id-type="pmid">8928006</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1537">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Guyer</surname>
                        <given-names>M</given-names>
                      </name>
                    </person-group>
                    <article-title>Statement on the rapid release of genomic DNA sequence</article-title>
                    <source>Genome Res</source>
                    <year>1998</year>
                    <volume>8</volume>
                    <fpage>413</fpage>
                    <pub-id pub-id-type="pmid">9582183</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1538">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Jang</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Chen</surname>
                        <given-names>HC</given-names>
                      </name>
                      <name>
                        <surname>Sicotte</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                    </person-group>
                    <article-title>Making effective use of human genomic sequence data</article-title>
                    <source>Trends Genet</source>
                    <year>1999</year>
                    <volume>15</volume>
                    <fpage>284</fpage>
                    <lpage>286</lpage>
                    <pub-id pub-id-type="pmid">10390628</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1539">
                  <label>4</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Altschul</surname>
                        <given-names>SF</given-names>
                      </name>
                      <name>
                        <surname>Madden</surname>
                        <given-names>TL</given-names>
                      </name>
                      <name>
                        <surname>Schaffer</surname>
                        <given-names>AA</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Miller</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                    </person-group>
                    <article-title>Gapped BLAST and PSI-BLAST: a new generation of protein database search programs</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>1997</year>
                    <volume>25</volume>
                    <fpage>3389</fpage>
                    <lpage>3402</lpage>
                    <pub-id pub-id-type="pmid">9254694</pub-id>
                    <pub-id pub-id-type="other">146917</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1540">
                  <label>5</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Zhao</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Malek</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Mahairas</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Fu</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Nierman</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Venter</surname>
                        <given-names>JC</given-names>
                      </name>
                      <name>
                        <surname>Adams</surname>
                        <given-names>MD</given-names>
                      </name>
                    </person-group>
                    <article-title>Human BAC ends quality assessment and sequence analyses</article-title>
                    <source>Genomics</source>
                    <year>2000</year>
                    <volume>63</volume>
                    <fpage>321</fpage>
                    <lpage>332</lpage>
                    <pub-id pub-id-type="pmid">10704280</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1541">
                  <label>6</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Mahairas</surname>
                        <given-names>GG</given-names>
                      </name>
                      <name>
                        <surname>Wallace</surname>
                        <given-names>JC</given-names>
                      </name>
                      <name>
                        <surname>Smith</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Swartzell</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Holzman</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Keller</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Shaker</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Furlong</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Young</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Zhao</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Adams</surname>
                        <given-names>MD</given-names>
                      </name>
                      <name>
                        <surname>Hood</surname>
                        <given-names>L</given-names>
                      </name>
                    </person-group>
                    <article-title>Sequence-tagged connectors: a sequence approach to mapping and scanning the human genome</article-title>
                    <source>Proc Natl Acad Sci U S A</source>
                    <year>1999</year>
                    <volume>96</volume>
                    <fpage>9739</fpage>
                    <lpage>9744</lpage>
                    <pub-id pub-id-type="pmid">10449764</pub-id>
                    <pub-id pub-id-type="other">22280</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1542">
                  <label>7</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Lander</surname>
                        <given-names>ES</given-names>
                      </name>
                      <name>
                        <surname>Linton</surname>
                        <given-names>LM</given-names>
                      </name>
                      <name>
                        <surname>Birren</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Nusbaum</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Zody</surname>
                        <given-names>MC</given-names>
                      </name>
                      <name>
                        <surname>Baldwin</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Devon</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Dewar</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Doyle</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>FitzHugh</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Funke</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Gage</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Harris</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Heaford</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Howland</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Kann</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Lehoczky</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>LeVine</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>McEwan</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>McKernan</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Meldrim</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Mesirov</surname>
                        <given-names>JP</given-names>
                      </name>
                      <name>
                        <surname>Miranda</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Morris</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Naylor</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Raymond</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Rosetti</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Santos</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Sheridan</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Sougnez</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Stange-Thomann</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Stojanovic</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Subramanian</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Wyman</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Rogers</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Sulston</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Ainscough</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Beck</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Bentley</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Burton</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Clee</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Carter</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Coulson</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Deadman</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Deloukas</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Dunham</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Dunham</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Durbin</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>French</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Grafham</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Gregory</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Hubbard</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Humphray</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Hunt</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Jones</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Lloyd</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>McMurray</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Matthews</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Mercer</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Milne</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Mullikin</surname>
                        <given-names>JC</given-names>
                      </name>
                      <name>
                        <surname>Mungall</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Plumb</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Ross</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Shownkeen</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Sims</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Waterston</surname>
                        <given-names>RH</given-names>
                      </name>
                      <name>
                        <surname>Wilson</surname>
                        <given-names>RK</given-names>
                      </name>
                      <name>
                        <surname>Hillier</surname>
                        <given-names>LW</given-names>
                      </name>
                      <name>
                        <surname>McPherson</surname>
                        <given-names>JD</given-names>
                      </name>
                      <name>
                        <surname>Marra</surname>
                        <given-names>MA</given-names>
                      </name>
                      <name>
                        <surname>Mardis</surname>
                        <given-names>ER</given-names>
                      </name>
                      <name>
                        <surname>Fulton</surname>
                        <given-names>LA</given-names>
                      </name>
                      <name>
                        <surname>Chinwalla</surname>
                        <given-names>AT</given-names>
                      </name>
                      <name>
                        <surname>Pepin</surname>
                        <given-names>KH</given-names>
                      </name>
                      <name>
                        <surname>Gish</surname>
                        <given-names>WR</given-names>
                      </name>
                      <name>
                        <surname>Chissoe</surname>
                        <given-names>SL</given-names>
                      </name>
                      <name>
                        <surname>Wendl</surname>
                        <given-names>MC</given-names>
                      </name>
                      <name>
                        <surname>Delehaunty</surname>
                        <given-names>KD</given-names>
                      </name>
                      <name>
                        <surname>Miner</surname>
                        <given-names>TL</given-names>
                      </name>
                      <name>
                        <surname>Delehaunty</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Kramer</surname>
                        <given-names>JB</given-names>
                      </name>
                      <name>
                        <surname>Cook</surname>
                        <given-names>LL</given-names>
                      </name>
                      <name>
                        <surname>Fulton</surname>
                        <given-names>RS</given-names>
                      </name>
                      <name>
                        <surname>Johnson</surname>
                        <given-names>DL</given-names>
                      </name>
                      <name>
                        <surname>Minx</surname>
                        <given-names>PJ</given-names>
                      </name>
                      <name>
                        <surname>Clifton</surname>
                        <given-names>SW</given-names>
                      </name>
                      <name>
                        <surname>Hawkins</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Branscomb</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Predki</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Richardson</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Wenning</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Slezak</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Doggett</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Cheng</surname>
                        <given-names>JF</given-names>
                      </name>
                      <name>
                        <surname>Olsen</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Lucas</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Elkin</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Uberbacher</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Frazier</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Gibbs</surname>
                        <given-names>RA</given-names>
                      </name>
                      <name>
                        <surname>Muzny</surname>
                        <given-names>DM</given-names>
                      </name>
                      <name>
                        <surname>Scherer</surname>
                        <given-names>SE</given-names>
                      </name>
                      <name>
                        <surname>Bouck</surname>
                        <given-names>JB</given-names>
                      </name>
                      <name>
                        <surname>Sodergren</surname>
                        <given-names>EJ</given-names>
                      </name>
                      <name>
                        <surname>Worley</surname>
                        <given-names>KC</given-names>
                      </name>
                      <name>
                        <surname>Rives</surname>
                        <given-names>CM</given-names>
                      </name>
                      <name>
                        <surname>Gorrell</surname>
                        <given-names>JH</given-names>
                      </name>
                      <name>
                        <surname>Metzker</surname>
                        <given-names>ML</given-names>
                      </name>
                      <name>
                        <surname>Naylor</surname>
                        <given-names>SL</given-names>
                      </name>
                      <name>
                        <surname>Kucherlapati</surname>
                        <given-names>RS</given-names>
                      </name>
                      <name>
                        <surname>Nelson</surname>
                        <given-names>DL</given-names>
                      </name>
                      <name>
                        <surname>Weinstock</surname>
                        <given-names>GM</given-names>
                      </name>
                      <name>
                        <surname>Sakaki</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Fujiyama</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Hattori</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Yada</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Toyoda</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Itoh</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Kawagoe</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Watanabe</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Totoki</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Taylor</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Weissenbach</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Heilig</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Saurin</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Artiguenave</surname>
                        <given-names>F</given-names>
                      </name>
                      <name>
                        <surname>Brottier</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Bruls</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Pelletier</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Robert</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Wincker</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Smith</surname>
                        <given-names>DR</given-names>
                      </name>
                      <name>
                        <surname>Doucette-Stamm</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Rubenfield</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Weinstock</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Lee</surname>
                        <given-names>HM</given-names>
                      </name>
                      <name>
                        <surname>Dubois</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Rosenthal</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Platzer</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Nyakatura</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Taudien</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Rump</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Yang</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Yu</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Wang</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Huang</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Gu</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Hood</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Rowen</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Madan</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Qin</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Davis</surname>
                        <given-names>RW</given-names>
                      </name>
                      <name>
                        <surname>Federspiel</surname>
                        <given-names>NA</given-names>
                      </name>
                      <name>
                        <surname>Abola</surname>
                        <given-names>AP</given-names>
                      </name>
                      <name>
                        <surname>Proctor</surname>
                        <given-names>MJ</given-names>
                      </name>
                      <name>
                        <surname>Myers</surname>
                        <given-names>RM</given-names>
                      </name>
                      <name>
                        <surname>Schmutz</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Dickson</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Grimwood</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Cox</surname>
                        <given-names>DR</given-names>
                      </name>
                      <name>
                        <surname>Olson</surname>
                        <given-names>MV</given-names>
                      </name>
                      <name>
                        <surname>Kaul</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Raymond</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Shimizu</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Kawasaki</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Minoshima</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Evans</surname>
                        <given-names>GA</given-names>
                      </name>
                      <name>
                        <surname>Athanasiou</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Schultz</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Roe</surname>
                        <given-names>BA</given-names>
                      </name>
                      <name>
                        <surname>Chen</surname>
                        <given-names>F</given-names>
                      </name>
                      <name>
                        <surname>Pan</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Ramser</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Lehrach</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Reinhardt</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>McCombie</surname>
                        <given-names>WR</given-names>
                      </name>
                      <name>
                        <surname>de la Bastide</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Dedhia</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Blocker</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Hornischer</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Nordsiek</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Agarwala</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Aravind</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Bailey</surname>
                        <given-names>JA</given-names>
                      </name>
                      <name>
                        <surname>Bateman</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Batzoglou</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Birney</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Bork</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Brown</surname>
                        <given-names>DG</given-names>
                      </name>
                      <name>
                        <surname>Burge</surname>
                        <given-names>CB</given-names>
                      </name>
                      <name>
                        <surname>Cerutti</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Chen</surname>
                        <given-names>HC</given-names>
                      </name>
                      <name>
                        <surname>Church</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Clamp</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Copley</surname>
                        <given-names>RR</given-names>
                      </name>
                      <name>
                        <surname>Doerks</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Eddy</surname>
                        <given-names>SR</given-names>
                      </name>
                      <name>
                        <surname>Eichler</surname>
                        <given-names>EE</given-names>
                      </name>
                      <name>
                        <surname>Furey</surname>
                        <given-names>TS</given-names>
                      </name>
                      <name>
                        <surname>Galagan</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Gilbert</surname>
                        <given-names>JG</given-names>
                      </name>
                      <name>
                        <surname>Harmon</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Hayashizaki</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Haussler</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Hermjakob</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Hokamp</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Jang</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Johnson</surname>
                        <given-names>LS</given-names>
                      </name>
                      <name>
                        <surname>Jones</surname>
                        <given-names>TA</given-names>
                      </name>
                      <name>
                        <surname>Kasif</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Kaspryzk</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Kennedy</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Kent</surname>
                        <given-names>WJ</given-names>
                      </name>
                      <name>
                        <surname>Kitts</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Koonin</surname>
                        <given-names>EV</given-names>
                      </name>
                      <name>
                        <surname>Korf</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Kulp</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Lancet</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Lowe</surname>
                        <given-names>TM</given-names>
                      </name>
                      <name>
                        <surname>McLysaght</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Mikkelsen</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Moran</surname>
                        <given-names>JV</given-names>
                      </name>
                      <name>
                        <surname>Mulder</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Pollara</surname>
                        <given-names>VJ</given-names>
                      </name>
                      <name>
                        <surname>Ponting</surname>
                        <given-names>CP</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Schultz</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Slater</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Smit</surname>
                        <given-names>AF</given-names>
                      </name>
                      <name>
                        <surname>Stupka</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Szustakowski</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Thierry-Mieg</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Thierry-Mieg</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Wagner</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Wallis</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Wheeler</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Williams</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Wolf</surname>
                        <given-names>YI</given-names>
                      </name>
                      <name>
                        <surname>Wolfe</surname>
                        <given-names>KH</given-names>
                      </name>
                      <name>
                        <surname>Yang</surname>
                        <given-names>SP</given-names>
                      </name>
                      <name>
                        <surname>Yeh</surname>
                        <given-names>RF</given-names>
                      </name>
                      <name>
                        <surname>Collins</surname>
                        <given-names>F</given-names>
                      </name>
                      <name>
                        <surname>Guyer</surname>
                        <given-names>MS</given-names>
                      </name>
                      <name>
                        <surname>Peterson</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Felsenfeld</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Wetterstrand</surname>
                        <given-names>KA</given-names>
                      </name>
                      <name>
                        <surname>Patrinos</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Morgan</surname>
                        <given-names>MJ</given-names>
                      </name>
                      <name>
                        <surname>Szustakowki</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>de Jong</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Catanese</surname>
                        <given-names>JJ</given-names>
                      </name>
                      <name>
                        <surname>Osoegawa</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Shizuya</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Choi</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Chen</surname>
                        <given-names>YJ</given-names>
                      </name>
                    </person-group>
                    <article-title>Initial sequencing and analysis of the human genome</article-title>
                    <source>Nature</source>
                    <year>2001</year>
                    <volume>409</volume>
                    <fpage>860</fpage>
                    <lpage>921</lpage>
                    <pub-id pub-id-type="pmid">11237011</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1543">
                  <label>8</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Pruitt</surname>
                        <given-names>KD</given-names>
                      </name>
                      <name>
                        <surname>Katz</surname>
                        <given-names>KS</given-names>
                      </name>
                      <name>
                        <surname>Sicotte</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Maglott</surname>
                        <given-names>DR</given-names>
                      </name>
                    </person-group>
                    <article-title>Introducing RefSeq and LocusLink: curated human genome resources at the NCBI</article-title>
                    <source>Trends Genet</source>
                    <year>2000</year>
                    <volume>16</volume>
                    <fpage>44</fpage>
                    <lpage>47</lpage>
                    <pub-id pub-id-type="pmid">10637631</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1544">
                  <label>9</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Pruitt</surname>
                        <given-names>KD</given-names>
                      </name>
                      <name>
                        <surname>Maglott</surname>
                        <given-names>DR</given-names>
                      </name>
                    </person-group>
                    <article-title>RefSeq and LocusLink: NCBI gene-centered resources</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2001</year>
                    <volume>29</volume>
                    <fpage>137</fpage>
                    <lpage>140</lpage>
                    <pub-id pub-id-type="pmid">11125071</pub-id>
                    <pub-id pub-id-type="other">29787</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1545">
                  <label>10</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Schwartz</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Wagner</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Miller</surname>
                        <given-names>W</given-names>
                      </name>
                    </person-group>
                    <article-title>A GREEDY algorithm for aligning DNA sequences</article-title>
                    <source>J Comput Biol</source>
                    <year>2000</year>
                    <volume>7</volume>
                    <fpage>203</fpage>
                    <lpage>214</lpage>
                    <pub-id pub-id-type="pmid">10890397</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1546">
                  <label>11</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Jurka</surname>
                        <given-names>J</given-names>
                      </name>
                    </person-group>
                    <article-title>Repeats in genomic DNA: mining and meaning</article-title>
                    <source>Curr Opin Struct Biol</source>
                    <year>1998</year>
                    <volume>8</volume>
                    <fpage>333</fpage>
                    <lpage>337</lpage>
                    <pub-id pub-id-type="pmid">9666329</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1547">
                  <label>12</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Smit</surname>
                        <given-names>AF</given-names>
                      </name>
                    </person-group>
                    <article-title>Interspersed repeats and other mementos of transposable elements in mammalian genomes</article-title>
                    <source>Curr Opin Genet Dev</source>
                    <year>1999</year>
                    <volume>9</volume>
                    <fpage>657</fpage>
                    <lpage>663</lpage>
                    <pub-id pub-id-type="pmid">10607616</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1548">
                  <label>13</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                    </person-group>
                    <article-title>Sequence mapping by electronic PCR</article-title>
                    <source>Genome Res</source>
                    <year>1997</year>
                    <volume>7</volume>
                    <fpage>541</fpage>
                    <lpage>550</lpage>
                    <pub-id pub-id-type="pmid">9149949</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1549">
                  <label>14</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Kirsch</surname>
                        <given-names>IR</given-names>
                      </name>
                      <name>
                        <surname>Green</surname>
                        <given-names>ED</given-names>
                      </name>
                      <name>
                        <surname>Yonescu</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Strausberg</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Carter</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Bentley</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Leversha</surname>
                        <given-names>MA</given-names>
                      </name>
                      <name>
                        <surname>Dunham</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Braden</surname>
                        <given-names>VV</given-names>
                      </name>
                      <name>
                        <surname>Hilgenfeld</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Lash</surname>
                        <given-names>AE</given-names>
                      </name>
                      <name>
                        <surname>Shen</surname>
                        <given-names>GL</given-names>
                      </name>
                      <name>
                        <surname>Martelli</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Kuehl</surname>
                        <given-names>WM</given-names>
                      </name>
                      <name>
                        <surname>Klausner</surname>
                        <given-names>RD</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                    </person-group>
                    <article-title>A systematic, high-resolution linkage of the cytogenetic and physical maps of the human genome</article-title>
                    <source>Nat Genet</source>
                    <year>2000</year>
                    <volume>24</volume>
                    <fpage>339</fpage>
                    <lpage>340</lpage>
                    <pub-id pub-id-type="pmid">10742091</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1550">
                  <label>15</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Cheung</surname>
                        <given-names>VG</given-names>
                      </name>
                      <name>
                        <surname>Nowak</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Jang</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Kirsch</surname>
                        <given-names>IR</given-names>
                      </name>
                      <name>
                        <surname>Zhao</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Chen</surname>
                        <given-names>XN</given-names>
                      </name>
                      <name>
                        <surname>Furey</surname>
                        <given-names>TS</given-names>
                      </name>
                      <name>
                        <surname>Kim</surname>
                        <given-names>UJ</given-names>
                      </name>
                      <name>
                        <surname>Kuo</surname>
                        <given-names>WL</given-names>
                      </name>
                      <name>
                        <surname>Olivier</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Conroy</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Kasprzyk</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Massa</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Yonescu</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Sait</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Thoreen</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Snijders</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Lemyre</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Bailey</surname>
                        <given-names>JA</given-names>
                      </name>
                      <name>
                        <surname>Bruzel</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Burrill</surname>
                        <given-names>WD</given-names>
                      </name>
                      <name>
                        <surname>Clegg</surname>
                        <given-names>SM</given-names>
                      </name>
                      <name>
                        <surname>Collins</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Dhami</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Friedman</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Han</surname>
                        <given-names>CS</given-names>
                      </name>
                      <name>
                        <surname>Herrick</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Lee</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Ligon</surname>
                        <given-names>AH</given-names>
                      </name>
                      <name>
                        <surname>Lowry</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Morley</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Narasimhan</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Osoegawa</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Peng</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Plajzer-Frick</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Quade</surname>
                        <given-names>BJ</given-names>
                      </name>
                      <name>
                        <surname>Scott</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Sirotkin</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Thorpe</surname>
                        <given-names>AA</given-names>
                      </name>
                      <name>
                        <surname>Gray</surname>
                        <given-names>JW</given-names>
                      </name>
                      <name>
                        <surname>Hudson</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Pinkel</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Ried</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Rowen</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Shen-Ong</surname>
                        <given-names>GL</given-names>
                      </name>
                      <name>
                        <surname>Strausberg</surname>
                        <given-names>RL</given-names>
                      </name>
                      <name>
                        <surname>Birney</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Callen</surname>
                        <given-names>DF</given-names>
                      </name>
                      <name>
                        <surname>Cheng</surname>
                        <given-names>JF</given-names>
                      </name>
                      <name>
                        <surname>Cox</surname>
                        <given-names>DR</given-names>
                      </name>
                      <name>
                        <surname>Doggett</surname>
                        <given-names>NA</given-names>
                      </name>
                      <name>
                        <surname>Carter</surname>
                        <given-names>NP</given-names>
                      </name>
                      <name>
                        <surname>Eichler</surname>
                        <given-names>EE</given-names>
                      </name>
                      <name>
                        <surname>Haussler</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Korenberg</surname>
                        <given-names>JR</given-names>
                      </name>
                      <name>
                        <surname>Morton</surname>
                        <given-names>CC</given-names>
                      </name>
                      <name>
                        <surname>Albertson</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>de Jong</surname>
                        <given-names>PJ</given-names>
                      </name>
                      <name>
                        <surname>Trask</surname>
                        <given-names>BJ</given-names>
                      </name>
                    </person-group>
                    <article-title>Integration of cytogenetic landmarks into the draft sequence of the human genome</article-title>
                    <source>Nature</source>
                    <year>2001</year>
                    <volume>409</volume>
                    <fpage>953</fpage>
                    <lpage>958</lpage>
                    <pub-id pub-id-type="pmid">11237021</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1551">
                  <label>16</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Marchler-Bauer</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Panchenko</surname>
                        <given-names>AR</given-names>
                      </name>
                      <name>
                        <surname>Shoemaker</surname>
                        <given-names>BA</given-names>
                      </name>
                      <name>
                        <surname>Thiessen</surname>
                        <given-names>PA</given-names>
                      </name>
                      <name>
                        <surname>Geer</surname>
                        <given-names>LY</given-names>
                      </name>
                      <name>
                        <surname>Bryant</surname>
                        <given-names>SH</given-names>
                      </name>
                    </person-group>
                    <article-title>CDD: a database of conserved domain alignments with links to domain three-dimensional structure</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>2002</year>
                    <volume>30</volume>
                    <fpage>281</fpage>
                    <lpage>283</lpage>
                    <pub-id pub-id-type="pmid">11752315</pub-id>
                    <pub-id pub-id-type="other">99109</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1552">
                  <label>17</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Yeh</surname>
                        <given-names>RF</given-names>
                      </name>
                      <name>
                        <surname>Lim</surname>
                        <given-names>LP</given-names>
                      </name>
                      <name>
                        <surname>Burge</surname>
                        <given-names>CB</given-names>
                      </name>
                    </person-group>
                    <article-title>Computational inference of homologous gene structures in the human genome</article-title>
                    <source>Genome Res</source>
                    <year>2001</year>
                    <volume>11</volume>
                    <fpage>803</fpage>
                    <lpage>816</lpage>
                    <pub-id pub-id-type="pmid">11337476</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1553">
                  <label>18</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Bairoch</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Apweiler</surname>
                        <given-names>R</given-names>
                      </name>
                    </person-group>
                    <article-title>The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998</article-title>
                    <source>Nucleic Acids Res</source>
                    <year>1998</year>
                    <volume>26</volume>
                    <fpage>38</fpage>
                    <lpage>42</lpage>
                    <pub-id pub-id-type="pmid">9399796</pub-id>
                    <pub-id pub-id-type="other">147215</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1554">
                  <label>19</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Sherry</surname>
                        <given-names>ST</given-names>
                      </name>
                      <name>
                        <surname>Ward</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Sirotkin</surname>
                        <given-names>K</given-names>
                      </name>
                    </person-group>
                    <article-title>dbSNP&mdash;database for single nucleotide polymorphisms and other classes of minor genetic variation</article-title>
                    <source>Genome Res</source>
                    <year>1999</year>
                    <volume>9</volume>
                    <fpage>677</fpage>
                    <lpage>679</lpage>
                    <pub-id pub-id-type="pmid">10447503</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1555">
                  <label>20</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Dib</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Faure</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Fizames</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Samson</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Drouot</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Vignal</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Millasseau</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Marc</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Hazan</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Seboun</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Lathrop</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Gyapay</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Morissette</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Weissenbach</surname>
                        <given-names>J</given-names>
                      </name>
                    </person-group>
                    <article-title>A comprehensive genetic map of the human genome based on 5,264 microsatellites</article-title>
                    <source>Nature</source>
                    <year>1996</year>
                    <volume>380</volume>
                    <fpage>152</fpage>
                    <lpage>154</lpage>
                    <pub-id pub-id-type="pmid">8600387</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1556">
                  <label>21</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Broman</surname>
                        <given-names>KW</given-names>
                      </name>
                      <name>
                        <surname>Murray</surname>
                        <given-names>JC</given-names>
                      </name>
                      <name>
                        <surname>Sheffield</surname>
                        <given-names>VC</given-names>
                      </name>
                      <name>
                        <surname>White</surname>
                        <given-names>RL</given-names>
                      </name>
                      <name>
                        <surname>Weber</surname>
                        <given-names>JL</given-names>
                      </name>
                    </person-group>
                    <article-title>Comprehensive human genetic maps: individual and sex-specific variation in recombination</article-title>
                    <source>Am J Hum Genet</source>
                    <year>1998</year>
                    <volume>63</volume>
                    <fpage>861</fpage>
                    <lpage>869</lpage>
                    <pub-id pub-id-type="pmid">9718341</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1557">
                  <label>22</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Kong</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Gudbjartsson</surname>
                        <given-names>DF</given-names>
                      </name>
                      <name>
                        <surname>Sainz</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Jonsdottir</surname>
                        <given-names>GM</given-names>
                      </name>
                      <name>
                        <surname>Gudjonsson</surname>
                        <given-names>SA</given-names>
                      </name>
                      <name>
                        <surname>Richardsson</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Sigurdardottir</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Barnard</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Hallbeck</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Masson</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Shlien</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Palsson</surname>
                        <given-names>ST</given-names>
                      </name>
                      <name>
                        <surname>Frigge</surname>
                        <given-names>ML</given-names>
                      </name>
                      <name>
                        <surname>Thorgeirsson</surname>
                        <given-names>TE</given-names>
                      </name>
                      <name>
                        <surname>Gulcher</surname>
                        <given-names>JR</given-names>
                      </name>
                      <name>
                        <surname>Stefansson</surname>
                        <given-names>K</given-names>
                      </name>
                    </person-group>
                    <article-title>A high-resolution recombination map of the human genome</article-title>
                    <source>Nat Genet</source>
                    <year>2002</year>
                    <volume>31</volume>
                    <fpage>241</fpage>
                    <lpage>247</lpage>
                    <pub-id pub-id-type="pmid">12053178</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1558">
                  <label>23</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                      <name>
                        <surname>Boguski</surname>
                        <given-names>MS</given-names>
                      </name>
                      <name>
                        <surname>Stewart</surname>
                        <given-names>EA</given-names>
                      </name>
                      <name>
                        <surname>Stein</surname>
                        <given-names>LD</given-names>
                      </name>
                      <name>
                        <surname>Gyapay</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Rice</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>White</surname>
                        <given-names>RE</given-names>
                      </name>
                      <name>
                        <surname>Rodriguez-Tome</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Aggarwal</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Bajorek</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Bentolila</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Birren</surname>
                        <given-names>BB</given-names>
                      </name>
                      <name>
                        <surname>Butler</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Castle</surname>
                        <given-names>AB</given-names>
                      </name>
                      <name>
                        <surname>Chiannilkulchai</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Chu</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Clee</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Cowles</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Day</surname>
                        <given-names>PJ</given-names>
                      </name>
                      <name>
                        <surname>Dibling</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Drouot</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Dunham</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Duprat</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>East</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Hudson</surname>
                        <given-names>TJ</given-names>
                      </name>
                    </person-group>
                    <article-title>A gene map of the human genome</article-title>
                    <source>Science</source>
                    <year>1996</year>
                    <volume>274</volume>
                    <fpage>540</fpage>
                    <lpage>546</lpage>
                    <pub-id pub-id-type="pmid">8849440</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1559">
                  <label>24</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Deloukas</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                      <name>
                        <surname>Gyapay</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Beasley</surname>
                        <given-names>EM</given-names>
                      </name>
                      <name>
                        <surname>Soderlund</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Rodriguez-Tome</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Hui</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Matise</surname>
                        <given-names>TC</given-names>
                      </name>
                      <name>
                        <surname>McKusick</surname>
                        <given-names>KB</given-names>
                      </name>
                      <name>
                        <surname>Beckmann</surname>
                        <given-names>JS</given-names>
                      </name>
                      <name>
                        <surname>Bentolila</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Bihoreau</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Birren</surname>
                        <given-names>BB</given-names>
                      </name>
                      <name>
                        <surname>Browne</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Butler</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Castle</surname>
                        <given-names>AB</given-names>
                      </name>
                      <name>
                        <surname>Chiannilkulchai</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Clee</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Day</surname>
                        <given-names>PJ</given-names>
                      </name>
                      <name>
                        <surname>Dehejia</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Dibling</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Drouot</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Duprat</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Fizames</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Bentley</surname>
                        <given-names>DR</given-names>
                      </name>
                    </person-group>
                    <article-title>A physical map of 30,000 human genes</article-title>
                    <source>Science</source>
                    <year>1998</year>
                    <volume>282</volume>
                    <fpage>744</fpage>
                    <lpage>746</lpage>
                    <pub-id pub-id-type="pmid">9784132</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1560">
                  <label>25</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Agarwala</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Applegate</surname>
                        <given-names>DL</given-names>
                      </name>
                      <name>
                        <surname>Maglott</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                      <name>
                        <surname>Schaffer</surname>
                        <given-names>AA</given-names>
                      </name>
                    </person-group>
                    <article-title>A fast and scalable radiation hybrid map construction and integration strategy</article-title>
                    <source>Genome Res</source>
                    <year>2000</year>
                    <volume>10</volume>
                    <fpage>350</fpage>
                    <lpage>364</lpage>
                    <pub-id pub-id-type="pmid">10720576</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1561">
                  <label>26</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Olivier</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Aggarwal</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Allen</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Almendras</surname>
                        <given-names>AA</given-names>
                      </name>
                      <name>
                        <surname>Bajorek</surname>
                        <given-names>ES</given-names>
                      </name>
                      <name>
                        <surname>Beasley</surname>
                        <given-names>EM</given-names>
                      </name>
                      <name>
                        <surname>Brady</surname>
                        <given-names>SD</given-names>
                      </name>
                      <name>
                        <surname>Bushard</surname>
                        <given-names>JM</given-names>
                      </name>
                      <name>
                        <surname>Bustos</surname>
                        <given-names>VI</given-names>
                      </name>
                      <name>
                        <surname>Chu</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Chung</surname>
                        <given-names>TR</given-names>
                      </name>
                      <name>
                        <surname>De Witte</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Denys</surname>
                        <given-names>ME</given-names>
                      </name>
                      <name>
                        <surname>Dominguez</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Fang</surname>
                        <given-names>NY</given-names>
                      </name>
                      <name>
                        <surname>Foster</surname>
                        <given-names>BD</given-names>
                      </name>
                      <name>
                        <surname>Freudenberg</surname>
                        <given-names>RW</given-names>
                      </name>
                      <name>
                        <surname>Hadley</surname>
                        <given-names>D</given-names>
                      </name>
                      <name>
                        <surname>Hamilton</surname>
                        <given-names>LR</given-names>
                      </name>
                      <name>
                        <surname>Jeffrey</surname>
                        <given-names>TJ</given-names>
                      </name>
                      <name>
                        <surname>Kelly</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Lazzeroni</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Levy</surname>
                        <given-names>MR</given-names>
                      </name>
                      <name>
                        <surname>Lewis</surname>
                        <given-names>SC</given-names>
                      </name>
                      <name>
                        <surname>Liu</surname>
                        <given-names>X</given-names>
                      </name>
                      <name>
                        <surname>Lopez</surname>
                        <given-names>FJ</given-names>
                      </name>
                      <name>
                        <surname>Louie</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Marquis</surname>
                        <given-names>JP</given-names>
                      </name>
                      <name>
                        <surname>Martinez</surname>
                        <given-names>RA</given-names>
                      </name>
                      <name>
                        <surname>Matsuura</surname>
                        <given-names>MK</given-names>
                      </name>
                      <name>
                        <surname>Misherghi</surname>
                        <given-names>NS</given-names>
                      </name>
                      <name>
                        <surname>Norton</surname>
                        <given-names>JA</given-names>
                      </name>
                      <name>
                        <surname>Olshen</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Perkins</surname>
                        <given-names>SM</given-names>
                      </name>
                      <name>
                        <surname>Perou</surname>
                        <given-names>AJ</given-names>
                      </name>
                      <name>
                        <surname>Piercy</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Piercy</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Qin</surname>
                        <given-names>F</given-names>
                      </name>
                      <name>
                        <surname>Reif</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Sheppard</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Shokoohi</surname>
                        <given-names>V</given-names>
                      </name>
                      <name>
                        <surname>Smick</surname>
                        <given-names>GA</given-names>
                      </name>
                      <name>
                        <surname>Sun</surname>
                        <given-names>WL</given-names>
                      </name>
                      <name>
                        <surname>Stewart</surname>
                        <given-names>EA</given-names>
                      </name>
                      <name>
                        <surname>Fernando</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Tejeda</surname>
                      </name>
                      <name>
                        <surname>Tran</surname>
                        <given-names>NM</given-names>
                      </name>
                      <name>
                        <surname>Trejo</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Vo</surname>
                        <given-names>NT</given-names>
                      </name>
                      <name>
                        <surname>Yan</surname>
                        <given-names>SC</given-names>
                      </name>
                      <name>
                        <surname>Zierten</surname>
                        <given-names>DL</given-names>
                      </name>
                      <name>
                        <surname>Zhao</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Sachidanandam</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Trask</surname>
                        <given-names>BJ</given-names>
                      </name>
                      <name>
                        <surname>Myers</surname>
                        <given-names>RM</given-names>
                      </name>
                      <name>
                        <surname>Cox</surname>
                        <given-names>DR</given-names>
                      </name>
                    </person-group>
                    <article-title>A high-resolution radiation hybrid map of the human genome draft sequence</article-title>
                    <source>Science</source>
                    <year>2001</year>
                    <volume>291</volume>
                    <fpage>1298</fpage>
                    <lpage>1302</lpage>
                    <pub-id pub-id-type="pmid">11181994</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
      </body>
    </book-part>
    <book-part id="bid.587" book-part-type="part" book-part-number="Part 3">
      <book-part-meta>
        <title-group>
          <title>Querying and Linking the Data</title>
        </title-group>
      </book-part-meta>
      <body>
        <book-part id="bid.588" book-part-type="chapter" book-part-number="15">
          <book-part-meta>
            <title-group>
              <title>The Entrez Search and Retrieval System</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Ostell</surname>
                  <given-names>Jim</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch15d1"/>
            <abstract>
              <title>Summary</title>
              <p>Entrez is the text-based search and retrieval system used at NCBI for all of the major databases, including PubMed, Nucleotide and Protein Sequences, Protein Structures, Complete Genomes, Taxonomy, OMIM, and many others. Entrez is at once an indexing and retrieval system, a collection of data from many sources, and an organizing principle for biomedical information. These general concepts are the focus of this chapter. Other chapters cover the details of a specific Entrez database (e.g., PubMed in <xref ref-type="other" rid="bid.42">Chapter 2</xref>) or a specific source of data (e.g., GenBank in <xref ref-type="other" rid="bid.2">Chapter 1</xref>).</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.589">
              <title>Entrez Design Principles</title>
              <sec id="bid.590">
                <title>History</title>
                <p>The first version of Entrez was distributed by NCBI in 1991 on CD-ROM. At that time, it consisted of nucleotide sequences from GenBank and PDB; protein sequences from translated GenBank, PIR, SWISS-PROT, PDB, and PRF; and associated citations and abstracts from MEDLINE (now PubMed and referred to as PubMed below). We will use this first design to illustrate the principles behind Entrez.</p>
              </sec>
              <sec id="bid.591">
                <title>Entrez Nodes Represent Data</title>
                <p>An Entrez &ldquo;node&rdquo; is a collection of data that is grouped together and indexed together. It is usually referred to as an Entrez database. In the first version of Entrez, there were three nodes: published articles, nucleotide sequences, and protein sequences. Each node represents specific data objects of the same type, e.g., protein sequences, which are each given a unique ID (UID) within that logical Entrez Proteins node. Records in a node may come from a single source (e.g., all published articles are from PubMed) or many sources (e.g., proteins are from translated GenBank sequences, <xref ref-type="other" rid="bid.1412">SWISS-PROT</xref>, or <xref ref-type="other" rid="bid.1374">PIR</xref>) (<xref ref-type="fig" rid="bid.592">Figure 1</xref>).<fig id="bid.592"><label>1</label><caption><title>The original version of Entrez had just 3 nodes: nucleotides, proteins, and PubMed abstracts. Entrez has now grown to nearly 20 nodes.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch15f1" mime-subtype="gif"/></fig></p>
                <p>Note that the UID identifies a single, well-defined object (i.e., a particular protein sequence or PubMed citation). There may be other information about objects in nodes, such as protein names or <xref ref-type="other" rid="bid.1280">EC number</xref>s, that may be used as index terms to find the record, but these pieces of information are not the central organizing principle of the node. Each data object represents a stable, objective observation of data as much as possible, rather than interpretations of the data, which are subject to change or confusion over time or across disciplines. For example, barring experimental error, a particular mRNA sequence report is not likely to change over the years; however, the given name, position on the chromosome, or function of the protein product may well change as our knowledge develops. Even a published article is a stable observation. The fact that the article was published at a certain time and contained certain words will not change over time, although the importance of the article topic may change many times.</p>
              </sec>
              <sec id="bid.593">
                <title>Entrez Nodes Are Intended for Linking</title>
                <p>Another criterion for selecting a particular data type to be an Entrez node is to enable linking to other Entrez nodes in a useful and reliable way. For example, given a protein sequence, it is very useful to quickly find the nucleotide sequence that encodes it. Or given a research article, it is useful to find the sequences it describes, if any.</p>
                <sec id="bid.594">
                  <title>
                    <bold>Links between Nodes</bold>
                  </title>
                  <p>One way to achieve this is to put all of the information into one record. For example, many GenBank records contain pertinent article citations. However, PubMed also contains the article abstract and additional index terms (e.g., MeSH terms); furthermore, the bibliographic information is also more carefully curated than the citation in a GenBank entry. It therefore makes much more sense to search for articles in PubMed rather than in GenBank.</p>
                  <p>When a subset of articles has been retrieved from PubMed, it may be useful to link to sequence information associated with the abstracts. The article citation in the GenBank record can be used to establish the link to PubMed and, conversely, to make the reciprocal link from the PubMed article back to the GenBank record. Treating each Entrez node separately but enabling linking between related data in different nodes means that the retrieval characteristics for each node can be optimized for the characteristics and strengths of that node, whereas related data can be reached in nodes with different strengths.</p>
                  <p>This approach also means that new connections between data can be made. In the example above, the GenBank record cited the published article, but there was no link from that article in PubMed to the sequence until Entrez made the reciprocal link from PubMed. Now, when searching articles in PubMed, it is possible to find this sequence, although no PubMed records have been changed. Because of this design principle, the Entrez system is richly interconnected, although any particular association may originate from only one record in one node.</p>
                </sec>
                <sec id="bid.595">
                  <title>
                    <bold>Links within Nodes</bold>
                  </title>
                  <p>Another type of linking in Entrez is between records of the same type, often called &ldquo;neighbors&rdquo;, in sequence and structure nodes. Most often these associations are computed at NCBI. For example, in Entrez Proteins, all of the protein sequences are &ldquo;<xref ref-type="other" rid="bid.1246">BLAST</xref>ed&rdquo; against each other, and the highest-scoring hits are stored as indexes within the node. This means that each protein record has associated with it a list of highly similar sequences, or neighbors.</p>
                  <p>Again, associations that may not be present in the original records can be made. For example, a well-annotated SWISS-PROT record for a particular protein may have fields that describe other protein or GenBank records from which it was derived. At a later date, a closely related protein may appear in GenBank that will not be referenced by the SWISS-PROT record. However, if a scientist finds an article in PubMed that has a link to the new GenBank record, that person can look at the protein and then use the BLAST-computed neighbors to find the SWISS-PROT record (as well as many others), although neither the SWISS-PROT record nor the new GenBank record refers to each other anywhere.</p>
                </sec>
              </sec>
              <sec id="bid.596">
                <title>Entrez Nodes Are Intended for Computation</title>
                <p>There are many advantages to establishing new associations by computational methods (as in the GenBank&ndash;SWISS-PROT example above), especially for large, rapidly changing data sets such as those in biomedicine.</p>
                <p>As computers get faster and cheaper, this type of association can be made more efficiently. As data sets get bigger, the problem remains tractable or may even improve because of better statistics. If a new algorithm or approach is found to be an improvement, it is possible to apply it over the whole data set within a practical timescale and by using a reasonable number of resources. Any associations that require human curation, such as the application of controlled vocabularies, do not scale well with rapidly growing sets of data or evolving data interpretations. Although these manual kinds of approaches certainly add value, computational approaches can often produce good results more objectively and efficiently.</p>
              </sec>
            </sec>
            <sec id="bid.597">
              <title>Entrez Is a Discovery System</title>
              <p>A data-retrieval system succeeds when you can retrieve the same data you put in. A discovery system is intended to let you find more information than appears in the original data. By making links between selected nodes and making computed associations within the same node, Entrez is designed to infer relationships between different data that may suggest future experiments or assist in interpretation of the available information, although it may come from different sources.</p>
              <p>The ability to compare genotype information across a huge range of organisms is a powerful tool for molecular biologists. For example, this technique was used in the discovery of a gene associated with hereditary nonpolyposis colon cancer (HNPCC). The tumor cells from most familial cases of HNPCC had altered, short, repeated DNA sequences, suggesting that DNA replication errors had occurred during tumor development. This information caused a group of investigators to look for human homologs of the well-characterized <italic>Escherichia coli</italic> DNA mismatch repair enzyme, MutS (<xref ref-type="bibr" rid="bid.609">1</xref>). Mutants in a <italic>MutS</italic> homolog in yeast, <italic>MSH2</italic>, showed expansion and contraction of dinucleotide repeats similar to the mutation found in the human tumor cells. By comparing the protein sequences between the yeast MSH2, the <italic>E. coli</italic> MutS, and a human gene product isolated and cloned from HNPCC colorectal tumor, the researchers could show that the amino acid sequences of all three proteins were very similar. From this, they inferred that the human gene, which they called <italic>hMSH2</italic>, may also play a role in repairing DNA, and that the mutation found in tumors negatively affects this function, leading to tumor development.</p>
              <p>The researchers could connect the functional data about the yeast and bacterial genes with the genetic mapping and clinical phenotype information in humans. Entrez is designed to support this kind of process when the underlying data are available electronically. In PubMed, the research paper about the discovery of <italic>hMSH2</italic> (<xref ref-type="bibr" rid="bid.609">1</xref>) has links to the protein sequence, which in turn has links to &ldquo;neighbors&rdquo; (related sequences). There are lots of records for this protein and its relatives in many organisms, but among them are the proteins from yeast and <italic>E. coli</italic> that prompted the study. From those records there are links back to the PubMed abstracts of articles that reported these proteins. PubMed also has a &ldquo;neighbor&rdquo; function, <named-content content-type="book-spec-emph">Related articles</named-content>, that represents other articles that contain words and phrases in common with the current record. Because phrases such as &ldquo;<italic>Escherichia coli</italic>&rdquo;, &ldquo;mismatch repair&rdquo;, and &ldquo;MutS&rdquo; all occur in the current article, many of the articles most related to this one describe studies on the <italic>E. coli</italic> mismatch repair system. These articles may not be directly linked to any sequence themselves and may not contain the words &ldquo;human&rdquo; or &ldquo;colon cancer&rdquo; but are relevant to HNPCC nonetheless, because of what the bacterial system may tell us.</p>
            </sec>
            <sec id="bid.598">
              <title>Entrez Is Growing</title>
              <p>The original three-node Entrez system has evolved over the past 10 years to include more nodes (<xref ref-type="fig" rid="bid.592">Figure 1</xref>). These include:
                <list list-type="arabic">
                  <list-item id="bid.599">
                    <label>1</label>
                    <p>Taxonomy, which is organized around the names and phylogenetic relationships of organisms</p>
                  </list-item>
                  <list-item id="bid.600">
                    <label>2</label>
                    <p>Structure, organized around the three-dimensional structures of proteins and nucleic acids</p>
                  </list-item>
                  <list-item id="bid.601">
                    <label>3</label>
                    <p>Genomes, in which each record represents a chromosome of an organism</p>
                  </list-item>
                  <list-item id="bid.602">
                    <label>4</label>
                    <p><italic>Online Mendelian Inheritance in Man</italic> (OMIM), a text-based resource organized around human genes and their phenotypes</p>
                  </list-item>
                  <list-item id="bid.603">
                    <label>5</label>
                    <p>PopSet, consisting of collections of aligned sequences from a single population study</p>
                  </list-item>
                  <list-item id="bid.604">
                    <label>6</label>
                    <p>Books, representing published books in biomedicine</p>
                  </list-item>
                </list>
              </p>
              <p>More nodes are planned for addition in the near future. Each one of these nodes is richly connected to others. Each offers unique information and unique new relationships among its members. The combination of new links and new relationships increases the chances for discovery. The addition of each new node creates different paths through the data that may lead to new connections, without more work on the old nodes.</p>
            </sec>
            <sec id="bid.605">
              <title>How Entrez Works</title>
              <p>Entrez integrates data from a large number of sources, formats, and databases into a uniform information model and retrieval system. The actual databases from which records are retrieved and on which the Entrez indexes are based have different designs, based on the type of data, and reside on different machines. These will be referred to as the &ldquo;source databases&rdquo;. A common theme in the implementation of Entrez is that some functions are unique to each source database, whereas others are common to all Entrez databases.</p>
              <sec id="bid.606">
                <title>A Division of Labor: Basic Principles</title>
                <p>Some of the common routines and formats for every Entrez node include the term lists and posting files (i.e., the retrieval engine) used for Boolean queries, the links within and between nodes, and the summary format used for listing search results in which each record is called a DocSum. Generally, an Entrez query is a <xref ref-type="other" rid="bid.1253">Boolean</xref> expression that is evaluated by the common Entrez engine and yields a list of unique ID numbers (UIDs), which identify records in an Entrez node. Given one or more UIDs, Entrez can retrieve the DocSum(s) very quickly. The links made within or between Entrez nodes from one or more UIDs is also a function across all Entrez source databases.</p>
                <p>The software that tracks the addition of new or updated records or identifies those that should be deleted from Entrez may be unique for each source database. Each database must also have accompanying software to gather index terms, DocSums, and links from the source data and present them to the common Entrez indexer. This can be achieved through either a set of C++ libraries or by generating an XML document in a specific DTD that contains the terms, DocSums, and links. Although the common engine retrieves a DocSum(s) given a UID(s), the retrieval of a full, formatted record is directed to the source database, where software unique to that database is used to format the record correctly. All of this software is written by the NCBI group that runs the database.</p>
                <p>This combination of database-specific software and a common set of Entrez routines and applications allows code sharing and common large-retrieval server administration but enables flexibility and simplicity for a wide variety of data sources.</p>
              </sec>
              <sec id="bid.607">
                <title>Software</title>
                <p>Although the basic principles of Entrez have remained the same for almost a decade, the software implementation has been through at least three major redesigns and many minor ones.</p>
                <p>Currently, Entrez is written using the NCBI C++ Toolkit. The indexing fields (which for PubMed, for example, would be Title, Author, Publication Date, Journal, Abstract, and so on) and DocSum fields (which for PubMed are Author, Title, Journal, Publication Date, Volume, and Page Number) for each node are defined in a configuration file; but for performance at runtime, the configuration files are used to automatically generate base classes for each database. These are the basic pieces of information used by Entrez that can also be inherited and used by more database-specific, hand-coded features. The term indexes are based on the Indexed Sequential-Access Method (ISAM) and are in large, shared, memory-mapped files. The postings are large bitmaps, with one bit per document in the node. Depending on how sparsely populated the posting is, the bit array is adaptively compressed on disk using one of four possible schemes. Boolean operations are performed by using AND or OR postings of bit arrays into a result bit array. DocSums are small, fielded data structures stored on the same machines as the postings to support rapid retrieval.</p>
                <p>The Web-based Entrez retrieval program, called <italic>query</italic>, is a fast cgi application that uses the Web application framework from the NCBI C++ Toolkit. One aspect of this framework is a set of classes that represents an HTML page. These classes allow the combination of static template pages, on the fly, with callbacks to class methods at tagged parts of the template. The Web page generated in an Entrez session contains elements from static templates and elements generated dynamically from common Entrez classes and from classes unique to one or a few Entrez nodes. Again, this design supports a common core of robust, common functionality maintained by one group, with support for customizations by diverse groups within NCBI.</p>
                <p>Boolean query processing, DocSum retrieval, and other common functions are supported on a number of load-balanced &ldquo;front-end&rdquo; UNIX machines. Because Entrez can support session context (for example, in the use of query history, NCBI Cubby, Filters, etc.), a &ldquo;history server&rdquo; has been implemented on the front-end machines so that if a user is sent to machine &ldquo;A&rdquo; by the load balancer for their first query but to machine &ldquo;B&rdquo; for the second query, Entrez can quickly locate the user's query history and obtain it from machine &ldquo;A&rdquo;. Other than that, the front-end machines are completely independent of each other and can be added and removed readily from <italic>query</italic> support. Retrieval of full documents comes from a variety of &ldquo;back-end&rdquo; databases, depending on the node. These might be Sybase or Microsoft SQL Server relational databases of a variety of schemas or text files of various formats. Links are supported using the Sybase IQ database product.</p>
              </sec>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.609">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Fishel</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Lescoe</surname>
                        <given-names>MK</given-names>
                      </name>
                      <name>
                        <surname>Rao</surname>
                        <given-names>MR</given-names>
                      </name>
                      <name>
                        <surname>Copeland</surname>
                        <given-names>NG</given-names>
                      </name>
                      <name>
                        <surname>Jenkins</surname>
                        <given-names>NA</given-names>
                      </name>
                      <name>
                        <surname>Garber</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Kane</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Kolodner</surname>
                        <given-names>R</given-names>
                      </name>
                    </person-group>
                    <article-title>The human mutator gene homolog MSH2 and its association with hereditary nonpolyposis colon cancer</article-title>
                    <source>Cell</source>
                    <volume>75</volume>
                    <fpage>1027</fpage>
                    <lpage>1038</lpage>
                    <year>1993</year>
                    <pub-id pub-id-type="pmid">8252616</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.610" book-part-type="chapter" book-part-number="16">
          <book-part-meta>
            <title-group>
              <title>The BLAST Sequence Analysis Tool</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Madden</surname>
                  <given-names>Tom</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch16d1"/>
            <abstract>
              <title>Summary</title>
              <p>The comparison of nucleotide or protein sequences from the same or different organisms is a very powerful tool in molecular biology. By finding similarities between sequences, scientists can infer the function of newly sequenced genes, predict new members of gene families, and explore evolutionary relationships. Now that whole genomes are being sequenced, sequence similarity searching can be used to predict the location and function of protein-coding and transcription-regulation regions in genomic DNA.</p>
              <p>Basic Local Alignment Search Tool (BLAST) (<xref ref-type="bibr" rid="bid.648">1</xref>, <xref ref-type="bibr" rid="bid.649">2</xref>) is the tool most frequently used for calculating sequence similarity. BLAST comes in variations for use with different query sequences against different databases. All BLAST applications, as well as information on which BLAST program to use and other help documentation, are listed on the BLAST <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/">homepage</ext-link>. This chapter will focus more on how BLAST works, its output, and how both the output and program itself can be further manipulated or customized, rather than on how to use <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/information3.html">BLAST</ext-link> or interpret BLAST results.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.611">
              <title>Introduction</title>
              <p>The way most people use BLAST is to input a nucleotide or protein sequence as a query against all (or a subset of) the public sequence databases, pasting the sequence into the textbox on one of the BLAST <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/">Web pages</ext-link>. This sends the query over the Internet, the search is performed on the NCBI databases and servers, and the results are posted back to the person's browser in the chosen display format. However, many biotech companies, genome scientists, and bioinformatics personnel may want to use &ldquo;stand-alone&rdquo; BLAST to query their own, local databases or want to customize BLAST in some way to make it better suit their needs. Stand-alone BLAST comes in two forms: the executables that can be run from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/blast_overview.html#executables">command line</ext-link>; or the Standalone WWW <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/blast_overview.html#wwwserver">BLAST Server</ext-link>, which allows users to set up their own in-house versions of the BLAST Web pages.</p>
              <p>There are many different <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/blast/html/BLASThomehelp.html">variations</ext-link> of BLAST available to use for different sequence comparisons, e.g., a DNA query to a DNA database, a protein query to a protein database, and a DNA query, translated in all six reading frames, to a protein sequence database. Other <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/producttable.html">adaptations</ext-link> of BLAST, such as PSI-BLAST (for iterative protein sequence similarity searches using a position-specific score matrix) and RPS-BLAST (for searching for protein domains in the Conserved Domains Database, <xref ref-type="other" rid="bid.94">Chapter 3</xref>) perform comparisons against sequence profiles.</p>
              <p>This chapter will first describe the BLAST architecture&mdash;how it works at the NCBI site&mdash;and then go on to describe the various BLAST outputs. The best known of these outputs is the default display from BLAST Web pages, the so-called &ldquo;traditional report&rdquo;. As well as obtaining BLAST results in the traditional report, results can also be delivered in structured output, such as a hit table (see below), XML, or ASN.1. The optimal choice of output format depends upon the application. The final part of the chapter discusses stand-alone BLAST and describes possibilities for customization. There are many interfaces to BLAST that are often not exploited by users but can lead to more efficient and robust applications.</p>
            </sec>
            <sec id="bid.612">
              <title>How BLAST Works: The Basics</title>
              <p>The BLAST algorithm is a heuristic program, which means that it relies on some smart shortcuts to perform the search faster. BLAST performs &ldquo;local&rdquo; alignments. Most proteins are modular in nature, with functional domains often being repeated within the same protein as well as across different proteins from different species. The BLAST algorithm is tuned to find these domains or shorter stretches of sequence similarity. The local alignment approach also means that a mRNA can be aligned with a piece of genomic DNA, as is frequently required in genome assembly and analysis. If instead BLAST started out by attempting to align two sequences over their entire lengths (known as a global alignment), fewer similarities would be detected, especially with respect to domains and motifs.</p>
              <p>When a query is submitted via one of the BLAST Web pages, the sequence, plus any other input information such as the database to be searched, word size, expect value, and so on, are fed to the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/BLAST_algorithm.html">algorithm</ext-link> on the BLAST server. BLAST works by first making a look-up table of all the &ldquo;words&rdquo; (short subsequences, which for proteins the default is three letters) and &ldquo;neighboring words&rdquo;, i.e., similar words in the query sequence. The sequence database is then scanned for these &ldquo;hot spots&rdquo;. When a match is identified, it is used to initiate <xref ref-type="other" rid="bid.1296">gap</xref>-free and gapped extensions of the &ldquo;word&rdquo;.</p>
              <p>BLAST does not search GenBank flatfiles (or any subset of GenBank flatfiles) directly. Rather, sequences are made into BLAST databases. Each entry is split, and two files are formed, one containing just the header information and one containing just the sequence information. These are the data that the algorithm uses. If BLAST is to be run in &ldquo;stand-alone&rdquo; mode, the data file could consist of local, private data, downloaded NCBI BLAST databases, or a combination of the two.</p>
              <p>After the algorithm has looked up all possible &ldquo;words&rdquo; from the query sequence and extended them maximally, it assembles the best alignment for each query&ndash;sequence pair and writes this information to an SeqAlign data structure (in <xref ref-type="other" rid="bid.1242">ASN.1</xref> ; also used by Sequin, see <xref ref-type="other" rid="bid.538">Chapter 12</xref>). The SeqAlign structure in itself does not contain the sequence information; rather, it refers to the sequences in the BLAST database (<xref ref-type="fig" rid="bid.613">Figure 1</xref>).<fig id="bid.613"><label>1</label><caption><title>How the BLAST results Web pages are assembled.</title><p>The QBLAST system located on the BLAST server executes the search, writing information about the sequence alignment in ASN.1. The results can then be formatted by fetching the ASN.1 (<italic>fetch ASN.1</italic>) and fetching the sequences (<italic>fetch sequence</italic>) from the BLAST databases. Because the execution of the search algorithm is decoupled from the formatting, the results can be delivered in a variety of formats without re-running the search.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch16f1" mime-subtype="gif"/></fig></p>
              <p>The BLAST Formatter, which sits on the BLAST server, can use the information in the SeqAlign to retrieve the similar sequences found and display them in a variety of ways. Thus, once a query has been completed, the results can be reformatted without having to re-execute the search. This is possible because of the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/blast_overview.html#blastq">QBLAST</ext-link> system.</p>
            </sec>
            <sec id="bid.614">
              <title>BLAST Scores and Statistics</title>
              <p>Once BLAST has found a similar sequence to the query in the database, it is helpful to have some idea of whether the alignment is &ldquo;good&rdquo; and whether it portrays a possible biological relationship, or whether the similarity observed is attributable to chance alone. BLAST uses <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/tutorial/Altschul-1.html">statistical theory</ext-link> to produce a <xref ref-type="other" rid="bid.1245">bit score</xref> and expect value (<xref ref-type="other" rid="bid.1279">E-value</xref>) for each alignment pair (query to hit).</p>
              <p>The bit score gives an indication of how good the alignment is; the higher the score, the better the alignment. In general terms, this score is calculated from a formula that takes into account the alignment of similar or identical residues, as well as any gaps introduced to align the sequences. A key element in this calculation is the &ldquo;<xref ref-type="other" rid="bid.1411">substitution matrix</xref> &rdquo;, which assigns a score for aligning any possible pair of residues. The <xref ref-type="other" rid="bid.1252">BLOSUM62</xref> matrix is the default for most BLAST programs, the exceptions being blastn and MegaBLAST (programs that perform nucleotide&ndash;nucleotide comparisons and hence do not use protein-specific matrices). Bit scores are normalized, which means that the bit scores from different alignments can be compared, even if different scoring matrices have been used.</p>
              <p>The E-value gives an indication of the statistical significance of a given pairwise alignment and reflects the size of the database and the scoring system used. The lower the E-value, the more significant the hit. A sequence alignment that has an E-value of 0.05 means that this similarity has a 5 in 100 (1 in 20) chance of occurring by chance alone. Although a statistician might consider this to be significant, it still may not represent a biologically meaningful result, and analysis of the alignments (see below) is required to determine &ldquo;biological&rdquo; significance.</p>
            </sec>
            <sec id="bid.615">
              <title>BLAST Output: 1. The Traditional Report</title>
              <p>Most BLAST users are familiar with the so-called &ldquo;traditional&rdquo; BLAST report. The report consists of three major sections: (1) the header, which contains information about the query sequence, the database searched (<xref ref-type="fig" rid="bid.616">Figure 2</xref>). On the Web, there is also a graphical overview (<xref ref-type="fig" rid="bid.617">Figure 3</xref>); (2) the one-line descriptions of each database sequence found to match the query sequence; these provide a quick overview for browsing (<xref ref-type="fig" rid="bid.618">Figure 4</xref>); (3) the alignments for each database sequence matched (<xref ref-type="fig" rid="bid.619">Figure 5</xref>) (there may be more than one alignment for a database sequence it matches).<fig id="bid.616"><label>2</label><caption><title>The BLAST report header.</title><p>The <italic>top line</italic> gives information about the type of program (in this case, <italic>BLASTP</italic>), the version (<italic>2.2.1</italic>), and a version release date. The research paper that describes BLAST is then cited, followed by the request ID (issued by QBLAST), the query sequence <xref ref-type="other" rid="bid.1273">definition line</xref>, and a summary of the database searched. The <named-content content-type="book-spec-emph">Taxonomy reports</named-content> link displays this BLAST result on the basis of information in the Taxonomy database (<xref ref-type="other" rid="bid.246">Chapter 4</xref>).</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch16f2" mime-subtype="gif"/></fig><fig id="bid.617"><label>3</label><caption><title>Graphical overview of BLAST results.</title><p>The query sequence is represented by the <italic>numbered red bar</italic> at the <italic>top</italic> of the figure. Database hits are shown aligned to the query, <italic>below</italic> the red bar. Of the aligned sequences, the most similar are shown closest to the query. In this case, there are three high-scoring database matches that align to most of the query sequence. The next twelve bars represent lower-scoring matches that align to two regions of the query, from about residues 3&ndash;60 and residues 220&ndash;500. The <italic>cross-hatched parts</italic> of the these bars indicate that the two regions of similarity are on the same protein, but that this intervening region does not match. The remaining bars show lower-scoring alignments. Mousing over the bars displays the definition line for that sequence to be shown in the window above the graphic.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch16f3" mime-subtype="gif"/></fig><fig id="bid.618"><label>4</label><caption><title>One-line descriptions in the BLAST report.</title><p>Each line is composed of four fields: (<italic>a</italic>) the gi number, database designation, Accession number, and locus name for the matched sequence, separated by vertical bars (Appendix 1); (<italic>b</italic>) a brief textual description of the sequence, the definition. This usually includes information on the organism from which the sequence was derived, the type of sequence (e.g., mRNA or DNA), and some information about function or phenotype. The definition line is often truncated in the one-line descriptions to keep the display compact; (<italic>c</italic>) the alignment score in bits. Higher scoring hits are found at the top of the list; and (<italic>d</italic>) the E-value, which provides an estimate of statistical significance. For the first hit in the list, the gi number is <italic>116365</italic>, the database designation is <italic>sp</italic> (for SWISS-PROT), the Accession number is <italic>P26374</italic>, the locus name is <italic>RAE2_HUMAN</italic>, the definition line is <italic>Rab proteins</italic>, the score is <italic>1216</italic>, and the E-value is <italic>0.0</italic>. Note that the first 17 hits have very low E-values (much less than 1) and are either RAB proteins or GDP dissociation inhibitors. The other database matches have much higher E-values, 0.5 and above, which means that these sequences may have been matched by chance alone.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch16f4" mime-subtype="gif"/></fig><fig id="bid.619"><label>5</label><caption><title>A pairwise sequence alignment from a BLAST report.</title><p>The alignment is preceded by the sequence identifier, the full definition line, and the length of the matched sequence, in amino acids. Next comes the bit score (the raw score is in <italic>parentheses</italic>) and then the E-value. The following line contains information on the number of identical residues in this alignment (<italic>Identities</italic>), the number of conservative substitutions (<italic>Positives</italic>), and if applicable, the number of gaps in the alignment. Finally, the actual alignment is shown, with the query on <italic>top</italic>, and the database match is labeled as <italic>Sbjct</italic>, below. The numbers at <italic>left</italic> and <italic>right</italic> refer to the position in the amino acid sequence. One or more dashes (&ndash;) within a sequence indicate insertions or deletions. Amino acid residues in the query sequence that have been masked because of low complexity are replaced by Xs (see, for example, the <italic>fourth</italic> and <italic>last</italic> blocks). The line between the two sequences indicates the similarities between the sequences. If the query and the subject have the same amino acid at a given location, the residue itself is shown. Conservative substitutions, as judged by the substitution matrix, are indicated with +.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch16f5" mime-subtype="gif"/></fig></p>
              <p>The traditional report is really designed for human readability, as opposed to being parsed by a program. For example, the one-line descriptions are useful for people to get a quick overview of their search results, but they are rarely complete descriptors because of limited space. Also, for convenience, there are several pieces of information that are displayed in both the one-line descriptions and alignments (for example, the E-values, scores, and descriptions); therefore, the person viewing the search output does not need to move back and forth between sections.</p>
              <p>New features may be added to the report, e.g., the addition of links to LocusLink records (<xref ref-type="other" rid="bid.788">Chapter 19</xref>) from sequence hits, which result in a change of output format. These are easy for people to pick up on and take advantage of but can trip programs that parse this BLAST output.</p>
              <p>By default, a maximum of 500 sequence matches are displayed, which can be changed on the advanced BLAST page with the <named-content content-type="book-spec-emph">Alignments</named-content> option. Many components of the BLAST results display via the Internet and are hyperlinked to the same information at different places in the page, to additional information including help documentation, and to the Entrez sequence records of matched sequences. These records provide more information about the sequence, including links to relevant research abstracts in PubMed.</p>
            </sec>
            <sec id="bid.620">
              <title>BLAST Output: 2. The Hit Table</title>
              <p>Although the traditional report is ideal for investigating the characteristics of one gene or protein, often scientists want to make a large number of BLAST runs for a specialized purpose and need only a subset of the information contained in the traditional BLAST report. Furthermore, in cases where the BLAST output will be processed further, it can be unreliable to parse the traditional report. The traditional report is merely a display format with no formal structure or rules, and improvements may be made at any time, changing the underlying <xref ref-type="other" rid="bid.1312">HTML</xref>. The hit table format provides a simple and clean alternative (<xref ref-type="fig" rid="bid.621">Figure 6</xref>).<fig id="bid.621"><label>6</label><caption><title>BLAST output in hit table format.</title><p>This shows the results of a search of an <italic>E. coli</italic> database using a human sequence as a query. The lines starting with a # sign should be considered comments and ignored. The <italic>last comment line</italic> lists the fields in the table.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch16f6" mime-subtype="gif"/></fig></p>
              <p>The screening of many newly sequenced human Expressed Sequence Tags (<xref ref-type="other" rid="bid.1283">EST</xref>s) for contamination by the <italic>Escherichia coli</italic> cloning vector is a good example of when it is preferable to use the hit table output over the traditional report. In this case, a strict, high E-value threshold would be applied to differentiate between contaminating <italic>E. coli</italic> sequence and the human sequence. Those human ESTs that find very strong, near-exact <italic>E.coli</italic> sequence matches can be discarded without further examination. (Borderline cases may require further examination by a scientist.)</p>
              <p>For these purposes, the hit table output is more useful than the traditional report; it contains only the information required in a more formal structure. The hit table output contains no sequences or definition lines, but for each sequence matched, it lists the sequence identifier, the start and stop points for stretches of sequence similarity (offset by one residue), the percent identity of the match, and the E-value.</p>
            </sec>
            <sec id="bid.622">
              <title>BLAST Output: 3. Structured Output</title>
              <p>There are drawbacks to parsing both the BLAST report and even the simpler hit table. There is no way to automatically check for truncated or otherwise corrupted output in cases when a large number of sequences are being screened. (This may happen if the disk is full, for example.) Also, there is no rigorous check for syntax changes in the output, such as the addition of new features, which can lead to erroneous parsing. Structured output allows for automatic and rigorous checks for syntax errors and changes. Both <xref ref-type="other" rid="bid.1435">XML</xref> and <xref ref-type="other" rid="bid.1242">ASN.1</xref> are examples of structured output in which there are built-in checks for correct and complete syntax and structure. (In the case of XML, for example, this is ensured by the necessity for matching tags and the DTD.) For text reports, there is often no specification, but perhaps a (incomplete) description of the file is written afterward.</p>
              <sec id="bid.623">
                <title>ASN.1 Is Used by the BLAST Server</title>
                <p>As well as the hit table and traditional report shown in HTML, BLAST results can also be formatted in plain text, XML, and ASN.1 (<xref ref-type="fig" rid="bid.624">Figure 7</xref>), and what's more, the format for a given BLAST result can be changed without re-executing the search.<fig id="bid.624"><label>7</label><caption><title>The different output formats that can be produced from ASN.1.</title><p>Note that some nodes can be viewed as both HTML and text. XML is also structured output but can be produced from ASN.1 because it has equivalent information.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch16f7" mime-subtype="gif"/></fig></p>
                <p>A change in BLAST format without re-executing the search is possible because when a scientist looks at a Web page of BLAST results at NCBI, the HTML that makes that page has been created from ASN.1 (<xref ref-type="fig" rid="bid.624">Figure 7</xref>). Although the formatted results are requested from the server, the information about the alignments is fetched from a disk in ASN.1, as are the corresponding sequences from the BLAST databases (see <xref ref-type="fig" rid="bid.613">Figure 1</xref>). The formatter on the BLAST server then puts these results together as a BLAST report. The BLAST search itself has been uncoupled from the way the result is formatted, thus allowing different output formats from the same search. The strict internal validation of ASN.1 ensures that these output formats can always be produced reliably.</p>
              </sec>
              <sec id="bid.625">
                <title>Information about the Alignment Is Contained within a SeqAlign</title>
                <p>SeqAlign is the ASN.1 object that contains the alignment information about the BLAST search. The SeqAlign does not contain the actual sequence that was found in the match but does contain the start, stop, and gap information, as well as scores, E-values, sequence identifiers, and (DNA) strand information.</p>
                <p>As mentioned above, the actual database sequences are fetched from the BLAST databases when needed. This means that an identifier must uniquely identify a sequence in the database. Furthermore, the query sequence cannot have the same identifier as any sequence in the database unless the query sequence itself is in the database. If one is using stand-alone BLAST with a custom database, it is possible to specify that every sequence is uniquely identified by using the <named-content content-type="book-spec-emph">&ndash;O</named-content> option with formatdb (the program that converts FASTA files to BLAST database format). This also indexes the entries by identifier. Similarly, the <named-content content-type="book-spec-emph">&ndash;J</named-content> option in the (stand-alone) programs blastall, blastpgp, megablast, or rpsblast certifies that the query does not use an identifier already in the database for a different sequence. If the <named-content content-type="book-spec-emph">&ndash;O</named-content> and <named-content content-type="book-spec-emph">&ndash;J</named-content> options are not used, BLAST assigns unique identifiers (for that run) to all sequences and shields the user from this knowledge.</p>
                <p>Any BLAST database or FASTA file from the NCBI Web site that contains gi numbers already satisfies the uniqueness criterion. Unique identifiers are normally a problem only when custom databases are produced and care is not taken in assigning identifiers. The identifier for a FASTA entry is the first token (meaning the letters up to the first space) after the &gt; sign on the definition line. The simplest case is to simply have a unique token (e.g., 1, 2, and so on), but it is possible to construct more complicated identifiers that might, for example, describe the data source. For the FASTA identifiers to be reliably parsed, it is necessary for them to follow a specific syntax (see Appendix 1).</p>
                <p>More information on the SeqAlign produced by BLAST can be found <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="presentation in html">here</ext-link> or be downloaded as a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href=" ftp://ftp.ncbi.nih.gov/blast/demo/blast_programming.ppt">PowerPoint presentation</ext-link>, as well as from the NCBI Toolkit Software Developer's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/IEB/ToolBox/SDKDOCS/INDEX.HTML">handbook</ext-link>.
			</p>
              </sec>
              <sec id="bid.626">
                <title>XML</title>
                <p>XML and ASN.1 are both structured languages and can express the same information; therefore, it is possible to produce a SeqAlign in XML. Some users do not find the format of the information in the SeqAlign to be convenient because it does not contain actual sequence information, and when the sequence is fetched from the BLAST database, it is packed two or four bases per byte. Typically, these users are familiar with the BLAST report and want something similar but in a format that can be parsed reliably. The XML produced by BLAST meets this need, containing the query and database sequences, sequence definition lines, the start and stop points of the alignments (one offset), as well as scores, E-values, and percent identity. There is a public <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/blast/documents/NCBI_BlastOutput.mod">DTD</ext-link> for this XML output.</p>
              </sec>
            </sec>
            <sec id="bid.627">
              <title>BLAST Code</title>
              <p>The BLAST code is part of the NCBI Toolkit, which has many low-level functions to make it platform independent; the Toolkit is supported under Linux and many varieties of UNIX, NT, and MacOS. To use the Toolkit, developers should write a function &ldquo;Main&rdquo;, which is called by the Toolkit &ldquo;main&rdquo;. The BLAST code is contained mostly in the tools directory (see Appendix 2 for an example).</p>
              <p>The BLAST code has a modular design. For example, the Application Programming Interface (API) for retrieval from the BLAST databases is independent of the compute engine. The compute engine is independent from the formatter; therefore, it is possible (as mentioned above) to compute results once but view them in many different modes.</p>
              <sec id="bid.628">
                <title>Readdb API</title>
                <p>The readdb API can be used to easily extract information from the BLAST databases. Among the data available are the date the database was produced, the title, the number of letters, number of sequences, and the longest sequence. Also available are the sequence and description of any entry. The latest version of the BLAST databases also contains a taxid (an integer specifying some node of the NCBI taxonomy tree; see <xref ref-type="other" rid="bid.246">Chapter 4</xref>). Users are strongly encouraged to use the readdb API rather than reading the files associated with the database, because the the files are subject to change. The API, on the other hand, will support the newest version, and an attempt will be made to support older versions. See Appendix 2 for an example of a simple program (db2fasta.c) that demonstrates the use of the readdb API.</p>
              </sec>
              <sec id="bid.629">
                <title>Performing a BLAST Search with C Function Calls</title>
                <p>Only a few function calls are needed to perform a BLAST search. Appendix 3 shows an excerpt from a Demonstration Program doblast.c.</p>
              </sec>
              <sec id="bid.630">
                <title>Formatting a SeqAlign</title>
                <p>MySeqAlignPrint (called in the example in Appendix 3) is a simple function to print a view of a SeqAlign (see Appendix 4).</p>
              </sec>
            </sec>
            <sec id="bid.631">
              <title>Appendix 1. FASTA identifiers.</title>
              <table-wrap id="bid.632">
                <label>1</label>
                <caption>
                  <title>Database identifiers in FASTA definition lines.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Database name</th>
                    <th align="left" valign="top">Identifier syntax</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GenBank</td>
                    <td align="left" valign="top">gb|accession|locus</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">EMBL Data Library</td>
                    <td align="left" valign="top">emb|accession|locus</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">DDBJ, DNA Database of Japan</td>
                    <td align="left" valign="top">dbj|accession|locus</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NBRF PIR</td>
                    <td align="left" valign="top">pir||entry</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Protein Research Foundation</td>
                    <td align="left" valign="top">prf||name</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">SWISS-PROT</td>
                    <td align="left" valign="top">sp|accession|entry name</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Brookhaven Protein Data Bank</td>
                    <td align="left" valign="top">pdb|entry|chain</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Patents</td>
                    <td align="left" valign="top">pat|country|number</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">GenInfo Backbone Id</td>
                    <td align="left" valign="top">bbs|number</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">General database identifier<xref ref-type="fn" rid="mul14"><sup><italic>a</italic></sup></xref></td>
                    <td align="left" valign="top">gnl|database|identifier</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NCBI Reference Sequence</td>
                    <td align="left" valign="top">ref|accession|locus</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Local Sequence identifier</td>
                    <td align="left" valign="top">lcl|identifier</td>
                  </tr>
                </table>
                <table-wrap-foot>
                  <fn symbol="a" id="mul14">
                    <p>gnl allows databases not included in this list to use the same identifying syntax. This is used for sequences in the <bold>trace databases</bold>, e.g., gnl|ti|53185177. The combination of the second and third fields should be unique.</p>
                  </fn>
                </table-wrap-foot>
              </table-wrap>
              <p>The syntax of the FASTA definition lines used in the NCBI BLAST databases depends upon the database from which each sequence was obtained (see <xref ref-type="other" rid="bid.2">Chapter 1</xref> on GenBank). <xref ref-type="table" rid="bid.632">Table 1</xref> shows how the sequence source databases are identified.</p>
              <p>For example, if the identifier of a sequence in a BLAST result is gb|M73307|AGMA13GT, the gb tag indicates that sequence is from GenBank, M73307 is the GenBank Accession number, and AGMA13GT is the GenBank locus.</p>
              <p>The bar (|) separates different fields. In some cases, a field is left empty, although the original specification called for including this field. To make these identifiers backwards-compatible for older parsers, the empty field is denoted by an additional bar (||).</p>
              <p>A gi identifier has been assigned to each sequence in NCBI's sequence databases. If the sequence is from an NCBI database, then the gi number appears at the beginning of the identifier in a traditional report. For example, gi|16760827|ref|NP_456444.1 indicates an NCBI reference sequence with the gi number 16760827 and Accession number NP_456444.1. (In stand-alone BLAST, or when running BLAST from the command line, the <named-content content-type="book-spec-emph">&ndash;I</named-content> option should be used to display the gi number.)</p>
              <p>The reason for adding the gi identifier is to provide a uniform, stable naming convention. If a nucleotide or protein sequence changes (for example, if it is edited by the original submitter of the sequence), a new gi identifier is assigned, but the Accession number of the record remains unchanged. Thus, the gi identifier provides a mechanism for identifying the exact sequence that was used or retrieved in a given search. This is also useful when creating crosslinks between different Entrez databases (<xref ref-type="other" rid="bid.588">Chapter 15</xref>).</p>
            </sec>
            <sec id="bid.633">
              <title>Appendix 2. Readdb API.</title>
              <p>A simple program (db2fasta.c) that demonstrates the use of the readdb API.
                <preformat>
Int2 Main (void)
{
    BioseqPtr bsp;
    Boolean is_prot;
    ReadDBFILEPtr rdfp;
    FILE *fp;
    Int4 index;
if (! GetArgs ("db2fasta", NUMARG, myargs))
{
    return (1);
}
    if (myargs[1].intvalue)
    is_prot = TRUE;
    else
    is_prot = FALSE;
    fp = FileOpen("stdout", "w");
    rdfp = readdb_new(myargs[0].strvalue, is_prot);
    index = readdb_acc2fasta(rdfp, myargs[2].strvalue);
    bsp = readdb_get_bioseq(rdfp, index);
    BioseqRawToFasta(bsp, fp, !is_prot);
    bsp = BioseqFree(bsp);
    rdfp = readdb_destruct(rdfp);
    return 0;
}</preformat>
              </p>
              <p>Note that:
                <list list-type="arabic">
                  <list-item id="bid.634">
                    <label>1</label>
                    <p>Readdb_new allocates an object for reading the database.</p>
                  </list-item>
                  <list-item id="bid.635">
                    <label>2</label>
                    <p>Readdb_acc2fasta fetches the ordinal number (zero offset) of the record given a FASTA identifier (e.g., gb|AAH06776.1|AAH0676).</p>
                  </list-item>
                  <list-item id="bid.636">
                    <label>3</label>
                    <p>Readdb_get_bioseq fetches the BioseqPtr (which contains the sequence, description, and identifiers) for this record.</p>
                  </list-item>
                  <list-item id="bid.637">
                    <label>4</label>
                    <p>BioseqRawToFasta dumps the sequence as FASTA.</p>
                  </list-item>
                </list>
              </p>
              <p>Note also that Main is called, rather than &ldquo;main&rdquo;, and a call to GetArgs is used to get the command-line arguments. db2fasta.c is contained in the tar archive ftp://ftp.ncbi.nih.gov/blast/demo/blast_demo.tar.gz.</p>
            </sec>
            <sec id="bid.638">
              <title>Appendix 3. Excerpt from a demonstration program doblast.c.</title>
              <table-wrap id="bid.642">
                <label>2</label>
                <caption>
                  <title>The most frequently used BLAST options in the BLASTOptionBlk structure.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Type<xref ref-type="fn" rid="mul15"><sup><italic>a</italic></sup></xref></th>
                    <th align="left" valign="top">Element</th>
                    <th align="left" valign="top">Description</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Nlm_FloatHi</td>
                    <td align="left" valign="top">expect_value</td>
                    <td align="left" valign="top">Expect value cutoff</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Int2</td>
                    <td align="left" valign="top">wordsize</td>
                    <td align="left" valign="top">Number of letters used in making words for lookup table</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Int2</td>
                    <td align="left" valign="top">penalty</td>
                    <td align="left" valign="top">Mismatch penalty (only blastn and MegaBLAST)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Int2</td>
                    <td align="left" valign="top">reward</td>
                    <td align="left" valign="top">Match reward (only blastn and MegaBLAST)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">CharPtr</td>
                    <td align="left" valign="top">matrix</td>
                    <td align="left" valign="top">Matrix used for comparison (not blastn or MegaBLAST)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Int4</td>
                    <td align="left" valign="top">gap_open</td>
                    <td align="left" valign="top">Cost for gap existence</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Int4</td>
                    <td align="left" valign="top">gap_extend</td>
                    <td align="left" valign="top">Cost to extend a gap one more letter (including first)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">CharPtr</td>
                    <td align="left" valign="top">filter_string</td>
                    <td align="left" valign="top">Filtering options (e.g., L, mL)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Int4</td>
                    <td align="left" valign="top">hitlist_size</td>
                    <td align="left" valign="top">Number of database sequences to save hits for</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Int2</td>
                    <td align="left" valign="top">number_of_cpus</td>
                    <td align="left" valign="top">Number of CPUs to use</td>
                  </tr>
                </table>
                <table-wrap-foot>
                  <fn symbol="a" id="mul15">
                    <p>The types are given in terms of those in the NCBI Toolkit. Nlm_FloatHi is a double, Int2/Int4 are 2- or 4-byte integers, and CharPtr is just char*.</p>
                  </fn>
                </table-wrap-foot>
              </table-wrap>
              
                <preformat>
/* Get default options. */
options = BLASTOptionNew(blast_program, TRUE);
if (options == NULL)
    return 5;
    
options-&gt;expect_value = (Nlm_FloatHi) myargs [3].floatvalue;

/* Perform the actual search. */
seqalign = BioseqBlastEngine(query_bsp, blast_program, blast_database, options,

    NULL, NULL, NULL);
    
/* Do something with the SeqAlign... */
MySeqAlignPrint(seqalign, outfp);

/* clean up. */
seqalign = SeqAlignSetFree(seqalign);
options = BLASTOptionDelete(options);

sep = SeqEntryFree(sep);
FileClose(infp);
FileClose(outfp);</preformat>
              
              <p>The main steps here are:
                <list list-type="arabic">
                  <list-item id="bid.639">
                    <label>1</label>
                    <p>BLASTOptionNew allocates a BLASTOptionBlk with default values for the specified program (e.g., blastp); the Boolean argument specifies a gapped search.</p>
                  </list-item>
                  <list-item id="bid.640">
                    <label>2</label>
                    <p>The expect_value member of the BLASTOptionBlk is changed to a non-default value specified on the command-line.</p>
                  </list-item>
                  <list-item id="bid.641">
                    <label>3</label>
                    <p>BioseqBlastEngine performs the search of the BioseqPtr (query_bsp). The BioseqPtr could have been obtained from the BLAST databases, Entrez, or from FASTA using the function call FastaToSeqEntry.</p>
                  </list-item>
                </list>
              </p>
              <p>The BLASTOptionBlk structure contains a large number of members. The most useful ones and a brief description for each are listed in <xref ref-type="table" rid="bid.642">Table 2</xref>.</p>
            </sec>
            <sec id="bid.643">
              <title>Appendix 4. A function to print a view of a SeqAlign: MySeqAlignPrint.</title>
              
                <preformat>
#define BUFFER_LEN 50

/*
    Print a report on hits with start/stop. Zero-offset is used.
*/
static void MySeqAlignPrint(SeqAlignPtr seqalign, FILE *outfp)
{
    Char query_id_buf[BUFFER_LEN+1], target_id_buf[BUFFER_LEN+1];
    SeqIdPtr query_id, target_id;
    while (seqalign)
    {
        query_id = SeqAlignId(seqalign, 0);
        SeqIdWrite(query_id, query_id_buf, PRINTID_FASTA_LONG, BUFFER_LEN);
        
        target_id = SeqAlignId(seqalign, 1);
        SeqIdWrite(target_id, target_id_buf, PRINTID_FASTA_LONG, BUFFER_LEN);
        
        fprintf(outfp, "%s:%ld-%ld\t%s:%ld-%ld\n",
            query_id_buf, (long) SeqAlignStart(seqalign, 0), (long) SeqAlignStop(seqalign, 0),
            target_id_buf, (long) SeqAlignStart(seqalign, 1), (long)     SeqAlignStop(seqalign, 1));
        seqalign = seqalign-&gt;next;
    }
return;
}</preformat>
              
              <p>Note that:
              <list list-type="arabic">
                <list-item id="bid.644">
                  <label>1</label>
                  <p>SeqAlignId gets the sequence identifier for the zero-th identifier (zero offset). This is actually a C structure.</p>
                </list-item>
                <list-item id="bid.645">
                  <label>2</label>
                  <p>SeqIdWrite formats the information in query_id into a FASTA identifier (e.g., gi|129295) and places it into query_buf.</p>
                </list-item>
                <list-item id="bid.646">
                  <label>3</label>
                  <p>SeqAlignStart and SeqAlignStop return the start values of the zero-th and first sequences (or first and second).</p>
                </list-item>
              </list>
              </p>
              <p>All of this is done by high-level function calls, and it is not necessary to write low-level function calls to parse the ASN.1.</p>
            </sec>
 </body>
 <back>
              <ref-list>
              <title>References</title>
                <ref id="bid.648">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Altschul</surname>
                        <given-names>SF</given-names>
                      </name>
                      <name>
                        <surname>Gish</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Miller</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Myers</surname>
                        <given-names>EW</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                    </person-group>
                    <article-title>Basic Local Alignment Search Tool</article-title>
                    <source>J Mol Biol</source>
                    <volume>215</volume>
                    <fpage>403</fpage>
                    <lpage>410</lpage>
                    <year>1990</year>
                    <pub-id pub-id-type="pmid">2231712</pub-id>
                  </citation>
                </ref>
                <ref id="bid.649">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Altschul</surname>
                        <given-names>SF</given-names>
                      </name>
                      <name>
                        <surname>Madden</surname>
                        <given-names>TL</given-names>
                      </name>
                      <name>
                        <surname>Schaffer</surname>
                        <given-names>AA</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Miller</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                    </person-group>
                    <article-title>Gapped BLAST and PSI-BLAST: a new generation of protein database search programs</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>25</volume>
                    <fpage>3389</fpage>
                    <lpage>3402</lpage>
                    <year>1997</year>
                    <pub-id pub-id-type="pmid">9254694</pub-id>
                    <pub-id pub-id-type="other">146917</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.650" book-part-type="chapter" book-part-number="17">
          <book-part-meta>
            <title-group>
              <title>LinkOut: Linking to External Resources from Entrez Databases</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Kwan</surname>
                  <given-names>Kathy</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch17d1"/>
            <abstract>
              <title>Summary</title>
              <p>The power of linking is one of the most important developments that the World Wide Web offers to the scientific and research community. By providing a convenient and effective means for sharing ideas, linking helps scientists and scholars promote their research goals.</p>
              <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/">LinkOut</ext-link> is a powerful linking feature of the Entrez search and retrieval system (<xref ref-type="other" rid="bid.588">Chapter 15</xref>). It is designed to provide Entrez users with links from database records to a wide variety of relevant online resources, including full-text publications, biological databases, consumer health information, and research tools. (See <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/linksexample.html">Sample Links</ext-link> for examples of LinkOut resources.) The goal of LinkOut is to facilitate access to relevant online resources beyond the Entrez system to extend, clarify, or supplement information found in the Entrez databases. By branching out to relevant resources on the Web, LinkOut expands on the theme of Entrez as an information discovery system.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.651">
              <title>How Is LinkOut Represented in Entrez?</title>
              <p>Any Entrez database record, e.g., a nucleotide sequence, a taxonomic record, a protein structure, or a PubMed abstract, can be linked to Web resources external to NCBI via LinkOut. The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ncbi.nlm.nih.gov/entrez/linkout/">LinkOut homepage</ext-link> contains up-to-date documentation about LinkOut. The page that lists LinkOut resources associated with a given record can be accessed in a variety of ways (Figures <xref ref-type="fig" rid="bid.652">1</xref>&ndash;<xref ref-type="fig" rid="bid.654">3</xref>). In the case of PubMed, the full-text article and other resources related to the abstract being viewed may be accessed directly by icon buttons above the abstract (<xref ref-type="fig" rid="bid.653">Figure 2</xref>). The LinkOut display can be customized by using the LinkOut <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/pmhelp.html#LinkOutPreferences">preferences</ext-link> in the <xref ref-type="other" rid="bid.1270">Cubby</xref>.<fig id="bid.652"><label>1</label><caption><title>The <named-content content-type="book-spec-emph">LinkOut</named-content> display can be accessed by selecting <named-content content-type="book-spec-emph">LinkOut</named-content> from a PubMed record (<italic>top panel</italic>), from other Entrez databases (<italic>middle panel</italic>), or from the <named-content content-type="book-spec-emph">Display</named-content> list (<italic>lower panel</italic>).</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch17f1a" mime-subtype="gif"/><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch17f1b" mime-subtype="gif"/><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch17f1c" mime-subtype="gif"/></fig><fig id="bid.653"><label>2</label><caption><title>From PubMed, the links to the full text of research articles are also managed by LinkOut and can be accessed through an icon from PubMed Abstracts, highlighted here in <italic>purple</italic>, as well as from the associated list of LinkOut resources in <xref ref-type="fig" rid="bid.654">Figure 3</xref>.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch17f2" mime-subtype="gif"/></fig><fig id="bid.654"><label>3</label><caption><title>Links to external resources are listed in the LinkOut <named-content content-type="book-spec-emph">Display</named-content> of an Entrez record.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch17f3" mime-subtype="gif"/></fig></p>
            </sec>
            <sec id="bid.655">
              <title>How Does LinkOut Work?</title>
              <sec id="bid.656">
                <title>Design Overview</title>
                <p>The URLs to LinkOut resources are all provided by the person or organization that owns or created the resource. Links can be provided in any URL syntax, and providers of links may choose as much or as little access to their resource as they wish. Providers use one format to submit links to all Entrez databases.</p>
                <p>LinkOut is in itself an Entrez database that holds all the linking information to external resources. The separation of the data records (e.g., PubMed abstracts) from the external linking information (e.g., URLs to journal articles on a publisher's Web site) enables both the external link providers and NCBI to manage linking in a flexible manner. This means that if links to external resources change, such as in the case of a Web site redesign, this will not affect the Entrez database records, and linking information can be updated as frequently as necessary.</p>
                <p>The LinkOut database contains information on the relationship between a link and all of the applicable unique Entrez ID numbers (UIDs). By taking advantage of the interconnectivity among Entrez nodes, the linking information is presented seamlessly and efficiently.</p>
              </sec>
              <sec id="bid.657">
                <title>LinkOut DTD and XML Files</title>
                <boxed-text id="bid.658" position="margin">
                  <label>1</label>
                  <title>Example of an identity file.</title>
                  
                    <preformat>
&lt;?xml version="1.0"?&gt;
&lt;!DOCTYPE Provider PUBLIC "-//NLM//DTD LinkOut 1.0//EN" "LinkOut.dtd"&gt;
&lt;Provider&gt;
    &lt;ProviderId&gt;777&lt;/ProviderId&gt;
    &lt;Name&gt;WebDatabase Co.&lt;/Name&gt;
    &lt;NameAbbr&gt;WebDB&lt;/NameAbbr&gt;
    &lt;SubjectType&gt;gene/protein/disease-specific&lt;/SubjectType&gt;
    &lt;Attribute&gt;registration required&lt;/Attribute&gt;
    &lt;Url&gt;http://www.webdatabase.com&lt;/Url&gt;
    &lt;IconUrl&gt;http://www.webdatabase.com/images/webdb.gif&lt;/IconUrl&gt;
    &lt;Brief&gt;On-line publisher of biomedical databases and other Web resources&lt;/Brief&gt;
&lt;/Provider&gt;</preformat>
                  
                  <sec id="bid.659">
                    <title>Identity File Elements</title>
                    <p><bold>Provider</bold>: root element of the identity file.</p>
                    <p><bold>ProviderId</bold>: unique ID assigned by NCBI.</p>
                    <p><bold>Name</bold>: full name of the resource provider.</p>
                    <p><bold>NameAbbr</bold>: short, one-word name of the provider assigned by NCBI. May only include alpha and numeric characters; spaces and special characters such as hyphens are not allowed.</p>
                    <p><bold>SubjectType, Attribute</bold>: descriptions of the resources and relationship of the provider to the resources listed in the resource file. SubjectType and Attribute values appearing in the identity file will apply to all of the resources listed by that provider.</p>
                    <p><bold>Url</bold>: URL of the provider's Web site, used in the LinkOut Providers list in <xref ref-type="other" rid="bid.1270">Cubby</xref>.</p>
                    <p><bold>IconUrl</bold>: logo of the provider, used to display the link from Entrez records.</p>
                    <p><bold>Brief</bold>: short (up to 256 characters) description of the provider.</p>
                  </sec>
                </boxed-text>
                <boxed-text id="bid.660" position="margin">
                  <label>2</label>
                  <title>Example of a resource file.</title>
                  
                    <preformat>
&lt;?xml version="1.0"?&gt;
&lt;!DOCTYPE LinkSet PUBLIC "-//NLM//DTD LinkOut 1.0//EN" "LinkOut.dtd"
[&lt;!ENTITY icon.url "http://www.webdatabase.com/images/webdb.gif"&gt;
&lt;!ENTITY base.url "http://www.webdatabase.com/cgi-bin/elegans?"&gt;]&gt;
&lt;LinkSet&gt;
    &lt;Link&gt;
        &lt;LinkId&gt;1&lt;/LinkId&gt;
        &lt;ProviderId&gt;777&lt;/ProviderId&gt;
        &lt;IconUrl&gt;&amp;icon.url;&lt;/IconUrl&gt;
        &lt;ObjectSelector&gt;
            &lt;Database&gt;Nucleotide&lt;/Database&gt;
            &lt;ObjectList&gt;
                &lt;Query&gt;Caenorhabditis elegans [orgn]&lt;/Query&gt;
            &lt;/ObjectList&gt;
        &lt;/ObjectSelector&gt;
        &lt;ObjectUrl&gt;
            &lt;Base&gt;&amp;base.url;&lt;/Base&gt;
            &lt;Rule&gt;an_lookup=&amp;lo.pacc;&lt;/Rule&gt;
            &lt;UrlName&gt;Caenorhabditis elegans&lt;/UrlName&gt;
            &lt;SubjectType&gt;organism-specific&lt;/SubjectType&gt;
        &lt;/ObjectUrl&gt;
    &lt;/Link&gt;
&lt;/LinkSet&gt;</preformat>
                  
                  <sec id="bid.661">
                    <title>Resource File Elements</title>
                    <p><bold>LinkSet</bold>: the root element of the resource file.</p>
                    <p><bold>Link</bold>: an element that describes a specific set of resources grouped together by access characteristics or for convenience. A resource file may have multiple Link elements.</p>
                    <p><bold>LinkId</bold>: an identifier assigned by the provider for its own reference. It may be any character string. Each Link should have a unique LinkId within each LinkSet or file.</p>
                    <p><bold>ProviderId</bold>: the identifier number assigned to the provider by NCBI and listed in the providerinfo.xml file.</p>
                    <p><bold>IconUrl</bold>: the URL to the icon that will be displayed on the PubMed Citation and Abstract Displays.</p>
                    <p><bold>ObjectSelector</bold>: an element containing sub-elements in which providers will specify which Entrez records are being linked from by a &lt;Link&gt; element.</p>
                    <p><bold>Database</bold>: a sub-element of &lt;ObjectSelector&gt;. Databases available for linking include: PubMed, Protein, Nucleotide, Genome, Structure, PopSet, Taxonomy, and OMIM.</p>
                    <p><bold>ObjectList</bold>: a sub-element of &lt;ObjectSelector&gt; containing either the &lt;Query&gt; or &lt;ObjectID&gt; that specifies the Entrez records from which the resource will be linked.</p>
                    <p><bold>Query</bold>: a sub-element of &lt;ObjectList&gt; that contains any valid Entrez search, used to select the Entrez records being linked from.</p>
                    <p><bold>ObjId</bold>: a sub-element of &lt;ObjectList&gt; that contains an Entrez record unique identifier (UID).</p>
                    <p><bold>ObjUrl</bold>: an element that contains the necessary information for the Entrez system to construct URLs to link to the provider's resources.</p>
                    <p><bold>Base</bold>: a sub-element of &lt;ObjUrl&gt; that is the base of the URL for the provider's records.</p>
                    <p><bold>Rule</bold>: a sub-element of &lt;ObjUrl&gt; that specifies the construction of the remainder of the URL, based upon the specification of systems where the resources reside.</p>
                    <p><bold>UrlName</bold>: a short (two- or three-word) description of the link. This may be used when multiple links are available for a single Entrez record. This may also be used if the allowed terms in SubjectType and Attribute cannot meet the need of a provider.</p>
                    <p><bold>SubjectType, Attribute</bold>: sub-elements of &lt;ObjectUrl&gt;, used to describe the subject(s) of the provider's resources, barriers (if any) to using the resources, and relationship of the provider to the resources listed in the resource file. The SubjectType(s) and Attribute(s) will be applied to the resources provided within a &lt;Link&gt;.</p>
                  </sec>
                </boxed-text>
                <p>LinkOut information is submitted in XML, defined by the LinkOut Document Type Definition (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/LinkOut.dtd">DTD</ext-link>).</p>
                <p>Linking information is supplied in two elements: the Provider element, which specifies information about a link provider; and the LinkSet element, which describes information about the link. Each element should be submitted to NCBI in a separate file. Identity files contain the Provider element, and Resource files contain the LinkSet element.</p>
                <p>The Identity file is always called providerinfo.xml. It describes the identity of a provider, including an ID (ProviderId) and an abbreviated name (NameAbbr) assigned by NCBI, the provider's name, and other general information about the provider. There should be only one providerinfo.xml file for each provider (see <xref ref-type="boxed-text" rid="bid.658">Box 1</xref> for an example of an Identity file).</p>
                <p>The Resource file, which contains the LinkSet information, specifies a set of Entrez records with a valid Entrez query, a specific rule to build the link to an external resource, and description of the resource using the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/subjecttypes.html">SubjectType, Attribute, and UrlName</ext-link> fields. There is no standard for naming the LinkSet files, except that they must use the .xml extension. There may be any number of LinkSet files associated with a ProviderId. (See <xref ref-type="boxed-text" rid="bid.660">Box 2</xref> for an example of a Resource file.)</p>
                <p>Terms used in SubjectType and Attribute elements are controlled to describe LinkOut resources in a systematic manner. This is because resources are presented to users by SubjectType on the LinkOut display page (and within the Cubby system), making it easier to browse and access available resources. Attributes can be used to describe the nature of a LinkOut resource (i.e., whether the resource requires a subscription or registration to access the content). A short text string may be used in the UrlName element to describe a resource. UrlName is typically used when the allowed SubjectType and Attribute terms cannot describe the resource adequately or when multiple links are available from one provider for a single Entrez record.</p>
              </sec>
              <sec id="bid.662">
                <title>XML File Processing and Indexing</title>
                <p>All links from Entrez are generated on a daily basis so that new or modified Entrez records will have accurate LinkOut resources connected to them. Once a day, all LinkOut files are parsed according to the LinkOut DTD, and the LinkOut database is rebuilt, relating the Entrez UIDs with the link information specified in the LinkSet XML files.</p>
                <p>A LinkOut record consists of a link and the associated information, including its URL and all descriptive terms (SubjectType, Attribute, and UrlName) pertaining to the link. The Entrez UIDs applicable to the link are indexed to associate this information to the corresponding Entrez databases. As explained in <xref ref-type="other" rid="bid.588">Chapter 15</xref>, LinkOut information is interconnected with all related Entrez records.</p>
              </sec>
              <sec id="bid.663">
                <title>LinkOut Filters</title>
                <p>To facilitate search and retrieval of LinkOut resources, there are a number of filters in the LinkOut-enabled Entrez databases. These filters, although not part of the LinkOut database, use the result generated in the LinkOut indexing process.</p>
                <p>The filters are all prefixed with <bold>lo</bold>. Filters are available for all allowable SubjectType and Attribute terms and the NameAbbr of a provider. Some examples include:
                  <list list-type="bullet">
                    <list-item id="bid.664">
                      <p><bold>loprov</bold>  LinkOut Provider</p>
                    </list-item>
                    <list-item id="bid.665">
                      <p><bold>loattr</bold>  LinkOut Attribute</p>
                    </list-item>
                    <list-item id="bid.666">
                      <p><bold>losubj</bold>  LinkOut SubjectType</p>
                    </list-item>
                    <list-item id="bid.667">
                      <p><bold>loall</bold>  all LinkOut resources in an Entrez database</p>
                    </list-item>
                  </list>
                </p>
                <p>To use these filters to retrieve a set of Entrez records with LinkOut resources, the filter term can be entered as a search. For example, in PubMed, searching <preformat>&ldquo;loattrfull text online&rdquo;[Filter]</preformat> will retrieve all records with LinkOut resources that have an attribute &ldquo;full-text online&rdquo;. The <named-content content-type="book-spec-emph">Preview/Index</named-content> section in PubMed can also be used to select LinkOut filters by first selecting <named-content content-type="book-spec-emph">Filter</named-content> and then typing in &ldquo;lo&rdquo; and selecting <named-content content-type="book-spec-emph">Index</named-content> to browse through all of the filters related to LinkOut.</p>
              </sec>
            </sec>
            <sec id="bid.668">
              <title>Guides for LinkOut Providers</title>
              <p>LinkOut resources should be directly relevant to specific subjects of the Entrez records to which they will be linked, thus providing further research resources for Entrez users. The information and its delivery system should be of high quality and must not, through typographic or factual errors, omissions, or other flaws or inconsistencies, mislead, hinder, or frustrate the research efforts of Entrez users. The resources should be easy to use and navigate. Resources from professional societies, government agencies, educational institutions, or individuals and organizations that have received grants from major funding organizations are preferred.</p>
              <p>Participation in LinkOut is voluntary. Providers need to submit two types of files to describe the LinkOut resources, Identity files and Resource files (see Boxes <xref ref-type="boxed-text" rid="bid.658">1</xref> and <xref ref-type="boxed-text" rid="bid.660">2</xref>). These files include the necessary information for the Entrez system to construct an appropriate URL to access specific resources.</p>
              <p>A list of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/linkoutfaq.html">Frequently Asked Questions</ext-link> is available to address questions that potential LinkOut providers may have. Current lists of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/journals/active_providers.html">LinkOut providers</ext-link> can also be browsed.</p>
              <sec id="bid.669">
                <title>Submission Procedures</title>
                <p><bold>Step 1. Initial Contact.</bold> A prospective provider can write to <email>linkout@ncbi.nlm.nih.gov</email>, indicating interest in creating links from Entrez records to the providers' Web-accessible online resources. Please include the name, email address, and phone number of an individual who will act as a designated contact. The email should also include a LinkOut Identity file (providerinfo.xml) based on the specifications described above.</p>
                <p><bold>Step 2. File Evaluation.</bold> NCBI staff will evaluate the resources before a ProviderId and NameAbbr are assigned. NCBI will also provide assistance with setting up an appropriate Resource file to describe the LinkOut resources.</p>
                <p><bold>Step 3. File Submission.</bold> An FTP account will be assigned to a provider for submission. Files must have been validated by the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/validate.html">LinkOut Validation</ext-link> utility before uploading. Providers may transfer new versions of current files or add new Resource files at their own discretion. Providers are responsible for keeping their files current and valid. Links in Entrez databases are regenerated each day based on the files in each provider's directory; therefore, providers must delete obsolete files from their holdings directory.</p>
                <p><bold>Step 4. Representation in Entrez.</bold> Once a provider's LinkOut files are processed, the resources described in the file will be available in the LinkOut display of a relevant Entrez record as described in the above section, <italic>How Is LinkOut Represented in Entrez?</italic>. In PubMed, publishers of the abstract can choose to display a &ldquo;button&rdquo; on the Abstract and Citation displays of the PubMed record by adding the parameter &ldquo;holding=NameAbbr&rdquo; to the basic PubMed URL. For example, to activate the icon of WebDatabase Co, the URL would be provided as: <preformat>http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?holding=WebDB</preformat>
			Furthermore, multiple NameAbbr parameters may be used in a URL to activate more than one icon. For example, to display icons for both WebDB and MyDB, the following URL should be provided: <preformat>http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?holding=WebDB,MyDB</preformat>
			A provider's icon can be activated if the provider is selected from the LinkOut Preferences in Cubby.</p>
                <p>All access restrictions will still apply. For example, if access to a database is limited by the user IP address, access will only be allowed via computers within an approved IP range; if access is password protected, the password must still be entered.</p>
              </sec>
              <sec id="bid.670">
                <title>Detailed Guides</title>
                <p>Interested parties can consult the following guides for more details:
                  <list list-type="bullet">
                    <list-item id="bid.671">
                      <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/nonbiblinkout.html">LinkOut and Non-Bibliographic Resources</ext-link> &ndash; written for providers of general LinkOut resources in all Entrez databases.</p>
                    </list-item>
                    <list-item id="bid.672">
                      <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/publinkout.html">LinkOut and Publisher Holdings</ext-link> &ndash; written for publishers and others that provide full-text links to PubMed records.</p>
                    </list-item>
                    <list-item id="bid.673">
                      <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/liblinkout.html">LinkOut and Library Holdings</ext-link> &ndash; written for libraries to indicate information about their electronic full-text subscription and print holdings.</p>
                    </list-item>
                  </list>
                </p>
              </sec>
              <sec id="bid.674">
                <title>Auxiliary Tools</title>
                <p>A number of tools are available to facilitate participation in LinkOut:
                  <list list-type="bullet">
                    <list-item id="bid.675">
                      <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/lbsub-i.html">Library LinkOut Files Submission Utility</ext-link> utility developed for Libraries to generate and manage their LinkOut files. Libraries simply check off their electronic journal collections from a list of journals that participate in LinkOut. With this utility, libraries can provide correct holdings information easily, and staff do not have to construct LinkOut files by hand.</p>
                    </list-item>
                    <list-item id="bid.676">
                      <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/validate.html">LinkOut File Validation</ext-link> utility to be used by providers of links to parse their LinkOut files, ensuring the accuracy of the files before submission. Besides validating the file syntax against the LinkOut DTD, this tool will ensure that only allowable SubjectType and Attribute terms have been provided.</p>
                    </list-item>
                  </list>
                </p>
                <p>Additional tools are being developed to assist other groups of providers. Interested parties can subscribe to announcement lists described in <italic>Communicating with LinkOut Providers</italic> (next section) to be informed of new developments.</p>
              </sec>
            </sec>
            <sec id="bid.677">
              <title>Communicating with LinkOut Providers</title>
              <p>LinkOut resource providers can communicate with NCBI's LinkOut team in a number of ways. Users and providers can write to <email>linkout@ncbi.nlm.nih.gov</email> to ask questions about LinkOut. There are also three announcement lists where development related to LinkOut will be communicated to link providers:
              <list list-type="arabic">
                <list-item id="bid.678">
                  <label>1</label>
                  <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mailman/listinfo/linkout-news">Linkout-news</ext-link> is for general announcements on LinkOut.</p>
                </list-item>
                <list-item id="bid.679">
                  <label>2</label>
                  <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mailman/listinfo/library-linkout">Library-linkout</ext-link> is for announcements on development related to library LinkOut participants.</p>
                </list-item>
                <list-item id="bid.680">
                  <label>3</label>
                  <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mailman/listinfo/tax-linkout">Tax-linkout</ext-link> is for announcements relevant to linking to taxonomic resources on the Web.</p>
                </list-item>
              </list>
              </p>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.681" book-part-type="chapter" book-part-number="18">
          <book-part-meta>
            <title-group>
              <title>The Reference Sequence (RefSeq) Project</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Pruitt</surname>
                  <given-names>Kim D</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Tatusova</surname>
                  <given-names>Tatiana</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Ostell</surname>
                  <given-names>James M</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch18d1"/>
            <abstract>
              <title>Summary</title>
              <p>The Reference Sequence (RefSeq) database provides a biologically non-redundant collection of DNA, RNA, and protein sequences. Each RefSeq represents a single, naturally occurring molecule from a particular organism. RefSeqs are frequently based on GenBank records but differ in that each RefSeq is a synthesis of information, not a piece of a primary research data in itself. Similar to a review article in the literature, a RefSeq is an interpretation by a particular group at a particular time. RefSeqs can be retrieved in several different ways: by searching the Entrez Nucleotide or Protein database, by BLAST searching, by FTP, or through links from other NCBI resources.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.682">
              <title>Introduction</title>
              <p>The goal of the NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/RefSeq/">RefSeq</ext-link> project is to provide the best non-redundant and comprehensive collection of naturally occurring DNA, RNA, and protein molecules for major organisms. The collection explicitly links the nucleotide and protein sequences. Ideally, all molecule types will be available for each well-studied organism, but the current database collection pragmatically includes those molecules and organisms that are most readily identified, either by collaborators or NCBI processing of public sequence data. Thus, different amounts of information are available for different organisms at any given time. Intermediate records are provided for some organisms when the genome sequence is not yet finished.</p>
              <p>As a non-redundant collection of sequences, RefSeq offers a significant advantage during database searches or when identifying sequences (whether by BLAST, text, or accession queries, or inclusion in a local custom database). RefSeq represents an objective and experimentally verifiable definition of non-redundancy by providing one example of each natural biological molecule per organism. For some organisms, the RefSeq collection includes alternatively spliced transcripts that share some identical exons, or identical proteins expressed from these alternatively spliced transcripts, or close paralogs or homologs. RefSeq provides the substrate for a variety of objective conclusions about non-redundancy based on clustering identical sequences or families of related sequences.</p>
              <p>RefSeq is unique in providing a large, multi-species, curated sequence database that explicitly links chromosome, transcript, and protein information; it establishes a baseline for integrating sequence, genetic, expression, and functional information into a single, consistent framework. RefSeq is substantially based on GenBank sequence records (<xref ref-type="other" rid="bid.2">Chapter 1</xref>). Hereafter, the shorter term &ldquo;GenBank&rdquo; is used to indicate the full set of archival sequence data that is submitted to, and redistributed by, the three collaborating databases <xref ref-type="other" rid="bid.1272">DDBJ</xref>, <xref ref-type="other" rid="bid.1281">EMBL</xref>, and <xref ref-type="other" rid="bid.1299">GenBank</xref>. RefSeq records include attribution to the original sequence data; however, RefSeq differs from GenBank in the same way that a review article differs from the relevant collection of primary research articles on the same subject. RefSeq represents a synthesis and summary of information by a person or group based on the primary information that was gathered by others. Other organizing principles or standards of judgment are possible, which is why such a work should be attributed to the synthesizing &ldquo;editors&rdquo;. RefSeq does not exclude other syntheses based on the same primary information. But, similar to a review article, it allows the comparison of many different observations taken over time. RefSeq also has the advantage of organizing a large body of diverse data into a single consistent framework with a uniform set of conventions and standards.</p>
              <p>Note that although based upon GenBank, RefSeq is distinct from GenBank. GenBank represents the sequence and annotations that are supplied by the original authors and is never changed by others. GenBank remains the primary sequence repository. RefSeq is one of many possible &ldquo;review articles&rdquo; based on that essential archive.</p>
              <p>The RefSeq collection establishes a consistent baseline and clear model of the central dogma. RefSeq standards support genome annotation, gene characterization, mutation analysis, expression studies, and polymorphism discovery. The RefSeq collection offers advantages in:
                <list list-type="bullet">
                  <list-item id="bid.683">
                    <p> Facile identification of sequence standards for genomes, transcripts, or proteins</p>
                  </list-item>
                  <list-item id="bid.684">
                    <p> Genome annotation</p>
                  </list-item>
                  <list-item id="bid.685">
                    <p> Comparative genomics</p>
                  </list-item>
                  <list-item id="bid.686">
                    <p> Reduction of redundancy in clustering approaches</p>
                  </list-item>
                  <list-item id="bid.688">
                    <p> Providing a foundation for unambiguous association of functional information (supports navigation)</p>
                  </list-item>
                </list>
              </p>
            </sec>
            <sec id="bid.690">
              <title>Database Content: Background</title>
              <table-wrap id="bid.691">
                <label>1</label>
                <caption>
                  <title>The RefSeq Accession number format and molecule types.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Accession prefix</th>
                    <th align="left" valign="top">Molecule type</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NC_</td>
                    <td align="left" valign="top">Complete genomic molecule</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NG_</td>
                    <td align="left" valign="top">Genomic region</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NM_</td>
                    <td align="left" valign="top">mRNA</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NP_</td>
                    <td align="left" valign="top">Protein</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NR_</td>
                    <td align="left" valign="top">RNA</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NT_<xref ref-type="fn" rid="mul16"><sup><italic>a</italic></sup></xref></td>
                    <td align="left" valign="top">Genomic contig</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NW_<xref ref-type="fn" rid="mul16"><sup><italic>a</italic></sup></xref></td>
                    <td align="left" valign="top">Genomic contig (WGS<xref ref-type="fn" rid="mul17"><sup><italic>b</italic></sup></xref>)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">XM_<xref ref-type="fn" rid="mul16"><sup><italic>a</italic></sup></xref></td>
                    <td align="left" valign="top">mRNA</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">XP_<xref ref-type="fn" rid="mul16"><sup><italic>a</italic></sup></xref></td>
                    <td align="left" valign="top">Protein</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">XR_<xref ref-type="fn" rid="mul16"><sup><italic>a</italic></sup></xref></td>
                    <td align="left" valign="top">RNA</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">NZ_<xref ref-type="fn" rid="mul18"><sup><italic>c</italic></sup></xref></td>
                    <td align="left" valign="top">Genomic (WGS)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">ZP_<xref ref-type="fn" rid="mul16"><sup><italic>a</italic></sup></xref></td>
                    <td align="left" valign="top">Protein, on NZ_</td>
                  </tr>
                </table>
                <table-wrap-foot>
                  <fn symbol="a" id="mul16">
                    <p>Computed.</p>
                  </fn>
                  <fn symbol="b" id="mul17">
                    <p>Assembly of Whole Genome Shotgun (WGS) sequence data.</p>
                  </fn>
                  <fn symbol="c" id="mul18">
                    <p>An ordered collection of WGS for a genome.</p>
                  </fn>
                </table-wrap-foot>
              </table-wrap>
              <p>The current RefSeq collection includes sequences from over 2,000 distinct taxonomic identifiers, ranging from viruses to bacteria to eukaryotes. It represents chromosomes, organelles, plasmids, transcripts, and over 700,000 proteins. Every sequence is assigned a stable Accession number, version number, and gi number; older versions remain available if a sequence is updated over time. <xref ref-type="table" rid="bid.691">Table 1</xref> indicates the types of sequence molecules and the corresponding RefSeq Accession number formats. Also see the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/RefSeq/">RefSeq</ext-link> Web site.</p>
              <sec id="bid.692">
                <title>Updates</title>
                <p>RefSeq updates are provided on a daily basis, as needed. New records may be added to the collection, or existing records may be updated to reflect sequence or annotation changes or as part of a bulk update from a collaborator. New and updated records are made available in Entrez as soon as possible. The FTP site also provides daily update information (see below).</p>
              </sec>
              <sec id="bid.693">
                <title>Flat File Format and Annotated Features</title>
                <p>RefSeq records appear similar in format to the GenBank records from which they are derived; distinguishing features include a unique accession prefix that includes an underscore, which is never present in a GenBank accession (<xref ref-type="table" rid="bid.691">Table 1</xref>), and a COMMENT field that indicates the RefSeq status and the source of the sequence information (<xref ref-type="fig" rid="bid.694">Figure 1</xref>). Some RefSeq records include feature annotation that is not present on the underlying GenBank record; new annotation is provided by computation as well as manual curation. For example, nucleotide variation features are computed using the data available in the dbSNP database (<xref ref-type="other" rid="bid.1143">Chapter 5</xref>), and protein domains are computed using NCBI's Conserved Domain Database (<xref ref-type="other" rid="bid.94">Chapter 3</xref>). Further nucleotide and protein features, publications, and comments may be added by collaborating groups or NCBI staff (see <xref ref-type="boxed-text" rid="bid.720">Box 2</xref>).<fig id="bid.694"><label>1</label><caption><title>Features of a RefSeq record.</title><p><italic>Dialog balloons</italic> indicate distinguishing features including the Accession number format, the COMMENT text indicating the status and source of the sequence information, the (optional) gene description (Summary text), the (optional) description of transcript variants, and links to additional information. Links may be provided to external sources of information (e.g., OMIM and the ExPASy Enzyme Commission Web site) as well as to other NCBI pages (e.g., the protein record and LocusLink), as seen in this example.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch18f1" mime-subtype="gif"/></fig></p>
              </sec>
            </sec>
            <sec id="bid.695">
              <title>Assembling and Maintaining the RefSeq Collection</title>
              <sec id="bid.696">
                <title>Summary</title>
                <table-wrap id="bid.697">
                  <label>2</label>
                  <caption>
                    <title>RefSeq status codes.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Code</th>
                      <th align="left" valign="top">Description</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">GENOME ANNOTATION</td>
                      <td align="left" valign="top">The RefSeq record is provided via automated processing and is not subject to individual review or revision between builds.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">INFERRED</td>
                      <td align="left" valign="top">The RefSeq record has been predicted by genome sequence analysis, but it is not yet supported by experimental evidence. The record may be partially supported by homology data.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">PREDICTED</td>
                      <td align="left" valign="top">The RefSeq record has not yet been subject to individual review, and some aspect of the RefSeq record is predicted.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">PROVISIONAL</td>
                      <td align="left" valign="top">The RefSeq record has not yet been subject to individual review. The initial sequence-to-gene name associations have been established by outside collaborators or NCBI staff.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">REVIEWED</td>
                      <td align="left" valign="top">The RefSeq record has been reviewed by NCBI staff or by a collaborator. The NCBI review process includes assessing available sequence data and the literature. Some RefSeq records may incorporate expanded sequence and annotation information.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">VALIDATED</td>
                      <td align="left" valign="top">The RefSeq record has undergone an initial review to provide the preferred sequence standard. The record has not yet been subject to final review, at which time additional functional information may be provided.</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">WGS</td>
                      <td align="left" valign="top">The RefSeq record is provided to represent a collection of whole genome shotgun sequences. These records are not subject to individual review or revisions between genome updates.</td>
                    </tr>
                  </table>
                </table-wrap>
                <p>The RefSeq database is compiled through several processes including collaboration, extraction from GenBank, and computation. Each molecule is annotated as accurately as possible with the correct organism name, correct gene symbol for that organism, and informative protein names whenever possible. Collaborations with authoritative groups outside of NCBI provide a variety of information ranging from curated sequence data, nomenclature, feature annotations, and links to external organism-specific resources. If a collaboration has not been established, then NCBI staff assembles the data from GenBank. Each record has a tag indicating the level of curation it has received (<xref ref-type="table" rid="bid.697">Table 2</xref>), and the collaborating group is attributed. Thus, the RefSeq record may be an essentially unchanged validated copy of the original GenBank record, or it may include corrected or additional information that has been added by collaborators or experts at NCBI.</p>
                <p>In cases when a molecule is represented by multiple sequences for an organism in GenBank, an effort is made by NCBI staff to select the &ldquo;best&rdquo; sequence to instantiate as a RefSeq. The goal is to avoid known mutations, sequencing errors, cloning artifacts, and erroneous annotation; should an existing RefSeq be identified with a problem of this type, it is corrected. Sequences are validated to confirm that the genomic sequence corresponding to an annotated mRNA feature matches the mRNA sequence record, and that coding region features really can be translated into the corresponding protein sequence.</p>
                <p>RefSeq records may be added to the collection, or existing records may be updated, on a daily basis. Separate working groups that use distinct process pipelines compile the RefSeq collection for different organisms. RefSeq records are provided by collaboration and the following three pipelines:
                  <list list-type="bullet">
                    <list-item id="bid.698">
                      <p> Genome Annotation pipeline</p>
                    </list-item>
                    <list-item id="bid.699">
                      <p> LocusLink pipeline</p>
                    </list-item>
                    <list-item id="bid.700">
                      <p> Entrez Genomes pipeline</p>
                    </list-item>
                  </list>
                </p>
              </sec>
              <sec id="bid.701">
                <title>Collaboration</title>
                <table-wrap id="bid.702">
                  <label>3</label>
                  <caption>
                    <title>Selected examples of collaborator-contributed RefSeq records.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Organism</th>
                      <th align="left" valign="top">Collaborator</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <italic>Saccharomyces cerevisiae</italic>
                      </td>
                      <td align="left" valign="top">Saccharomyces Genome Database (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://genome-www.stanford.edu/Saccharomyces/">SGD</ext-link>)</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <italic>Arabidopsis thaliana</italic>
                      </td>
                      <td align="left" valign="top">The Institute for Genomic Research (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.tigr.org/">TIGR</ext-link>)</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <italic>Pseudomonas aeruginosa</italic>
                      </td>
                      <td align="left" valign="top"><italic>Pseudomonas aeruginosa</italic> Community Annotation Project (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pseudomonas.com/">PseudoCAP</ext-link>)</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <italic>Drosophila melanogaster</italic>
                      </td>
                      <td align="left" valign="top">
								Drosophila Sequencing Consortium, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://flybase.bio.indiana.edu/">FlyBase</ext-link>, and NCBI staff</td>
                    </tr>
                  </table>
                </table-wrap>
                <p>We welcome collaborations whenever authoritative groups outside NCBI are willing to provide sequences, nomenclature, annotations, or links to phenotypic or organism-specific resources. For some species, the RefSeq collection is curated entirely by a collaborating authoritative group that provides both the sequences and annotations. Other species may be provided via varying levels of collaborative efforts. For example, a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/viradvisors.html">Viral Genome Advisory group</ext-link> has been established to support curation of the viral RefSeq collection. For some species, some information is provided by the external authoritative source, and some information is provided by NCBI. This process may be automated in that data are periodically downloaded and subject to validation to detect errors and apply annotation in a more uniform way; NCBI does not otherwise carry out additional curation to add annotation or make sequence changes to records supplied by collaborators. As the collaborating group supplies updates, changes are reflected in the RefSeq collection for that organism. Collaborator-supplied records have a REVIEWED status, and the collaborating group is identified. <xref ref-type="table" rid="bid.702">Table 3</xref> lists examples of curated and annotated sequence records provided wholly or in part through collaboration.</p>
              </sec>
              <sec id="bid.703">
                <title>Genome Annotation Pipeline</title>
                <p>NCBI is providing annotation of genomic sequence data for some genomes including some microbial species, human, and mouse. These pipelines are automated and yield genomic, transcript, and protein RefSeq records (records that are provided vary by organism). Data are refreshed periodically, and for the eukaryotic annotation pipeline, records are not subject to individual incremental updates or manual curation (see <xref ref-type="table" rid="bid.691">Table 1</xref>; see <xref ref-type="other" rid="bid.1440">Chapter 14</xref> for more information on the eukaryotic genome annotation pipeline).</p>
              </sec>
              <sec id="bid.704">
                <title>LocusLink Pipeline</title>
                <boxed-text id="bid.705" position="margin">
                  <label>1</label>
                  <title>Approaches to associate sequence data with LocusLink loci.</title>
                  <p>Collaboration with authoritative groups including:
                    <list list-type="bullet">
                      <list-item id="bid.706">
                        <p> FlyBase</p>
                      </list-item>
                      <list-item id="bid.707">
                        <p> Human Gene Nomenclature Committee (HGNC)</p>
                      </list-item>
                      <list-item id="bid.708">
                        <p> Mouse Genome Informatics (MGI)</p>
                      </list-item>
                      <list-item id="bid.709">
                        <p> OMIM</p>
                      </list-item>
                      <list-item id="bid.710">
                        <p> RATMAP</p>
                      </list-item>
                      <list-item id="bid.711">
                        <p> Rat Genome Database (RGD)</p>
                      </list-item>
                      <list-item id="bid.712">
                        <p> WormBase</p>
                      </list-item>
                      <list-item id="bid.713">
                        <p> Zebrafish Information Network (ZFIN)</p>
                      </list-item>
                      <list-item id="bid.1771">
                        <p> Gene Family Authorities</p>
                      </list-item>
                    </list>
                  </p>
                  <p>In-House Processing:
                  <list list-type="bullet">
                    <list-item id="bid.1743">
                      <p> Extraction from GenBank</p>
                    </list-item>
                    <list-item id="bid.1744">
                      <p> Genome Annotation pipeline</p>
                    </list-item>
                    <list-item id="bid.1745">
                      <p> Homology analysis</p>
                    </list-item>
                    <list-item id="bid.1746">
                      <p> In-house curation</p>
                    </list-item>
                    <list-item id="bid.1747">
                      <p> UniGene analysis</p>
                    </list-item>
                  </list>
                  </p>
                </boxed-text>
                <boxed-text id="bid.720" position="margin">
                  <label>2</label>
                  <title>RefSeq curation: error correction and data addition.</title>
                  <p>The RefSeq curation process may result in correction of errors as well as provision of additional sequence information and feature annotation, as indicated below.</p>
                  <sec id="bid.721">
                    <title>Curation of Sequences</title>
                    <sec id="bid.722">
                      <title>
                        <bold>Error Corrections</bold>
                      </title>
                      
                        <list list-type="arabic">
                          <list-item id="bid.723">
                            <label>1</label>
                            <p>Remove chimeric sequences.</p>
                          </list-item>
                          <list-item id="bid.724">
                            <label>2</label>
                            <p>Remove vector and linker sequence.</p>
                          </list-item>
                          <list-item id="bid.725">
                            <label>3</label>
                            <p>Remove sequences annotated with incorrect organism information.</p>
                          </list-item>
                          <list-item id="bid.726">
                            <label>4</label>
                            <p>Resolve sequence-to-locus mis-associations.</p>
                          </list-item>
                          <list-item id="bid.727">
                            <label>5</label>
                            <p>Correct apparent sequence errors as identified through sequence analysis and personal communication; an attempt is made to reconcile sequence differences to finished genomic sequence.</p>
                          </list-item>
                          <list-item id="bid.728">
                            <label>6</label>
                            <p>Modify the extent of the original GenBank CDS annotation, as determined through sequence analysis (including homology considerations), literature review, and personal communication.</p>
                          </list-item>
                        </list>
                    
                    </sec>
                    <sec id="bid.729">
                      <title>
                        <bold>Provide Additional Information</bold>
                      </title>
                      
                        <list list-type="arabic">
                          <list-item id="bid.730">
                            <label>1</label>
                            <p>Extend UTRs based on publicly available sequence data or literature review.</p>
                          </list-item>
                          <list-item id="bid.731">
                            <label>2</label>
                            <p>Provide RefSeq with full-length CDS from a series of overlapping partial GenBank sequences.</p>
                          </list-item>
                          <list-item id="bid.732">
                            <label>3</label>
                            <p>Provide splice variants and corresponding protein variants.</p>
                          </list-item>
                        </list>
                      
                    </sec>
                  </sec>
                  <sec id="bid.733">
                    <title>Curation of Annotations (information added)</title>
                    <sec id="bid.734">
                      <title>
                        <bold>Nucleotide and Protein</bold>
                      </title>
                      
                        <list list-type="arabic">
                          <list-item id="bid.735">
                            <label>1</label>
                            <p>Alias symbols and names</p>
                          </list-item>
                          <list-item id="bid.736">
                            <label>2</label>
                            <p>Brief description of transcript and protein variants</p>
                          </list-item>
                          <list-item id="bid.737">
                            <label>3</label>
                            <p>Publications</p>
                          </list-item>
                          <list-item id="bid.738">
                            <label>4</label>
                            <p>Summary of gene, protein function</p>
                          </list-item>
                        </list>
                      
                    </sec>
                    <sec id="bid.739">
                      <title>
                        <bold>Protein Only</bold>
                      </title>
                      
                        <list list-type="arabic">
                          <list-item id="bid.740">
                            <label>1</label>
                            <p>Alternate translation start sites</p>
                          </list-item>
                          <list-item id="bid.741">
                            <label>2</label>
                            <p>Enzyme Commission number</p>
                          </list-item>
                          <list-item id="bid.742">
                            <label>3</label>
                            <p>Mature peptide products</p>
                          </list-item>
                          <list-item id="bid.743">
                            <label>4</label>
                            <p>Non-AUG translation start sites</p>
                          </list-item>
                          <list-item id="bid.744">
                            <label>5</label>
                            <p>Protein domains</p>
                          </list-item>
                          <list-item id="bid.745">
                            <label>6</label>
                            <p>Selenocysteine proteins</p>
                          </list-item>
                        </list>
                     
                    </sec>
                    <sec id="bid.746">
                      <title>
                        <bold>Nucleotide Only</bold>
                      </title>
                      
                        <list list-type="arabic">
                          <list-item id="bid.747">
                            <label>1</label>
                            <p>Indication of transcript completeness, as known</p>
                          </list-item>
                          <list-item id="bid.748">
                            <label>2</label>
                            <p>Poly(A) signal, site</p>
                          </list-item>
                          <list-item id="bid.749">
                            <label>3</label>
                            <p>RNA editing</p>
                          </list-item>
                          <list-item id="bid.750">
                            <label>4</label>
                            <p>Variation features</p>
                          </list-item>
                        </list>
                      
                    </sec>
                  </sec>
                </boxed-text>
                <p>LocusLink supports the generation of RefSeq records for human, mouse, rat, fly, zebrafish, and cow; the RefSeq process for these organisms takes advantage of the descriptive information available in LocusLink. Multiple collaborations support the collection of this descriptive information (<xref ref-type="boxed-text" rid="bid.705">Box 1</xref>; see also <xref ref-type="other" rid="bid.788">Chapter 19</xref>). The <italic>Drosophila</italic> RefSeq collection is provided in collaboration with FlyBase, and numerous additional collaborations have occurred for individual genes and gene families, primarily for the human RefSeq collection.</p>
                <p>This data set consists of genomic regions, transcripts, and protein. Records representing genomic regions are provided primarily to support more comprehensive genome-level annotation and may represent gene clusters or single genes or pseudogenes. Records are annotated with the level of curation it has undergone; records may have an INFERRED, PREDICTED, PROVISIONAL, REVIEWED, or VALIDATED status (<xref ref-type="table" rid="bid.697">Table 2</xref>).</p>
                <p>Sequences in LocusLink records enter RefSeq by a mixture of computational analysis process, collaboration, and in-house curation. As illustrated in <xref ref-type="fig" rid="bid.714">Figure 2</xref>, generation of the initial RefSeq record is dependent on first identifying a representative sequence for a gene in the LocusLink database. New genes are identified and added to the collection by collaborators or in-house processing that mines information available in UniGene, the Genome Annotation pipeline, and new GenBank submissions (<xref ref-type="boxed-text" rid="bid.705">Box 1</xref>). Once sequence data are associated with a LocusID, it is used as a query sequence for automated BLAST analysis. The longest mRNA that meets our stringent matching criteria is taken from the BLAST results; this is either a cDNA sequence or a mRNA feature annotated on a DNA sequence record. The identified sequence is subsequently used to provide the initial RefSeq record. This stage of the process includes detection of conflicts and problems including:
                  <list list-type="bullet">
                    <list-item id="bid.715">
                      <p> Sequence-to-locus association conflicts (e.g., close paralogs)</p>
                    </list-item>
                    <list-item id="bid.716">
                      <p> Vector contamination</p>
                    </list-item>
                    <list-item id="bid.717">
                      <p> The GenBank record used is a genomic record with &ldquo;not experimental&rdquo; annotation</p>
                    </list-item>
                    <list-item id="bid.718">
                      <p> The protein sequence is annotated as partial</p>
                    </list-item>
                    <list-item id="bid.719">
                      <p> The RefSeq transcript would be suspiciously long</p>
                    </list-item>
                  </list>
                  <fig id="bid.714">
                    <label>2</label>
                    <caption>
                      <title>LocusLink-supported RefSeq pipeline.</title>
                      <p>Once a gene is defined and associated with sufficient sequence information in LocusLink, it can be pushed into the RefSeq pipeline. New genes are added to LocusLink by collaborators and in-house review. The RefSeq process is initiated by an automated BLAST step, which uses the sequence information in LocusLink as a query against GenBank to identify the longest mRNA for each locus. This sequence is represented as a provisional, predicted, or inferred RefSeq record. Subsequent review and curation may result in a sequence or annotation update (as described in <xref ref-type="boxed-text" rid="bid.720">Box 2</xref>). Records are refreshed if the underlying GenBank Accession number is updated or if an official nomenclature or cytogenetic map location is updated in LocusLink.</p>
                    </caption>
                    <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch18f2" mime-subtype="gif"/>
                  </fig>
                </p>
                <p>Conflicts and problems of these types must be resolved before the RefSeq can become public. Records are subject to validation to correct annotation errors and to provide annotation in a more consistent format. LocusLink descriptive information including official nomenclature, alias symbols, alternate descriptive names, map location, and additional citations are applied to the records. Records at this stage have a PROVISIONAL, PREDICTED, or INFERRED status.</p>
                <p>These initial records then enter into the in-house review pipeline, where additional manual curation may be applied. The review process prioritizes known genes, gene families, and problem cases that are identified through user input or analysis. The curation process includes analysis of a suite of precomputed BLAST results, literature review, review of available database and Web resources (both in-house and external), and collaboration to curate the nucleotide and protein sequence data and to apply additional annotation and descriptive information to both the RefSeq sequence record and to the LocusLink database. <xref ref-type="boxed-text" rid="bid.720">Box 2</xref> lists more detailed information concerning the type of errors corrected and information added by the manual curation process. Sequence review is carried out primarily by NCBI staff, but some sequences and annotations are provided by collaboration including a portion of the <italic>Drosophila melanogaster</italic> RefSeq collection. The curation process also provides additional sequence records to represent splice variants when sufficient information about their full-length composition is available. Records that have undergone manual curation, either by NCBI staff or a collaborating group, have a VALIDATED or REVIEWED status. Note that for many genes, intermediate levels of manual curation occur to correct sequence mis-associations, to use a more optimal GenBank record, and to provide additional data to LocusLink before full review of a RefSeq sequence record.</p>
                <p>Additional ongoing review is applied to identified problem sets; for example, periodic analysis may be carried out to identify sequences that include repeats, that have poor-quality splice sites, are very short or very long, or that are extremely similar to sequences associated with a different LocusID. Review of problem sets may result in discontinuing a RefSeq, LocusID, or both. A RefSeq is suppressed if it is found to represent a transcribed repeat element, to be derived from the wrong organism (i.e., the GenBank sequence it was based on does not have accurate organism annotation), or to not represent a &ldquo;gene&rdquo;. An Entrez query will still retrieve a suppressed record, with a disclaimer appearing on the query result document summary (<xref ref-type="fig" rid="bid.751">Figure 3a</xref>), but the record is not included in the BLAST databases nor in the calculation of related sequences or the BLink display (precomputed protein BLAST results). If a RefSeq is found to be redundant with another public RefSeq, then one is retained and the other becomes a secondary Accession number (<xref ref-type="fig" rid="bid.751">Figure 3b</xref>). If the sequences were associated with two different LocusIDs, then the LocusIDs are merged so that in LocusLink, a query with either of the original LocusIDs will retrieve the remaining single record.<fig id="bid.751"><label>3</label><caption><title>How to recognize suppressed and redundant RefSeq records.</title><p>(<italic>a</italic>) A standard text statement is included on the Entrez document summary for suppressed RefSeq records (<italic>red arrow</italic>). (<italic>b</italic>) If redundant RefSeq records are merged, then both Accession numbers appear on the flat file ACCESSION line (<italic>green arrow</italic>). The first Accession number listed is the primary identifier, and all other listed accessions are &ldquo;secondary&rdquo; Accession numbers.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch18f3" mime-subtype="gif"/></fig></p>
                <p>Input from the research community is welcome to further improve the quality of this data set. Interested parties are invited to contact us by sending an email to the NCBI Help Desk (<email>info@ncbi.nlm.nih.gov</email>). See also the RefSeq Web site at http://www.ncbi.nih.gov/RefSeq/.</p>
              </sec>
              <sec id="bid.752">
                <title>Entrez Genomes Pipeline</title>
                <table-wrap id="bid.1772">
                  <label>4</label>
                  <caption>
                    <title>Selected Entrez Genomes Resources.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Web Page</th>
                      <th align="left" valign="top">Web Site</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Entrez Genomes Home</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Genome">http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Genome</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Prominent Organisms</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/org.html">http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/org.html</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Microbes</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/micr.html">http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/micr.html</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Viral Genomes</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/viruses.html">http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/viruses.html</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Organelle Genome Resources</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/organelles.html">http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/organelles.html</ext-link>
                      </td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">Plant Genomes Central</td>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/PlantList.html">http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/PlantList.html</ext-link>
                      </td>
                    </tr>
                  </table>
                </table-wrap>
                <p>The Entrez
				<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Genome">Genomes</ext-link> database represents a collection of complete, or nearly complete, genomes and chromosomes. It is divided into six large taxonomic groups: Archaea, Eubacteria, Eukaryotae, Viroids, Viruses, and Plasmids. See the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/org.html">Prominent Organisms</ext-link> page for a complete list of organisms included in this database; this list is automatically updated daily. Entrez Genomes RefSeq records include genomic, transcript, and protein records and are provided by collaboration, in-house automatic processing, and in-house curation. The Entrez Genomes Web site includes custom displays, analysis, and tools for some genomes (see <xref ref-type="table" rid="bid.1772">Table 4</xref>).</p>
                <p>In general, these RefSeq records undergo an initial automated validation step before being released. The resulting record is a copy of a GenBank entry, but validation may make some corrections and provides more consistent feature annotation. Records provided via collaboration have a status of REVIEWED and are attributed to the collaborating group. Records provided by in-house processing have a PROVISIONAL, VALIDATED, or REVIEWED status.</p>
                <p>Entrez Genomes record processing falls into four primary categories: chromosomes, microbial genomes, small complete genomes, and viruses.</p>
                <sec id="bid.753">
                  <title>
                    <bold>Chromosomes</bold>
                  </title>
                  <p>RefSeq records in this category are usually submitted directly to Entrez Genomes as a complete chromosome sequence representing an assembly of individual clones that are themselves available in GenBank. Examples of this type of RefSeq include <italic>Arabidopsis thaliana, Saccharomyces cerevisiae,</italic> and <italic>Caenorhabditis elegans.</italic> In addition, RefSeq records may be available for some genomes that are not yet fully complete but for which complete sequence is available for individual chromosomes; for example, a RefSeq chromosome 1 record is provided for <italic>Leishmania major</italic>.</p>
                  <p>These records are curated by the organism-specific collaborating group and undergo NCBI validation before being released. The validation process checks for logical conflicts in the annotation (which are reported back to the submitting group) and makes small modifications to format the submission as a RefSeq record. This category of RefSeq record is displayed using the NCBI Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>) in addition to being available in the main Entrez sequence databases. An organism-specific genome BLAST service is provided for these genomic, transcript, and protein records.</p>
                </sec>
                <sec id="bid.754">
                  <title>
                    <bold>Microbial Genomes</bold>
                  </title>
                  <table-wrap id="bid.1773">
                    <label>5</label>
                    <caption>
                      <title>Selected Examples of REVIEWED Microbial Genomes.</title>
                    </caption>
                    <table>
                      <tr>
                        <th align="left" valign="top">RefSeq</th>
                        <th align="left" valign="top">Organism</th>
                      </tr>
                      <tr>
                        <td align="left" valign="top">
                          <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov:80/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=NC_003450">NC_003450</ext-link>
                        </td>
                        <td align="left" valign="top">
                          <italic>Corynebacterium glutamincum</italic>
                        </td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">
                          <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov:80/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=NC_001802">NC_001802</ext-link>
                        </td>
                        <td align="left" valign="top">
                          <italic>Human immunodeficiency virus 1</italic>
                        </td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">
                          <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov:80/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=NC_004193">NC_004193</ext-link>
                        </td>
                        <td align="left" valign="top">
                          <italic>Oceanobacillus iheyensis</italic>
                        </td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">
                          <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov:80/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=NC_002689">NC_002689</ext-link>
                        </td>
                        <td align="left" valign="top">
                          <italic>Thermoplasma volcanium</italic>
                        </td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">
                          <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov:80/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=NC_004718">NC_004718</ext-link>
                        </td>
                        <td align="left" valign="top">
                          <italic>SARS coronavirus</italic>
                        </td>
                      </tr>
                    </table>
                  </table-wrap>
                  <p>Microbial complete genomes are submitted to GenBank, but because of the GenBank/EMBL/DDBJ collaboration agreement, which limits the size to 350 kb, they are made available in GenBank as a series of Accession numbers. RefSeq does not need to adhere to this upper limit, and therefore complete genome sequences are available as RefSeq records in Entrez Genomes.</p>
                  <p>These records are subject to additional automated validation and computational analysis; the computational analysis results are then manually reviewed, resulting in more complete annotation of the RefSeq record (see <xref ref-type="table" rid="bid.1773">Table 5</xref>).</p>
                </sec>
                <sec id="bid.755">
                  <title>
                    <bold>Small Complete Genomes</bold>
                  </title>
                  <p>Smaller complete genomic sequences, including organelles, plasmids, and viruses, are based on single GenBank records. Automatic processing scans GenBank daily for complete genome updates and new submissions; identified records are candidates for a complete genome RefSeq. These records are manually evaluated to make the final decision; if more than one genomic sequence is available for the genome, then only one is selected to become the RefSeq standard. This selection takes into account various factors including the level of annotation, strain information, and community input.</p>
                  <p>Some of these RefSeq records undergo manual curation and have a REVIEWED status. For example, viral genomes are re-annotated using GeneMarkS in collaboration with Mark Borodovsky's Bioinformatics Group at the Georgia Institute of Technology. Following <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://opal.biology.gatech.edu/GeneMark/">GeneMarkS</ext-link> analysis, viral genomes are subject to manual review by NCBI staff. Viral annotation also relies on an established <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/viradvisors.html">Viral Genome Advisors</ext-link> group. For example, the HIV-1 RefSeq was curated by NCBI Staff in collaboration with the authors of the book <italic><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="/books/bv.fcgi?call=bv.View..ShowTOC&amp;rid=rv.TOC">Retroviruses</ext-link></italic>, and NCBI curators expanded the mature protein annotation for several viruses, including the poliovirus and hepatitis C, based on observations reported in the literature that have not been included in a GenBank submission.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.756">
              <title>Access and Retrieval</title>
              <p>RefSeq records can be accessed by direct query, BLAST and FTP download, or indirectly through links provided from several NCBI resources. In addition, RefSeqs are included in some computed resources, and therefore links may be found from those pages to individual RefSeq records. For example, the RefSeq collection is included in Clusters of Orthologous Groups of proteins (COG; <xref ref-type="other" rid="bid.885">Chapter 22</xref>) analysis and in Conserved Domain Database (CDD; <xref ref-type="other" rid="bid.94">Chapter 3</xref>) analysis to identify proteins with similar domain architecture. Some links to RefSeq records are based on LocusLink associations (e.g., links from OMIM; <xref ref-type="other" rid="bid.405">Chapter 7</xref>), whereas others are based on sequence similarity. Links to RefSeq records may be found in the following resources:
                <list list-type="bullet">
                  <list-item id="bid.1662">
                    <p> BLAST results (<xref ref-type="other" rid="bid.610">Chapter 16</xref>)</p>
                  </list-item>
                  <list-item id="bid.759">
                    <p> BLink (precomputed protein blast results)</p>
                  </list-item>
                  <list-item id="bid.1663">
                    <p>CDD</p>
                  </list-item>
                  <list-item id="bid.757">
                    <p> dbSNP (<xref ref-type="other" rid="bid.1143">Chapter 5</xref>)</p>
                  </list-item>
                  <list-item id="bid.758">
                    <p> Entrez (<xref ref-type="other" rid="bid.588">Chapter 15</xref>)</p>
                  </list-item>
                  <list-item id="bid.1664">
                    <p> LocusLink (<xref ref-type="other" rid="bid.788">Chapter 19</xref>)</p>
                  </list-item>
                  <list-item id="bid.762">
                    <p> Map Viewer (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>)</p>
                  </list-item>
                  <list-item id="bid.1665">
                    <p> OMIM (<xref ref-type="other" rid="bid.405">Chapter 7</xref></p>
                  </list-item>
                  <list-item id="bid.761">
                    <p> UniGene (<xref ref-type="other" rid="bid.857">Chapter 21</xref>)</p>
                  </list-item>
                  <list-item id="bid.760">
                    <p> UniSTS</p>
                  </list-item>
                </list>
              </p>
              <p>The distinct Accession number format used for RefSeq records (<xref ref-type="table" rid="bid.691">Table 1</xref>) makes it easy to spot links to RefSeq records from these and other NCBI resources. Several approaches to access and retrieve RefSeq records are described below.</p>
              <sec id="bid.763">
                <title>Entrez Query Access</title>
                <table-wrap id="bid.765">
                  <label>6</label>
                  <caption>
                    <title>Entrez queries to retrieve sets of RefSeq records.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">Query</th>
                      <th align="left" valign="top">Accession prefix</th>
                      <th align="left" valign="top">RefSeq category retrieved</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq%5Bprop%5D">srcdb_refseq[prop]</ext-link>
                      </td>
                      <td align="left" valign="top">NC_, NG_, NT_, NW_, NZ_, NM_, NR_, XM, XR_, NP_, XP_, ZP_</td>
                      <td align="left" valign="top">All</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq_known%5Bprop%5D">srcdb_refseq_known[prop]</ext-link>
                      </td>
                      <td align="left" valign="top">NC_, NG_, NM_, NR_, NP_</td>
                      <td align="left" valign="top">REVIEWED, PROVISIONAL, PREDICTED, INFERRED, and VALIDATED</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq_reviewed%5Bprop%5D">srcdb_refseq_reviewed[prop]</ext-link>
                      </td>
                      <td align="left" valign="top">NC_, NG_, NM_, NR_, NP_</td>
                      <td align="left" valign="top">REVIEWED records</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq_validated%5Bprop%5D">srcdb_refseq_validated[prop]</ext-link>
                      </td>
                      <td align="left" valign="top">NC_,NM_,NR_,NP_</td>
                      <td align="left" valign="top">VALIDATED records</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq_provisional%5Bprop%5D">srcdb_refseq_provisional[prop]</ext-link>
                      </td>
                      <td align="left" valign="top">NC_, NG_, NM_, NR_, NP_</td>
                      <td align="left" valign="top">PROVISIONAL records</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq_predicted%5Bprop%5D">srcdb_refseq_predicted[prop]</ext-link>
                      </td>
                      <td align="left" valign="top">NM_, NR_, NP_</td>
                      <td align="left" valign="top">PREDICTED records</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq_inferred%5Bprop%5D">srcdb_refseq_inferred[prop]</ext-link>
                      </td>
                      <td align="left" valign="top">NM_,NR_,NP_</td>
                      <td align="left" valign="top">INFERRED records</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">
                        <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=PureSearch&amp;db=nucleotide&amp;details_term=srcdb_refseq_model%5Bprop%5D">srcdb_refseq_model[prop]</ext-link>
                        <xref ref-type="fn" rid="mul19"><sup>
                          <italic>a</italic>
                        </sup></xref>
                      </td>
                      <td align="left" valign="top">NT_, NW_, XM_, XR_, XP_, ZP_</td>
                      <td align="left" valign="top">Genome annotation model records</td>
                    </tr>
                  </table>
                  <table-wrap-foot>
                    <fn symbol="a" id="mul19">
                      <p>Retrieves those NT_ and NW_ records that have gene annotation.</p>
                    </fn>
                  </table-wrap-foot>
                </table-wrap>
                <p>RefSeq records can be retrieved by querying various databases in the Entrez system (<xref ref-type="other" rid="bid.588">Chapter 15</xref>). All RefSeqs can be found in the Entrez nucleotide or protein databases, whereas queries against the Entrez Genomes database retrieve a subset of the RefSeq collection. Together, Entrez Genomes and LocusLink represent the entire RefSeq collection.</p>
                <p>Queries against the nucleotide and protein databases can be restricted to the RefSeq collection by using the <named-content content-type="book-spec-emph">Limited to:</named-content> settings (<xref ref-type="fig" rid="bid.764">Figure 4</xref>) or by entering a query restriction using the property fields listed in <xref ref-type="table" rid="bid.765">Table 6</xref>. In addition, both the <named-content content-type="book-spec-emph">Limited to:</named-content> settings and property field terms can also be used to restrict the query by type of molecule (e.g., DNA <italic>versus</italic> mRNA) and other parameters. See Entrez <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query/static/help/helpdoc.html">Help</ext-link> and the RefSeq <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/RefSeq/RSfaq.html">FAQs</ext-link> for further details. RefSeq records in Entrez Genomes can be retrieved using an Accession number or text term including organism name, gene name, and protein name.<fig id="bid.764"><label>4</label><caption><title>Using Entrez <named-content content-type="book-spec-emph">Limits</named-content> to restrict a query to RefSeq.</title><p>The <italic>red arrows</italic> indicate the menu boxes used to restrict the query to a genomic or mRNA sequence or to restrict the query to the RefSeq collection.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch18f4" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.766">
                <title>LocusLink Query Access</title>
                <p>A subset of the RefSeq collection is represented in LocusLink, a gene-centered database, including human, mouse, rat, Drosophila, zebrafish, cow, and HIV-1 (<xref ref-type="other" rid="bid.788">Chapter 19</xref>); additional organisms continue to be added to LocusLink and RefSeq over time. LocusLink-supported RefSeq records include a link to the LocusLink gene report page (note the LocusID link on the gene and CDS features in <xref ref-type="fig" rid="bid.694">Figure 1</xref>). Loci with associated RefSeq records can be retrieved using the controlled query term &ldquo;has_refseq&rdquo;. LocusLink query results and gene reports indicate when a RefSeq is available, with links provided to the nucleotide and protein sequences and to related resources, including the Map Viewer and BLink (precomputed protein alignments) and Conserved Domain Database (CDD). The process of RefSeq curation also expands the data available in LocusLink by providing a range of information including:
                  <list list-type="bullet">
                    <list-item id="bid.767">
                      <p> Alternate names</p>
                    </list-item>
                    <list-item id="bid.769">
                      <p> Enzyme Committee numbers</p>
                    </list-item>
                    <list-item id="bid.771">
                      <p> Gene summaries</p>
                    </list-item>
                    <list-item id="bid.768">
                      <p> Publications</p>
                    </list-item>
                    <list-item id="bid.770">
                      <p> Related GenBank accessions</p>
                    </list-item>
                    <list-item id="bid.772">
                      <p> Transcript variant descriptions</p>
                    </list-item>
                  </list>
                </p>
              </sec>
              <sec id="bid.773">
                <title>BLAST</title>
                <p>RefSeq transcript and protein records are included in the non-redundant nucleotide and protein BLAST databases, and genomic sequences are included in the &ldquo;chromosome&rdquo; database; therefore, when a query sequence matches a RefSeq record, the hit is included in the BLAST results (see <xref ref-type="fig" rid="bid.774">Figure 5</xref>). Additional organism-specific BLAST pages provide access to specific custom databases, which vary by organism as needed. For example, the human genome BLAST page provides access to query the RefSeq contigs, transcripts, or proteins in addition to (human) sequences in the GenBank High-Throughput Genomic (HTG) division.<fig id="bid.774"><label>5</label><caption><title>RefSeq records are included in NCBI BLAST databases.</title><p>In a BLAST summary list of results, the abbreviation <italic>ref</italic> identifies records that are provided by the RefSeq collection.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch18f5" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.786">
                <title>FTP</title>
                <p>Currently, RefSeq records generated by different pipelines are available in different FTP areas [<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/refseq/">RefSeq LocusLink pipeline</ext-link>; <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/genomes/">Entrez Genomes RefSeqs</ext-link>; <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/genomes/">Models</ext-link> (Genome Annotation pipeline)]. In the future, the entire collection will be made available in a regular release cycle similar to GenBank releases. The RefSeq release will be available in the &ldquo;refseq&rdquo; FTP directory. Separate files are provided for nucleotide and protein records in a variety of formats. See the available README files for additional information.</p>
              </sec>
            </sec>
            <sec id="bid.787">
              <title>Related Reading</title>
              <list list-type="bullet">
                <list-item id="bid.1774">
                  <p>Besemer J, Lomsadze A, Borodovski M. GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions. Nucleic Acids Res 29:2607-2618; 2001.</p>
                </list-item>
                <list-item id="bid.1666">
                  <p>Blake JA, Eppig JT, Richardson JE, Davisson MT. The Mouse Genome Database (MGD): expanding genetic and genomic resources for the laboratory mouse. The Mouse Genome Database Group. Nucleic Acids Res 28:108-111; 2000.</p>
                </list-item>
                <list-item id="bid.1667">
                  <p>Boguski MS, Schuler GD. ESTablishing a human transcript map. Nat Genet 10:369-371; 1995.</p>
                </list-item>
                <list-item id="bid.1668">
                  <p>Coffin JM, Hughes SH, Varmus E. Retroviruses. Plainview (NY): Cold Spring Harbor Laboratory Press; 1997.</p>
                </list-item>
                <list-item id="bid.1669">
                  <p>FlyBase Consortium. The FlyBase database of the Drosophila Genome Projects and community literature. The FlyBase Consortium. Nucleic Acids Res 27:85-88; 1999.</p>
                </list-item>
                <list-item id="bid.1670">
                  <p>Hamosh A, Scott AF, Amberger J, Valle D, McKusick VA. Online Mendelian Inheritance in Man (OMIM). Hum Mutat 15:57-61; 2000.</p>
                </list-item>
                <list-item id="bid.1775">
                  <p>Lukashin A, Borodovski M. GeneMark.hmm new solutions for gene finding. Nucleic Acids Res 26:1107-1115; 1998.</p>
                </list-item>
                <list-item id="bid.1671">
                  <p>Marchler-Bauer A, Panchenko AR, Shoemaker BA, Thiesse PA, Geer LY, Bryant SH. CDD: a database of conserved domain alignments with links to domain three-dimensional structure. Nucleic Acids Res 30:281-283; 2002.</p>
                </list-item>
                <list-item id="bid.1672">
                  <p>Pruitt KD, Katz KS, Sicotte H, Maglott DR. Introducing RefSeq and LocusLink: curated human genome resources at the NCBI. Trends Genet 16:44-47; 2000.</p>
                </list-item>
                <list-item id="bid.1673">
                  <p>Tatusova TA, Karsch-Mizrachi I, Ostell JA. Complete genomes in WWW Entrez: data representation and analysis. Bioinformatics 15:536-543; 1999.</p>
                </list-item>
                <list-item id="bid.1674">
                  <p>Twigger S, Lu J, Shimoyama M, Chen D, Pasko D, Long H, Ginster J, Chen CF, Nigam R, Kwitek A, Eppig J, Maltais L, Maglott D, Schuler G, Jacob H, Tonellato PJ. Rat Genome Database (RGD): mapping disease into the genome. Nucleic Acids Res 30:125-128; 2002.</p>
                </list-item>
                <list-item id="bid.1675">
                  <p>Westerfield M, Doerry E, Kirkpatrick AE, Douglas SA. Zebrafish informatics and the ZFIN database. Methods Cell Biol 60:339-355; 1999.</p>
                </list-item>
                <list-item id="bid.1676">
                  <p>White JA, McAlpine PJ, Antonarakis S, Cann H, Eppig JT, Frazer K, Frezal J, Lancet D, Nahmias J, Pearson P, Peters J, Scott A, Scott H, Spurr N, Talbot C Jr, Povey S. Guidelines for human gene nomenclature (1997). HUGO Nomenclature Committee. Genomics 45:468-471; 1997.</p>
                </list-item>
              </list>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.788" book-part-type="chapter" book-part-number="19">
          <book-part-meta>
            <title-group>
              <title>LocusLink: A Directory of Genes</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Maglott</surname>
                  <given-names>Donna</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>12</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch19d1"/>
            <abstract>
              <title>Summary</title>
              <p>LocusLink organizes information from collaborating public databases and from other groups within NCBI to provide a locus-centered view of genomic information from human, mouse, rat, zebrafish, <italic>Drosophila melanogaster</italic>, and human immunodeficiency virus 1 (HIV-1). Each LocusLink record (one for each locus) consists of a collection of links to more information about a locus. The intent, therefore, is not to copy all information about a locus into one record; the intent is to provide a sufficient description of each linked item so that LocusLink can be searched and to provide enough connections for users to find related resources. As part of this process, a unique integer is assigned to each locus. This identifier can then be used by other resources to connect back to LocusLink.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.789">
              <title>Overview</title>
              <p>LocusLink serves as a hub of information for <xref ref-type="other" rid="bid.1332">loci</xref> from several model organisms: human, mouse, rat, fruit fly, and human immunodeficiency virus 1 (HIV-1). On a regular basis, &ldquo;source&rdquo; databases are checked for novel information about genes or other loci. If the record already exists in LocusLink, the novel information is added. Otherwise, a new record is created. Because many of the contributing databases are curated and because the LocusLink records are reviewed by NCBI staff, LocusLink can be considered a curated resource.</p>
              <p>LocusLink gathers information from databases both within and external to NCBI. These are scanned using a combination of automatic and manual techniques. LocusLink collaborates with organism-specific and nomenclature databases, from which a combination of sequence data and names are used to initiate new LocusLink records. LocusLink also finds new records by a weekly review of submissions to GenBank and new versions of UniGene (which are released periodically). Each new LocusLink record is assigned a unique identifying number that is tracked, a LocusID.</p>
              <p>LocusLink records are used in turn by UniGene, dbSNP, and the organism-specific and nomenclature databases. For example, a new LocusLink record created on the basis of a sequence submitted to GenBank can be the basis for a new entry in an organism-specific database. This new defining sequence of the LocusLink record is &ldquo;<xref ref-type="other" rid="bid.1246">BLAST</xref>ed&rdquo; against sequences from the same or other genomes, which identifies other <xref ref-type="other" rid="bid.1299">GenBank</xref> records for the same gene or its homologs. When a related sequence is identified in a genome within the scope of LocusLink and that sequence has not yet been included in the appropriate genome-specific database, LocusLink makes the connection and reports it for others to use.</p>
              <p>LocusLink also feeds new sequences into the Reference Sequence (RefSeq) project. The prerequisite for this process is that the sequence encodes a complete protein. For more details about how eukaryotic RefSeq mRNA and protein records are initiated, see <xref ref-type="other" rid="bid.681">Chapter 18</xref> in this Handbook.</p>
              <p>Although historically many of the GenBank sequences used to initiate LocusLink records represented characterized genes sequenced by individual research labs, an increasing number of records in LocusLink are generated as a part of NCBI's genome annotation pipeline (see <xref ref-type="other" rid="bid.1440">Chapter 14</xref>). The IDs assigned to these records are termed &ldquo;interim IDs&rdquo; and are not tracked within the LocusLink database until they have been reviewed or a preponderance of evidence confirms that the gene prediction is real.</p>
              <p>In summary, no matter whether the data are stored as a result of curation or computation, the central function of LocusLink is to establish an accurate connection between the defining sequence for a locus and other descriptors for that locus. With such a connection in place, it is possible to:
                <list list-type="bullet">
                  <list-item id="bid.790">
                    <p> Establish a RefSeq for that locus.</p>
                  </list-item>
                  <list-item id="bid.791">
                    <p> Identify or validate putative <xref ref-type="other" rid="bid.1362">ortholog</xref>s (genes with a common ancestor) based on a combination of sequence similarity and conserved <xref ref-type="other" rid="bid.1414">synteny</xref> (a stretch of chromosome in which the gene order is conserved across species).</p>
                  </list-item>
                  <list-item id="bid.792">
                    <p> Support the NCBI annotation pipeline based on mRNA sequence alignment.</p>
                  </list-item>
                </list>
              </p>
              <p>As will be discussed in more detail in the following sections, other NCBI resources make use of the LocusLink LocusID-to-sequence connection to provide appropriate nomenclature and other identifiers for sequences within their scope [see also <xref ref-type="other" rid="bid.1143">Chapter 5</xref> (<!--<glossref>-->dbSNP<!--</glossref>-->), <xref ref-type="other" rid="bid.857">Chapter 21</xref> (UniGene and HomoloGene), and <xref ref-type="other" rid="bid.1562">Chapter 20</xref> (Map Viewer)].</p>
            </sec>
            <sec id="bid.793">
              <title>How to Query LocusLink</title>
              <table-wrap id="bid.794">
                <label>1</label>
                <caption>
                  <title>LocusLink query terms: field restrictions.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Field</th>
                    <th align="left" valign="top">Meaning</th>
                    <th align="left" valign="top">Example of search term<xref ref-type="fn" rid="mul20"><sup><italic>a</italic></sup></xref></th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">[chr]</td>
                    <td align="left" valign="top">Chromosome number</td>
                    <td align="left" valign="top">21[chr]</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">[loc]</td>
                    <td align="left" valign="top">LocusLink ID</td>
                    <td align="left" valign="top">4292[loc]</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">[mim]</td>
                    <td align="left" valign="top">OMIM number</td>
                    <td align="left" valign="top">300200[mim]</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">[sym]</td>
                    <td align="left" valign="top">Gene symbol</td>
                    <td align="left" valign="top">abc&ast;[sym]</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">[pm]</td>
                    <td align="left" valign="top">PubMed ID</td>
                    <td align="left" valign="top">123456[pm]</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">[ngi]</td>
                    <td align="left" valign="top">Nucleotide gi number</td>
                    <td align="left" valign="top">223344[ngi]</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">[pgi]</td>
                    <td align="left" valign="top">Protein gi number</td>
                    <td align="left" valign="top">1234567[pgi]</td>
                  </tr>
                </table>
                <table-wrap-foot>
                  <fn symbol="a" id="mul20">
                    <p>No space is allowed between the value and the field name.</p>
                  </fn>
                </table-wrap-foot>
              </table-wrap>
              <table-wrap id="bid.795">
                <label>2</label>
                <caption>
                  <title>LocusLink query terms: controlled terms.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Term</th>
                    <th align="left" valign="top">Meaning</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">disease_known</td>
                    <td align="left" valign="top">Human locus associated with a phenotype defined by a MIM number (may be only that phenotype)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">has_homol</td>
                    <td align="left" valign="middle">Associated with a HomoloGene record</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">has_omim</td>
                    <td align="left" valign="top">Associated with an OMIM record</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">has_refseq</td>
                    <td align="left" valign="top">Associated with a RefSeq</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">has_seq</td>
                    <td align="left" valign="top">Associated with nucleotide sequence</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">has_snp</td>
                    <td align="left" valign="top">Associated with a dbSNP record</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">seq_map</td>
                    <td align="left" valign="top">Either the gene or an STS in this record has been localized to the sequence-based map available from Map Viewer</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">type_dseg</td>
                    <td align="left" valign="top">A DNA segment. May include BAC or YAC ends.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">type_gene_other</td>
                    <td align="left" valign="top">A gene that does not fall into the category type_gene_protein</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">type_gene_protein</td>
                    <td align="left" valign="top">A gene that encodes a protein product. Does not include unreviewed, putative genes based only on modeling; named genes that encode only part of a protein product, such as immunoglobulin variable, diversity, joining, or constant regions; or genes that exhibit somatic rearrangement.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">type_pheno</td>
                    <td align="left" valign="top">Characterized as a mapped phenotype only</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">type_pseudo</td>
                    <td align="left" valign="top">A pseudogene</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">type_qtl</td>
                    <td align="left" valign="top">A phenotype characterized as a QTL only</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">type_region</td>
                    <td align="left" valign="top">A region on the genome. Examples are named gene clusters and viral integration sites.</td>
                  </tr>
                </table>
              </table-wrap>
              <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/index.html">LocusLink homepage</ext-link> gives a brief introduction to the resource as well as information on new features. At the top of this and all LocusLink pages is a query bar that allows users to search not only LocusLink but also a selection of other resources within NCBI. Thus, if a query against LocusLink does not return a result, a different resource can be selected and searched without retyping the query term. Currently, these other resources include: OMIM, PubMed, Entrez Nucleotide, Entrez Protein, Human Map Viewer, UniGene, and UniSTS. The default for LocusLink queries is to search all of the available organisms, although a specific organism can be selected by using the pull-down menu in the search bar.</p>
              <p>LocusLink supports several types of queries, including gene names, GenBank Accession numbers, and other resource-specific ID numbers such as <xref ref-type="other" rid="bid.1424">UniGene cluster</xref> or <xref ref-type="other" rid="bid.1410">STS</xref> marker IDs. Queries are not case sensitive, and retrievals are based on a word index, not a phrase index. Punctuation is not removed; thus, entering the query <italic>beta actin</italic> will retrieve any record that contains the word &ldquo;beta&rdquo; and the word &ldquo;actin&rdquo; (processed as &ldquo;beta AND actin&rdquo;). <italic>&ldquo;beta actin&rdquo;</italic> will retrieve No Results, because the quotes are not removed, and look-up fails on <italic>&ldquo;beta&rdquo;</italic> or <italic>actin?</italic>. More advanced searches can limit queries to specific fields (<xref ref-type="table" rid="bid.794">Table 1</xref>) or by controlled terms (<xref ref-type="table" rid="bid.795">Table 2</xref>). Further information on options for formulating queries are documented in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/help.html">Help</ext-link> pages.</p>
              <p>A query on the LocusLink homepage returns a report page that is organized in a tabular format, as shown in <xref ref-type="fig" rid="bid.796">Figure 1</xref>. The complete record is viewed by selecting <named-content content-type="book-spec-emph">Locus ID</named-content> or <named-content content-type="book-spec-emph">More</named-content> (for interim IDs from the annotation pipeline).
			<fig id="bid.796"><label>1</label><caption><title>Example of a LocusLink query results page.</title><p><italic>LocusID</italic> field: NCBI assigns a unique, stable LocusID to each locus. Selecting the number retrieves the detailed LocusLink report page. As part of NCBI's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/guide/build.html">Genome Annotation Project</ext-link>, some LocusLink records are generated that are likely to be temporary and which, therefore, are not represented by a stable ID. Such records are indicated by <named-content content-type="book-spec-emph">More</named-content>, which can be selected to display the report page. <italic>Org</italic> field: a two-letter abbreviation for the organism in which this locus is described. The symbols (<italic>Hs</italic> for <italic>Homo sapiens</italic>, <italic>Mm</italic> for <italic>Mus musculus</italic>, and <italic>Rn</italic> for <italic>Rattus norvegicus</italic>) match the abbreviations used by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/">UniGene</ext-link>. <italic>Symbol</italic> field: the official or alias symbol. In some cases, a symbol may be used for multiple loci. <italic>Description</italic> field: the gene name. <italic>Position</italic> field: the cytogenetic (human or rat) or genetic (mouse) location. <italic>Links</italic> field: small, color-coded <italic>boxes</italic> indicate when links are available for PubMed (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_P" mime-subtype="gif"/>), OMIM (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_O" mime-subtype="gif"/>), RefSeq (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_R" mime-subtype="gif"/>), GenBank nucleotide (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_G" mime-subtype="gif"/>), Protein (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_Pr" mime-subtype="gif"/>), HomoloGene (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_H" mime-subtype="gif"/>), UniGene (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_U" mime-subtype="gif"/>), and variation data (<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19_V" mime-subtype="gif"/>). Note: neither PubMed nor GenBank links are comprehensive. Use the <named-content content-type="book-spec-emph">Related Articles</named-content> or <named-content content-type="book-spec-emph">Related Sequences</named-content> in the <named-content content-type="book-spec-emph">Links</named-content> menu to retrieve more records.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19f1" mime-subtype="gif"/></fig></p>
            </sec>
            <sec id="bid.797">
              <title>A LocusLink Report: The Details</title>
              <p>This section is organized according to the features and subdivisions seen on a LocusLink report page, from top to bottom. We will illustrate the descriptions of the features by using screen shots from the LocusLink report of BRCA2 and CDKN1A interacting protein (BCCIP; LocusID <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/LocRpt.cgi?l=56647">56647</ext-link>).
		</p>
              <p>On each LocusLink report page, the first set of links is a table of contents for the sections included in the report being displayed, which are hyperlinked to the appropriate section. <named-content content-type="book-spec-emph">Top of page</named-content> links to the top. The latter is useful if you want to gain access to the query bar to enter another query.</p>
              <sec id="bid.798">
                <title>mRNA&ndash;Genomic Alignments Diagram</title>
                <p>As part of the genome annotation process at NCBI, GenBank and RefSeq mRNAs are aligned to genomic <xref ref-type="other" rid="bid.1267">contig</xref>s to determine the exon/intron organization of genes. The results of these calculations can be displayed using NCBI's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Entrez/Genome/evvdoc.html">Evidence Viewer</ext-link>.</p>
                <p>The diagram at the top of a LocusLink report represents the intron/exon structure of the gene as determined by the Genome Annotation Pipeline (see <xref ref-type="other" rid="bid.1440">Chapter 14</xref>). For example, the annotation for BCCIP looks like this (the details may change with each reannotation of the genome): <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19df1" mime-subtype="gif"/></p>
                <p>This abbreviated output displays the longest alignment, where tick marks indicate the positions of the exons. Genes with a single exon are represented as a solid thick bar with flanking horizontal lines. The diagram is hyperlinked to the Evidence Viewer, which in turn is linked to the appropriate Entrez Nucleotide records by the Accession number of each sequence. Mousing over the individual entries in the Evidence Viewer displays a definition line from that sequence record.</p>
              </sec>
              <sec id="bid.799">
                <title>Quick Link Buttons</title>
                <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/help.html#reports">Buttons</ext-link> are provided at the top of each report page to indicate what information may be available from other resources. In the case of BCCIP, the buttons are: <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19df2" mime-subtype="gif"/></p>
                <p>where (from left to right) they represent links to: PubMed publications, UniGene, Map Viewer, dbSNP variation database, the Genome Database, Ensembl, and UCSC Human Genome Browser. The buttons are used for connecting to related resources within NCBI or to external genomic databases. Others in this category include Ace Viewer, OMIM, MGC, MGI, and RGD. More links for any record may be listed at the end of the report page in the <named-content content-type="book-spec-emph">Additional Links</named-content> section.</p>
              </sec>
              <sec id="bid.800">
                <title>Nomenclature</title>
                <p>
                  <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19df3" mime-subtype="gif"/>
                </p>
                <p>LocusLink uses symbols and gene names from official authority lists when available. If no connection to official nomenclature can be made, symbols and names are selected as available from the defining sequence record. If sequence and positional homology (synteny) suggest that a locus not named officially in one species is orthologous to a named gene in another species, the symbol from the ortholog may be included in the LocusLink record. If no symbol can be identified for a new locus, the letters LOC are prepended to the LocusID. Once an official or meaningful symbol has been identified, that LOC symbol is discontinued (because the record will still be searchable and identified by the LocusID itself).</p>
                <sec id="bid.801">
                  <title>
                    <bold>Sources of Nomenclature</bold>
                  </title>
                  <p>Information on the nomenclature authorities that collaborate with LocusLink can be viewed by following the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/LLnomen.html">Nomenclature</ext-link> link on the LocusLink homepage. Except for the human genome, data from these authorities are imported automatically and used either to correct existing LocusLink names or to create new records. For the human genome, the Human Gene Nomenclature Committee has direct access to the LocusLink database and inserts and updates records interactively.</p>
                </sec>
                <sec id="bid.802">
                  <title>
                    <bold>Information Processing</bold>
                  </title>
                  <p/>
                  <p><bold>Rules.</bold> In addition to official symbols and full names, LocusLink provides other symbols and names seen in publications and sequence records appropriate to the locus. These alternative names are not meant to be comprehensive and are usually reviewed only when the RefSeq is being reviewed.</p>
                  <p><bold>Stability and Tracking.</bold> Although LocusLink does maintain some nomenclature history, the appropriate nomenclature committee performs this function more comprehensively. For individual LocusLink records, links to those committees are provided from the LocusLink report.</p>
                  <p><bold>Update Frequency.</bold> Official nomenclature is modified when changes are available, ranging from daily (human, mouse) to weekly (zebrafish), or longer (fly, rat). Names provided by NCBI staff are available daily.</p>
                  <p><bold>Methods.</bold> Data files are imported by FTP and analyzed through a combination of shell and Perl scripts, in conjunction with relational database tables used to hold the input data. Records with no data conflicts are updated automatically. If a flag is raised because of uncertainty about the relationships among a name, a sequence, and a record identifier, then an expert reviews that record. During review, either the record is corrected or the appropriate nomenclature committee(s) is contacted to resolve the discrepancies.</p>
                </sec>
                <sec id="bid.807">
                  <title>
                    <bold>Interactions with Other Resources at NCBI</bold>
                  </title>
                  <table-wrap id="bid.808">
                    <label>3</label>
                    <caption>
                      <title>Nomenclature interactions with NCBI resources.</title>
                    </caption>
                    <table>
                      <tr>
                        <th align="left" valign="top">Resource</th>
                        <th align="left" valign="top">Keys into LocusLink</th>
                        <th align="left" valign="top">Method of matching resource to LocusID</th>
                      </tr>
                      <tr>
                        <td align="left" valign="top">HomoloGene</td>
                        <td align="left" valign="top">mRNA</td>
                        <td align="left" valign="top">Alignment</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">Map Viewer</td>
                        <td align="left" valign="top">mRNA, gene feature</td>
                        <td align="left" valign="top">Alignment</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">UniGene</td>
                        <td align="left" valign="top">mRNA, protein gi</td>
                        <td align="left" valign="top">Clustering</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top">UniSTS</td>
                        <td align="left" valign="top">mRNA</td>
                        <td align="left" valign="top">e-PCR</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top"/>
                        <td align="left" valign="top">Genomic annotation</td>
                        <td align="left" valign="top">e-PCR</td>
                      </tr>
                      <tr>
                        <td align="left" valign="top"/>
                        <td align="left" valign="top">Marker name</td>
                        <td align="left" valign="top">Publications</td>
                      </tr>
                    </table>
                  </table-wrap>
                  <p>Several NCBI databases use the nomenclature maintained by LocusLink. These names are incorporated into other databases based primarily on name&ndash;LocusID&ndash;sequence connections, i.e., sequence comparisons identify similar sequences in LocusLink, all of which have a LocusID; from this, the associated nomenclature can be extracted and applied to the original sequence from the collaborating database (<xref ref-type="table" rid="bid.808">Table 3</xref>). Nomenclature data can be extracted from files available on the FTP site.</p>
                </sec>
              </sec>
              <sec id="bid.809">
                <title>Overview</title>
                <p>This section of the LocusLink report page may include any or all of the following categories of information.<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19df4" mime-subtype="gif"/></p>
                <p>The <bold>Summary</bold> is written by RefSeq staff and/or by external contributors such as OMIM, Proteome, or Protein Reviews on the Web (<xref ref-type="other" rid="bid.1383">PROW</xref>). These summaries provide a quick synopsis of what is known about the gene, the function of its encoded protein or RNA products, disease associations, spatial and temporal distribution, and so on.</p>
                <p><bold>Locus Type</bold> indicates the type of molecule that is the defining sequence for the Locus Link record. It is selected from the options listed in <xref ref-type="boxed-text" rid="bid.810">Box 1</xref>.</p><boxed-text id="bid.810" position="margin"><label>1</label><title>List of possible locus types for LocusLink records.</title><list list-type="bullet"><list-item id="bid.811"><p> D segment</p></list-item><list-item id="bid.812"><p> RNA, ribosomal</p></list-item><list-item id="bid.813"><p> RNA, small cytoplasmic</p></list-item><list-item id="bid.814"><p> RNA, small nuclear</p></list-item><list-item id="bid.815"><p> RNA, small nucleolar</p></list-item><list-item id="bid.816"><p> RNA, transfer</p></list-item><list-item id="bid.817"><p> gene with no protein product</p></list-item><list-item id="bid.818"><p> gene with protein product, demonstrates somatic rearrangement</p></list-item><list-item id="bid.819"><p> gene with protein product, function known or inferred</p></list-item><list-item id="bid.820"><p> gene with protein product, function unknown</p></list-item><list-item id="bid.821"><p> gene, segment</p></list-item><list-item id="bid.822"><p> model, <italic>ab initio</italic></p></list-item><list-item id="bid.823"><p> model, <italic>ab initio</italic>, with EST support</p></list-item><list-item id="bid.824"><p> model, supported by EST alignments</p></list-item><list-item id="bid.825"><p> model, supported by mRNA alignments</p></list-item><list-item id="bid.826"><p> model, supported by mRNA and EST alignments</p></list-item><list-item id="bid.827"><p> phenotype only</p></list-item><list-item id="bid.828"><p> pseudogene</p></list-item><list-item id="bid.829"><p> pseudogene, transcribed</p></list-item><list-item id="bid.830"><p> quantitative trait locus (QTL)</p></list-item><list-item id="bid.831"><p> region</p></list-item><list-item id="bid.832"><p> regulatory element</p></list-item><list-item id="bid.833"><p> repetitive element</p></list-item><list-item id="bid.834"><p> unknown</p></list-item></list></boxed-text>
                <p><bold>Product</bold> lists the known names of proteins associated with the locus. This is not an exhaustive list. The intent is to support data retrieval and to document usage; therefore, only some of the names from sequence databases or published literature are reported.</p>
                <p><bold>Alternate Symbols</bold> indicates other symbols or abbreviations. These symbols may be for the gene, the protein, or a disease phenotype.</p>
                <p><bold>Alias</bold> indicates other names.</p>
              </sec>
              <sec id="bid.835">
                <title>Function</title>
                <p>Information about the function of a gene and its RNA or protein products is gathered from several sources.<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19df5" mime-subtype="gif"/></p>
                <p>For human genes, links to OMIM, if available, are given under the heading Phenotype. For all genomes, links to the published literature are provided under the heading Gene References into Function (GeneRIF). Any user can submit a reference to a paper they think is important for a locus, but beginning in February 2002, these links are also supplied by <xref ref-type="other" rid="bid.1341">MeSH</xref> indexers at the National Library of Medicine.</p>
                <p>Increasingly, the Gene Ontology&trade; (GO) vocabulary terms are incorporated from GO's <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.geneontology.org/pub/go">FTP site</ext-link>, and each term is linked to <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.godatabase.org/cgi-bin/go.cgi">AmiGO</ext-link>, the GO database, which can be browsed or searched. When the literature citation(s) supporting the GO term is available, a link to PubMed (<named-content content-type="book-spec-emph">pm</named-content>) is provided.</p>
              </sec>
              <sec id="bid.836">
                <title>Relationships</title>
                <p>This section reports other loci, and/or the proteins that they encode, that have a defined relationship to the locus being displayed.<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19df6" mime-subtype="gif"/></p>
                <p>At present, this section includes: (<italic>a</italic>) reports of how interim loci (i.e., those identified by genome annotation processes) are related to other loci, based on the Accession number of the mRNA that was used to define the intron/exon organization; and (<italic>b</italic>) reports of human&ndash;mouse homologs, which are linked to the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Homology/">Human&ndash;Mouse Homology Map</ext-link>.</p>
                <p>We plan to expand this section to include other types of pairwise relationships. For example, &ldquo;overlap&rdquo;, &ldquo;interspersed&rdquo;, and &ldquo;protein binds&rdquo; will refer to genes that overlap, that are contained in or contain another gene, and for which the encoded proteins interact, respectively. Although primarily for relationships within a species, when proteins of one species interact with those of another (e.g., infectious agents), such a relationship will be reported in this section as well.</p>
              </sec>
              <sec id="bid.837">
                <title>Map Information</title>
                <p>This section reports map data for the locus. The location listed is the same as that on the LocusLink query result page, but if there are any conflicts, additional locations may be reported, along with the source of the conflicting data and a link to that resource.<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch19df7" mime-subtype="gif"/></p>
                <p>For the human and mouse genomes, if no published map location has been identified and if the gene has been aligned to NCBI's genome assembly, the cytogenetic position is recalculated with each build. The conversion files used in the process are <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nlm.nih.gov/genomes/H_sapiens/maps/mapview/ISCN800_abc">ISCN800_abc</ext-link> for human. This location is also displayed from <xref ref-type="other" rid="bid.1336">Map Viewer</xref>. With each assembly of the genome, genes reported to be on one chromosome that appear to have better alignments on a different chromosome are identified. If marker and other data are consistent but distinct from the published map location, then the LocusLink report will be modified to conform to the evidence gained from analysis of the human genome.</p>
                <p>Genetic and physical map positions are incorporated from the published maps used in Map Viewer. Rather than report all position data for any locus in any coordinate system, links are provided to Map Viewer via the <named-content content-type="book-spec-emph">mv</named-content> link in the map section or indirectly through the marker names, which are linked to the UniSTS record.</p>
                <sec id="bid.838">
                  <title>
                    <bold>Marker Data</bold>
                  </title>
                  <p>In LocusLink, markers are defined as sequence tagged sites (STSs) associated with the locus. LocusLink reports markers either as the locus itself or as a marker that has some relationship with a gene. LocusLink does not store all of the markers available for a genome, which is the function of <xref ref-type="other" rid="bid.1425">UniSTS</xref>. Some LocusLink records contain only markers. These are considered historical records; new records are added only if a contributing database reports a marker to represent a gene.</p>
                  <sec id="bid.839">
                    <title>Information Sources</title>
                    <p>The marker data that LocusLink reports come primarily from any of the following paths: (1) a report from a genome-specific database that states that a marker is within a gene or locus; (2) for genes, a caluculation based on <xref ref-type="other" rid="bid.1284">e-PCR</xref> using mRNA as the electronic template, that the marker is within a mRNA defined to be associated with genes; and (3) for genes and in genomes for which NCBI is making an assembly and/or providing annotation e-PCR based on placement of a marker within the range beginning 2.0 kb upstream of the most 5&prime; mRNA alignment of a gene feature and ending 0.5 kb downstream of the most 3&prime; mRNA alignment.</p>
                  </sec>
                  <sec id="bid.841">
                    <title>Information Processing</title>
                    <p/>
                    <p><bold>Rules.</bold> Markers are included in a locus report if they can be assayed by PCR, i.e., if a marker is a STS, and if, according to current computation, they are detected at no more than two locations in the genome being reported.</p>
                    <p><bold>Stability and Tracking.</bold> Relationships between STS and gene records are not archived or tracked. New markers may be added to a gene report if more sequence is identified as being valid for a gene and the new marker &ldquo;detects&rdquo; that sequence by e-PCR. A marker may be removed from a report if the sequence that it detected is no longer considered valid for that gene.</p>
                    <p><bold>Update Frequency.</bold> The LocusID-to-UniSTS marker relationship that is based on e-PCR from mRNA templates is calculated daily. The LocusID-to-UniSTS marker relationship that is based on genome annotation is recalculated with each genome build.</p>
                  </sec>
                </sec>
              </sec>
              <sec id="bid.845">
                <title>Sequence Information</title>
                <p>A LocusLink record includes several categories of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/LocRpt.cgi?l=56647#refseq">sequence information</ext-link>. The RefSeq section provides information on accessions (nucleotide and protein), the GenBank accessions used to create the RefSeq, and the status (REVIEWED, PROVISIONAL, VALIDATED, PREDICTED). If alternate splice variants have been reviewed, a brief description of each is included. If there is a protein and it demonstrates significant sequence matches to domains defined in the Conserved Domain Database (<xref ref-type="other" rid="bid.1257">CDD</xref>), then the domains are listed by name, with the score for the domain match. A link is also provided to the CDD browser. To facilitate identification of related proteins and structures, <named-content content-type="book-spec-emph">BL</named-content> provides a link to <xref ref-type="other" rid="bid.1250">BLink</xref>, a resource of precalculated BLAST searches for any given protein.</p>
                <p>A second category in the RefSeq section reports sequences generated or annotated as part of an NCBI genome annotation project. The first Accession number(s) displayed is the genomic record (contig), as well as links to the genomic sequence containing the gene (<named-content content-type="book-spec-emph">gb</named-content>), a graphic sequence view (<named-content content-type="book-spec-emph">sv</named-content>), the Map Viewer display for that contig (<named-content content-type="book-spec-emph">mv</named-content>), the Evidence Viewer (<named-content content-type="book-spec-emph">ev</named-content>), and Model Maker (<named-content content-type="book-spec-emph">mm</named-content>. Please note that the <named-content content-type="book-spec-emph">mv</named-content> link in this section is contig based, whereas the <named-content content-type="book-spec-emph">mv</named-content> link in the map section is based on the locus itself. These links are provided to make it easier to: retrieve a gene-specific region of a large chromosomal sequence, rather than finding that gene in a record that may be several megabases (<named-content content-type="book-spec-emph">gb</named-content>); review the alignment evidence for a gene that has been annotated on a genomic assembly (<named-content content-type="book-spec-emph">ev</named-content>); and construct your own mRNA model(s) in a genomic region based on mRNA and ESTs that align there (<named-content content-type="book-spec-emph">mm</named-content>). The next Accession numbers listed are for sequences annotated on the genomic sequence. They may be the same as the curated RefSeq in the previous category, or they may be models, distinguished by the format of their Accession numbers (XM_000000, XP_000000, and XR_000000 refer to models). CDD domain matches and BLink links are also provided for model proteins.</p>
                <p>The <italic>GenBank Sequences</italic> section provides a list of representative nucleotide and protein Accession numbers for the locus, each Accession number of which is linked to an Entrez sequence display. Proteins listed in this section also provide a link to BLink.</p>
              </sec>
              <sec id="bid.846">
                <title>Additional Links</title>
                <p>This section provides a list of sites that may have additional information. The OMIM and UniGene links are redundant, with the button links at the top of the page, but provide the respective ID numbers that are hidden in the button links. Other links shown here, such as <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://bioinfo.weizmann.ac.il/cards/">GeneCards</ext-link> or <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.geneclinics.org/">GeneTests &middot; GeneClinics</ext-link>, a resource containing genetic testing and disease information, are generated automatically based on links from OMIM provided through LinkOut (see <xref ref-type="other" rid="bid.650">Chapter 17</xref>). Other links are added by RefSeq curators as they review a gene or after suggestions by other authorities.</p>
              </sec>
            </sec>
            <sec id="bid.847">
              <title>Maintenance and Reporting</title>
              <table-wrap id="bid.848">
                <label>4</label>
                <caption>
                  <title>LocusLink record tracking.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Maintenance</th>
                    <th align="left" valign="top">Type of locus</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Tracked</td>
                    <td align="left" valign="top">Officially named genes and pseudogenes, for both nuclear and mitochondrial genomes, whether or not the final gene product is known to be a protein or a RNA.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Tracked</td>
                    <td align="left" valign="top">Probable protein-coding genes, defined by one or more mRNAs. The function of the encoded protein is not necessarily known.</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Tracked</td>
                    <td align="left" valign="top">Mapped phenotypes, such as disease susceptibility loci or QTLs</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Tracked</td>
                    <td align="left" valign="top">Gene segments (such as coding regions for variable regions of immunoglobulin or T-cell receptors)</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Not tracked</td>
                    <td align="left" valign="top">Gene predictions from NCBI's genome annotation pipeline.</td>
                  </tr>
                </table>
              </table-wrap>
              <p>LocusLink records can be categorized by two major criteria: the type of locus being described and the tracking maintained for the record (<xref ref-type="table" rid="bid.848">Table 4</xref>). All but the interim loci (ascribed to LocusLink records that are based on gene-prediction software only) are tracked, which means that if the record is discontinued or determined to be redundant with another, queries based on the ID representing the discontinued or secondary record will return a result. For interim loci, the interim identifier is not reused, but a query based on that identifier number will not return a result if that gene model is no longer being annotated. It is anticipated that the number of interim loci will decrease as more curation is applied to the records of the genome.</p>
              <sec id="bid.849">
                <title>Retrieving Historical Data</title>
                <p>LocusLink supports retrieval of inactive records in the following ways:
                  <list list-type="arabic">
                    <list-item id="bid.850">
                      <label>1</label>
                      <p>If a non-interim record has been discontinued and the record appeared to have been created in error (i.e., could not be merged into or made secondary to another record), then the withdrawn record can be retrieved. These records are clearly noted with the term &ldquo;withdrawn&rdquo; on the query result table and the report page. The report page also includes an explanation for why the record was discontinued.</p>
                    </list-item>
                    <list-item id="bid.851">
                      <label>2</label>
                      <p>If a locus record has been made secondary to another record, a query on the secondary ID will take you to the current record.</p>
                    </list-item>
                  </list>
                </p>
                <p>Merges are reported in the FTP file LocusID_history, and the status of all queryable LocusIDs is reported in the file LL_tmpl.gz (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/refseq/LocusLink/">ftp://ftp.ncbi.nih.gov/refseq/LocusLink/</ext-link>).</p>
              </sec>
              <sec id="bid.1748">
                <title>New Records</title>
                <p>Records are added to LocusLink in several ways:
		<list list-type="bullet"><list-item id="bid.1749"><p> External resources and collaborators provide information on new, officially named genes and the sequences that define them.</p></list-item><list-item id="bid.1750"><p> New sequences are released from GenBank/DDBJ/EMBL and are identified as being from a gene not yet described in LocusLink. If sequence alignments indicate that a new sequence matches an interim locus annotated in the Annotation pipeline, then the interim locus is converted to a &ldquo;curated&rdquo; one, and the sequence accession is added to that record.</p></list-item><list-item id="bid.1751"><p> Communications from the public. LocusLink provides three mechanisms for users: (<italic>a</italic>) at the bottom of each report is a mail link to the NCBI Service Desk. It is helpful when the mail message contains a sequence accession and a published citation in addition to the description of the request; (<italic>b</italic>) The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/GeneRIFhelp.html">GeneRIF</ext-link> link within the blue function bar on any LocusLink report page can be used. For loci not yet in the database, there are generic <named-content content-type="book-spec-emph">NEWENTRY</named-content> records established for human, mouse, rat, zebrafish, and fly. Thus, the user can use <named-content content-type="book-spec-emph">NEWENTRY</named-content> as a query term, select the species that needs a new entry, and add the GeneRIF. Please note that GeneRIFs require a published citation with a PubMed ID; and (<italic>c</italic>) The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/update.cgi">Update Submission Form</ext-link> accessed from the LocusLink homepage. This form should be used to notify RefSeq, LocusLink, and OMIM staff of new genes, to suggest updates, or to report errors.</p></list-item></list></p>
              </sec>
              <sec id="bid.852">
                <title>FTP Site</title>
                <table-wrap id="bid.853">
                  <label>5</label>
                  <caption>
                    <title>The LocusLink FTP site.</title>
                  </caption>
                  <table>
                    <tr>
                      <th align="left" valign="top">File or directory</th>
                      <th align="left" valign="top">Description</th>
                    </tr>
                    <tr>
                      <td align="left" valign="top">README</td>
                      <td align="left" valign="top">Documentation for the directory</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">HomologyMaps</td>
                      <td align="left" valign="top">Directory of data files used in the human/mouse comparative map</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LL.out.gz</td>
                      <td align="left" valign="top">Tab-delimited file of descriptors for current LocusLink records</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LL.out_dm.gz</td>
                      <td align="left" valign="top">Drosophila subset of LL.out</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LL.out_dr.gz</td>
                      <td align="left" valign="top">Zebrafish subset of LL.out</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LL.out_hs.gz</td>
                      <td align="left" valign="top">Human subset of LL.out</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LL.out_mm.gz</td>
                      <td align="left" valign="top">Mouse subset of LL.out</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LL.out_rn.gz</td>
                      <td align="left" valign="top">Rat subset of LL.out</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LL_tmpl.gz</td>
                      <td align="left" valign="top">Current text file for displays on the LocusLink site</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">LocusID_history</td>
                      <td align="left" valign="top">Report of locus_id merges</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">homol_seq_pairs.gz</td>
                      <td align="left" valign="top">mRNA accession pairs determined by MegaBlast</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">loc2UG</td>
                      <td align="left" valign="top">Current LocusID/UniGene cluster conversion table</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">loc2acc</td>
                      <td align="left" valign="top">Current LocusID/GenBank accession report</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">loc2cit</td>
                      <td align="left" valign="top">Current LocusID/PubMed ID/MedLine ID report</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">loc2ref</td>
                      <td align="left" valign="top">Current LocusID/RefSeq accession report</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">loc2sts</td>
                      <td align="left" valign="top">Current LocusID/UniSTS ID relationships</td>
                    </tr>
                    <tr>
                      <td align="left" valign="top">mim2loc</td>
                      <td align="left" valign="top">Current LocusID/MIM ID relationships</td>
                    </tr>
                  </table>
                </table-wrap>
                <p>LocusLink can be obtained by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/refseq/LocusLink">FTP</ext-link> (<xref ref-type="table" rid="bid.853">Table 5</xref>). These files are refreshed when new data are available, which for most files is daily (weekdays).</p>
                <p>The file LL_tmpl.gz represents the text file for displays of the complete LocusLink site and is in a semistructured tag:value format, which makes it challenging to parse. A subset of these data is available in tab-delimited format, either for all species covered by LocusLink (LL.out) or in species-specific files (e.g., LL.out.hs and LL.out.dm). These files report the LocusID, symbols, names, map location, and identity of the contributing database.</p>
                <p>The file loc2UG (the LocusID/UniGene cluster conversion table) is refreshed with each UniGene build, and homology reports are refreshed with new genome annotation builds. These files can be used to obtain names and sequences connected to a locus.</p>
              </sec>
            </sec>
            <sec id="bid.854">
              <title>Integration with Other Resources</title>
              <table-wrap id="bid.855">
                <label>6</label>
                <caption>
                  <title>Connections between LocusID and other NCBI resources.</title>
                </caption>
                <table>
                  <tr>
                    <th align="left" valign="top">Resource</th>
                    <th align="left" valign="top">Connection made</th>
                    <th align="left" valign="top">Basis</th>
                  </tr>
                  <tr>
                    <td align="left" valign="top">dbSNP</td>
                    <td align="left" valign="top">LocusID to Reference SNP ID</td>
                    <td align="left" valign="top">dbSNP accessions aligned to mRNAs, or within the boundary of 2000 nt upstream through 500 nt downstream of known exons</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">Map Viewer</td>
                    <td align="left" valign="top">LocusID to cytogenetic, genetic, or sequence position</td>
                    <td align="left" valign="top">Reports of cytogenetic position or calculated from sequence assembly; reports of genetic position, connecting LocusIDs to sequence position based on alignment of accessions associated with LocusIDs</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">RefSeq mRNAs and proteins</td>
                    <td align="left" valign="top">LocusID to mRNA and protein accessions</td>
                    <td align="left" valign="top">Calculated and curated LocusID/mRNA or protein_id relationships</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">UniGene</td>
                    <td align="left" valign="top">LocusID to UniGene cluster ID</td>
                    <td align="left" valign="top">Based on the LocusID/mRNA or protein_id relationships, identifying the cluster ID and requiring that the LocusID/UniGene cluster ID relationship be 1:1</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">UniSTS</td>
                    <td align="left" valign="top">LocusID to UniSTS sts_id</td>
                    <td align="left" valign="top">Based on the LocusID/mRNA relationships or overlap in the genomic assembly, after positioning the STS by e-PCR</td>
                  </tr>
                </table>
              </table-wrap>
              <p>The database supporting LocusLink houses more than only the unique loci identifiers for the genomes in its scope. It also records, whenever possible, the public sequence Accession numbers that define these loci and, along with its collaborators, applies several tests for data consistency. Thus, the relationships between LocusID and sequence Accession numbers are used by other databases at NCBI to convert information about mRNA or protein sequences to the LocusID and all other information associated with that LocusID (name, database cross-references, protein product, and others). These relationships are outlined in <xref ref-type="table" rid="bid.855">Table 6</xref>.</p>
            </sec>
            <sec id="bid.856">
              <title>More Information on LocusLink</title>
              <p>On LocusLink's static HTML pages, such as the homepage, there are links to general resources. Specific sites provide information for <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/LLfaq.html">LocusLink FAQs</ext-link> and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/RSstatistics.html">RefSeq statistics</ext-link>.</p>
            </sec>
          </body>
        </book-part>
        <book-part id="bid.1562" book-part-type="chapter" book-part-number="20">
          <book-part-meta>
            <title-group>
              <title>Using the Map Viewer to Explore Genomes</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Dombrowski</surname>
                  <given-names>Susan M.</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Maglott</surname>
                  <given-names>Donna</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch20d1"/>
            <abstract>
              <title>Summary</title>
              <p>There are many different approaches to starting a genomic analysis. These include literature searching, searching databases for gene names and other genomic features, performing sequence comparisons, or using map data to find gene information by position relative to other landmarks. The NCBI Map Viewer has been developed to facilitate this latter approach.</p>
              <p>The purpose of this chapter is to provide a foundation for gaining maximum benefit from using the Map Viewer and related resources at NCBI. It is important to note that in this document, the term &ldquo;map&rdquo; refers to a position of a particular type of object in a particular coordinate system. This means, for example, that there is not one sequence map but a set of maps in sequence coordinates. Readers interested in precisely how sequence-based maps are annotated and assembled should refer to <xref ref-type="other" rid="bid.1440">Chapter 14</xref>.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.1677">
              <title>Introduction</title>
              <p>First launched with the release of the sequence of <italic>Drosophila melanogaster</italic> in March 2000, Map Viewer is now used to present genetic, radiation hybrid (RH), cytogenetic, breakpoint, sequence-based, and clone maps for many genomes. The availability of whole genome sequences means that objects such as genes, markers, clones, sites of variation, and clone boundaries can be positioned by aligning defining sequence from these objects against the genomic sequence. This position information can then be compared to information about order obtained by other means, such as genetic or physical mapping. The results of sequence-based queries (e.g., BLAST) can also be viewed in genomic context. Our view of the genomes of a variety of organisms is constantly being improved through the increase in underlying data.</p>
              <p>Map Viewer integrates map and sequence data from a variety of sources. The basic architecture and principle of Map Viewer can be applied to any complete or incomplete genome as long as map data exist to support it. Map Viewer is a powerful tool because it provides: (1) a mechanism to compare maps in different coordinate systems; (2) a robust query interface; (3) diverse options for configuring the display; (4) multiple functions to report and download maps and annotated information; (5) tools to manipulate nucleotide sequence such as <!--<glossref>-->ModelMaker<!--</glossref>--> (for constructing mRNAs from putative exon sequences); (6) connections to comprehensive data files for transfer by FTP; and (7) detailed descriptions of the objects displayed on the maps.</p>
            </sec>
            <sec id="bid.1563">
              <title>Maintenance of Data</title>
              <sec id="bid.1564">
                <title>Data Sources</title>
                <p><bold>Non-Sequence-based Maps.</bold> Sources of maps that are not based directly on sequence include published maps in genetic, radiation hybrid, cytogenetic, and ordinal coordinate systems (where ordinal refers to clone order). The primary sources of each map are described in the online help documentation of each genome-specific Map Viewer. We are indebted to the researchers who make their mapping results so freely available. When a new version of any map becomes available, the data are also updated in the appropriate NCBI database.</p>
                <p><bold>Sequence-based Maps.</bold> The sequence-based maps shown through Map Viewer can be supplied by external sources and/or supplied from features computed within NCBI. For example, when the annotated sequence for a complete genome is submitted to the sequence databases (GenBank/EMBL/DDBJ), a copy of the data may also be accessioned as Reference Sequences (RefSeqs; see <xref ref-type="other" rid="bid.681">Chapter 18</xref>). The gene, transcript, and other feature annotations of the submitted complete genome are processed for display in the Map Viewer. NCBI staff may then calculate and display the position of other types of features, such as marker position or points of variation, as separate maps (<xref ref-type="table" rid="bid.1565">Table 1</xref>).</p>
                <table-wrap id="bid.1565"><label>1</label><caption><title>Types of Map Viewer annotation provided by NCBI.</title></caption><table><tr><th align="left" valign="top">Feature<xref ref-type="fn" rid="mul21"><sup><italic>a</italic></sup></xref></th><th align="left" valign="top">Coordinate system<xref ref-type="fn" rid="mul22"><sup><italic>b</italic></sup></xref></th><th align="left" valign="top">Representative maps<xref ref-type="fn" rid="mul23"><sup><italic>c</italic></sup></xref></th></tr><tr><td align="left" valign="top"><bold>STS</bold></td><td align="left" valign="top">Sequence (Mb), Radiation hybrid (cRay), Genetic (cM), Clone content (ordinal), Cytogenetic</td><td align="left" valign="top">STS, STSnw, G3, GM4, GeneMap'99, TNG, Marshfield, Genethon, deCode, Whitehead YAC, phenotype maps such as Quantitative Trait Loci (QTL)</td></tr><tr><td align="left" valign="top"><bold>Clones</bold></td><td align="left" valign="top">Sequence, Cytogenetic</td><td align="left" valign="top">Clone, BES, Components</td></tr><tr><td align="left" valign="top">&emsp;Expression</td><td align="left" valign="top">Sequence</td><td align="left" valign="top">SAGE tag, UniGene</td></tr><tr><td align="left" valign="top"><bold>Genes</bold></td><td align="left" valign="top">Sequence (Mb), Cytogenetic (band names)</td><td align="left" valign="top">Genes_seq, Genes_cyto</td></tr><tr><td align="left" valign="top">&emsp;Gene-related</td><td align="left" valign="top">Sequence, Cytogenetic</td><td align="left" valign="top">UniGene, GenomeScan, Mitelman recurrent breakpoint, morbid</td></tr><tr><td align="left" valign="top"><bold>Variation</bold></td><td align="left" valign="top">Sequence (Mb)</td><td align="left" valign="top">Variation</td></tr><tr><td align="left" valign="top">&emsp;Published accessions</td><td align="left" valign="top">Sequence (Mb)</td><td align="left" valign="top">GenBank</td></tr><tr><td align="left" valign="top">&emsp;Phenotype</td><td align="left" valign="top">Cytogenetic, Cytogenetic (abnormalities), Sequence</td><td align="left" valign="top">OMIM's morbid map, Mitelman's recurrent breakpoint, QTL (in progress)</td></tr><tr><td align="left" valign="top"><bold>Source clones</bold></td><td align="left" valign="top">Sequence (Mb)</td><td align="left" valign="top">Component</td></tr><tr><td align="left" valign="top">&emsp;Homology</td><td align="left" valign="top">Sequence (Mb)</td><td align="left" valign="top">Indirectly via LocusLink or UniGene. For mouse and human, through the homology (hm) link to the mouse&ndash;human homology map</td></tr></table><table-wrap-foot><fn symbol="a" id="mul21"><p>The feature column lists the types of objects annotated on maps seen in Map Viewer. Those features in bold type are annotated on the RefSeqs; the rest are provided only from the Map Viewer, and the files are available for FTP transfer.</p></fn><fn symbol="b" id="mul22"><p>The different map types and coordinate systems that may contain a particular type of feature.</p></fn><fn symbol="c" id="mul23"><p>A partial enumeration of named maps that represents positions of this feature type.</p></fn></table-wrap-foot></table-wrap>
                <p>Some of the annotation of genomic sequence carried out by NCBI is included in the genomic reference sequences (NC, NT, and NW Accession number format); however, other annotation is represented only in the Map Viewer and in the associated reports (<xref ref-type="table" rid="bid.1565">Table 1</xref>). This latter type of annotation is based on information in several NCBI databases (<xref ref-type="table" rid="bid.1566">Table 2</xref>) and is particularly important for attaching biological information to sequence data. Links to these resources are provided in Map Viewer to provide further information about each annotated object. It should be noted, however, that although sequence features may be placed in a genomic context automatically, there are curation steps that affect the final displays. For example, for the human and mouse genomes, sequences defining genes and pseudogenes are reviewed by collaborators and NCBI staff and, whenever possible, used as the basis of RefSeq records (NG, NM, and NR Accession number format).</p>
                <table-wrap id="bid.1566"><label>2</label><caption><title>NCBI data resources used in NCBI-generated annotation.</title></caption><table><tr><th align="left" valign="top">Resource</th><th align="left" valign="top">Description</th></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/clone/">Clone Registry</ext-link></td><td align="left" valign="top">Clone sequencing sequence status, STS content, and availability</td></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/SNP/">dbSNP</ext-link></td><td align="left" valign="top">Single Nucleotide Polymorphisms (SNPs), polymorphisms, small-scale insertions/deletions, polymorphic repetitive elements</td></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genomes/">Genome Guides</ext-link></td><td align="left" valign="top">Directory of key resources for the genome, with links to related resources and tutorials. The directory to guide pages is available from Genomic Biology.</td></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/">LocusLink</ext-link></td><td align="left" valign="top">Locus-specific data for a subset of organisms with extensive links to related resources and sequence data</td></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=OMIM">OMIM</ext-link></td><td align="left" valign="top">Human genes and Mendelian disorders</td></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/RefSeq.html">RefSeq</ext-link></td><td align="left" valign="top">NCBI's curated, non-redundant RefSeqs</td></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/index.html">UniGene</ext-link></td><td align="left" valign="top">Computed clusters of cDNA and Expressed Sequence Tag (EST) sequences from the same gene, with tissue expression information and links to related resources</td></tr><tr><td align="left" valign="top"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/sts/">UniSTS</ext-link></td><td align="left" valign="top">Unified, nonredundant database of sequence tagged sites (STSs)</td></tr></table></table-wrap>
                <p>Feature annotation is computed primarily in two ways: (1) by alignment of the defining sequence to the genome; or (2) for sequence tagged sites (STSs), by e-PCR (<xref ref-type="bibr" rid="bid.1625">1</xref>). In some genomes, gene placement is based primarily on the alignment of mRNA [Expressed Sequence Tags (ESTs) and cDNAs], but only when an encoded protein is predicted. In other cases, where transcription evidence is weaker, more weight is given to identification of protein-coding regions. Gene identification is also constrained in that a known gene cannot be placed more than once in a haplotype (except for pseuodo-autosomal regions) or on an incorrect chromosome. Thus, if any reference haplotype retains inappropriately redundant sequence that encodes a gene, only one copy will be annotated as that gene. Others will be assigned interim IDs (see <xref ref-type="other" rid="bid.1440">Chapter 14</xref>). Some <italic>ab initio</italic> methods may also be used for gene prediction. The predicted genes, as well as the mRNAs, are supplied as separate maps (gene, RNA, or <xref ref-type="other" rid="bid.1301">GenomeScan</xref> maps).</p>
                <p>In some cases, the position of these features may suggest the location of other genomic regions of interest. For example, the position of STS markers can help define the position of phenotypes such as quantitative trait loci (QTL). Although the best annotation of a gene or region is always through annotation by an expert researcher, automated annotation of genomes and comparison to that provided by experts can provide significant useful information. Experts interested in analyzing or assisting with genome annotation should contact us at <email>info@ncbi.nlm.nih.gov</email>.</p>
              </sec>
              <sec id="bid.1567">
                <title>Relationships among Coordinate Systems</title>
                <p>In addition to supporting the display of multiple maps in the same coordinate system (e.g., multiple sequence-based maps), Map Viewer also displays maps in different coordinate systems by calculating the correspondances among them (e.g., sequence to genetic). This is accomplished by: (<italic>a</italic>) identifying features that have been placed on maps in different coordinate systems; and (<italic>b</italic>) using general conversion factors. In the first case, placement of STSs on the genome is critical for the integration of sequence data with other, non-sequence-based maps, such as genetic and RH maps. The integration of cytogenetic data with sequence data is achieved through alignment of sequence from clones that have been placed cytogentically, such as the human fluorescence <italic>in situ</italic> hybridization (FISH)-mapped clones from the Bacterial Artificial Chromosome (BAC) Resource Consortium (<xref ref-type="bibr" rid="bid.1626">2</xref>). The integration of non-sequence-based maps with the sequence provides a powerful mechanism to access portions of sequence on the basis of marker or cytogenetic data. Many features, such as Single Nucleotide Polymorphisms (SNPs), ESTs, mRNAs, whole genome shotgun reads, and clones can be placed on the genome assembly by using standard DNA sequence alignment methods such as BLAST.</p>
                <p>The identification of known genes within the genome assembly provides critical landmarks and functional context to the sequence data, which in turn makes it easier to traverse to other rich sources of gene and protein information, including publications, OMIM, RefSeq, Conserved Domain Database (CDD), and LocusLink.</p>
                <p>The power of calculating correspondances between coordinate systems may be more apparent when considering a common application of Map Viewer, i.e., identifying candidate genes within a region defined by genetic markers. When markers are palced on both genetic and sequence maps, it is then possible to use the gene-related maps (gene, UniGene/EST, or <italic>ab initio</italic> predictions) to identify possible genes of interest. For more details on how to do this, see the Map Viewer Exercises in <xref ref-type="other" rid="bid.1679">Chapter 23</xref>.</p>
              </sec>
              <sec id="bid.1568">
                <title>A Work in Progress</title>
                <p>For many genomes, identifying and positioning chromosomes and genes within sequence blocks is an ongoing process. In those cases, the Map Viewer can be used to evaluate the evidence that supports the current representation of the sequence and visualize possible conflicts. Inconsistencies in map order or in the placement of any object can be seen in the Map Viewer; this is assisted in some cases by the use of color coding (Figures <xref ref-type="fig" rid="bid.1569">1</xref> and <xref ref-type="fig" rid="bid.1570">2</xref>).<fig id="bid.1569"><label>1</label><caption><title>Evaluation of a chromosome sequence (STS) map.</title><p>Potential inconsistencies in the order or orientation of sequence blocks can be investigated by displaying a genetic map (<italic>Marshfield</italic>), radiation hybrid map (<italic>TNG</italic>), and sequence map (<italic>STS</italic>) together and checking the <named-content content-type="book-spec-emph">Show connections</named-content> box in the <named-content content-type="book-spec-emph">Maps &amp; Options</named-content> window. Note that some of the <italic>gray lines</italic> (connecting the same marker on different maps) are <italic>crossed</italic>, indicating that either the placement is incorrect on a map or the chromosome sequence is not ordered and oriented consistently with all map data.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f1" mime-subtype="gif"/></fig><fig id="bid.1570"><label>2</label><caption><title>Evaluation of gene localization and annotation.</title><p>A comparison of cDNA alignments (UniGene, RNA) and gene predictions (GenomeScan) to the genomic contig annotation can be achieved by displaying three maps simultaneously. The genomic contig (NT_024981.9) annotation is shown in the <italic>Genes_seq</italic> map and is displayed with the GenomeScan predictions (the <italic>GScan</italic> map) and the EST/mRNA alignments labeled by human UniGene clusters (the <italic>UniG_Hs</italic> map). Note that in this case, there are two sequence objects not included in the contig annotation: one is an <italic>ab initio</italic> prediction (the last model in the GScan map) (<italic>a</italic>); and the other is either some small gene or an alternative 3&prime; exon for PIK3C3 from the UniG_Hs map (<italic>b</italic>). This approach is especially useful when reviewing BLAST results in a genomic context.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f2" mime-subtype="gif"/></fig></p>
                <p>For some genomes, the color-coded contig map displays whether the annotation is based on sequence assembled from draft or finished clones (blue, finished; green, whole genome shotgun; orange, draft). This is helpful when evaluating the level of confidence in the completeness of the annotation of a gene and/or its coding region.</p>
                <p>Map Viewer also uses color coding or diagrams to represent the level of confidence in the placement of any mapped object. For example, SNPs or STSs that are placed at more that one position in a given map are noted by color (yellow) in the detailed labels (<xref ref-type="fig" rid="bid.1571">Figure 3a</xref>). Annotated genes are shown in different colors, based on the source and level of confidence in the annotation or the model (<xref ref-type="fig" rid="bid.1571">Figure 3b</xref>).<fig id="bid.1571"><label>3</label><caption><title>Representation of ambiguity.</title><p>
							(<italic>a</italic>) The marker D1S2894 is found on several maps. Note that for the first map (<italic>STS</italic>), the circle is diagonally split with two colors. The diagonal means that the marker has been placed more than once; the two colors mean that the placements are not on the same chromosome. (<italic>b</italic>) A Map Viewer display of a region of chromosome 16. SNPs that are placed more than once on the chromosome are designated by a <italic>yellow triangle</italic>. From the <italic>Contig</italic> map, it appears that at least one of these SNPs (rs3220808) is placed both on draft sequence (<italic>orange</italic>) and on finished sequence (<italic>blue</italic>). This may be an artifact resulting from misassembly or perhaps a region of segmental duplication. This diagram also illustrates the use of color to indicate the source and level of confidence in annotated genes. <italic>Blue</italic> indicates a confirmed gene with no conflicts; <italic>light green</italic> indicates EST evidence only; <italic>dark brown</italic> indicates a GenomeScan prediction with protein homology; <italic>orange</italic> means that there is a conflict between the annotated gene and the mRNA evidence. (<italic>Ab initio</italic> predictions from GenomeScan are categorized into two types, based on presence or absence of sequence similarity to vertebrate proteins or protein domains.)</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f3" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.1572">
                <title>Frequency of Updates</title>
                <p>Although maps provided from external sources are updated when new data are available, the maps dependent on NCBI's annotation process are updated periodically in versions called &ldquo;builds&rdquo;. Thus, mRNA or other supporting evidence that becomes available after the data &ldquo;freeze&rdquo; date for one build will not be incorporated into the display until the next build. However, some of the supporting databases linked from the Map Viewer may have more updated information. For example, UniSTS may provide more recent e-PCR results, or LocusLink may show a newer name or additional sequence data. dbSNP may make major data releases between builds; in this case, the variation map is updated.</p>
              </sec>
            </sec>
            <sec id="bid.1573">
              <title>Methods of Access</title>
              <p>Although most of this chapter discusses the human genome Map Viewer, there is a growing number of organisms for which there is Map Viewer access to the genome. To identify the taxa that have Map Viewer access to the genome, query the taxonomy database by typing &ldquo;loprovmapviewer&rdquo;[filter] into the query box on the Entrez Taxonomy <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Taxonomy">homepage</ext-link>; or more simply, review the options provided on the Map Viewer <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/">homepage</ext-link>.</p>
              <sec id="bid.1574">
                <title>Links from NCBI Resources</title>
                <p>Many NCBI databases are now integrated into Map Viewer (<xref ref-type="table" rid="bid.1566">Table 2</xref>); therefore, database records are often linked to Map Viewer displays. If a sequence in the public databases was released before the date of the current Map Viewer data freeze, then the position of this sequence may be displayed within Map Viewer. For example, Entrez Nucleotide, UniGene, UniSTS, and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ll" xlink:href="793">LocusLink</ext-link> records for sequences annotated on the human genome provide links directly to the appropriate region of the genome via links called <named-content content-type="book-spec-emph">Map Viewer</named-content> (in the <named-content content-type="book-spec-emph">Links</named-content> menu), <named-content content-type="book-spec-emph">Nucleotide</named-content>, <named-content content-type="book-spec-emph">Map View</named-content>, or <named-content content-type="book-spec-emph">mv</named-content> links, respectively (<xref ref-type="fig" rid="bid.1575">Figure 4</xref>). It should be noted that such links are only precomputed if at least 50% of the sequence aligns with an identity of greater than 90%.<fig id="bid.1575"><label>4</label><caption><title>Connecting to Map Viewer from Entrez Nucleotide, using the <named-content content-type="book-spec-emph">Links</named-content> menu in Entrez Nucleotide to connect from a record to Map Viewer.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f6" mime-subtype="gif"/></fig></p>
                <p>Genome-specific resource pages also support queries via chromosome diagrams (<xref ref-type="fig" rid="bid.1576">Figure 5</xref>).<fig id="bid.1576"><label>5</label><caption><title>Example of a genome-specific resource page supporting queries to Map Viewer.</title><p>Note that there are two ways to connect: (1) by selecting <named-content content-type="book-spec-emph">Maps</named-content> from the pull-down menu in the <italic>gray bar</italic> at the <italic>top</italic> of the page and entering a query term in the <named-content content-type="book-spec-emph">Search</named-content> box; or (2) by selecting a chromosome in the genome diagram on the <italic>right</italic> of the page (<italic>yellow</italic> background).</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f4" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.1577">
                <title>Sequence Similarity Searches</title>
                <p>Genome-specific BLAST pages that restrict a search to a specific genome are provided for several organisms and allow the results of the search to be displayed in a genomic context (provided by Map Viewer). Genome-specific BLAST searches can be accessed from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/">BLAST</ext-link> homepage, the Map Viewer pages of individual organisms (e.g., <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/Entrez/map_search">human</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/Entrez/map_srchdb?chr=mouse_chr.inf">mouse</ext-link>), and the genome-specific resource pages of individual organisms. If the reference genome (the default) is selected as the database to be searched, the <named-content content-type="book-spec-emph">Genome View</named-content> button (<xref ref-type="fig" rid="bid.1578">Figure 6</xref>) will appear on the BLAST results display page.<fig id="bid.1578"><label>6</label><caption><title>Accessing the Map Viewer display from a genome-specific BLAST results page.</title><p>Selecting the <named-content content-type="book-spec-emph">Genome View</named-content> button shows all of the BLAST hits on the genome. <!--See Exercises (Gene Family).--></p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f5" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.1579">
                <title>Direct Query</title>
                <sec id="bid.1580">
                  <title>
                    <bold>Simple Searches</bold>
                  </title>
                  <p>When already at a genome-specific Map Viewer page, any combination of query terms can be entered into a Map Viewer <named-content content-type="book-spec-emph">Search for</named-content> box (<xref ref-type="fig" rid="bid.1581">Figure 7</xref>). Boolean operators (AND, OR, and NOT) and the use of <bold>*</bold> as a wild card (applied to the right of any term) are supported. The <named-content content-type="book-spec-emph">Search for</named-content> and <named-content content-type="book-spec-emph">Help</named-content> document hyperlinks provide current details about query options. An advanced search is available for some genomes.<fig id="bid.1581"><label>7</label><caption><title>Representative of species-specific homepage.</title><p>Note the links to the help documentation and related resources. Also note the check boxes to use the advanced query page and/or to display objects calculated to have links to any object returned by the query. A link to the genome-specific BLAST site is also provided at the <italic>top</italic> of the form.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f7" mime-subtype="gif"/></fig></p>
                  <p>Queries may include any unique identifier for a database record, e.g., a sequence Accession number or OMIM (MIM) number, or a text term or phrase, e.g., a gene symbol (BRCA2) or descriptor (p53-binding), or disease name (lung cancer). The Boolean AND operator is used automatically if multiple terms are entered. Therefore, a query for &ldquo;fanconi anemia&rdquo; will automatically be interpreted as &ldquo;fanconi AND anemia&rdquo;. The wildcard operator (*) provides a convenient mechanism to retrieve genes that share a common symbol or name, as is often found for gene families. For example, a query for ABC* will return matches to the ATP-binding cassette superfamily.</p>
                  <p>The advanced query page, accessed by checking the <named-content content-type="book-spec-emph">Advanced search</named-content> box, provides additional options to refine a query. These additional options, which may vary from genome to genome, are useful for restricting queries to a particular search field or map type. The advanced query page also includes predefined search options to restrict the search to data with certain properties, e.g., to only find genes associated with a known disease or with sequence variation (SNPs). Additional refinements to queries against the variation map can also be made, for example, to search for variation markers known to be in a gene or coding region.</p>
                  <p>The same options for wild cards and Boolean operators for your query term(s) apply when starting at the Map Viewer <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview">homepage</ext-link>. At present, however, you must select a genome to which to restrict your search. An option to query across multiple genomes is under development.</p>
                </sec>
                <sec id="bid.1582">
                  <title>
                    <bold>Position-based Access</bold>
                  </title>
                  <p>To use Map Viewer to display a particular section of a genome by using a range of positions as a query, it is first necessary to select a particular chromosome for display from either a genome-specific Map Viewer page or a Genome Guide page.</p>
                  <p>Once a single chromosome is displayed, position-based queries can be defined by: (1) entering a value into the <named-content content-type="book-spec-emph">Region Shown</named-content> box. This could be a numerical range (base pairs are the default if no units are entered), the names of clones, genes, markers, SNPs, or any combination. The screen will be refreshed with only that region shown. If the first entry cannot be resolved, the display will extend to the top of the map; if the second entry cannot be resolved, the display will extend to the bottom of the map. Both of these navigational aids are found on the left of the page; and (2) using the <named-content content-type="book-spec-emph">Maps &amp; Options</named-content> controls. One of the options in this menu is to define the region shown. Here it may be clearer that the region selected will be in the coordinates of the rightmost, or Master, map, which may also be adjusted in this menu. The values that can be used to specify the range are the same as those described in (1), above. (See <italic>Customizing the Display</italic> for more details on fine-tuning.)</p>
                  <p>Tutorials in Chpater 23, particularly #2, provide more examples of querying Map Viewer by position.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.1583">
              <title>Interpreting the Display</title>
              <sec id="bid.1584">
                <title>Map Viewer Summary Results</title>
                <p>The results from a query are displayed both graphically and in a summary table (<xref ref-type="fig" rid="bid.1585">Figure 8</xref>). When the query is executed by BLAST, the graphical view is color-coded according to the BLAST score, and the table summarizes the scores and the RefSeq Accession numbers that have matches. <!--Exercises (Gene Family)-->Clicking on the RefSeq Accession number (i.e., those beginning with NT_ or NW_) displays that BLAST result in the Map Viewer.<fig id="bid.1585"><label>8</label><caption><title>Map Viewer query results.</title><p>
							(<italic>a</italic>) Note that an OMIM (MIM) number was entered as a query, and the <named-content content-type="book-spec-emph">Show linked entries</named-content> box was checked. (<italic>b</italic>) The <italic>red tick marks</italic> next to the chromosome diagram indicate where the results appear to be placed on the chromosome. The left/right placement does not indicate strand; it allows more resolution between tick marks. (<italic>c</italic>) When a result is returned as the result of a link, the data in the <named-content content-type="book-spec-emph">Match</named-content> column of the table begin with <named-content content-type="book-spec-emph">linked to</named-content>. The complete results for a particular chromosome can be displayed by selecting the name of the chromosome on the graphical portion of the display or selecting <named-content content-type="book-spec-emph">all matches</named-content> in the summary table. The number of matches per chromosome is reported under each chromosome in the graphical overview, and the total number of results returned is indicated below that. Additional pages are provided if the query returns over 100 results. The table indicates the chromosomal location, the match found, the map element returned, the type of match found, and the specific map(s) that contains the query match. Only the first 40 characters are shown in the <italic>Match</italic> column, and therefore the portion of text that matches the query may not be displayed. All maps that contain any object are viewed by selecting the name of the object in the <italic>Map element</italic> column. To see only one map, select the name of that map. In some cases, the resultant display will contain related maps in the same sequence coordinates. For example, selecting a sequence-based gene map may result in the display of mRNA alignments, labeled with UniGene cluster designations, and <italic>ab initio</italic> predictions.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f8" mime-subtype="gif"/></fig></p>
              </sec>
              <sec id="bid.1586">
                <title>Viewing the Maps</title>
                <sec id="bid.1587">
                  <title>
                    <bold>The Graphical Display</bold>
                  </title>
                  <sec id="bid.1588">
                    <title>Text or Position Queries</title>
                    <p>General information on the chromosome being viewed is summarized at the top of the map page: the species and chromosome currently being viewed, the query term, and the name of the focal map, termed the Master Map (<xref ref-type="fig" rid="bid.1752">Figure 9c</xref>).<fig id="bid.1752"><label>9</label><caption><title>Representative map display.</title><p>(<italic>a</italic>) Search and find options. (<italic>b</italic>) <named-content content-type="book-spec-emph">Compress Map</named-content> allows you to select whether to display labels on maps other than the rightmost one. The <named-content content-type="book-spec-emph">Region Shown:</named-content> and <named-content content-type="book-spec-emph">zoom</named-content> graphic allow you to reset the range displayed either explicitly or by chromosome fractions, respectively. (<italic>c</italic>) Explanation and tools. This section reminds you of your chitin* query and provides statistics about the elements mapped on the Master Map for that chromosome. In this example, 3978 genes are annotated on chromosome IV, and 2 of 2 in the region are being displayed. The range shown is summarized, and links are provided to tools to download a region of the sequence and use ModelMaker and Evidence Viewer. The labels at the <italic>top</italic> of each map are connected to the documentation for that map, the -&gt; makes the map become the master, and the <named-content content-type="book-spec-emph">X</named-content> removes the map from the viewer. (<italic>d</italic>) Map. This section contains a representation of the order of elements in the range displayed and provides information about each, either by mousing over the label (uncompressed) or circle (compressed) of the non-master map, or by clicking on the links provided on the labels for the Master Map. This example shows that there is an option to display a ruler for any map. As described in the text and other figures, color is used in this display to convey information as well.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f9" mime-subtype="gif"/></fig></p>
                    <p>The summary also includes the following statistics concerning the number of objects on the Master Map, which are:
                      <list list-type="bullet">
                        <list-item id="bid.1589">
                          <p> the number of objects localized (positioned) on the chromosome</p>
                        </list-item>
                        <list-item id="bid.1590">
                          <p> the number of objects not localized but present on the chromosome</p>
                        </list-item>
                        <list-item id="bid.1591">
                          <p> the number of objects localized in the region displayed (i.e., the number decreases as you zoom in)</p>
                        </list-item>
                        <list-item id="bid.1592">
                          <p> the number of objects for which text descriptions are shown (dependent on user-defined page length)</p>
                        </list-item>
                      </list>
                    </p>
                    <p>A thumbnail map on the left of the page provides a coarse indication of the region displayed; by default, this is a cytogenetic map, although the Master Map can be selected (<xref ref-type="fig" rid="bid.1752">Figure 9b</xref>).</p>
                    <p>Maps are displayed vertically, with the name of each map hyperlinked to a description of it (<xref ref-type="fig" rid="bid.1752">Figure 9d</xref>). Features displayed on the Master Map have brief descriptive labels; information on features on the non-Master Maps can be found by mousing over an object. The labels on the Master Map depend on the type of object and genome being explored but can provide: (<italic>a</italic>) links to resources defining the mapped element, some of which may not be at NCBI; (<italic>b</italic>) indicators of the confidence in the placement or naming or sequence in the region; (<italic>c</italic>) biological features of the element (for SNPs, this includes position in a gene or effects on the coding region); (<italic>d</italic>) direction of transcription for genes; and (<italic>e</italic>) links to tools to facilitate reviewing of the sequence (<named-content content-type="book-spec-emph">sv</named-content>), downloading a subsequence of interest (<named-content content-type="book-spec-emph">seq</named-content>), the mRNA alignments in a region (<named-content content-type="book-spec-emph">ev</named-content>), homology maps (<named-content content-type="book-spec-emph">hm</named-content>), or to create cDNA sequences in real time (<named-content content-type="book-spec-emph">mm</named-content>). (See the section on <italic>Associated Tools</italic> for more information.).</p>
                  </sec>
                  <sec id="bid.1593">
                    <title>Sequence (BLAST) Queries</title>
                    <p>The positions of BLAST hits are highlighted on the Contig map, and a text summary of the BLAST hit is provided with links to regional alignment reports. All of the options described previously for configuring your display are still available. Thus, it is possible to evaluate the sequence match by the location (possible intron/exon structure, percent identity) as well as to determine whether the matching genomic region contains all of the query sequence in the expected order. Adding other maps to the display using the <named-content content-type="book-spec-emph">Maps&amp;Options</named-content> window provides a powerful mechanism to determine how the query sequence corresponds to existing annotation, such as genes, gene predictions, STS markers, or SNPs. For more hints, see the tutorial section on querying the human genome by sequence.</p>
                  </sec>
                </sec>
                <sec id="bid.1594">
                  <title>
                    <bold>The Tabular Display (View Data as Table/Download)</bold>
                  </title>
                  <p>A tabular report of the region and maps being displayed can be generated by selecting the <named-content content-type="book-spec-emph">Data as Table View</named-content> link (<xref ref-type="fig" rid="bid.1752">Figure 9b</xref>). The default report is restricted to maps that were in the previous graphical display. Tables indicating the object name, or other identifier, and chromosome coordinates are provided for each map, along with many of the links seen in the graphical display. If the region being displayed on the map includes more than 1000 features per map, a warning message is displayed that points to the FTP site as an alternative for large-scale access.</p>
                  <p>If any of the maps are in sequence coordinates, an option is presented to report data for any sequence map in the region. Note: Links are provided for downloading tab-delimited files for any or all maps.</p>
                </sec>
              </sec>
            </sec>
            <sec id="bid.1595">
              <title>Customizing the Display</title>
              <p>The Map Viewer display can be customized with regard to the region shown, the number and coordinate systems of maps, the number of objects labeled on the Master Map, and whether to show connections between objects. Each of these will be described in this section.</p>
              <sec id="bid.1596">
                <title>Selecting the Region to Display</title>
                <p>The Map Viewer provides zoom, navigation, and other map display controls. These can be found on the display page itself and in the <named-content content-type="book-spec-emph">Maps&amp;Options</named-content> window (<xref ref-type="fig" rid="bid.1678">Figure 10</xref>).<fig id="bid.1678"><label>10</label><caption><title>Representative Maps&amp;Options window.</title><p>The menus and boxes in this window allow definition of the range of the chromosome to display (<named-content content-type="book-spec-emph">Region Shown:</named-content>), selection of the new map(s) to add to the display (<named-content content-type="book-spec-emph">Available Maps:</named-content>), establishment of the order and ruler options [<named-content content-type="book-spec-emph">Maps Displayed (left to right):</named-content>], control of the display of lines connected to related objects on different maps (<named-content content-type="book-spec-emph">Show connections</named-content> box), control of the length of the label on the rightmost (Master) map (<named-content content-type="book-spec-emph">Verbose Mode</named-content> box), compression of the graphic (<named-content content-type="book-spec-emph">Compressed Graphics</named-content> box), control of the number of labels on the Master map (<named-content content-type="book-spec-emph">Page Length:</named-content> box), and control of the diagram in the <named-content content-type="book-spec-emph">Thumbnail View:</named-content>. To add map(s) to the display, select the map name(s) and then click on <named-content content-type="book-spec-emph">ADD&gt;&gt;</named-content>. To remove map(s) from the display, select the map name(s) in the <named-content content-type="book-spec-emph">Maps Displayed</named-content> box and then click on <named-content content-type="book-spec-emph">&lt;&lt;REMOVE</named-content>. Please note that this is an example of a configuration that might be useful in displaying gene-related information, i.e., maps of UniGene, Gene, and GenomeScan.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch20f10" mime-subtype="gif"/></fig></p>
                <p>As the resolution of a view is changed, the chromosome diagram is updated. The view automatically centers on a highlighted query term, or on the middle of the chromosome if browsing only. The chromosome view can be moved up and down or zoomed in and out. Zooming can be achieved in several ways: (1) by using the zoom control, located in the left column; (2) by providing a range or bounding markers in the <named-content content-type="book-spec-emph">Region Shown</named-content> text boxes; or (3) by selecting part of the map to display a menu with predefined zoom levels. Most menu-based zooms should be carried out in two or three steps to avoid missing the region of interest. It is also possible to scroll or reposition the display by selecting <named-content content-type="book-spec-emph">recenter</named-content> from the menu that pops up when you click on the chromosome diagram at the left or on a map or by clicking on the arrows at the top and bottom of a map.</p>
              </sec>
              <sec id="bid.1597">
                <title>Selecting the Maps (Tracks) to Display</title>
                <p>Maps are categorized by the coordinate system as well as type of feature. The maps available for a genome can be seen by scrolling through the <named-content content-type="book-spec-emph">Maps</named-content> menu in the <named-content content-type="book-spec-emph">Maps&amp;Options</named-content> window or in the genome-specific help documentation. For display and query purposes, different types of features annotated on sequence coordinates are treated as different maps. The maps in sequence coordinates are comparable because all of the sequence maps are based on the reference to a standard genome assembly. Thus, one can display the SNP map (at high zoom level) next to the Gene, UniGene, or GenomeScan map to ascertain the number and location of polymorphisms in a region.</p>
                <p>Some basic map controls are available directly on the display including removal of a map from the display by clicking on the <named-content content-type="book-spec-emph">X</named-content> over the map and moving a secondary map to the Master Map position by clicking on the arrow next to the map label.</p>
                <p>The <named-content content-type="book-spec-emph">Maps&amp;Options</named-content> window provides advanced options to: (<italic>a</italic>) add a ruler to any map; (<italic>b</italic>) reset the page length to display more (or less) information; (<italic>c</italic>) define region to display by providing coordinates or marker name in <named-content content-type="book-spec-emph">Region Shown</named-content> boxes (also available directly on the Map Viewer display); (<italic>d</italic>) display direct connections between maps by checking the <named-content content-type="book-spec-emph">Show connections</named-content> box; (<italic>e</italic>) optionally view text in <named-content content-type="book-spec-emph">Verbose</named-content> or <named-content content-type="book-spec-emph">Condensed</named-content> mode by selecting the checkbox. These user-defined preferences will be maintained for additional queries on different regions or chromosomes, until reset.</p>
                <p>There has been considerable effort to integrate data on the sequence-based maps with data from non-sequence-based maps. Map connections provide a unique and powerful mechanism to identify features in a relevant region of the sequence map when starting with information from a different coordinate system (see <italic>Relationships among Coordinate Systems</italic>).</p>
                <p>The features that are available with Map Viewer are summarized in <xref ref-type="boxed-text" rid="bid.1598">Box 1</xref>.</p>

<boxed-text id="bid.1598" position="margin"><label>1</label><title>Map Viewer-associated functions.</title><sec id="bid.1599"><title/><p><bold>Query:</bold></p><p>&emsp;Text</p><p>&emsp;Text, advanced</p><p>&emsp;Nucleotide query (by alignment or Accession number)</p><p>&emsp;Protein query (by alignment)</p><p>&emsp;By position in genome</p><p><bold>Display Data:</bold></p><p>&emsp;Graphical</p><p>&emsp;Tabular</p><p>&emsp;Assembled sequence</p><p>&emsp;Annotated feature sequence</p><p><bold>Download:</bold></p><p>&emsp;Sequence region</p><p>&emsp;Other map data for region</p><p>&emsp;Custom model (Model Maker)</p><p><bold>Change Display Configuration:</bold></p><p>&emsp;Zoom</p><p>&emsp;Scroll along chromosome</p><p>&emsp;Add/Remove tracks/maps</p><p>&emsp;Scalebar (ruler)</p><p>&emsp;Change order of track/map</p><p>&emsp;Specify coordinates to view</p><p>&emsp;Jump to different chromosome</p><p>&emsp;Show links</p><p>&emsp;Alter number of rows displayed</p><p><bold>FTP:</bold></p><p>&emsp;Assembled sequence</p><p>&emsp;Model mRNA sequence</p><p>&emsp;Model protein sequence</p><p>&emsp;Contig/chromosome conversion tables</p><p>&emsp;Map location, sequence-based</p><p>&emsp;Map location, non-sequence-based</p><p><bold>Links:</bold></p><p>&emsp;Help documentation</p><p>&emsp;Statistics</p><p>&emsp;FAQs</p></sec></boxed-text>
              </sec>
            </sec>
            <sec id="bid.1612">
              <title>Associated Tools</title>
              <p>Map Viewer provides links to several tools to display, download, or manipulate the sequence in a user-defined region. Whenever a sequence-based map is the master (the one at the right), the link <named-content content-type="book-spec-emph">Download/View Sequence/Evidence</named-content> is provided above the map display. This opens a window that provides access to the <named-content content-type="book-spec-emph">seq</named-content>, <named-content content-type="book-spec-emph">ev</named-content>, and <named-content content-type="book-spec-emph">mm</named-content> tools described below. In addition, when the annotated object is a gene (sequence or cytogenetic maps) or the species-specific UniGene cluster, the label may include these links.</p>
              <p>The Evidence Viewer (<named-content content-type="book-spec-emph">ev</named-content>) displays graphically the GenBank and RefSeq cDNAs that align to the genome in a particular region, along with a density plot for ESTs. The positions of any mismatches or insertions/deletions are marked, the multiple pairwise sequence alignments are provided, and computed translations are shown.</p>
              <p>The Sequence Viewer (<named-content content-type="book-spec-emph">sv</named-content>) is the Entrez graphical display option for any nucleotide sequence, focused on the gene indicated. By default, a 2-kb section of sequence is shown below the representation of the features, but that limit can be increased at the bottom of the page. It is also possible to zoom and navigate in the display.</p>
              <p>Sequence Download (<named-content content-type="book-spec-emph">seq</named-content>) provides the same function as the <named-content content-type="book-spec-emph">Download/View Sequence</named-content> link provided at the top of the Maps page. The scope of the sequence passed to the tool corresponds to what is being viewed on the page. When connected to a gene feature, the scope corresponds to that gene. The tool allows the user to alter the sequence scope and to select a report format (e.g., FASTA, GenBank, ASN.1). For the human and mouse genomes, a link is also provided to the Human&ndash;Mouse Homology Map (<named-content content-type="book-spec-emph">hm</named-content>).</p>
              <p>Model Maker (<named-content content-type="book-spec-emph">mm</named-content>) displays the evidence for exons in a genomic region by diagramming the exons predicted from the alignment of cDNAs, from <italic>ab initio</italic> models (the default), and from alignment of ESTs (after an explicit selection). To facilitate construction of your own model transcript or transcripts, the splice junctions and the exons they connect are displayed, and the coding potential of any combination of exons can quickly be evaluated using ORFfinder. The sequence can also be edited, and the results can be saved or downloaded.</p>
            </sec>
            <sec id="bid.1613">
              <title>Technical Details</title>
              <sec id="bid.1614">
                <title>Data Access</title>
                <p>The data displayed in Map Viewer are freely available. In addition to the view-specific reports, all of the data are available by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/genomes/[genome_name]/maps/mapview/">FTP</ext-link>. README files document the content and format of each file. Genomic data are also available by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp.ncbi.nih.gov/genomes/[genome name]/">chromosome</ext-link>; this includes genomic contigs (NT_ or NW_ Accession numbers) built from finished and unfinished sequence data. The contig data are available in various formats, including ASN.1, FASTA, GenBank, and GenPept. Also available in this directory are the RNAs (NM_, XM_ , and XR_ Accession numbers) and proteins (NP_, XP_).</p>
              </sec>
              <sec id="bid.1615">
                <title>Constructing URLs to Generate Specific Displays</title>
                <boxed-text id="bid.1616" position="margin">
                  <label>2</label>
                  <title>Examples of URL construction.</title>
                  <p><bold>(a)</bold> Find the neighborhood (zoom=2) of the <italic>HIRA</italic> gene (chromosome 22) on all gene-containing human maps plus an ideogram, with the sequence map (loc) as the master map. Provide the detailed description of the genes (verbose=on). URL: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=hum&amp;chr=22&amp;query=HIRA&amp;zoom=2&amp;maps=ideogr,morbid,gene,loc&amp;verbose=on">http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=hum&amp;chr=22&amp;query=HIRA&amp;zoom=2&amp;maps=ideogr,morbid,gene,loc&amp;verbose=on</ext-link></p>
                  <p><bold>(b)</bold> Find human FISH-mapped clones (fish) in a cytogenetic region (coordinates are added to define the region) and also on the sequence map (clone). URL: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=hum&amp;chr=1&amp;maps=clone,fish[1pter-p31]">http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=hum&amp;chr=1&amp;maps=clone,fish[1pter-p31]</ext-link></p>
                  <p><bold>(c)</bold> Show comparable regions on the human contig (cntg), component (comp), gene (gene,loc), and STS (sts) maps between the markers D7S726 and D7S2686. Show the ruler for the STS map (-r) and highlight the query terms. URL: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=hum&amp;chr=7&amp;maps=cntg,comp,gene,loc,sts[D7S726:D7S2686]-r&amp;query=D7S726+D7S2686">http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=hum&amp;chr=7&amp;maps=cntg,comp,gene,loc,sts[D7S726:D7S2686]-r&amp;query=D7S726+D7S2686</ext-link></p>
                  <p><bold>(d)</bold> Show potential genes (RefSeqs, ESTs, GenomeScan models) on a human genomic contig (NT_), with corresponding GenBank Accession numbers used to build the contig (displayed on the comp map) and FISH-mapped clones (on the clone map). URL: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?ORG=hum&amp;gnl=NT_023567&amp;maps=cntg-r,clone,comp,scan,est,loc&amp;query=NT_023567&amp;cmd=focus">http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?ORG=hum&amp;gnl=NT_023567&amp;maps=cntg-r,clone,comp,scan,est,loc&amp;query=NT_023567&amp;cmd=focus</ext-link></p>
                  <p><bold>(e)</bold> Display mouse chromosome 6 on the radiation hybrid (rh) and genetic (mgi, wigen) maps and highlight the query term, D6Mit113. Zoom into 30% of the chromosome, with A2m in the center of that region. URL: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=mouse&amp;chr=6&amp;maps=rh,mgi,wigen&amp;query=D6Mit113&amp;zoom=30">http://www.ncbi.nlm.nih.gov/mapview/maps.cgi?org=mouse&amp;chr=6&amp;maps=rh,mgi,wigen&amp;query=D6Mit113&amp;zoom=30</ext-link></p>
                </boxed-text>
                <p>Dynamic links to Map Viewer can be generated by constructing URLs with arguments that define the species, chromosome, range, types of maps (with or without units), display order, number of labels, query string, how to center a display around a query result, and the type of label for the display. The most current documentation is provided in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/MapViewerHelp.html#URLs">online help</ext-link>. The examples in <xref ref-type="boxed-text" rid="bid.1616">Box 2</xref>, however, may illustrate the flexibility of the approach. Please note that the argument of the map in the URL is processed as an ordered list, with the order in the list controlling the left-to-right order in the display. Additional qualifiers control the display of a ruler and the range on the chromosome. If a query term is included as a part of the URL and that value cannot be identified on any of the maps in the list, that map will not be displayed.</p>
              </sec>
              <sec id="bid.1617">
                <title>Implementation</title>
                <p>Query terms are indexed for retrieval using the Entrez system. Thus, wild cards, Boolean operators, filters, and properties are managed as for other Entrez databases.</p>
                <p>Each distinct object on the map is assigned a unique identifier that is specific to a particular build. Each object may have other secondary identifiers, such as IDs, in the sequence, Clone Repository, dbSNP, LocusLink, UniGene, or UniSTS databases. All descriptors are indexed as text. In addition, some are indexed by specific field values or by pre-identified properties, such as genes with associated diseases, SNPs with heterozygosity values in pre-defined ranges, or evidence type for genes. These field names or properties can be applied to restrict a query either in the Web-based query form or within a URL. The complete listings of current implementations for field qualifiers and properties are provided in the online help documentation.</p>
                <p>Data for each map are retrieved for display from a relational database based on the IDs returned from the Entrez query. The database is used only to support display; it is refreshed with each NCBI build or update of any other map but not to track changes from build to build. Data from previous builds are archived at NCBI, but direct access is not currently supported.</p>
              </sec>
            </sec>
            <sec id="bid.1618">
              <title>Caveats for Using Evolving Data</title>
              <table-wrap id="bid.1623">
                <label>3</label>
                <caption>
                  <title>Web sites of interest.</title>
                </caption>
                <table>
                  <tr>
                    <td align="left" valign="top">
                      <bold>Map Viewers</bold>
                    </td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Ensembl</td>
                    <td align="left" valign="top">www.ensembl.org</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;NCBI MapViewer</td>
                    <td align="left" valign="top">www.ncbi.nlm.nih.gov/mapview</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;UCSC Genome Browser</td>
                    <td align="left" valign="top">www. genome.ucsc.edu</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Sequencing Information</td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;NHGRI Sequencing Information</td>
                    <td align="left" valign="top">www.nhgri.nih.gov/Data/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Celera Genomics</td>
                    <td align="left" valign="top">www.celera.com</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">
                      <bold>Analysis tools</bold>
                    </td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;BLAT</td>
                    <td align="left" valign="top">http://genome.ucsc.edu/cgi-bin/hgBlat?command=start</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;BLAST</td>
                    <td align="left" valign="top">http://www.ncbi.nlm.nih.gov/BLAST/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;e-PCR</td>
                    <td align="left" valign="top">http://www.ncbi.nlm.nih.gov/ sts/epcr.cgi</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Sim4 (mRNA to genomic alignment&nbsp;tool)</td>
                    <td align="left" valign="top">http://globin.cse.psu.edu/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Spidey (mRNA to genomic alignment&nbsp;tool)</td>
                    <td align="left" valign="top">http://www.ncbi.nlm.nih.gov/spidey</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;SSAHA</td>
                    <td align="left" valign="top">http://www.sanger.ac.uk/Software/analysis/SSAHA/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;RepeatMasker</td>
                    <td align="left" valign="top">http://repeatmasker.genome.washington.edu/cgi-bin/RepeatMasker</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Maps</td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;BAC FingerPrint Map</td>
                    <td align="left" valign="top">http://genome.wustl.edu/gsc/human/Mapping/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Other Annotation Sources and Viewers</td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Celera Genomics</td>
                    <td align="left" valign="top">http://www.celera.com</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;DAS</td>
                    <td align="left" valign="top">http://www.biodas.org</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;DoubleTwist</td>
                    <td align="left" valign="top">http://www.doubletwist.com</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;The Genome Channel</td>
                    <td align="left" valign="top">http://compbio.ornl.gov/channel/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Incyte Genomics</td>
                    <td align="left" valign="top">http://www.incyte.com</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;FTP Sites</td>
                    <td align="left" valign="top"/>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;Ensembl</td>
                    <td align="left" valign="top">ftp.ensembl.org/pub/current/data/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;NCBI</td>
                    <td align="left" valign="top">ftp.ncbi.nlm.nih.nih.gov/genomes/H_sapiens/</td>
                  </tr>
                  <tr>
                    <td align="left" valign="top">&emsp;UCSC</td>
                    <td align="left" valign="top">ftp.genome.cse.ucsc.edu/goldenPath</td>
                  </tr>
                </table>
              </table-wrap>
              <p>Map Viewer displays represent the current synthesis of information available at the time of the data freeze (<xref ref-type="table" rid="bid.1623">Table 3</xref>). It is important to understand that the underlying data may change from build to build, as our view of a genome becomes more refined. The data presented should always be critically reviewed, with a view to assessing the reliability of the assembly and annotation.</p>
              <p>Means of reviewing reliability include: (<italic>a</italic>) noting the color coding of the contigs according to whether the sequence is draft or finished (this primarily applies to the human sequence); (<italic>b</italic>) noting the descriptions of the genes, STS, or SNPs to determine whether the element has been placed more than once; (<italic>c</italic>) checking that the STS order is the same on different maps; and (<italic>d</italic>) viewing features from different coordinate systems on the same map, e.g., showing STS features on the sequence (nucleotide coordinates), RH (cRay coordinates), and genetic maps (centiMorgan coordinates) to check for ambiguities. For more information, see the Pipeline FAQ /genome/guide/BuildFAQ.html.</p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.1625">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                    </person-group>
                    <year>1998</year>
                    <article-title>Electronic PCR: bridging the gap between genome mapping and genome sequencing</article-title>
                    <source>Trends Biotechnol</source>
                    <volume>16</volume>
                    <issue>11</issue>
                    <fpage>456</fpage>
                    <lpage>459</lpage>
                    <pub-id pub-id-type="pmid">9830153</pub-id>
                  </citation>
                </ref>
                <ref id="bid.1626">
                  <label>2</label>
                  <citation>
                    <person-group/>
                    <collab>The BAC Resource Consortium</collab>
                    <year>2001</year>
                    <article-title>Integration of cytogenetic landmarks into the draft sequence of the human genome</article-title>
                    <source>Nature</source>
                    <volume>409</volume>
                    <fpage>953</fpage>
                    <lpage>958</lpage>
                    <pub-id pub-id-type="pmid">11237021</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.857" book-part-type="chapter" book-part-number="21">
          <book-part-meta>
            <title-group>
              <title>UniGene: A Unified View of the Transcriptome</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Pontius</surname>
                  <given-names>Joan U.</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Wagner</surname>
                  <given-names>Lukas</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Schuler</surname>
                  <given-names>Gregory D.</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch21d1"/>
            <abstract>
              <title>Summary</title>
              <p>The task of assembling an inventory of all genes of <italic>Homo sapiens</italic> and other organisms began more than a decade ago with large-scale survey sequencing of transcribed sequences. The resulting Expressed Sequence Tags (ESTs) were a gold mine of novel gene sequences that provided an infrastructure for additional large-scale projects, such as gene maps, expression systems, and full-length cDNA projects. In addition, untold numbers of targeted gene-hunting projects have benefited from the availability of these sequences and the physical clone reagents. However, the high level of redundancy found among transcribed sequences, not to mention a variety of common experimental artifacts, made it difficult for many people to make effective use of the data. This problem was the motivation for the development of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/">UniGene</ext-link>, a largely automated analytical system for producing an organized view of the transcriptome. In this chapter, we discuss the properties of the input sequences, the process by which they are analyzed in UniGene, and some pointers on how to use the resource.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.858">
              <title>Expressed Sequence Tags (ESTs)</title>
              <p>At a time when the genomes of many species have been sequenced completely, a fundamental resource expected by many researchers is a simple list of all of an organism's genes. A gene list, together with associated physical reagents and electronic information, allows one to begin to investigate the ways in which many genes interact in the complex system of the organism. However, many species of medical and agricultural importance have not yet been prioritized for genomic sequencing, and expressed cDNAs have provided the primary source of gene sequences. Furthermore, when the genomic sequence of an organism becomes available, a collection of cDNA sequences provides the best tool for identifying genes within the DNA sequence. Thus, we can anticipate that the sequencing of transcribed products will remain a significant area of interest well into the future.</p>
              <p>The era of high-throughput cDNA sequencing was initiated in 1991 by a landmark study from Venter and his colleagues (<xref ref-type="bibr" rid="bid.869">1</xref>). The basic strategy involves selecting cDNA clones at random and performing a single, automated, sequencing read from one or both ends of their inserts. They introduced the term EST to refer to this new class of sequence, which is characterized by being short (typically about 400&ndash;600 bases) and relatively inaccurate (around 2% error). The use of single-pass sequencing was an important aspect of making the approach cost effective. In most cases, there is no initial attempt to identify or characterize the clones. Instead, they are identified using only the small bit of sequence data obtained, comparing it to the sequences of known genes and other ESTs. It is fully expected that many clones will be redundant with others already sampled and that a smaller number will represent various sorts of contaminants or cloning artifacts. There is little point in incurring the expense of high-quality sequencing until later in the process, when clones can be validated and a non-redundant set selected.</p>
              <p>Despite their fragmentary and inaccurate nature, ESTs were found to be an invaluable resource for the discovery of new genes, particularly those involved in human disease processes (<xref ref-type="bibr" rid="bid.870">2</xref>, <xref ref-type="bibr" rid="bid.871">3</xref>). After the initial demonstration of the utility and cost effectiveness of the EST approach, many similar projects were initiated, resulting in an ever-increasing number of human ESTs (<xref ref-type="bibr" rid="bid.872">4</xref>&ndash;<xref ref-type="bibr" rid="bid.876">8</xref>). In addition, large-scale EST projects were launched for several other organisms of experimental interest. In 1992, a database called dbEST (<xref ref-type="bibr" rid="bid.877">9</xref>) was established to serve as a collection point for ESTs, which are then distributed to the scientific community as the EST division of GenBank (<xref ref-type="bibr" rid="bid.878">10</xref>). The EST division continues to dominate GenBank, accounting for roughly two-thirds of all submissions. The 20 organisms with the largest numbers of ESTs in the public database (as of March 7, 2002) are shown in <xref ref-type="table" rid="bid.859">Table 1</xref>.</p>
              <table-wrap id="bid.859"><label>1</label><caption><title>Top 20 organisms in dbEST (as of March 7, 2002).</title></caption><table><tr><th align="left" valign="top">Organism</th><th align="left" valign="top">ESTs</th></tr><tr><td align="left" valign="top"><italic>Homo sapiens</italic> (human)</td><td align="left" valign="top">4,070,035</td></tr><tr><td align="left" valign="top"><italic>Mus musculus</italic> (mouse)</td><td align="left" valign="top">2,522,776</td></tr><tr><td align="left" valign="top"><italic>Rattus norvegicus</italic> (rat)</td><td align="left" valign="top">326,707</td></tr><tr><td align="left" valign="top"><italic>Drosophila melanogaster</italic> (fruit fly)</td><td align="left" valign="top">255,456</td></tr><tr><td align="left" valign="top"><italic>Glycine max</italic> (soybean)</td><td align="left" valign="top">234,900</td></tr><tr><td align="left" valign="top"><italic>Bos taurus</italic> (cow)</td><td align="left" valign="top">230,256</td></tr><tr><td align="left" valign="top"><italic>Danio rerio</italic> (zebrafish)</td><td align="left" valign="top">197,630</td></tr><tr><td align="left" valign="top"><italic>Xenopus laevis</italic> (African clawed frog)</td><td align="left" valign="top">197,565</td></tr><tr><td align="left" valign="top"><italic>Caenorhabditis elegans</italic> (nematode)</td><td align="left" valign="top">191,268</td></tr><tr><td align="left" valign="top"><italic>Lycopersicon esculentum</italic> (tomato)</td><td align="left" valign="top">148,338</td></tr><tr><td align="left" valign="top"><italic>Zea mays</italic> (maize)</td><td align="left" valign="top">147,658</td></tr><tr><td align="left" valign="top"><italic>Medicago truncatula</italic> (barrel medic)</td><td align="left" valign="top">137,588</td></tr><tr><td align="left" valign="top"><italic>Arabidopsis thaliana</italic> (thale cress)</td><td align="left" valign="top">113,330</td></tr><tr><td align="left" valign="top"><italic>Chlamydomonas reinhardtii</italic></td><td align="left" valign="top">112,489</td></tr><tr><td align="left" valign="top"><italic>Hordeum vulgare</italic> (barley)</td><td align="left" valign="top">104,803</td></tr><tr><td align="left" valign="top"><italic>Oryza sativa</italic> (rice)</td><td align="left" valign="top">104,284</td></tr><tr><td align="left" valign="top"><italic>Sus scrofa</italic> (pig)</td><td align="left" valign="top">103,321</td></tr><tr><td align="left" valign="top"><italic>Anopheles gambiae</italic> (mosquito)</td><td align="left" valign="top">88,963</td></tr><tr><td align="left" valign="top"><italic>Ciona intestinalis</italic> (sea squirt)</td><td align="left" valign="top">88,742</td></tr><tr><td align="left" valign="top"><italic>Sorghum bicolor</italic> (sorghum)</td><td align="left" valign="top">84,712</td></tr></table></table-wrap>
              <p>One avenue to gene discovery is to use a database search tool, such as BLAST (<xref ref-type="bibr" rid="bid.879">11</xref>), to perform a sequence similarity search against dbEST. The query for such a search would be a gene or protein sequence, perhaps from a model organism, that is expected to be related to the human gene of interest. Because clone identifiers are carried with the sequence tags, it is possible to obtain the original material to generate a more accurate sequence or to use as an experimental reagent. For many EST projects, the IMAGE consortium (<xref ref-type="bibr" rid="bid.880">12</xref>) has been particularly instrumental in collecting the cDNA libraries, arraying the clones, and making the clones available for sequencing and redistribution.</p>
              <p>For EST sequencing to be maximally productive, certain details of the library construction require some attention. For example, normalization procedures have been used to reduce the abundance of highly expressed genes so as to favor the sampling of rarer transcripts (<xref ref-type="bibr" rid="bid.881">13</xref>). More recently, subtraction techniques have been used to construct libraries depleted of clones already subjected to EST sampling (<xref ref-type="bibr" rid="bid.882">14</xref>). Although these techniques make it more efficient to find transcripts that are at low abundance in a particular tissue, it is possible that a small number of genes will still be missed because they are simply not expressed in tissues, cell types, and developmental stages that have been sampled.</p>
              <p>Although ESTs are a useful way to identify clones of interest and provide guidance in identifying gene structure, a full-insert sequence of cDNA clones is preferable for both purposes. High-throughput full-insert cDNA sequencing projects have been the source of over 80,000 sequence submissions accessioned to date (August 2002). The full-insert cDNA sequence can allow identification of the translation product of the sequenced transcript, as well as potentially providing evidence for gene structure. Moreover, for the investigator wanting to use the clone as a reagent, having the accurate and complete sequence of the clone's insert at hand makes complete resequencing unnecessary, if the full-insert cDNA sequencing project makes clones available. Verifying that the full-insert sequence corresponds to either the complete transcript of interest or to its complete, uncorrupted coding sequence is possible without committing laboratory resources and time to a clone that produced an EST. cDNA libraries do not generally include the entire transcript sequence; therefore, many full-insert sequences do not contain the entire transcription unit. Large transcripts (&gt;6 kb) are particularly difficult to obtain.</p>
            </sec>
            <sec id="bid.860">
              <title>Sequence Clusters</title>
              <p>The sheer number of transcribed sequences is extraordinary, indeed for most organisms much larger than the number of genes. A major challenge is to make putative gene assignments for these sequences, recognizing that many of these genes will be anonymous, defined only by the sequences themselves. Computationally, this can be thought of as a clustering problem in which the sequences are vertices that may be coalesced into clusters by establishing connections among them.</p>
              <p>Experience has shown that it is important to eliminate low-quality or apparently artifactual sequences before clustering because even a small level of noise can have a large corrupting effect on a result. Thus, procedures are in place to eliminate sequences of foreign origin (most commonly <italic>Escherichia coli</italic>) and identify regions that are derived from the cloning vector or artificial primers or linkers. At present, UniGene focuses on protein-coding genes of the nuclear genome; therefore, those identified as rRNA or mitochondrial sequence are eliminated. Through the NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Traces/trace.cgi?">Trace Archive</ext-link>, an increasing number of EST sequences now have base-level error probabilities that are used to identify the highest quality segment of each sequence. Repetitive sequences sometimes lead to false alignments and must be treated with caution. Simple repeats (low-complexity regions) are identified using a word-overrepresentation algorithm called DUST, and transposable repetitive elements are identified by comparison with a library of known repeats for each organism. Rather than eliminating them outright, subsequences classified as repetitive are &ldquo;soft-masked&rdquo;, which is to say that they are not allowed to initiate a sequence alignment, although they may participate in one that is triggered within a unique sequence. For a sequence to be included in UniGene, the clone insert must have at least 100 base pairs that are of high quality and not repetitive.</p>
              <p>With a given a set of sequences, a variety of different sources of information may be used as evidence that any pair of them is or is not derived from the same gene. The most obvious type of relationship would be one in which the sequences overlap and can form a near-perfect sequence alignment. One dilemma is that some level of mismatching should be tolerated because of known levels of base substitution errors in ESTs, whereas allowing too much mismatching will cause highly similar paralogous genes to cluster together. One way to improve the results is to require that alignments show an approximate &ldquo;dovetail&rdquo; relationship, which is to say that they extend about as far to the ends of the sequences as possible. Values of specific parameters governing acceptable sequence alignments are chosen by examining ratios of true to false connections in curated test sets. It is important to note that the resulting clusters may contain more than one alternative-splice form.</p>
              <p>Multiple incomplete but non-overlapping fragments of the same gene are frequently recognized in hindsight when the gene's complete sequence is submitted. To minimize the frequency of multiple clusters being identified for a single gene, UniGene clusters are required to contain at least one sequence carrying readily identifiable evidence of having reached the 3&prime; terminus. In other words, UniGene clusters must be anchored at the 3&prime; end of a transcription unit. This evidence can be either a canonical polyadenylation signal (<xref ref-type="bibr" rid="bid.883">15</xref>) or the presence of a poly(A) tail on the transcript, or the presence of at least two ESTs labeled as having been generated using the 3&prime; sequencing primer. Because some clusters do not contain such evidence (typically, they are single ESTs), not all uncontaminated sequences in dbEST appear in UniGene clusters. Of course, alternatively spliced terminal 3&prime; exons will appear as distinct clusters until sequence that spans the distinct splice forms is submitted. With the availability of genome sequence, a more stringent test of 3&prime; anchoring is possible, because internal priming can be recognized. Clusters that satisfy this more-stringent requirement can be identified by adding the term &ldquo;has_end&rdquo; to any query. Specific query possibilities such as this one are listed under the rubric <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/query.cgi">Query Tips</ext-link> on the UniGene homepage.</p>
              <p>The UniGene Web site allows the user to view UniGene information on a per cluster, per sequence, or per library basis. Each UniGene Web page (<xref ref-type="fig" rid="bid.861">Figure 1</xref>) includes a header with a query bar and a sidebar providing links to related online resources. UniGene is also the basis for three other NCBI resources: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/ProtEST/">ProtEST</ext-link>, a facility for browsing protein similarities; Digital Differential Display (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/ddd.cgi?">DDD</ext-link>), for comparison of EST-based expression profiles; and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/HomoloGene/">HomoloGene</ext-link>, which provides information about putative homology relationships.<fig id="bid.861"><label>1</label><caption><title>Cluster view.</title><p>A Web view of the UniGene cluster representing the human serine proteinase inhibitor gene <italic>SERPINF2</italic> is shown.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch21f1" mime-subtype="gif"/></fig></p>
            </sec>
            <sec id="bid.862">
              <title>UniGene Cluster Browser</title>
              <p>The UniGene Cluster page summarizes the sequences in the cluster and a variety of derived information that may be used to infer the identity of the gene. <xref ref-type="fig" rid="bid.861">Figure 1</xref> shows an example of such a view for the human <italic>SERPINF2</italic> gene. When available, links are provided to a corresponding entry in other NCBI resources (e.g., LocusLink, OMIM) or external databases [e.g., Mouse Genome Informatics (MGI) at the Jackson Laboratory and the Zebrafish Information Network (ZFIN) at the University of Oregon]. Additional sections on the page provide protein similarities, mapping data, expression information, and lists of the clustered sequences.</p>
              <p>Possible protein products for the gene are suggested by providing protein similarities between one representative sequence from the cluster and protein sequences from eight selected model organisms. For each organism, the protein with the highest degree of sequence similarity to the nucleotide sequence is listed, with its title and GenBank Accession number. The sequence alignment is described using the percent identity and length of the aligned region. Also provided is a link to ProtEST, which summarizes the UniGene protein similarities on a per protein basis.</p>
              <p>The next section summarizes information on the inferred map position of the gene. In some cases, chromosome assignments can be drawn from other databases, such as OMIM or MGI. In other cases, radiation hybrid (RH) maps have been constructed using Sequence Tagged Site (STS) markers derived from ESTs. In these cases, the UniGene cluster can be associated with a marker in the UniSTS database, and a map position can be assigned from the RH map. More recently, map positions have been derived by alignment of the cDNA sequences to the finished or draft genomic sequences present in the NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview">MapViewer</ext-link>. For example, the <italic>SERPINF2</italic> gene in <xref ref-type="fig" rid="bid.861">Figure 1</xref> has a link to human chromosome 17 in the Map Viewer. The map is initially shown with a few selected tracks that are likely to be of interest, but others may be added by the user.</p>
              <p>Although ESTs are a poor probe of gene expression, both the total number of ESTs and the tissues from which they originated are often useful. Both of these are displayed in the cluster browser. The tissues are listed under Expression Information, which includes the tissue source of libraries of the component sequences and, for human, links to the SAGE resource. Moreover, if genomic sequence is available, the UniGene map view displays expression for each exon (more precisely, for each portion of genome similar to a transcript; because incompletely processed mRNAs are not unheard of, the presence of a transcript is insufficient to identify an exon).</p>
              <p>The component sequences of the cluster are listed, with a brief description of each one and a link to its UniGene Sequence page. The Sequence page provides more detailed information about the individual sequence, and in the case of ESTs, includes a link to its corresponding UniGene Library page. On the Cluster page, the EST clones that are considered by the Mammalian Gene Collection (MGC) project to be putatively full length are listed at the top, whereas others follow in order of their reported insert length. At the bottom of the UniGene Cluster page is an option for the user to download the sequences of the cluster in FASTA format.</p>
            </sec>
            <sec id="bid.863">
              <title>Protein Similarity Analysis</title>
              <p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/ProtEST/">ProtEST</ext-link> section of UniGene allows the user to explore precomputed protein similarities for the cDNA sequences found in a cluster. The BLASTX program has been used to compare each sequence in UniGene to selected protein sequences drawn from eight model organisms: <italic>Homo sapiens</italic>, <italic>Mus musculus</italic>, <italic>Rattus norvegicus</italic>, <italic>Danio rerio</italic>, <italic>Drosophila melanogaster</italic>, <italic>Caenorhabditis elegans</italic>, <italic>Arabidopsis thaliana</italic>, and <italic>Saccharomyces cerevisiae</italic>. These species were chosen as spanning a variety of taxonomic classes, as well as being well represented in the protein databases. To exclude proteins that are strictly conceptual translations and models, the proteins used in ProtEST are those originating from the RefSeq, SWISS-PROT, PIR, PDB, or PRF databases.</p>
              <p>The ProtEST Web site has three features: information describing the amino acid sequence; information describing the nucleotide&ndash;protein alignments; and the ability for the user to modify various display options. The sequence alignments in ProtEST are summarized in tabular form (<xref ref-type="fig" rid="bid.864">Figure 2</xref>). The first column is a schematic representation of the nucleotide&ndash;protein alignment. The width of the column represents the entire length of the protein, whereas the unaligned nucleotide sequence is represented as a thin gray line and the aligned region is represented as a thick magenta bar. The alignment representation is a hyperlink to the full alignment regenerated on-the-fly using BLAST. Other information in the table includes the frame and strand of the alignment, a link to the corresponding trace as provided in the NCBI Trace Archive, the UniGene cluster ID, the GenBank Accession number, and columns that describe the aligned region and percent identity.<fig id="bid.864"><label>2</label><caption><title>ProtEST view.</title><p>A view of protein similarities for the human <italic>SERPINF2</italic> gene, found by BLASTX searching of a selected subset of the protein database, is shown.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch21f2" mime-subtype="gif"/></fig></p>
              <p>To further refine the view, the sequence alignments in the table can be sorted by: (<italic>a</italic>) percent identity; (<italic>b</italic>) alignment length; (<italic>c</italic>) beginning coordinate of the alignment; (<italic>d</italic>) ending coordinate of the alignment; (<italic>e</italic>) UniGene cluster ID; or (<italic>f</italic>) GenBank Accession number. It is also possible to omit various rows of the table by restricting the display to a chosen organism or by choosing a cut-off value for the percent identity of the alignment and the length of the alignment.</p>
            </sec>
            <sec id="bid.865">
              <title>Digital Differential Display (DDD)</title>
              <p>DDD is a tool for comparing EST-based expression profiles among the various libraries, or pools of libraries, represented in UniGene. These comparisons allow the identification of those genes that differ among libraries of different tissues, making it possible to determine which genes may be contributing to a cell's unique characteristics, e.g., those that make a muscle cell different from a skin or liver cell. Along similar lines, DDD can be used to try to identify genes for which the expression levels differ between normal, premalignant, and cancerous tissues or different stages of embryonic development.</p>
              <p>As in UniGene, the DDD resource is organism specific and is available from the UniGene <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/">Web site</ext-link> for that organism. For those libraries that have sequences in UniGene, DDD lists the title and tissue source and provides a link to the UniGene Library page, which gives additional information about the library. From the libraries listed, the user can select two for comparison. DDD then displays those genes for which the frequency of the transcript is significantly different between the two libraries. The output includes, for each gene, the frequency of its transcript in each library and the title of the gene's corresponding UniGene cluster. Results are sorted by significance, with the genes having the largest differences in frequencies displayed at the top. Libraries can be added sequentially to the analysis, and DDD will perform an analysis on each possible library&ndash;gene pair combination. Similarly, groups of libraries can be pooled together and compared with other pools or single libraries.</p>
              <p>DDD uses the Fisher Exact test to restrict the output to statistically significant differences (<italic>P</italic> &le;  0.05). The analysis is also restricted to deeply sequenced libraries; only those with over 1000 sequences in UniGene are included in DDD. These requirements place limitations on the capabilities of the analysis. Unless there are a large number of sequences in each pool, the frequencies of genes are generally not found to be statistically significant. Furthermore, the wide variety of tissue types, cell types, histology, and methods of generating the libraries can make it difficult to attribute significant differences to any one aspect of the libraries. These issues underscore the need for more libraries to be made public and the need for the comparisons to be made using proper controls. Libaries generated by the Cancer Genome Anatomy Project (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://cgap.nci.nih.gov/">CGAP</ext-link>) will become especially valuable to this end. This project has resulted in a plethora of human libraries made from a variety of tissue types and generated using a variety of methods.</p>
            </sec>
            <sec id="bid.866">
              <title>HomoloGene</title>
              <p>HomoloGene is a resource for exploring putative homology relationships among genes, bringing together curated homology information and results from automated sequence comparisons. UniGene clusters, supplemented by data from genome sequencing projects, have been used as a source of gene sequences for automated comparisons.</p>
              <p>Homology relationships, according to the experts who judge these, have been obtained from several sources. Collaborations with MGI and ZFIN at the University of Oregon have provided a large body of literature-derived data centered around <italic>M. musculus</italic> and <italic>D. rerio</italic>, respectively. Ortholog pairs involving sequences from <italic>H. sapiens</italic> and <italic>M. musculus</italic> have been imported from the NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Homology/">Human&ndash;Mouse Homology Map</ext-link>. Additional information has been extracted from the literature by NCBI staff specifically for the HomoloGene project.</p>
              <p>MegaBLAST (<xref ref-type="bibr" rid="bid.884">16</xref>) is used to perform cross-species sequence alignments and to identify those sequence pairs that share high degrees of nucleotide similarity. For each sequence, its best alignment with the sequences of the other organisms is retained. However, the best match for a sequence is not necessarily the best match for its partner sequence. For example, if there are several more sequences representing a particular gene in one organism than in the other organism, several sequences in one organism might have the same best match in the less well-represented organism. Similarly, if there are several paralogous genes in one species, they may find one identical homologous gene in another species. HomoloGene discriminates &ldquo;one-way best matches&rdquo; from cases where two sequences are each other's best match, or &ldquo;reciprocal best matches&rdquo;, and only these reciprocal best matches are used. These sequence pairs are then used to find cross-species homologies between UniGene clusters. When reciprocal best matches are consistent between three or more organisms, the pair is described as being part of a &ldquo;consistent triplet&rdquo;.</p>
              <p>The connections made by these methods result in a complex web of relationships. To simplify the Web view, it is useful to have each report page focus on an individual gene, called the &ldquo;key gene&rdquo;, and to show connections that follow from it. An example of the report for the <italic>M. musculus Serpinf2</italic> gene is shown in <xref ref-type="fig" rid="bid.867">Figure 3</xref>. The title of this key gene is shown at the top of the page, followed by genes from other species that show reciprocal best match relationships to the key gene. Each of these may have hypertext links to provide additional biological information about the gene. This is followed by a section providing the curated homology information (if any), with links to the source of the data. Reciprocal best-match relationships are listed in the next two sections, first those directly involving the key gene and then those from a second round of walking that may be of interest. In each case, the description includes the sequence identifiers and percent identity of the alignment, with a hyperlink to reproduce a full alignment using BLAST.<fig id="bid.867"><label>3</label><caption><title>HomoloGene view.</title><p>Homology information for the mouse <italic>Serpinf2</italic> gene, with curated homologies for mouse and computed homologies extending to rat, zebrafish, and cow, is shown.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch21f3" mime-subtype="gif"/></fig></p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.869">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Adams</surname>
                        <given-names>MD</given-names>
                      </name>
                      <name>
                        <surname>Kelley</surname>
                        <given-names>JM</given-names>
                      </name>
                      <name>
                        <surname>Gocayne</surname>
                        <given-names>JD</given-names>
                      </name>
                      <name>
                        <surname>Dubnick</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Polymeropoulos</surname>
                        <given-names>MH</given-names>
                      </name>
                      <name>
                        <surname>Xiao</surname>
                        <given-names>H</given-names>
                      </name>
                      <name>
                        <surname>Merril</surname>
                        <given-names>CR</given-names>
                      </name>
                      <name>
                        <surname>Wu</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Olde</surname>
                        <given-names>RF</given-names>
                      </name>
                      <name>
                        <surname>Moreno</surname>
                        <given-names>RF</given-names>
                      </name>
                    </person-group>
                    <year>1991</year>
                    <article-title>Complementary DNA sequencing: expressed sequence tags and human genome project</article-title>
                    <source>Science</source>
                    <volume>252</volume>
                    <issue>5013</issue>
                    <fpage>1651</fpage>
                    <lpage>1656</lpage>
                    <pub-id pub-id-type="pmid">2047873</pub-id>
                  </citation>
                </ref>
                <ref id="bid.870">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Sikela</surname>
                        <given-names>JM</given-names>
                      </name>
                      <name>
                        <surname>Auffray</surname>
                        <given-names>C</given-names>
                      </name>
                    </person-group>
                    <year>1993</year>
                    <article-title>Finding new genes faster than ever</article-title>
                    <source>Nat Genet</source>
                    <volume>3</volume>
                    <issue>3</issue>
                    <fpage>189</fpage>
                    <lpage>191</lpage>
                    <pub-id pub-id-type="pmid">8485571</pub-id>
                  </citation>
                </ref>
                <ref id="bid.871">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Boguski</surname>
                        <given-names>MS</given-names>
                      </name>
                      <name>
                        <surname>Tolstoshev</surname>
                        <given-names>CM</given-names>
                      </name>
                      <name>
                        <surname>Bassett</surname>
                        <given-names>DE Jr</given-names>
                      </name>
                    </person-group>
                    <year>1994</year>
                    <article-title>Gene discovery in dbEST</article-title>
                    <source>Science</source>
                    <volume>265</volume>
                    <issue>5181</issue>
                    <fpage>1993</fpage>
                    <lpage>1994</lpage>
                    <pub-id pub-id-type="pmid">8091218</pub-id>
                  </citation>
                </ref>
                <ref id="bid.872">
                  <label>4</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Okubo</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Hori</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Matoba</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Niiyama</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Fukushima</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Kojima</surname>
                        <given-names>Y</given-names>
                      </name>
                      <name>
                        <surname>Matsuba</surname>
                        <given-names>K</given-names>
                      </name>
                    </person-group>
                    <year>1992</year>
                    <article-title>Large scale cDNA sequencing for analysis of quantitative and qualitative aspects of gene expression</article-title>
                    <source>Nat Genet</source>
                    <volume>2</volume>
                    <issue>3</issue>
                    <fpage>173</fpage>
                    <lpage>179</lpage>
                    <pub-id pub-id-type="pmid">1345164</pub-id>
                  </citation>
                </ref>
                <ref id="bid.873">
                  <label>5</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Adams</surname>
                        <given-names>MD</given-names>
                      </name>
                      <name>
                        <surname>Soares</surname>
                        <given-names>MB</given-names>
                      </name>
                      <name>
                        <surname>Kerlavage</surname>
                        <given-names>AR</given-names>
                      </name>
                      <name>
                        <surname>Fields</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Venter</surname>
                        <given-names>JC</given-names>
                      </name>
                    </person-group>
                    <year>1993</year>
                    <article-title>Rapid cDNA sequencing (expressed sequence tags) from a directionally cloned human infant brain cDNA library</article-title>
                    <source>Nat Genet</source>
                    <volume>4</volume>
                    <issue>4</issue>
                    <fpage>373</fpage>
                    <lpage>380</lpage>
                    <pub-id pub-id-type="pmid">8401585</pub-id>
                  </citation>
                </ref>
                <ref id="bid.874">
                  <label>6</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Houlgatte</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Mariage-Samson</surname>
                        <given-names>R</given-names>
                      </name>
                      <name>
                        <surname>Duprat</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Tessier</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Bentolia</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Larry</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Auffray</surname>
                        <given-names>C</given-names>
                      </name>
                    </person-group>
                    <year>1995</year>
                    <article-title>The Genexpress Index: a resource for gene discovery and the genic map of the human genome</article-title>
                    <source>Genome Res</source>
                    <volume>5</volume>
                    <issue>3</issue>
                    <fpage>272</fpage>
                    <lpage>304</lpage>
                    <pub-id pub-id-type="pmid">8593614</pub-id>
                  </citation>
                </ref>
                <ref id="bid.875">
                  <label>7</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Hillier</surname>
                        <given-names>LD</given-names>
                      </name>
                      <name>
                        <surname>Lennon</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Becker</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Bonaldo</surname>
                        <given-names>MF</given-names>
                      </name>
                      <name>
                        <surname>Chiapelli</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Chissoe</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Dietrich</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>DuBuque</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Favello</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Gish</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Hawkins</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Hultman</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Kucaba</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Lacy</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Le</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Le</surname>
                        <given-names>N</given-names>
                      </name>
                      <name>
                        <surname>Mardis</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Moore</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Parsons</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Prange</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Rifkin</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Rohlfing</surname>
                        <given-names>T</given-names>
                      </name>
                      <name>
                        <surname>Schellenberg</surname>
                        <given-names>K</given-names>
                      </name>
                      <name>
                        <surname>Marra</surname>
                        <given-names>M</given-names>
                      </name>
                    </person-group>
                    <year>1996</year>
                    <article-title>Generation and analysis of 280,000 human expressed sequence tags</article-title>
                    <source>Genome Res</source>
                    <volume>6</volume>
                    <issue>9</issue>
                    <fpage>807</fpage>
                    <lpage>828</lpage>
                    <pub-id pub-id-type="pmid">8889549</pub-id>
                  </citation>
                </ref>
                <ref id="bid.876">
                  <label>8</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Krizman</surname>
                        <given-names>DB</given-names>
                      </name>
                      <name>
                        <surname>Wagner</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Lash</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Strausberg</surname>
                        <given-names>RL</given-names>
                      </name>
                      <name>
                        <surname>Emmert-Buck</surname>
                        <given-names>MR</given-names>
                      </name>
                    </person-group>
                    <year>1999</year>
                    <article-title>The Cancer Genome Anatomy Project: EST sequencing and the genetics of cancer progression</article-title>
                    <source>Neoplasia</source>
                    <volume>1</volume>
                    <issue>2</issue>
                    <fpage>101</fpage>
                    <lpage>106</lpage>
                    <pub-id pub-id-type="pmid">10933042</pub-id>
                  </citation>
                </ref>
                <ref id="bid.877">
                  <label>9</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Boguski</surname>
                        <given-names>MS</given-names>
                      </name>
                      <name>
                        <surname>Lowe</surname>
                        <given-names>TM</given-names>
                      </name>
                      <name>
                        <surname>Tolstoshev</surname>
                        <given-names>CM</given-names>
                      </name>
                    </person-group>
                    <year>1993</year>
                    <article-title>dbEST: database for &ldquo;expressed sequence tags&rdquo;</article-title>
                    <source>Nature Genet</source>
                    <volume>4</volume>
                    <fpage>332</fpage>
                    <lpage>333</lpage>
                    <pub-id pub-id-type="pmid">8401577</pub-id>
                  </citation>
                </ref>
                <ref id="bid.878">
                  <label>10</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Benson</surname>
                        <given-names>DA</given-names>
                      </name>
                      <name>
                        <surname>Karsch-Mizrachi</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                      <name>
                        <surname>Ostell</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Rapp</surname>
                        <given-names>BA</given-names>
                      </name>
                      <name>
                        <surname>Wheeler</surname>
                        <given-names>DL</given-names>
                      </name>
                    </person-group>
                    <year>2002</year>
                    <article-title>GenBank</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>30 </volume>
                    <issue>1</issue>
                    <fpage>17</fpage>
                    <lpage>20</lpage>
                    <pub-id pub-id-type="pmid">11752243</pub-id>
                    <pub-id pub-id-type="other">99127</pub-id>
                  </citation>
                </ref>
                <ref id="bid.879">
                  <label>11</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Altschul</surname>
                        <given-names>SF</given-names>
                      </name>
                      <name>
                        <surname>Madden</surname>
                        <given-names>TL</given-names>
                      </name>
                      <name>
                        <surname>Schaffer</surname>
                        <given-names>AA</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Miller</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                    </person-group>
                    <year>1997</year>
                    <article-title>Gapped BLAST and PSI-BLAST: a new generation of protein database search programs</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>25</volume>
                    <issue>17</issue>
                    <fpage>3389</fpage>
                    <lpage>3402</lpage>
                    <pub-id pub-id-type="pmid">9254694</pub-id>
                    <pub-id pub-id-type="other">146917</pub-id>
                  </citation>
                </ref>
                <ref id="bid.880">
                  <label>12</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Lennon</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Auffray</surname>
                        <given-names>C</given-names>
                      </name>
                      <name>
                        <surname>Polymeropoulos</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Soares</surname>
                        <given-names>MB</given-names>
                      </name>
                    </person-group>
                    <year>1996</year>
                    <article-title>The I.M.A.G.E. consortium: an integrated molecular analysis of genomes and their expression</article-title>
                    <source>Genomics</source>
                    <volume>33</volume>
                    <fpage>151</fpage>
                    <lpage>152</lpage>
                    <pub-id pub-id-type="pmid">8617505</pub-id>
                  </citation>
                </ref>
                <ref id="bid.881">
                  <label>13</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Soares</surname>
                        <given-names>MB</given-names>
                      </name>
                      <name>
                        <surname>Bonaldo</surname>
                        <given-names>MF</given-names>
                      </name>
                      <name>
                        <surname>Jelene</surname>
                        <given-names>P</given-names>
                      </name>
                      <name>
                        <surname>Su</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Lawton</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Efstratiadis</surname>
                        <given-names>A</given-names>
                      </name>
                    </person-group>
                    <year>1994</year>
                    <article-title>Construction and characterization of a normalized cDNA library</article-title>
                    <source>Proc Natl Acad Sci U S A</source>
                    <volume>91</volume>
                    <issue>20</issue>
                    <fpage>9228</fpage>
                    <lpage>9232</lpage>
                    <pub-id pub-id-type="pmid">7937745</pub-id>
                    <pub-id pub-id-type="other">44785</pub-id>
                  </citation>
                </ref>
                <ref id="bid.882">
                  <label>14</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Bonaldo</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Lennon</surname>
                        <given-names>G</given-names>
                      </name>
                      <name>
                        <surname>Soares</surname>
                        <given-names>MB</given-names>
                      </name>
                    </person-group>
                    <year>1996</year>
                    <article-title>Normalization and subtraction: two approaches to facilitate gene discovery</article-title>
                    <source>Genome Res</source>
                    <volume>6</volume>
                    <fpage>791</fpage>
                    <lpage>806</lpage>
                    <pub-id pub-id-type="pmid">8889548</pub-id>
                  </citation>
                </ref>
                <ref id="bid.883">
                  <label>15</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Wahle</surname>
                        <given-names>E</given-names>
                      </name>
                      <name>
                        <surname>Keller</surname>
                        <given-names>W</given-names>
                      </name>
                    </person-group>
                    <year>1996</year>
                    <article-title>The biochemistry of polyadenylation</article-title>
                    <source>Trends Biochem Sci</source>
                    <volume>21</volume>
                    <issue>7</issue>
                    <fpage>247</fpage>
                    <lpage>250</lpage>
                    <pub-id pub-id-type="pmid">8755245</pub-id>
                  </citation>
                </ref>
                <ref id="bid.884">
                  <label>16</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Schwartz</surname>
                        <given-names>S</given-names>
                      </name>
                      <name>
                        <surname>Wagner</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Miller</surname>
                        <given-names>W</given-names>
                      </name>
                    </person-group>
                    <year>2000</year>
                    <article-title>A greedy algorithm for aligning DNA sequences</article-title>
                    <source>J Comput Biol</source>
                    <volume>7</volume>
                    <issue>1-2</issue>
                    <fpage>203</fpage>
                    <lpage>214</lpage>
                    <pub-id pub-id-type="pmid">10890397</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.885" book-part-type="chapter" book-part-number="22">
          <book-part-meta>
            <title-group>
              <title>The Clusters of Orthologous Groups (COGs) Database: Phylogenetic Classification of Proteins from Complete Genomes</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Koonin</surname>
                  <given-names>Eugene V.</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch22d1"/>
            <abstract>
              <title>Summary</title>
              <p>The protein database of Clusters of Orthologous Groups (COGs) is an attempt to phylogenetically classify the complete complement of proteins (both predicted and characterized) encoded by complete genomes. Each COG is a group of three or more proteins that are inferred to be orthologs, i.e., they are direct evolutionary counterparts. The current release of the COGs database consists of 4,873 COGs, which include 136,711 proteins (~71% of all encoded proteins) from 50 bacterial genomes, 13 archaeal genomes, and 3 genomes of unicellular eukaryotes, the yeasts <italic>Saccharomyces cerevisiae</italic> and <italic>Schizosaccharomyces pombe</italic>, and the microsporidian <italic>Encephalitozoon cuniculi</italic>. The COG database is updated periodically as new genomes become available. The COGs for complete eukaryotic genomes are in preparation. The COGs can be applied to the task of functional annotation of newly sequenced genomes by using the COGnitor program, which is available on the COGs <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/COG/">homepage</ext-link>.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.886">
              <title>Introduction</title>
              <p>The recent progress in genome sequencing has led to a rapid enrichment of protein databases with an unprecedented variety of deduced protein sequences, most of them without a documented functional role. Computational biology strives to extract the maximal possible information from these sequences by classifying them according to their <xref ref-type="other" rid="bid.1310">homologous</xref> relationships, predicting their likely biochemical activities and/or cellular functions, three-dimensional structures, and evolutionary origin. This challenge is daunting, given that even in <italic>Escherichia coli</italic>, arguably the best-studied organism, only about 40% of the gene products have been characterized experimentally. However, computational analysis of complete microbial genomes has shown that prokaryotic proteins are, in general, highly conserved, with about 70% of them containing ancient conserved regions shared by homologs from distantly related species. This allows one to use functional information from experimentally characterized proteins to suggest function in their homologs from poorly studied organisms. For such functional predictions to be reliable, it is critical to infer orthologous relationships between genes from different species. <xref ref-type="other" rid="bid.1362">Ortholog</xref>s are evolutionary counterparts related by vertical descent (i.e., they have evolved from a common ancestor) as opposed to <!--<glossref>-->paralogs<!--</glossref>-->, which are genes related by duplication (<xref ref-type="bibr" rid="bid.902">1</xref>, <xref ref-type="bibr" rid="bid.903">2</xref>). Typically, orthologous proteins have the same domain architecture and the same function, although there are significant exceptions and complications to this generalization, particularly among multicellular eukaryotes.</p>
              <p>The COGs database has been designed as an attempt to classify proteins from completely sequenced genomes on the basis of the orthology concept (<xref ref-type="bibr" rid="bid.904">3</xref>&ndash;<xref ref-type="bibr" rid="bid.906">5</xref>). The COGs reflect one-to-many and many-to-many orthologous relationships as well as simple one-to-one relationships (hence, orthologous groups of proteins). In addition to the classification itself, the COGs Web site includes the COGnitor program, which assigns proteins from newly sequenced genomes to COGs that already exist and to several functionalities that allow the user to select and analyze various subsets of COGs.</p>
            </sec>
            <sec id="bid.887">
              <title>Construction of the COGs</title>
              <p>COGs have been identified on the basis of an all-against-all sequence comparison of the proteins encoded in complete genomes using the gapped BLAST program (<xref ref-type="bibr" rid="bid.907">6</xref>; see also <xref ref-type="other" rid="bid.610">Chapter 16</xref>) after masking low-complexity (<xref ref-type="bibr" rid="bid.908">7</xref>) and predicted coiled-coil (<xref ref-type="bibr" rid="bid.909">8</xref>) regions. The COG construction procedure is based on the simple notion that any group of at least three proteins from distant genomes that are more similar to each other than they are to any other proteins from the same genomes are most likely to form an orthologous set (<xref ref-type="fig" rid="bid.888">Figure 1</xref>). This prediction holds even if the absolute level of sequence similarity between the proteins in question is relatively low, and thus the COG approach accommodates both slow-evolving and fast-evolving genes.<fig id="bid.888"><label>1</label><caption><title>Example of a COG: monoamine oxidase.</title><p>The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/palox?COG1231">COG</ext-link> for monoamine oxidase currently contains eight proteins from seven different organisms: one each from <italic>Deinococcus radiodurans</italic> (DRA0274), <italic>Mycobacterium tuberculosis</italic> (Rv3170), <italic>Bacillus subtilis</italic> (BS_yobN), <italic>Synechocystis</italic> (slr0782), <italic>Pseudomonas aeruginosa</italic> (PA0421), and <italic>Mesorhizobium loti</italic> (mll3668), and two paralogs from <italic>Caulobacter crescentus</italic> (CC2793 and CC1091). This is the only COG in the COGs database that has this <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/phylox?A=0&amp;O=0&amp;M=0&amp;P=0&amp;K=0&amp;Z=0&amp;Y=0&amp;Q=0&amp;V=0&amp;D=1&amp;R=1&amp;B=1&amp;L=0&amp;C=1&amp;E=0&amp;F=1&amp;G=0&amp;H=0&amp;S=0&amp;N=0&amp;U=0&amp;J=1&amp;X=0&amp;I=0&amp;T=0&amp;W=0">phylogenetic pattern</ext-link>. In humans, monoamine oxidase is an enzyme of the mitochondrial outer membrane that seems to be involved in the metabolism of antibiotics and neurologically active agents and is a target for one class of antidepressant drugs.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch22f1" mime-subtype="gif"/></fig></p>
              <p>Briefly, COG construction includes the following steps:
                <list list-type="arabic">
                  <list-item id="bid.889">
                    <label>1</label>
                    <p>Perform the all-against-all protein sequence comparison.</p>
                  </list-item>
                  <list-item id="bid.890">
                    <label>2</label>
                    <p>Detect and collapse obvious paralogs, i.e., proteins from the same genome that are more similar to each other than to any proteins from other species.</p>
                  </list-item>
                  <list-item id="bid.891">
                    <label>3</label>
                    <p>Detect triangles of mutually consistent, genome-specific best hits (BeTs), taking into account the paralogous groups detected at step 2.</p>
                  </list-item>
                  <list-item id="bid.892">
                    <label>4</label>
                    <p>Merge triangles with a common side to form COGs.</p>
                  </list-item>
                  <list-item id="bid.893">
                    <label>5</label>
                    <p>Perform a case-by-case analysis of each COG. This analysis serves to eliminate false-positives and to identify groups that contain multidomain proteins by examining the pictorial representation of the BLAST search outputs. The sequences of detected multidomain proteins are split into single-domain segments, and steps 1&ndash;4 are repeated with the resulting shorter sequences, which assigns individual domains to COGs in accordance with their distinct evolutionary affinities.</p>
                  </list-item>
                  <list-item id="bid.894">
                    <label>6</label>
                    <p>Examine large COGs that include multiple members from all or several of the genomes using phylogenetic trees, cluster analysis, and visual inspection of alignments. As a result, some of these groups are split into two or more smaller ones that are included in the final set of COGs.</p>
                  </list-item>
                </list>
              </p>
              <p>By the design of this procedure, a minimal COG includes three genes from distinct phylogenetic lineages; protein sets from closely related species were merged before COG construction. The approach used for the construction of COGs does not supplant a comprehensive phylogenetic analysis. Nevertheless, it provides a fast and convenient shortcut to delineate a large number of families that most likely consist of orthologs.</p>
              <sec id="bid.895">
                <title>The COGnitor Program</title>
                <p>New proteins can be assigned to the COGs using the COGnitor program, the principal tool associated with the COGs database. COGnitor &ldquo;BLASTs&rdquo; the query sequences against all protein sequences encoded in the genomes that are classified in the current release of the COG system. To assign proteins to COGs, COGnitor applies the same principle that is embedded in the COG construction procedure, i.e., the consistency of genome-specific BeTs. For any given query protein, if the number of BeTs for a particular COG exceeds a predefined cut-off (three by default; the cut-off value can be changed by the user), the query protein is assigned to that COG; in cases where there are more than three BeTs to two different COGs, an ambiguous result is reported.</p>
              </sec>
              <sec id="bid.896">
                <title>The Current State of the COGs Database, Updates, and Additional Classification of the COGs</title>
                <p>Once the COGs have been identified using the above procedure, new members can be added using the COGnitor program. The assignments are further checked and curated by hand to eliminate potential false-positives. It has been shown that 95&ndash;97% of the COGnitor assignments typically require no correction (<xref ref-type="bibr" rid="bid.910">9</xref>). Once the proteins from a new genome are assigned to the appropriate pre-existing COGs through this combination of COGnitor and manual refinement, the remaining proteins from this genome are compared to the proteins from non-COG proteins from previously available genomes, and an attempt is made to construct new COGs using the original procedure. In addition, when new sequences are added to an exisiting COG, the COG is examined for the possibility of a split (isolation of a new COG) by inspecting BLAST search outputs for all COG members and, in some cases, phylogenetic tree analysis. Thus, the number of COGs continuously grows through the construction of new COGs that typically include just a small number of species, whereas the number of proteins in the COG system increases primarily through the addition of new members to pre-existing COGs.</p>
                <p>In bacterial and archaeal genomes, approximately 70% of the proteins typically belong to the COGs. Because each COG includes proteins from at least three distantly related species, this reveals the generally high level of evolutionary conservation of protein sequences, making the COGs a powerful tool for functional annotation of uncharacterized proteins. The COGs were classified into 18 functional categories that loosely follow those introduced by Riley (<xref ref-type="bibr" rid="bid.911">10</xref>) and also include a class for which only a general functional prediction (e.g., that of biochemical activity) was feasible, as well as a class of uncharacterized COGs. A significant majority of the COGs could be assigned to one of the well-defined functional categories, but the single largest class includes the functionally uncharacterized COGs. Additionally, the COGs were clustered according to the common metabolic pathways and macromolecular complexes.</p>
              </sec>
            </sec>
            <sec id="bid.897">
              <title>Phyletic Pattern Analysis in COGs</title>
              <p>A phyletic pattern is the pattern of species that are represented or not represented in a given COG; alternatively, phyletic patterns can be described in terms of the sets of COGs that are represented in a given range of species. The COGs show a broad diversity of phyletic patterns; only a small fraction are universal COGs, i.e., they are represented in all sequenced genomes, whereas COGs present in only three or four species are most abundant. This patchy distribution of phyletic patterns probably reflects the major role of horizontal gene transfer and lineage-specific gene loss in the evolution of prokaryotes, as well as the rapid evolution of certain genes in specific lineages, which may be linked to functional changes. Phyletic patterns are informative not only as indicators of probable evolutionary scenarios but also functionally; most often, different steps of the same pathway are associated with proteins that have the same phyletic pattern, whereas on some occasions, complementary patterns indicate that distinct (sometimes unrelated) proteins are responsible for the same function in different sets of species. The COG system includes a simple phyletic pattern search tool that allows the selection of COGs according to any given pattern of species. This tool effectively provides the functionality of &ldquo;differential genome display&rdquo; (for example, allowing the selection of all COGs that are present in one, but not the other, of a pair of genomes of interest) and can be helpful for delineating sets of candidate proteins for a particular range of functional features, e.g., virulence or hyperthermophily.</p>
            </sec>
            <sec id="bid.898">
              <title>Description of the COGs Website</title>
              <p>The main COGs <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/COG">Web page</ext-link> contains the following principal features: (<italic>a</italic>) a list of all COGs organized by the (predicted) <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/palox?fun=all">functional category</ext-link>; (<italic>b</italic>) separate lists of COGs for each functional category and for a variety of major pathways and functional <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/palox?sys=all">systems</ext-link>; (<italic>c</italic>) a table of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/palox?spec=all">co-occurrences</ext-link> of genomes in COGs; (<italic>d</italic>) a list of COGs organized by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/palox?pytall=a">phyletic patterns</ext-link>; (<italic>e</italic>) the phyletic patterns <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/COG/phylox.html">search tool</ext-link>; (<italic>f</italic>) the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/COG/xognitor.html">COGnitor</ext-link> program; (<italic>g</italic>) a search engine to search COGs for gene names, COG numbers, and arbitrary text; and (<italic>h</italic>) <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/COG/COGhelp.html#h_home_page">Help</ext-link>, which covers the principal subjects related to COGs.</p>
              <p>The individual COG pages can be reached from any of the COG lists mentioned above or by searching the site (see, for example, the COG for <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/palox?COG2925">exonuclease I</ext-link>). Each of the COG pages shows the respective phyletic pattern in a table that also gives the ID number for the contributing sequence(s), a cluster dendrogram generated using the BLAST scores as the measure of similarity between proteins, and a graphical representation of BeTs for the given COG (not shown for the largest COGs). Also, each of the COG pages is hyperlinked to: (<italic>a</italic>) pictorial representations of BLAST search outputs for each member of the COG, which also includes links to the respective GenBank and Entrez-Genomes entries (see, for example, the link from <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/blux?cog=XF2022&amp;XF2022">XF2022</ext-link>, the protein from <italic>Xyella fastidiosa</italic> in the exonuclease I COG); (<italic>b</italic>) a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/COG/aln/COG2925.aln">multiple alignment</ext-link> of the COG members produced automatically using the ClustalW program (<xref ref-type="bibr" rid="bid.912">11</xref>); (<italic>c</italic>) a FASTA library of the protein sequences that belong to the COG (represented by the floppy disc icon); (<italic>d</italic>) the respective functional category of COGs and pathway (functional system) if applicable (in this exonuclease I example, the functional category <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/palox?fun=L">L</ext-link> represents proteins involved in DNA replication, recombination, and repair); (<italic>e</italic>) a COG information page that includes functional, evolutionary, and structural information on the COG and its members (many of these pages are still under construction); (<italic>f</italic>) other COGs that include distinct domains of multidomain proteins that belong to the given COG through one of their domains; and (<italic>g</italic>) the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/cgi-bin/COG/coogtik?COG2925">Genome Context</ext-link> tool that shows the gene neighborhood around the given COG for all genomes that encode proteins of the given COG.</p>
              <p>The COG data set and the COGnitor program also are available by anonymous ftp at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/pub/COG">ftp://ftp.ncbi.nih.gov/pub/COG</ext-link>.</p>
            </sec>
            <sec id="bid.899">
              <title>Future Directions</title>
              <p>Substantial evolution of the COGs is expected in the near future in terms of both growth by adding more genomes and the addition of new functionalities and layers of presentation. Quantitatively, the main forthcoming addition is the COGs for eukaryotic genomes, which are expected to approximately double the size of the COG system. Many of the COGs include paralogous proteins, and this will be addressed by introducing hierarchical organization into the COG system, whereby related COGs will be unified at a higher level. In addition, partial integration of the COGs with the NCBI's Conserved Domains Database (CDD) is expected (<xref ref-type="other" rid="bid.94">Chapter 3</xref>), which will result in a more flexible and informative representation of the domain organization of proteins and of structural information that is available for COG members.</p>
            </sec>
            <sec id="bid.900">
              <title>The COG Team</title>
              <p>The COG system is developed and maintained by a team of programmers and expert biologists.</p>
              <p>Project leader: Eugene V. Koonin.</p>
              <p>The programming group: Roman L. Tatusov (group leader), Boris Kiryutin, Victor Smirnov, and Alexander Sverdlov (student)</p>
              <p>The annotation group: Darren A. Natale (group leader), Natalie Fedorova, Anastasia Nikolskaya, Aviva Jacobs, Jodie Yin, B. Sridhar Rao, Dmitri M. Krylov, Sergei Mekhedov, John Jackson, Raja Mazumder, and Sona Vasudevan</p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.902">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Fitch</surname>
                        <given-names>WM</given-names>
                      </name>
                    </person-group>
                    <year>1970</year>
                    <article-title>Distinguishing homologous from analogous proteins</article-title>
                    <source>Syst Zool</source>
                    <volume>19</volume>
                    <fpage>99</fpage>
                    <lpage>106</lpage>
                    <pub-id pub-id-type="pmid">5449325</pub-id>
                  </citation>
                </ref>
                <ref id="bid.903">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Fitch</surname>
                        <given-names>WM</given-names>
                      </name>
                    </person-group>
                    <year>2000</year>
                    <article-title>Homology: a personal view on some of the problems</article-title>
                    <source>Trends Genet</source>
                    <volume>16</volume>
                    <fpage>227</fpage>
                    <lpage>231</lpage>
                    <pub-id pub-id-type="pmid">10782117</pub-id>
                  </citation>
                </ref>
                <ref id="bid.904">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Tatusov</surname>
                        <given-names>RL</given-names>
                      </name>
                      <name>
                        <surname>Koonin</surname>
                        <given-names>EV</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                    </person-group>
                    <year>1997</year>
                    <article-title>A genomic perspective on protein families</article-title>
                    <source>Science</source>
                    <volume>278</volume>
                    <fpage>631</fpage>
                    <lpage>637</lpage>
                    <pub-id pub-id-type="pmid">9381173</pub-id>
                  </citation>
                </ref>
                <ref id="bid.905">
                  <label>4</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Tatusov</surname>
                        <given-names>RL</given-names>
                      </name>
                      <name>
                        <surname>Galperin</surname>
                        <given-names>MY</given-names>
                      </name>
                      <name>
                        <surname>Natale</surname>
                        <given-names>DA</given-names>
                      </name>
                      <name>
                        <surname>Koonin</surname>
                        <given-names>EV</given-names>
                      </name>
                    </person-group>
                    <year>2000</year>
                    <article-title>The COG database: a tool for genome-scale analysis of protein functions and evolution</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>28</volume>
                    <fpage>33</fpage>
                    <lpage>36</lpage>
                    <pub-id pub-id-type="pmid">10592175</pub-id>
                    <pub-id pub-id-type="other">102395</pub-id>
                  </citation>
                </ref>
                <ref id="bid.906">
                  <label>5</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Tatusov</surname>
                        <given-names>RL</given-names>
                      </name>
                      <name>
                        <surname>Natale</surname>
                        <given-names>DA</given-names>
                      </name>
                      <name>
                        <surname>Garkavtsev</surname>
                        <given-names>IV</given-names>
                      </name>
                      <name>
                        <surname>Tatusova</surname>
                        <given-names>TA</given-names>
                      </name>
                      <name>
                        <surname>Shankavaram</surname>
                        <given-names>UT</given-names>
                      </name>
                      <name>
                        <surname>Rao</surname>
                        <given-names>BS</given-names>
                      </name>
                      <name>
                        <surname>Kiryutin</surname>
                        <given-names>B</given-names>
                      </name>
                      <name>
                        <surname>Galperin</surname>
                        <given-names>MY</given-names>
                      </name>
                      <name>
                        <surname>Fedorova</surname>
                        <given-names>ND</given-names>
                      </name>
                      <name>
                        <surname>Koonin</surname>
                        <given-names>EV</given-names>
                      </name>
                    </person-group>
                    <year>2001</year>
                    <article-title>The COG database: new developments in phylogenetic classification of proteins from complete genomes</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>29</volume>
                    <fpage>22</fpage>
                    <lpage>28</lpage>
                    <pub-id pub-id-type="pmid">11125040</pub-id>
                    <pub-id pub-id-type="other">29819</pub-id>
                  </citation>
                </ref>
                <ref id="bid.907">
                  <label>6</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Altschul</surname>
                        <given-names>SF</given-names>
                      </name>
                      <name>
                        <surname>Madden</surname>
                        <given-names>TL</given-names>
                      </name>
                      <name>
                        <surname>Schaffer</surname>
                        <given-names>AA</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Zhang</surname>
                        <given-names>Z</given-names>
                      </name>
                      <name>
                        <surname>Miller</surname>
                        <given-names>W</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                    </person-group>
                    <year>1997</year>
                    <article-title>Gapped BLAST and PSI-BLAST: a new generation of protein database search programs</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>25</volume>
                    <fpage>3389</fpage>
                    <lpage>3402</lpage>
                    <pub-id pub-id-type="pmid">9254694</pub-id>
                    <pub-id pub-id-type="other">146917</pub-id>
                  </citation>
                </ref>
                <ref id="bid.908">
                  <label>7</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Wootton</surname>
                        <given-names>JC</given-names>
                      </name>
                      <name>
                        <surname>Federhen</surname>
                        <given-names>S</given-names>
                      </name>
                    </person-group>
                    <year>1996</year>
                    <article-title>Analysis of compositionally biased regions in sequence databases</article-title>
                    <source>Methods Enzymol</source>
                    <volume>266</volume>
                    <fpage>554</fpage>
                    <lpage>571</lpage>
                    <pub-id pub-id-type="pmid">8743706</pub-id>
                  </citation>
                </ref>
                <ref id="bid.909">
                  <label>8</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Lupas</surname>
                        <given-names>A</given-names>
                      </name>
                      <name>
                        <surname>Van Dyke</surname>
                        <given-names>M</given-names>
                      </name>
                      <name>
                        <surname>Stock</surname>
                        <given-names>J</given-names>
                      </name>
                    </person-group>
                    <year>1991</year>
                    <article-title>Predicting coiled coils from protein sequences</article-title>
                    <source>Science</source>
                    <volume>252</volume>
                    <fpage>1162</fpage>
                    <lpage>1164</lpage>
                    <pub-id pub-id-type="pmid">2031185</pub-id>
                  </citation>
                </ref>
                <ref id="bid.910">
                  <label>9</label>
                  <citation><person-group><name><surname>Natale</surname><given-names>DA</given-names></name><name><surname>Shankavaram</surname><given-names>UT</given-names></name><name><surname>Galperin</surname><given-names>MY</given-names></name><name><surname>Wolf</surname><given-names>YI</given-names></name><name><surname>Aravind</surname><given-names>L</given-names></name><name><surname>Koonin</surname><given-names>EV</given-names></name></person-group><article-title>Genome annotation using clusters of orthologous groups of proteins (COGs)&mdash;towards understanding the first genome of a Crenarchaeon</article-title><source>Genome Biol</source><italic>in press</italic>.
<pub-id pub-id-type="pmid">11178258</pub-id></citation>
                </ref>
                <ref id="bid.911">
                  <label>10</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Riley</surname>
                        <given-names>M</given-names>
                      </name>
                    </person-group>
                    <year>1993</year>
                    <article-title>Functions of the gene products of <italic>Escherichia coli</italic></article-title>
                    <source>Microbiol Rev</source>
                    <volume>57</volume>
                    <fpage>862</fpage>
                    <lpage>952</lpage>
                    <pub-id pub-id-type="pmid">7508076</pub-id>
                  </citation>
                </ref>
                <ref id="bid.912">
                  <label>11</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Thompson</surname>
                        <given-names>JD</given-names>
                      </name>
                      <name>
                        <surname>Higgins</surname>
                        <given-names>DG</given-names>
                      </name>
                      <name>
                        <surname>Gibson</surname>
                        <given-names>TJ</given-names>
                      </name>
                    </person-group>
                    <year>1994</year>
                    <article-title>CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>22</volume>
                    <fpage>4673</fpage>
                    <lpage>4680</lpage>
                    <pub-id pub-id-type="pmid">7984417</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
      </body>
    </book-part>
    <book-part id="bid.913" book-part-type="part" book-part-number="Part 4">
      <book-part-meta>
        <title-group>
          <title>User Support</title>
        </title-group>
      </book-part-meta>
      <body>
        <book-part id="bid.914" book-part-type="chapter" book-part-number="23">
          <book-part-meta>
            <title-group>
              <title>User Services: Helping You Find Your Way</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Wheeler</surname>
                  <given-names>David</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Rapp</surname>
                  <given-names>Barbara</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>9</day>
                <month>10</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch23d1"/>
            <abstract>
              <title>Summary</title>
              <p>The User Services team is the primary liaison between the public and the resources and data at NCBI. User Services disseminates information through outreach training programs and exhibits at scientific conferences and responds to incoming questions by email and telephone assistance. The team instructs people in the use of NCBI resources, responds to a wide range of questions, receives comments and suggestions, and coordinates with the NCBI resource developers to implement suggestions from users. In addition, User Services develops documentation, tutorials, and other support materials; produces the NCBI News; and publishes articles on NCBI resources.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.915">
              <title>The User Services Team</title>
              <p>User Services consists of a staff of scientists and information specialists with diverse backgrounds and experiences. Scientist members of the staff hold Masters or Ph.D. degrees in an area of molecular biology, biochemistry, or biotechnology. Information specialists have Masters degrees in Library and Information Science and extensive experience using online databases of scientific information.</p>
              <sec id="bid.916">
                <title>The Help Desk</title>
                <p>Help Desk assistance is available from 8:30 a.m. to 5:30 p.m. Eastern Time, Monday through Friday. Two email addresses are available, info@ncbi.nlm.nih.gov and blast-help@ncbi.nlm.nih.gov. The <email>info@ncbi.nlm.nih.gov</email> address is for any type of inquiry, including questions about services; how to get started using the NCBI tools for a particular research problem; reports of technical problems; press inquiries; and comments or suggestions about NCBI resources. The <email>blast-help@ncbi.nlm.nih.gov</email> address is for questions and comments regarding sequence similarity searches using BLAST tools and databases. The Help Desk phone number is 301-496-2475.</p>
                <p>email questions are answered as expeditiously as possible, usually within a day of receipt of the question. However, those that require extended investigation may take longer. Questions are usually handled directly by members of the User Services staff, although some are referred to a specific database development team for attention.</p>
                <p>Examples of question topics include: data submission protocols, including the use of <xref ref-type="other" rid="bid.1244">BankIt</xref> and <xref ref-type="other" rid="bid.1398">Sequin</xref> (<xref ref-type="other" rid="bid.538">Chapter 12</xref>); finding summary information about a gene or disease; using <xref ref-type="other" rid="bid.1282">Entrez</xref> (<xref ref-type="other" rid="bid.588">Chapter 15</xref>) to find sequence records for genes and proteins; choosing the best <xref ref-type="other" rid="bid.1246">BLAST</xref> service to use for a particular application (<xref ref-type="other" rid="bid.610">Chapter 16</xref>); how to interpret BLAST results; how to display and manipulate three-dimensional structures with <xref ref-type="other" rid="bid.1264">Cn3D</xref> (<xref ref-type="other" rid="bid.94">Chapter 3</xref>); how to display genome data using the <xref ref-type="other" rid="bid.1336">Map Viewer</xref> (<xref ref-type="other" rid="bid.1562">Chapter 20</xref>); how to print or save search output; how to set up sequence databases and install NCBI software locally; and which databases to use for a specific research question. We also accept reports of possible data errors, suggestions of new features or content to include in databases, and reports of system bugs or Web problems. The NCBI also receives a number of press inquiries about the research applications of our services, bioinformatics in general, and various genome projects. Occasionally, a high-school student will submit questions for a classroom assignment, providing a special outreach opportunity to young scientists.</p>
                <p>Because of the genetic focus of many of NCBI resources, we receive a number of questions from the general public regarding medical issues. The NCBI Help Desk staff can neither provide direct answers to medical questions nor give medical advice or guidance. However, we do provide suggestions on how to search our resources for information on the gene or condition of interest and refer users to the National Library of Medicine (NLM) customer service group for further assistance with PubMed (<xref ref-type="other" rid="bid.42">Chapter 2</xref>), <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/medlineplus/">MEDLINEplus</ext-link>, and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.clinicaltrials.gov/">ClinicalTrials.gov</ext-link>. We also refer them to outside organizations that can provide information on such topics as support groups and sources of medical advice.</p>
                <p>Questions about <xref ref-type="other" rid="bid.1387">PubMed</xref> are handled by a separate customer service group within the NLM. Their direct address is <email>custserv@nlm.nih.gov</email>, and their phone number is 1-888-FIND-NLM. PubMed questions that are received at the NCBI Help Desk are forwarded to NLM.</p>
              </sec>
            </sec>
            <sec id="bid.917">
              <title>Development of User Support Materials</title>
              <p>Because of its ongoing personal contact with our users, the User Services group plays an important role in communicating with database development and production teams, making suggestions, testing new releases and new features, and keeping them informed of problems that people are having with the services. The team also collaborates with developers in creating help documents, frequently asked questions (FAQs), tutorials, and workshop materials.</p>
              <sec id="bid.918">
                <title>Tutorials on the Web</title>
                <p>Web-based tutorials for <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/information3.html">BLAST</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Database/tut1.html">Entrez</ext-link>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/CN3D/cn3dtut.shtml">Cn3D</ext-link>, and <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/bsd/pubmed_tutorial/m1001.html">PubMed</ext-link> are currently available, with additional topics under development. Tutorials are produced on a collaborative basis by database development and User Services staff.</p>
              </sec>
              <sec id="bid.919">
                <title>About NCBI</title>
                <p>In keeping with the &ldquo;plain language initiative&rdquo; at NIH, the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/index.html"><italic>About NCBI</italic></ext-link> section of the NCBI Web site presents many fundamentals of NCBI's bioinformatics tools and databases, including a science primer covering such topics as molecular genetics, genome mapping, Single Nucleotide Polymorphisms (SNPs) (<xref ref-type="other" rid="bid.1143">Chapter 5</xref>), and microarray technology (<xref ref-type="other" rid="bid.337">Chapter 6</xref>). A model organism guide presents various model organisms and their uses in laboratory settings. As an introduction and orientation to NCBI's multifaceted Web site, the <italic>About NCBI</italic> section appeals to the general public, educators, and researchers alike.</p>
              </sec>
              <sec id="bid.920">
                <title>NCBI Site Map</title>
                <p>The NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sitemap/index.html">Site Map</ext-link> serves as a guide to NCBI resources. It provides a comprehensive, linked list of resources, along with a brief description of each resource. An effective way to locate a resource of interest within the Site Map is to perform a Find in Page search, a function that is built into all commonly used Web browsers.</p>
              </sec>
              <sec id="bid.921">
                <title>Publications</title>
                <p>The <italic>NCBI News</italic> is a quarterly newsletter that includes articles on new services, new features, and basic research at NCBI, as well as how to use selected resources for common applications. The newsletter is available free of charge and is offered online and by print subscription.</p>
                <p>The User Services group also prepares fact sheets, brochures, and other public information materials to describe and illustrate NCBI services. A <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/news/print.html">list</ext-link> of available materials is provided in the <italic>About NCBI</italic> section of the Web site, under News.</p>
                <p>Overview articles entitled <italic>GenBank</italic> and <italic>Database Resources of the NCBI</italic> have also been published recently in the annual database issues of <italic>Nucleic Acids Research</italic> (<xref ref-type="bibr" rid="bid.937">1</xref>-<xref ref-type="bibr" rid="bid.939">3</xref>).</p>
              </sec>
            </sec>
            <sec id="bid.922">
              <title>Outreach</title>
              <p>NCBI's continuing emphasis on outreach to the scientific community is evident in its multifaceted program that includes exhibiting its services at scientific meetings, offering a variety of training courses, and developing Web-based tutorials and workshops.</p>
              <sec id="bid.923">
                <title>Exhibits at Scientific Meetings</title>
                <p>NCBI exhibits at approximately 15 scientific meetings per year, providing an opportunity for a wide range of researchers, students, and teachers to see demonstrations of NCBI resources and interact directly with NCBI staff. The current exhibit <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/glance/index.html">schedule</ext-link> is posted on the NCBI Web site in the <italic>About NCBI</italic> section, under <italic>NCBI at a Glance</italic>.</p>
                <p>Workshops are offered at select scientific meetings and include the standing workshops described below in the Training section, but workshops also can be customized for particular audiences. Meeting organizers who would like to invite NCBI to offer a workshop are encouraged to do so.</p>
              </sec>
              <sec id="bid.924">
                <title>Training Courses</title>
                <p>NCBI has a growing training program consisting of full-day, half-day, and two-hour courses that are usually a combination of lecture and computer-based formats. There are also advanced courses that are given over a more extended time period. Each is described briefly below, and further information on the training programs can be found in the Education section of the Web site, under <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/outreach/courses.html"><italic>NCBI Courses</italic></ext-link>.</p>
                <sec id="bid.925">
                  <title>
                    <bold>A Field Guide to GenBank and Other NCBI Resources</bold>
                  </title>
                  <p><italic>A Field Guide to GenBank and NCBI Molecular Biology Resources</italic> is a training course offered in a lecture format, followed by hands-on computer sessions. It is designed as a basic but broad introduction to NCBI tools and resources.</p>
                  <p><italic>Field Guide</italic> topics include the following: description and scope of the primary database, GenBank (<xref ref-type="other" rid="bid.2">Chapter 1</xref>); derivative databases, such as UniGene (<xref ref-type="other" rid="bid.857">Chapter 21</xref>), LocusLink (<xref ref-type="other" rid="bid.788">Chapter 19</xref>), and Reference Sequence (RefSeq) (<xref ref-type="other" rid="bid.681">Chapter 18</xref>); effective database searching using Entrez; NCBI structure databases and the structure viewer, Cn3D; sequence similarity searching using the BLAST programs; the Conserved Domain Database (CDD) and associated search engine; and genome resources, including the NCBI assembly of the draft human genome, access to both finished and unfinished microbial genomes, and the genome Map Viewer.</p>
                  <p>The course is offered by invitation at academic institutions as well as at selected scientific conferences. If you are interested in hosting a course at your institution or conference, write to <email>info@ncbi.nlm.nih.gov</email>, and your request will be routed to the course coordinator.</p>
                  <p>The course is also offered four times a year at the NLM on the NIH campus in Bethesda, Maryland, and is free and open to anyone who would like to attend.</p>
                  <p>More information on <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Class/FieldGuide/">this course</ext-link> is available in the Education section of the NCBI Web site, under <italic>NCBI Courses</italic>. The Web site includes the course handout, slide presentations, and problem sets with answers. A schedule of planned courses at NLM and elsewhere is posted under <italic>Upcoming Courses</italic>.</p>
                </sec>
                <sec id="bid.926">
                  <title>
                    <bold>Molecular Biology Information Resources</bold>
                  </title>
                  <p>This course is designed primarily for medical and science librarians or other professionals who are providing support services for molecular biology information resources. It provides an introduction to four categories of molecular biology information available from NCBI: nucleotide sequences, protein sequences, three-dimensional structures, and complete genomes and maps. An overview of search systems available at the NCBI, particularly Entrez and BLAST, emphasizes how search skills related to other types of information resources also apply to molecular biology databases. The course concludes with a discussion of various levels of molecular biology information services provided by librarians.</p>
                  <p>The Medical Library Association approves this course for eight continuing education credit hours. The course has been given at 24 locations since May 1997. Because of the increase in NCBI services, courses are being revised and are not being scheduled at this time.</p>
                  <p>More information on <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/outreach/courses.html">this course</ext-link> is available in the Education section of the NCBI Web site, under <italic>NCBI Courses</italic>. The Web site includes the course materials used for the lecture and a set of exercises.</p>
                </sec>
                <sec id="bid.927">
                  <title>
                    <bold>NCBI Advanced Workshop for Bioinformatics Information Specialists</bold>
                  </title>
                  <p>A new 5-day advanced course on NCBI resources has been developed as part of a collaborative project with a group of scientists and librarians who currently provide bioinformatics support services at their universities. The course provides detailed descriptive information as well as hands-on experience with handling a wide range of user questions. The course is designed for bioinformatics support staff based in university medical libraries so that they can, in turn, assist students, faculty, staff, and clinicians at their institutions in the use of molecular biology information resources. Additional information on <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/About/outreach/index.html">this course</ext-link> is available in the Education section of the NCBI Web site.</p>
                </sec>
                <sec id="bid.928">
                  <title>
                    <bold>Specialized Mini-Courses</bold>
                  </title>
                  <p>The Service Desk staff also offers four mini-courses: BLAST QuickStart, Unmasking Genes in the Human Genome, Making Sense of DNA and Protein Sequences, and GenBank and PubMed Searching. Each is described briefly below. The purpose of the mini-courses is to focus on specific research application areas and address how to use multiple NCBI resources together to answer a research question. Additional problem-oriented mini-courses are under development.</p>
                  <p>The courses are 2 hours each in length. An overview is given during the first hour in lecture format, followed by a 1-hour hands-on session. Although primarily given on the NIH campus, NCBI is beginning to offer these workshops at outside institutions as well. Although the mini-courses were originally designed to be presented by an instructor, they are constructed in an online notebook format; therefore, it is possible to take the course on your own. Revisions to augment the online notebooks with lecture material and make the courses completely self-guided are currently under way.</p>
                  <sec id="bid.929">
                    <title>
                      <bold>BLAST QuickStart!</bold>
                    </title>
                    <p>This mini-course is a practical introduction to the BLAST family of sequence-similarity search programs. Exercises range from simple searches to creative uses of the BLAST programs.</p>
                  </sec>
                  <sec id="bid.930">
                    <title>
                      <bold>Unmasking Genes in the Human Genome</bold>
                    </title>
                    <p>This mini-course covers how to find genes, promoters, and transcription factor-binding sites in human DNA sequences. It is designed around a program developed within User Services called Greengene, which integrates the output of several gene-finding tools and allows a coding sequence and accompanying protein translation to be assembled from the exons detected by these programs. Because the output of several programs is integrated, there is increased reliability in exon selection.</p>
                  </sec>
                  <sec id="bid.931">
                    <title>
                      <bold>Making Sense of DNA and Protein Sequences</bold>
                    </title>
                    <p>In this course, participants find a gene within a eukaryotic DNA sequence. They then predict the function of the derived protein by seeking sequence similarities to proteins with documented function using BLAST and other tools. Finally, a 3D modeling template is located for the protein sequence using the Conserved Domain Search (CDD-Search).</p>
                    <p>During the first hour, an instructor walks the class through an analysis of an uncharacterized <italic>Drosophila melanogaster</italic> genomic sequence from a GenBank record. During the second hour, participants perform the same analysis independently, using a different genomic sequence.</p>
                  </sec>
                  <sec id="bid.932">
                    <title>
                      <bold>GenBank and PubMed Searching</bold>
                    </title>
                    <p>This mini-course provides an overview of literature searching and sequence retrieval using the PubMed and Entrez database search interfaces. Exercises illustrate advanced search tips for using Entrez, many of which explore the use of the <named-content content-type="book-spec-emph">Preview/Index</named-content> options for specifying parameters to limit the search results. The course also features 21 self-scoring exercises for GenBank.</p>
                  </sec>
                </sec>
              </sec>
              <sec id="bid.933">
                <title>The NCBI Learning Center</title>
                <p>In addition to communication by email and phone provided through the Help Desk, a regular research consultation service provides one-on-one support for researchers in the NIH community. The consults are available by appointment and are provided in 1-hour time slots at the NIH Library as well as the NCBI training facility. Because of the success of the program, this type of service may be offered by appointment at selected scientific meetings in the future.</p>
              </sec>
              <sec id="bid.934">
                <title>CoreBio</title>
                <p>An innovative training program that began in 2001 aims to train molecular biologists for a new type of career as bioinformatics specialists who provide institutional support for users of computational biology tools. The NCBI Core Bioinformatics Facility (referred to as the CoreBio program) currently functions to train and support a network of bioinformatics specialists serving individual Institutes at NIH. NCBI's CoreBio facility trains Core members identified by their respective institutes in the use of its bioinformatics tools. The Core members, in turn, support the use of NCBI tools and databases by researchers at their institutes.</p>
                <p>The training is provided over a 9-week period, with students attending lectures and completing practical exercises in the morning and returning to their regular workplace in the afternoon. The coursework centers on one major topic each week and follows the rough schedule given below:</p>
                <p>WEEK 1: Introduction to the Sequence Databases</p>
                <p>WEEK 2: BLAST</p>
                <p>WEEK 3: The Human Genome</p>
                <p>WEEK 4: Genomic Biology</p>
                <p>WEEK 5: Molecular Modeling</p>
                <p>WEEK 6: Web Page Development</p>
                <p>WEEK 7: Setting Up a BLAST Web Server</p>
                <p>WEEK 8: Interaction with Users</p>
                <p>WEEK 9: Practicum</p>
                <p>During week 9, the students pursue an institute-related project with the assistance of NCBI instructors. These projects run the gamut from the compilation of specialized datasets and data mining to the creation of novel BLAST interfaces and the construction of new data display tools. Students also develop a Web page to support the services they are developing for the respective Institutes at the NIH.</p>
                <p>Although currently a NIH-based program, other organizations are welcome to consider using the program as a model for development of similar initiatives to meet their bioinformatics support needs.</p>
              </sec>
            </sec>
            <sec id="bid.935">
              <title>Conclusion</title>
              <p>At NCBI, we encourage our users to contact us with questions, suggestions, and requests for training or presentations on NCBI services. We invite feedback on tutorials, FAQs, and other support materials and welcome suggestions regarding additional materials that would be useful in guiding users through the wide range of services offered by NCBI.</p>
            </sec>
</body>
<back>
              <ref-list>
              <title>References</title>
                <ref id="bid.937">
                  <label>1</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Benson</surname>
                        <given-names>DA</given-names>
                      </name>
                      <name>
                        <surname>Karsch-Mizrachi</surname>
                        <given-names>I</given-names>
                      </name>
                      <name>
                        <surname>Lipman</surname>
                        <given-names>DJ</given-names>
                      </name>
                      <name>
                        <surname>Ostell</surname>
                        <given-names>J</given-names>
                      </name>
                      <name>
                        <surname>Rapp</surname>
                        <given-names>BA</given-names>
                      </name>
                      <name>
                        <surname>Wheeler</surname>
                        <given-names>DL</given-names>
                      </name>
                    </person-group>
                    <article-title>GenBank</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>30</volume>
                    <fpage>17</fpage>
                    <lpage>20</lpage>
                    <year>2002</year>
                    <pub-id pub-id-type="pmid">11752243</pub-id>
                    <pub-id pub-id-type="other">99127</pub-id>
                  </citation>
                </ref>
                <ref id="bid.938">
                  <label>2</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Wheeler</surname>
                        <given-names>DL</given-names>
                      </name>
                      <name>
                        <surname>Church</surname>
                        <given-names>DM</given-names>
                      </name>
                      <name>
                        <surname>Lash</surname>
                        <given-names>AE</given-names>
                      </name>
                      <name>
                        <surname>Leipe</surname>
                        <given-names>DD</given-names>
                      </name>
                      <name>
                        <surname>Madden</surname>
                        <given-names>TL</given-names>
                      </name>
                      <name>
                        <surname>Pontius</surname>
                        <given-names>JU</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                      <name>
                        <surname>Schriml</surname>
                        <given-names>LM</given-names>
                      </name>
                      <name>
                        <surname>Tatusova</surname>
                        <given-names>TA</given-names>
                      </name>
                      <name>
                        <surname>Wagner</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Rapp</surname>
                        <given-names>BA</given-names>
                      </name>
                    </person-group>
                    <article-title>Database resources of the National Center for Biotechnology Information: 2002 update</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>30</volume>
                    <fpage>13</fpage>
                    <lpage>16</lpage>
                    <year>2002</year>
                    <pub-id pub-id-type="pmid">11752242</pub-id>
                    <pub-id pub-id-type="other">99094</pub-id>
                  </citation>
                </ref>
                <ref id="bid.939">
                  <label>3</label>
                  <citation>
                    <person-group>
                      <name>
                        <surname>Wheeler</surname>
                        <given-names>DL</given-names>
                      </name>
                      <name>
                        <surname>Church</surname>
                        <given-names>DM</given-names>
                      </name>
                      <name>
                        <surname>Lash</surname>
                        <given-names>AE</given-names>
                      </name>
                      <name>
                        <surname>Leipe</surname>
                        <given-names>DD</given-names>
                      </name>
                      <name>
                        <surname>Madden</surname>
                        <given-names>TL</given-names>
                      </name>
                      <name>
                        <surname>Pontius</surname>
                        <given-names>JU</given-names>
                      </name>
                      <name>
                        <surname>Schuler</surname>
                        <given-names>GD</given-names>
                      </name>
                      <name>
                        <surname>Schriml</surname>
                        <given-names>LM</given-names>
                      </name>
                      <name>
                        <surname>Tatusova</surname>
                        <given-names>TA</given-names>
                      </name>
                      <name>
                        <surname>Wagner</surname>
                        <given-names>L</given-names>
                      </name>
                      <name>
                        <surname>Rapp</surname>
                        <given-names>BA</given-names>
                      </name>
                    </person-group>
                    <article-title>Database resources of the National Center for Biotechnology Information</article-title>
                    <source>Nucleic Acids Res</source>
                    <volume>29</volume>
                    <fpage>11</fpage>
                    <lpage>16</lpage>
                    <year>2001</year>
                    <pub-id pub-id-type="pmid">11125038</pub-id>
                    <pub-id pub-id-type="other">29800</pub-id>
                  </citation>
                </ref>
              </ref-list>
          </back>
        </book-part>
        <book-part id="bid.1679" book-part-type="chapter" book-part-number="24">
          <book-part-meta>
            <title-group>
              <title>Exercises: Using Map Viewer</title>
            </title-group>
            <contrib-group>
              <contrib contrib-type="author">
                <name>
                  <surname>Wheeler</surname>
                  <given-names>David</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Pruitt</surname>
                  <given-names>Kim</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Maglott</surname>
                  <given-names>Donna</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Dombrowski</surname>
                  <given-names>Susan</given-names>
                </name>
              </contrib>
              <contrib contrib-type="author">
                <name>
                  <surname>Gabrelian</surname>
                  <given-names>Andrei</given-names>
                </name>
              </contrib>
            </contrib-group>
            <history>
              <date date-type="created">
                <day>4</day>
                <month>11</month>
                <year>2002</year>
              </date>
              <date date-type="updated">
                <day>13</day>
                <month>8</month>
                <year>2003</year>
              </date>
            </history>
            <alternate-form xmlns:xlink="http://www.w3.org/1999/xlink" alternate-form-type="pdf" xlink:href="ch24d1"/>
            <abstract>
              <title>Introduction</title>
              <p>This chapter contains tutorials for using Map Viewer. Step-by-step instructions are provided for several common biological research problems that can be addressed by exploiting the whole-genome and positional perspectives of Map Viewer. Please be aware that the examples in these tutorials may return different results when you execute them, because the underlying data may have been updated, but we hope that the framework for obtaining, interpreting, and processing your results will be sufficiently clear if that happens. Most of the examples are for human genes, but the same logic applies to other genomes as well.</p>
              <p>Please note that each of these tutorials is accompanied by a figure. If you are using this tutorial on the Web, we suggest that you open another browser window so that you can view the figure as you are reading the text. You are also encouraged to use Map Viewer interactively.</p>
              <p>We welcome any suggestions that you have for improving the existing tutorials or for adding new ones.</p>
            </abstract>
          </book-part-meta>
          <body>
            <sec id="bid.1680">
              <title>1. How Do I Obtain the Genomic Sequence around My Gene of Interest?</title>
              <p>There are many instances in molecular biological research when you may have only a cDNA sequence but need to have the nucleotide sequence that lies 5&prime; or 3&prime; to a gene or the introns for additional analyses. Because genomic sequence available from the public database may not have this annotation or may be so large as to make it difficult to retrieve only a region of interest, tools have been added to Map Viewer to make it easier to define, view, and download genomic sequence in multiple formats.</p>
              <sec id="bid.1681">
                <title>Direct Sequence Information via Map Viewer</title>
                <p>From the Map Viewer <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview">homepage</ext-link>, select <named-content content-type="book-spec-emph">Search</named-content> Homo sapiens (human) from the pull-down menu, enter FMR1 in the box labeled <named-content content-type="book-spec-emph">for</named-content>, and then select <named-content content-type="book-spec-emph">Go</named-content>. On the results page, four entries are returned (<italic>4 hits</italic>). <italic>FMR1</italic> has been mapped on three different maps: <italic>Genes_cyto, Genes_seq,</italic> and <italic>Morbid</italic>. Select the <named-content content-type="book-spec-emph">Genes_seq</named-content> map to see the structure of the <italic>FMR1</italic> gene. Two links to the right of the Genes_seq map are of use in retrieving the 5&prime; and 3&prime; flanking DNA, the <named-content content-type="book-spec-emph">sv</named-content> and <named-content content-type="book-spec-emph">seq</named-content> links (<italic>boxed</italic>). At the top of the page, <named-content content-type="book-spec-emph">Download/View Sequence/Evidence</named-content> can also be used.</p>
                <p>The most informative maps, for the purpose of this example, are the Contig, Component, and Genes_seq maps. Selecting the <named-content content-type="book-spec-emph">Genes_seq</named-content> map will display the current gene model. To start our search for flanking DNAs, display the <named-content content-type="book-spec-emph">Contig</named-content> and <named-content content-type="book-spec-emph">Component</named-content> maps. To do so, select <named-content content-type="book-spec-emph">Maps &amp; Options</named-content>, and from the available maps, choose the <named-content content-type="book-spec-emph">Contig</named-content> map by left-clicking with the mouse. Select <named-content content-type="book-spec-emph">ADD&gt;&gt;</named-content>. Now add the <named-content content-type="book-spec-emph">Component</named-content> map. Next, make the Gene_seq map the master map in the display by left-clicking on the <named-content content-type="book-spec-emph">Gene</named-content> map under the <named-content content-type="book-spec-emph">Maps Displayed</named-content> (<italic>left</italic> to <italic>right</italic>) and selecting <named-content content-type="book-spec-emph">Make Master/Move to Bottom:</named-content>. Finally, select the <named-content content-type="book-spec-emph">Contig</named-content> map and choose the <named-content content-type="book-spec-emph">Toggle Ruler</named-content> to add a ruler (<italic>[R]</italic>) to the display. This will guide you in finding your region of interest. Select <named-content content-type="book-spec-emph">Apply</named-content>. From <xref ref-type="fig" rid="bid.1682">Figure 1</xref>, we can see that FMR1 has been annotated on a finished contig, NT_011537.9 (<italic>c</italic>), and that the contig is built from two components in this region, AC016925.15 and L29074.1 [(<italic>a</italic>) and (<italic>b</italic>), respectively].<fig id="bid.1682"><label>1</label><caption><title>Making a master map.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f1" mime-subtype="gif"/></fig></p>
                <p>There are two links on the Map Viewer display that are used to view and download of the region of interest: the <named-content content-type="book-spec-emph">seq</named-content> link and the <named-content content-type="book-spec-emph">Download/View Sequence/Evidence</named-content> link (<italic>boxed</italic>). The <named-content content-type="book-spec-emph">seq</named-content> link is displayed only when the Gene_seq map is made the master map in the display. Selecting the <named-content content-type="book-spec-emph">seq</named-content> link will open a window, prompting the user to enter a region to retrieve. This region can then be refined further by adjusting the position, in kilobase pairs. Selecting <named-content content-type="book-spec-emph">Display Region</named-content> will change the start and stop positions on the contig, where the gene has been annotated. The region can then be displayed and saved locally in FASTA or GenBank formats.</p>
                <p>Selecting the <named-content content-type="book-spec-emph">Download/View Sequence/Evidence</named-content> link from the main display page will generate the same window, but the initial coordinates are those spanning the entire region displayed by the Map Viewer rather than the region of a particular gene, as in the case above.</p>
                <p>Let us assume that we would like to download 5.0 kb of upstream DNA and 1.0 kb of downstream DNA. To define this region, we will need to follow the <named-content content-type="book-spec-emph">seq</named-content> link, which will open a new display showing the chromosome coordinates for <italic>FMR1</italic> and the corresponding position on the contig. To adjust the region, simply enter the amount of desired upstream DNA and downstream DNA into the two <named-content content-type="book-spec-emph">adjust by:</named-content> input boxes provided, and select <named-content content-type="book-spec-emph">Change Region</named-content>. Notice that the corresponding region on the contig has adjusted to reflect this change in position. Now we can either display the data in GenBank or FASTA formats and save the data to a disk.</p>
              </sec>
              <sec id="bid.1683">
                <title>Using the Sequence Viewer</title>
                <p>Another tool for obtaining the desired sequence is the Sequence Viewer, available via the <named-content content-type="book-spec-emph">sv</named-content> link when the Gene_seq map is the master map. The Sequence Viewer presents a graphical view of the gene within the contig. The sequence is also annotated with the coding regions, RNA and gene features, Sequence Tagged Sites (STSs), and single nucleotide polymorphisms. <italic>Blue arrows</italic> at the <italic>top</italic> and <italic>bottom</italic> of the display allow the user to navigate upstream or downstream. The <named-content content-type="book-spec-emph">Get Subsequence</named-content> link (<italic>large, open arrow</italic>) at the <italic>top</italic> of the <named-content content-type="book-spec-emph">sv</named-content> page allows the user to change the sequence range on the contig and will also display the reverse complement of the sequence. The specified region can then be displayed and saved locally in FASTA format.</p>
              </sec>
            </sec>
            <sec id="bid.1684">
              <title>2. If I Have Physical and/or Genetic Mapping Data, How Do I Use the Map Viewer to Find a Candidate Disease Gene in That Region?</title>
              <p>In this example, we will use the Map Viewer to look for human candidate genes in a region. The types of queries that can be posted to the Map Viewer that will address this type of question are queries by genetic marker or STS.</p>
              <p>Please note that Map Viewer supports queries by any named object positioned on a map so that it is possible to query by gene symbol or GenBank Accession number or any other object that might define your range of interest.</p>
              <sec id="bid.1685">
                <title>Querying by STSs</title>
                <p>To refine our search,  we will enter the names of two STSs. In the text box, enter &ldquo;sWXD113 OR DXS52&rdquo; on chromosome X. Select <named-content content-type="book-spec-emph">Find</named-content>. We can see that these two STSs map to the distal region of the long arm of the X chromosome, Xq, by the <italic>red tick marks</italic> that appear alongside the schematic of the X chromosome. These two STS markers have been mapped on several other maps that are also represented in the results. Select the <named-content content-type="book-spec-emph">X</named-content> chromosome <italic>above</italic> the <italic>red 3</italic> to see both markers in the same display. The Map Viewer page now displays three different maps, each showing the physical location of these two markers.</p>
                <p>The maps that are displayed include the UniG_Hs, Genes_seq, and STS maps. The UniG_Hs map shows the density of ESTs and mRNAs that align to the current assembly of the human genome. The Genes_seq map displays known and predicted genes that are annotated on the genomic contigs. The <italic>rightmost</italic> map, the STS map in this case, is termed the master map and contains descriptive information about each map element (<xref ref-type="fig" rid="bid.1686">Figure 2</xref>). The two STSs for which we are searching (<italic>GDB:192503</italic> and <italic>DXS52</italic>) are highlighted in <italic>pink</italic>. To the <italic>right</italic> of the display is a grid indicating other maps upon which these STSs are located. The <italic>red dots</italic> (<italic>circled</italic>) to the <italic>left</italic> of each highlighted STS show the relative position of these STSs in the context of the two other maps. By default, a ruler is displayed alongside the STS map so that the region of interest can be localized further. Notice on the <italic>far left</italic> of the page that there is an area where you can enter the region that you would like to display [(<italic>a</italic>)].<fig id="bid.1686"><label>2</label><caption><title>Master Map: STS.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f2" mime-subtype="gif"/></fig></p>
                <p>At the current resolution, it is not possible to view all of the information displayed on the three maps. Therefore, some adjustment will be necessary.</p>
              </sec>
              <sec id="bid.1687">
                <title>Identifying a Candidate Gene</title>
                <p>Narrow the region further using the ruler adjacent to the STS map as a guide. It is not necessary to display a ruler alongside the other two maps because they are all on the same coordinate system. In instances where sequence, genetic, cytogenetic, or radiation hybrid maps are being displayed in the same view, it is advisable to display additional rulers because the different maps show the mapped element on different coordinate systems (Kbp, cM, banding position, or centiRays, respectively). Enter the range 133.0 M to147.0 M in the <italic>boxes</italic> to the <italic>left</italic> of the page, in the <named-content content-type="book-spec-emph">Region Shown</named-content> [(<italic>a</italic>)]. Select <named-content content-type="book-spec-emph">Go</named-content>. This is a slightly better view; however, there are many more genes in the region that can be displayed.</p>
                <p>To see all of the genes in the region defined by the two markers, select the <named-content content-type="book-spec-emph">Data as Table View</named-content> link (<italic>boxed</italic>). The table that is generated lists all 144 genes in this region in a format easily read by people or computers. The table also preserves the links to additional gene-related resources seen in the graphical display, as well as reporting other objects in your displayed region. Please note that links are also provided to make it easy for you to download reports, not only of the objects in your map display, but other objects within the region defined by your display. This feature is especially useful if you are looking for other gene markers in your region of interest.</p>
                <p>We can also change the page length under <named-content content-type="book-spec-emph">Maps &amp; Options</named-content> to a number large enough so that all genes are displayed graphically on a single page. By default, Map Viewer will show 20 map elements on a page. When the Genes_seq map is made the master map, there will be information at the <italic>top</italic> of the page, indicating how many genes have been labeled (<italic>n</italic> = 20) and how many genes are in the region (<italic>n</italic> = 144). To see all of the candidate genes in the specified region, go to <named-content content-type="book-spec-emph">Maps &amp; Options</named-content> and change the page length to the number of genes in the region (<italic>n</italic> = 144) and then select <named-content content-type="book-spec-emph">Apply</named-content>.</p>
              </sec>
              <sec id="bid.1688">
                <title>Interpreting Your Results</title>
                <p>At this point, you can now browse the description of the genes that are being displayed. Each gene or locus name is hyperlinked to LocusLink, where a detailed report about the gene or locus is provided. If the gene or locus of interest has supporting EST and mRNA data, then you can select the UniGene cluster number and link to UniGene, where more detailed information is provided about this gene, including its pattern of expression. LocusLink also provides connections to BLink and thus indirectly to reports of related proteins in the protein database and to viewers of protein structure, if your protein of interest is related to a protein for which the structure is known (see also Exercise 8 in this chapter).</p>
              </sec>
              <sec id="bid.1689">
                <title>Other Ways to Query</title>
                <p>The example above summarizes the approach taken when defining a region of interest by entering names of markers in the query box. Gene symbols, reference SNP names, and GenBank Accession numbers for ESTs could also be used. It should also be noted that when a chromosome is displayed, you may also submit a query using the <named-content content-type="book-spec-emph">Region Shown</named-content> boxes in the bar at the left side.</p>
              </sec>
            </sec>
            <sec id="bid.1690">
              <title>3. How Can I Find and Display a Gene with the Map Viewer?</title>
              <p>In this example, we will locate and display the human gene implicated in Fragile X syndrome using the Map Viewer. We can find the gene beginning with several types of data. Refer to <xref ref-type="fig" rid="bid.1691">Figure 3</xref>. For these examples, we assume you are starting from the human-specific Map Viewer. If you instead are starting from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/mapview/">homepage</ext-link>, please remember to select &ldquo;Homo sapiens (human)&rdquo; as the species.<fig id="bid.1691"><label>3</label><caption><title>Location and display of the human gene implicated in Fragile X syndrome.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f3" mime-subtype="gif"/></fig></p>
              <sec id="bid.1692">
                <title>By Gene Symbol</title>
                <p>If we are fortunate enough to know the official gene symbol for the Fragile X gene, <italic>FMR1</italic>, or an alternative symbol such as <italic>FRAXA</italic>, we can type the symbol into the search box at the <italic>top</italic> of the page and press the <named-content content-type="book-spec-emph">Find</named-content> button. This gene has been annotated on the genome, and the gene symbol <italic>FMR1</italic> appears on the <named-content content-type="book-spec-emph">Genes_seq</named-content>, <named-content content-type="book-spec-emph">Genes_cyto</named-content>, and the <named-content content-type="book-spec-emph">Morbid</named-content> maps. To generate a Map Viewer display that includes all three maps, select the link to <named-content content-type="book-spec-emph">all matches</named-content>.</p>
              </sec>
              <sec id="bid.1693">
                <title>By Linkage to a Disease</title>
                <p>Genes that are linked to a disease in Online Mendelian Inheritance in Man (OMIM) are referenced on the Morbid map and can be found by searching with a disease name or phenotype. In our case, &ldquo;fragile-X&rdquo; can be used. Using this query, we pick up hits to genes related to <italic>FMR1</italic>, as well as our intended gene. Selecting the <named-content content-type="book-spec-emph">FMR1</named-content> link generates a display of the gene.</p>
              </sec>
              <sec id="bid.1694">
                <title>By Physical Marker</title>
                <p>The <italic>FMR1</italic> gene contains the STS, STS-X69962 [(<italic>b</italic>)]; therefore, we can also use the name of this STS to find it, as in the case of a gene symbol. The search yields a table of hits showing that STS-X69962 appears on the STS map. Selecting the <named-content content-type="book-spec-emph">STS</named-content> link gives us a Map Viewer display of the STS we found but not of the gene we sought. To get the gene into the Map Viewer display, we can add the Genes_seq map to the display. Although <italic>FMR1</italic> is located on the Genes_seq map rather than the STS map, because the coordinate systems used in different maps are synchronized, we can see the <italic>FMR1</italic> gene if we ask for the Genes_seq track in the region corresponding to the hit on the STS map. To do this, we can select the <named-content content-type="book-spec-emph">Maps &amp; Options</named-content> link [(<italic>a</italic>)], highlight the <named-content content-type="book-spec-emph">Gene</named-content> map from the list of <named-content content-type="book-spec-emph">available maps</named-content> in the <italic>left-hand box</italic>, and select the <named-content content-type="book-spec-emph">ADD&gt;&gt;</named-content> button to add this map to the list of displayed maps. After selecting the <named-content content-type="book-spec-emph">Apply</named-content> button, the Gene map is added to the Map Viewer display. Note, however, that the view is limited to a very small 200-base pair portion of the gene. This is because we are still focused on the STS returned by our initial search. To see an expanded view, we must zoom out using the zoom control located directly over the thumbnail chromosome map in the <italic>blue sidebar</italic> [(<italic>d</italic>)]. Mousing over the control indicates that we are viewing 1/10,000th of chromosome X. We can click further up on the control to view 1/1,000th of the chromosome and see most of the <italic>FMR1</italic> gene. The <italic>FMR1</italic> gene is now centered in the Map Viewer display, and our STS hit is marked in <italic>red</italic>.</p>
              </sec>
              <sec id="bid.1695">
                <title>By Region</title>
                <p>Often, a gene is known to reside only in a particular region. Suppose that we know only that <italic>FMR1</italic> resides somewhere between markers DXS532 and DXS7389 [(<italic>e</italic>) and (<italic>c</italic>), respectively]. We can use a query containing a Boolean OR to force the Map Viewer to search for both markers simultaneously. In this case, a query to Map Viewer of &ldquo;DXS532 OR DXS7389&rdquo; generates a number of hits to various physical maps, all to a region on chromosome X, which is marked in <italic>red</italic> under the <named-content content-type="book-spec-emph">X</named-content> chromosome graphic. If we select the chromosome <named-content content-type="book-spec-emph">X</named-content> link under the chromosome graphic, the Map Viewer display shows marker hits on several physical maps, highlighted in <italic>red</italic>, over a fairly large sequence region from about 141 to 142 megabases. To generate a tabular listing of all the genes in this region, select the <named-content content-type="book-spec-emph">Data as Table View</named-content> link to the <italic>left</italic> of the Map Viewer display. With a genetic map, such as the Genethon map, as the master map, it is also possible to define or refine the region of display using coordinates in centimorgans. In the case of <italic>FMR1</italic>, entering a range of 176&ndash;198 into the <named-content content-type="book-spec-emph">Region Shown</named-content> boxes generates a display of the genes falling within that range.</p>
              </sec>
              <sec id="bid.1696">
                <title>By Sequence Homology</title>
                <p>Suppose that we have the sequence of the mouse homolog of the human <italic>FMR1</italic> gene and want to locate it on the human genome assembly. We might consult the Human/Mouse homology map at NCBI as the most direct approach for mapped genes, but let us assume that the human homolog of the <italic>FMR1</italic> gene is unmapped. In this case, we can perform a BLAST sequence similarity search with the mouse sequence to attempt to locate the corresponding human gene. We will use the mouse <italic>Fmr1</italic> mRNA sequence, taken from NCBI's LocusLink database (Accession number NM_002024), as our probe and follow the link to <named-content content-type="book-spec-emph">BLAST search the human genome</named-content> located at the <italic>top</italic> of the Map Viewer search page; type the above Accession number into the BLAST form, and press the <named-content content-type="book-spec-emph">Search</named-content> button. Such searches are extremely fast because they make use of an NCBI program called MegaBLAST, designed especially for this purpose. In the MegaBLAST results, the <named-content content-type="book-spec-emph">Genomic View</named-content> button near the <italic>top</italic> of the page provides an entry point into a Map Viewer display. The MegaBLAST hits are indicated by a <italic>red mark</italic> on chromosome X, and links to hits are provided in the table at the <italic>bottom</italic> of the page, as in the searches by gene symbol or marker. In the case of any MegaBLAST search, all hits are to sequences [(<italic>f</italic>)] on the <named-content content-type="book-spec-emph">Contig</named-content> map; therefore, it is the contig link (Accession number beginning with NT_) that we follow to view the hits in their genomic context. Following this link brings us to a display of the <italic>FMR1</italic> gene, with the BLAST hits indicated as highlighted &ldquo;hits&rdquo; on the contig map and <italic>colored ticks</italic> on the other maps (<italic>large, expanded view</italic>).</p>
              </sec>
              <sec id="bid.1753">
                <title>Other Ways to Query</title>
                <p>Please note that other resources within NCBI also support querying for genes. Consider also LocusLink, UniGene, and Entrez Nucleotide. When a record of interest has been retrieved, each of these provides links to Map Viewer.</p>
              </sec>
            </sec>
            <sec id="bid.1697">
              <title>4. How Can I Analyze a Gene Using the Map Viewer?</title>
              <p>We will analyze the <italic>FMR1</italic> gene. To find this gene in the human Map Viewer, enter FMR1 into the query box and select the <named-content content-type="book-spec-emph">Genes_seq</named-content> map link in the table of search results.</p>
              <p>To begin the analysis, we can select the link in the <italic>blue sidebar</italic> entitled <named-content content-type="book-spec-emph">Data as Table View</named-content> to see a tabular listing of the chromosomal coordinates of  the features visible in the current Map Viewer display. In the section of the table giving features on the Genes_seq map, we find that the <italic>FMR1</italic> gene extends from 141519781 to 141558823, a span of about 40,000 base pairs. We can now return  to the graphical Map Viewer display and limit our view to this region by entering this range into the <named-content content-type="book-spec-emph">Region Shown</named-content> boxes [<xref ref-type="fig" rid="bid.1698">Figure 4</xref>, (<italic>a</italic>)] in the <italic>blue sidebar</italic> and pressing the <named-content content-type="book-spec-emph">Go</named-content> button. The coordinate system operative in the <named-content content-type="book-spec-emph">Region Shown</named-content> boxes is that of the <italic>rightmost</italic> map in the Map Viewer display, called the master map; be sure that the Genes_seq map is the master before changing the coordinates for the region shown.</p>
              <p>We are now ready to select the maps to display. The maps displayed will depend on the sort of analysis intended; however, one useful set of maps includes the Genes_seq, Contig, Comp, GScan, UniG_Hs, RNA, and gbDNA maps. This set of maps can be selected for viewing using the panel invoked by the <named-content content-type="book-spec-emph">Maps &amp; Options</named-content> link (<italic>large, open arrow</italic>). A Map Viewer display of the <italic>FMR1</italic> gene using this set of maps is given in <xref ref-type="fig" rid="bid.1698">Figure 4</xref>.<fig id="bid.1698"><label>4</label><caption><title>Set of maps.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f4" mime-subtype="gif"/></fig></p>
              <sec id="bid.1699">
                <title>The Genes_seq Track</title>
                <p>In <xref ref-type="fig" rid="bid.1698">Figure 4</xref>, the Genes_seq map, which displays annotated genes, is the master map and is the first one we will examine. The <italic>FMR1</italic> gene comprises 18 exons, represented by <italic>thick lines</italic>, interspersed over about 40,000 base pairs of sequence. The gene is drawn to the <italic>right</italic> of the Genes_seq track line and therefore runs from the <italic>top</italic> of the display to the <italic>bottom</italic> and is coded on the &ldquo;plus&rdquo; strand in the human genome assembly. Genes located on the opposite strand run from the <italic>bottom</italic> of the display upward and are shown to the <italic>left</italic> of the Genes_seq track. The <italic>FMR1</italic> gene is displayed in <italic>orange</italic>, which indicates that the alignment between the genomic sequence and the <italic>FMR1</italic> transcript sequences that were used to produce the gene model was not perfect; <italic>blue</italic> alignments are of the best quality. The 3&prime;-most exon (<italic>bottom inset</italic>) of the gene is exceptionally large and probably includes a significant untranslated region. To verify this, we can select the <named-content content-type="book-spec-emph">sv</named-content> link (<italic>boxed</italic>) to the <italic>right</italic> of the Genes_seq map to invoke the Sequence Viewer, which shows the sequence of the <italic>FMR1</italic> gene. In the Sequence Viewer, we can navigate to the display for the last exon, number 18, to see that the coding sequence ends toward the beginning of the exon and that the majority of the exon is indeed untranslated.</p>
              </sec>
              <sec id="bid.1700">
                <title>The RNA Track</title>
                <p>The RNA or Transcript map shows the alignment of a single mRNA sequence to the genome, and in this case, the pattern of exons produced matches exactly that shown on the Genes_seq map. If additional splice variants are sequenced, multiple alignments will be shown on the RNA track, and the gene model given on the Genes_seq track will be a composite model made up of all the exons implied by these alignments.</p>
              </sec>
              <sec id="bid.1701">
                <title>The GScan Track</title>
                <p>The GenomeScan track shows gene predictions made using GenomeScan that are independent of supporting mRNA alignments. The GenomeScan model for <italic>FMR1</italic> is very similar to the model shown on the Genes_seq track; however, there are differences, and the alignment-based model shown on the Genes_seq track covers exons that are part of two different GenomeScan models (<italic>topmost inset</italic>). Note also that the GenomeScan model covers only the initial translated portion of the large 3&prime; exon of the transcript-based model (<italic>bottom inset</italic>).</p>
              </sec>
              <sec id="bid.1702">
                <title>The UniG_Hs Track</title>
                <p>Both the predicted model (GScan) and the alignment-based model (Genes_seq) can be compared to the mapping of ESTs on the UniG_Hs  track. The <italic>bars</italic> extending to the <italic>left</italic> of the UniG_Hs track line depict EST mapping density, whereas the <italic>lines</italic> to the <italic>right</italic> connect ESTs arising from common UniGene clusters. From the UniG_Hs track, it is clear that most of the exons arising from either of the two gene models have some EST support, and that most of the ESTs that map to these regions are members of UniGene cluster Hs.89764 (<italic>boxed, center</italic>). Selecting the <named-content content-type="book-spec-emph">Hs.89764</named-content> link leads to the <italic>FMR1</italic> UniGene cluster.</p>
              </sec>
              <sec id="bid.1754">
                <title>The UniG_Mm Track</title>
                <p>This map is comparable to the UniG_Hs track but is based instead on alignment of mouse cDNA sequences (conventional and EST). In this example, the exons suggested by the alignment of mouse cDNAs is comparable to that based on alignment of human cDNAs.</p>
              </sec>
              <sec id="bid.1755">
                <title>SAGE_tag</title>
                <p>SAGE_tag provides another view of expression levels and connections to more information about the tissue of origin of the expressed sequences. The SAGE_tag map also provides a histogram of expression, and each tag is connected to a tag-specific report page.</p>
              </sec>
              <sec id="bid.1703">
                <title>The Contig Track</title>
                <p>Looking across to the Contig track, we can see that the gene maps to a contig that is drawn in <italic>blue</italic>, indicating that the contig is derived from high quality, finished sequence. If we consult the Component map (<italic>Comp</italic>), we can see that the portion of the contig containing the <italic>FMR1</italic> gene is composed of two overlapping finished sequences, also drawn in <italic>blue</italic>. Because the sequence underlying the <italic>FMR1</italic> gene is finished, rather than draft sequence, the <italic>FMR1</italic> sequence and structure are likely to remain stable in future human genome assemblies.</p>
              </sec>
              <sec id="bid.1704">
                <title>The gbDNA Track</title>
                <p>Additional GenBank sequences that align to the genome but were not part of the assembly are shown on the gbDNA track. In this case, a number of short sequences for the individual exons of <italic>FMR1</italic>, the product of intense research on this gene, are aligned to the genome (Figure 4b).</p>
              </sec>
            </sec>
            <sec id="bid.1705">
              <title>5. How Can I Create My Own Transcript Models with the Map Viewer?</title>
              <p>The Map Viewer displays the alignment of transcripts, such as mRNA GenBank sequences and RefSeqs, to genomic sequence and shows the positions of predicted genes, but it does not stop there. By using a utility called the ModelMaker, it is possible to combine the alignment evidence with the results of gene prediction to construct novel transcripts.</p>
              <p>Beginning with the standard Map Viewer display for the gene <italic>FMR1</italic>, we can display the Gscan and Genes_seq maps in parallel, as shown in <xref ref-type="fig" rid="bid.1706">Figure 5</xref>. The GScan map shows gene predictions made by the GenomeScan program. The Genes_seq track shows the exons implied by the composite alignment of transcripts, such as NCBI mRNA RefSeqs, to the genome. There are two <italic>boxed regions</italic> in the Map Viewer display. The <italic>upper boxed</italic> region shows that the GenomeScan model for <italic>FMR1</italic> begins at a point that is part of an intron in the transcript-based model. Furthermore, there is a separate GenomScan model upstream of the first that overlaps with the initial exon of the transcript-based model. It would be of interest to investigate whether the two GenomeScan models could be fused to produce a longer transcript. In the <italic>lower boxed</italic> region, we see that the transcript-based model includes an exon lacking in the GenomeScan model. Perhaps we can create a model transcript based on the fusion of the two GenomeScan models that also includes the extra exon seen in the transcript-based model.<fig id="bid.1706"><label>5</label><caption><title>Use of ModelMaker to test alternative cDNA XSmodels based on GenomeScan predictions of mRNA alignments.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f5" mime-subtype="gif"/></fig></p>
              <p>To attempt this synthesis, we first select the <named-content content-type="book-spec-emph">mm</named-content> link [(<italic>a</italic>)] to the <italic>right</italic> of the <italic>FMR1</italic> gene link on the Genes_seq track to invoke the ModelMaker. A number of alignments between transcript or model sequences and the genomic contig upon which <italic>FMR1</italic> lies, NT_011537 [(<italic>b</italic>)], are given at the <italic>top</italic> of the ModelMaker display, and the implied exons resulting from these alignments are shown just <italic>below</italic>, numbered sequentially in the <named-content content-type="book-spec-emph">Putative exons</named-content> pallet. We may choose any of these exons for inclusion in our model by selecting it. We can also choose a complete set of exons from an existing alignment by selecting the <named-content content-type="book-spec-emph">set</named-content> link next to an alignment. Because we plan to begin with the GenomeScan model, we can select the <named-content content-type="book-spec-emph">set</named-content> link next to the second alignment from the <italic>bottom</italic> to start. This gives us an initial model identical to that of the GenomeScan prediction and yields, as its longest Open Reading Frame (ORF), an ORF of 600 amino acids. We can also see the second, small, GenomeScan-predicted model at the <italic>far left</italic> of the ModelMaker display. We hope to fuse this model with the larger model. The second model comprises three closely spaced exons, resembling a single exon in the display, numbered 2&ndash;4 in the <named-content content-type="book-spec-emph">Putative exons</named-content> pallet. Selecting exons 2 and 3 adds them to the model (<italic>first boxed pair</italic> in the ModelMaker display). At this point, the longest ORF detected in the transcript has increased to 762 amino acids. If we try to include exon 4, however, we drop back to 600 amino acids because of the introduction of an internal stop codon; therefore, we can remove exon 4 from our model with another click. Finally, we can select exon 22, which is the exon from the alignment-based model that we want to include, and notice that the longest ORF detected has risen to 823 amino acids [(<italic>c</italic>)]. Because long ORFs without stop codons occur rarely by chance, this transcript model is promising. To explore further, we can select the <named-content content-type="book-spec-emph">ORF Finder</named-content> link to generate a graphical view of all ORFs found in the transcript, including the longest ORF, and subject the translation of the latter to a BLAST search. In this case, we find that the majority of the predicted protein matches the <italic>FMR1</italic> gene product but that we have introduced some novel peptide sequence at the amino-terminal end as well as some near the carboxy terminus.</p>
              <p>In this example, there were multiple, putative, full-length mRNAs. Please note that ESTs can be added to the display by selecting <named-content content-type="book-spec-emph">add ESTs</named-content> (Figure 5, upper right-hand corner). Additional exons and splicing patterns may then be available to be considered in your model. This feature may be of particular importance if most of the evidence for splicing and exons comes from ESTs rather than complete mRNAs.</p>
            </sec>
            <sec id="bid.1707">
              <title>6. Using the Mouse Map Viewer</title>
              <p>This assembly can be displayed by using Map Viewer for the mouse. In this particular case, the official symbol for the human and mouse genes is the same; therefore, a query by symbol returns the expected result. If this were not the case, however, it is also possible to search the mouse genome by the human sequence using <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/seq/page.cgi?F=MmBlast.html">BLAST the Mouse Genome</ext-link>. Simulating this, you can see that the best match for the human sequence is for a gene labeled LOC207836. If you check this in LocusLink, you will see that this is now annotated as Fmr1.</p>
              <sec id="bid.1708">
                <title>By Human/Mouse Homology Map</title>
                <p>Consider an example in which we want to know whether there is a human/mouse synteny region that includes our favorite gene, <italic>FMR1</italic>. To begin, enter FMR1 into the Map Viewer search box and press the <named-content content-type="book-spec-emph">Find</named-content> button. Selecting the gene name in the results table leads us to a Map Viewer display of the <italic>FMR1</italic> gene with the Genes_seq map, UniG_Hs map, and Genes_cyto maps shown. In <xref ref-type="fig" rid="bid.1709">Figure 6</xref>, the Genes_cyto map has been replaced with the UniG_Mm map. Selecting the <named-content content-type="book-spec-emph">hm</named-content> link (<italic>boxed</italic>) to the <italic>right</italic> of the Genes_seq map leads to the human/mouse homology maps. One can choose between two slightly different variants of human genome assembly (NCBI and UCSC). Three sources of mouse mapping data are available: the Mouse Genome Database (MGD) map, Jackson Lab map, and the Whitehead/MRC Radiation Hybrid map. Either of the two human maps may be compared with either of the two mouse maps.<fig id="bid.1709"><label>6</label><caption><title>The human/mouse synteny region.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f6" mime-subtype="gif"/></fig></p>
                <p>Let us choose the NCBI <italic>versus</italic> MGD variant and then check for synteny in the region of the <italic>FMR1</italic> homologs in the two species. Because our primary interest is the human gene <italic>FMR1</italic>, we select NCBI <italic>versus</italic> MGD_Hs, and the comparative map will appear. The row corresponding to the <italic>FMR1</italic> gene is highlighted (<italic>inset</italic>). The mouse homolog, called <italic>Fmr1</italic>, is also positioned on chromosome X, and the order of genes surrounding it is conserved in both genomes. Available STS data are indicated by a <italic>green dot</italic> that is a hyperlink to UniSTS. Links to the Map Viewer (select <named-content content-type="book-spec-emph">Cytogenetic map position</named-content>) and LocusLink (select <named-content content-type="book-spec-emph">Gene Symbol</named-content>) are available for each pair of genes. The user can examine the pairwise BLAST alignment that is accessible by selecting a chromosome color-coded dot preceding the mouse gene name. A dot-matrix similarity plot on the pairwise alignment page makes it easy to visually estimate the differences in sequences of the two genes. It is interesting to see that the similarity between <italic>FMR1</italic> and <italic>Fmr1</italic> actually extends past the coding regions.</p>
              </sec>
              <sec id="bid.1710">
                <title>Using the Mouse UniGene Map</title>
                <p><xref ref-type="fig" rid="bid.1711">Figure 6</xref> shows a Map Viewer display of human <italic>FMR1</italic> using the Genes_seq map, the UniG_Mm map, and the UniG_Hs map. It is apparent that the EST mapping patterns on the two UniGene maps are similar for <italic>FMR1</italic>. This indicates a similarity at the expressed sequence level but does not indicate a similar gene structure because both sets of ESTs are being mapped to the same human genomic sequence. To investigate similarities in gene structure, it is necessary to use the mouse version of the Map Viewer.
     </p>
              </sec>
              <sec id="bid.1712">
                <title>Using the Mouse Map Viewer</title>
                <p>The most complete assembly of the mouse genome available at NCBI is the Mouse Genome Consortium Version 3 WGS assembly. This assembly can be displayed using the mouse version of the Map Viewer; however, the mouse homolog of <italic>FMR1</italic> is not labeled as such in this assembly; therefore, we cannot find it using a simple search by gene name. We can overcome this obstacle by searching the mouse genome with the human mRNA sequence using mouse genome BLAST. The result of such a search is shown in the <italic>leftmost inset</italic> in Figure 6 and indicates that a gene model labeled <named-content content-type="book-spec-emph">LOC207836</named-content> is the probable homolog. The <italic>double-arrow</italic> links the human and mouse Map Viewer displays of the corresponding genes in the two species; the structures are not identical, but they do share a rather large 3&prime; exon. If we zoom out a bit in the two displays (<italic>two right insets</italic>) to see the surrounding genes, we can observe that the organization of the two genomes is similar in the region of <italic>FMR1</italic> to the extent that a second gene called <italic>FMR2</italic> lies downstream of <italic>FMR1</italic>, whereas the mouse version, <italic>Fmr2</italic>, likewise lies downstream of the mouse <italic>Fmr1</italic> homolog.</p>
              </sec>
            </sec>
            <sec id="bid.1713">
              <title>7. How Can I Find Members of a Gene Family Using the Map Viewer?</title>
              <p>Finding members of a gene family is not straightforward by any means. However, the Map Viewer can be used to flag sets of genes that are related, either by nomenclature or by sequence similarity.</p>
              <sec id="bid.1714">
                <title>By Common Annotation</title>
                <p>Consider the gene <italic>FMR1</italic>. Let us assume we do not know much about genes from this family, but we suppose (recognizing that we may have cause to regret this supposition) that they all share the common root name FMR. We start our search from the main Map Viewer page by entering FMR* (the asterisk is a wild-card symbol) into the <named-content content-type="book-spec-emph">Search</named-content> box and pressing the <named-content content-type="book-spec-emph">Find</named-content> button. This search results in several hits, and some obviously do not belong to the <italic>FMR1</italic> family (cytoplasmic FMR1 interacting proteins 1 and 2, FMRFAL); but two genes on chromosome X (<italic>FMR2</italic> and <italic>FMR3</italic>) and one on chromosome 17 (<italic>FMR1L2</italic>)  look promising. By selecting the chromosome X <named-content content-type="book-spec-emph">all matches</named-content> link, we go to the graphical representation of the genomic region containing three FMR genes.</p>
                <p>Selecting the gene name (the rows for the FMR* query hits are highlighted) invokes a corresponding LocusLink page that serves as a portal to available information for the gene, including the precomputed results of a similarity search against the nr database. Go to the NCBI Reference Sequences section of the LocusLink page and select the BLAST Link (<named-content content-type="book-spec-emph">BL</named-content>) link. BLink displays a schematic representation of BLAST alignments with links to displays of the best hit from each organism, protein domains found in the query sequence, or sequences similar to the query that have known 3D structures.</p>
                <p>When we look at the BLAST summary for <italic>FMR1</italic>, we find neither <italic>FMR2</italic> nor <italic>FMR3</italic>. We did not expect <italic>FMR3</italic> to show up because it had not been mapped on the Genes_seq map, which suggested that its sequence is not yet known. However, why do we not see <italic>FMR2</italic>? If we use LocusLink to retrieve the Reference Sequences for the <italic>FMR1</italic> and <italic>FMR2</italic> gene products (NP_002015 and NP_002016) and perform a pairwise BLAST comparison, we find that there is no significant similarity between the two sequences. Apparently, the names of &ldquo;FMR&rdquo; genes do not reflect common sequence features but rather a physiological condition, &ldquo;fragile X mental retardation&rdquo; syndrome, associated with this gene. In this sense, the two are members of a group or family, but they show sequence similarity only in the pathological trinucleotide repeats (CGG)<sub>n</sub> that are often found upstream of their coding regions.</p>
              </sec>
              <sec id="bid.1715">
                <title>By Precomputed Sequence Similarity</title>
                <p>The BLink page lists two annotated human homologs of the <italic>FMR1</italic> gene, called &ldquo;fragile X mental retardation&rdquo;, autosomal homologs 1 and 2 (FXR1 and FXR2). They show a significant similarity to the FMR1 protein and possess the same functional domains (K-homology RNA-binding domains documented in the LocusLink reports). Selecting the <named-content content-type="book-spec-emph">Score column</named-content> link invokes a page with the results of pairwise BLAST, as well as a visual representation of similarity between the two proteins (a variant of the dot matrix similarity plot).</p>
              </sec>
              <sec id="bid.1716">
                <title>By Sequence Similarity to a Query</title>
                <p>To see whether there are undocumented homologs of <italic>FMR1</italic> in the genome, return to the Map Viewer maps page for the <italic>FMR1</italic> gene and select the <named-content content-type="book-spec-emph">BLAST the Human Genome</named-content> link. Enter the Accession number for the FMR1 protein taken from the LocusLink report (NP_002015) into the <named-content content-type="book-spec-emph">Search</named-content> window, select <named-content content-type="book-spec-emph">Genome</named-content> as the database, <named-content content-type="book-spec-emph">tblastn</named-content> as the BLAST program, and search. A tblastn search takes a protein sequence as a query and translates a nucleotide database in all reading frames to find any coding regions, documented or undocumented, that might code for a protein similar to the query. Such a search is very sensitive because it is tolerant of differences in codon usage as well as of insertions and deletions.</p>
                <p>The results of such a search are shown in <xref ref-type="fig" rid="bid.1717">Figure 7</xref>. BLAST hits to regions on four chromosomes are shown in the genomic overview, indicated by <italic>small tick marks</italic>. There are 14 hits to chromosome X, clustered near the end of the q-arm. These hits are to <italic>FMR1</italic>, which is located in band Xq27.3. There are also 5 hits apiece to chromosomes 3 and 17 [(<italic>a</italic>) and (<italic>b</italic>), respectively]. These  hits are to known homologs of <italic>FMR1, FXR1,</italic> and <italic>FXR2</italic>. Selecting the links below the chromosome graphic or on the appropriate contig link in the table below leads to a graphical display of the hits on the contig, as shown in the two <italic>insets</italic>. Note that the BLAST hits track the exons of the genes. The hit on chromosome 12 is to a hypothetical protein and may be worth further investigation.<fig id="bid.1711"><label>7</label><caption><title>Review of results of a tblastn query against the mouse genome using a human protein sequence.</title><p>The summary of significant BLAST hits is shown in the <italic>top</italic> graphic. Sections (<italic>a</italic>) and (<italic>b</italic>) show expanded views of hits to the related genes <italic>FXR1</italic> and <italic>FXR2</italic> on chromosomes 3 and 17, respectively.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f7" mime-subtype="gif"/></fig></p>
              </sec>
            </sec>
            <sec id="bid.1718">
              <title>8. How Can I Find Genes Encoding a Protein Domain Using the Map Viewer?</title>
              <p>The Map Viewer displays a graphic representation of genomic information with links to related resources that allow it to serve as a springboard for many types of analyses.</p>
              <sec id="bid.1719">
                <title>Finding Protein Domains in a Gene Product</title>
                <p>For example, one might be interested in the domain structure of the gene product. Let us use the human <italic>FMR1</italic> gene as an example. On the main Map Viewer page, enter FMR1 into the <named-content content-type="book-spec-emph">Search for</named-content> window and press the <named-content content-type="book-spec-emph">Find</named-content> button. Select the <named-content content-type="book-spec-emph">FMR1</named-content> link to display three maps (cyto, UniG_Hs, and Genes_seq) of the <italic>FMR1</italic> locus on chromosome X. Selecting the gene name next to the gene model (<italic>FMR1</italic>) leads to the LocusLink page, which is not only a compilation of genomic, genetic, and reference data related to the gene but also a gateway to the external resources and NCBI tools and databases. The LocusLink report for <italic>FMR1</italic> is linked to precomputed BLAST results for the <italic>FMR1</italic> gene product. Selecting <named-content content-type="book-spec-emph">BL</named-content> (BLink, for BLAST link) invokes a page with a schematic view of the BLAST comparison of the FMR1 protein against the non-redundant (nr) database. The BLink page lists database hits to the FMR1 protein, sorted according to their BLAST similarity scores. Results may be reformatted by selecting <named-content content-type="book-spec-emph">Sort by Taxonomy Proximity</named-content> to cluster hits from the same species.</p>
                <p>To see the functional domains that have been identified in the FMR1 protein, select the Conserved Domains Database (CDD) <named-content content-type="book-spec-emph">Search</named-content> button. The resulting page shows two hits to the KH-domain from the SMART database and one hit to the KH domain in the Pfam database. Notice that one of the SMART domain hits is a partial hit, as indicated by the <italic>jagged edge</italic> in the schematic representation.</p>
              </sec>
              <sec id="bid.1720">
                <title>Visualizing 3D structures</title>
                <p>We can easily see whether there exists a three-dimensional structure that includes these conserved domains. The <italic>pink dot</italic> preceding the domain name in the BLink <named-content content-type="book-spec-emph">Description of Alignments</named-content> section  leads to a display of the corresponding three-dimensional structure using Cn3D, the NCBI macromolecular viewer available over the Web. The Pfam and SMART domains are linked to two 3D structures (1K1G_A and 1J4W_A). The CDD page also lists other sequences that have the same domain and shows their multiple sequence alignments.</p>
              </sec>
              <sec id="bid.1721">
                <title>Searching for Similar Domains in Genomic Sequence</title>
                <p>Cut and paste the sequence of the KH-domain from the FMR1 protein and run a genome-specific BLAST search. We will use the tblastn program to compare the 44-amino acid sequence of the KH domain to the nucleotide sequence of the human genome. The results will show us other regions of the genome with the potential to code for this domain. Of course, we already know from the BLink page that there are some autosomal homologs, but we do not know whether they actually contain the KH RNA-binding domain. The obvious caveat is that some pseudogenes might contain the domain as well. Our tblastn search returns 4 hits, one being the <italic>FMR1</italic> gene itself on the X chromosome, and three others on chromosomes 3, 12, and 17. Select the <named-content content-type="book-spec-emph">Genome View</named-content> button in the Genome BLAST results to see the positions of the hits on the chromosomes. One may then select the names of sequences producing significant alignments to invoke a Map Viewer display that shows the corresponding maps and report the names of the loci. As expected, hits to chromosomes 3 and 17 correspond to the autosomal homologs FXR1 and FXR2. The hit to chromosome 12 corresponds to the hypothetical protein HSPC232, and the Map Viewer display for this BLAST hit is shown (<xref ref-type="fig" rid="bid.1717">Figure 8</xref>). The hit is to a segment of intronic sequence, rather than to an exon, and is without supporting human EST alignments. There are, however, mouse ESTs mapping to this region (<italic>lower inset</italic>), and there is also a GenomeScan model that covers the hit; therefore, the BLAST hit may indeed represent coding sequence. Selecting the <named-content content-type="book-spec-emph">Blast hit</named-content> link (<italic>boxed</italic>) leads to an alignment (<italic>upper inset</italic>) that indicates a good match between our 44-amino acid domain sequence and a protein translation of the genomic sequence. If we map this 44-amino acid sequence onto the structure of 1K1G using Cn3D, we see that it covers a module consisting of two alpha helices and a three-stranded beta sheet (<italic>leftmost inset</italic>). It appears reasonable that our domain hit may represent an exon of the <italic>HSPC232</italic> gene. To follow this line of analysis, the next step might be to produce a transcript model incorporating this new exon using the Model Maker (see the Model Maker exercise in this series).<fig id="bid.1717"><label>8</label><caption><title>BLAST hits.</title></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ch24f8" mime-subtype="gif"/></fig></p>
              </sec>
            </sec>
          </body>
        </book-part>
      </body>
</book-part><?mtl /hellip?>
</body>
<?mtl hellip?><back>
<glossary id="bid.1237">
      <title>Glossary</title>
      <def-list>
        <def-item>
          <term id="bid.1238">3-D or 3D</term>
          <def>
            <p>Three-dimensional.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1722">Accession number</term>
          <def>
            <p>An Accession number is a unique identifier given to a sequence when it is submitted to one of the DNA repositories (GenBank, EMBL, DDBJ). The initial deposition of a sequence record is referred to as version 1. If the sequence is updated, the version number is incremented, but the Accession number will remain constant.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1239">
            <italic>Alu</italic>
          </term>
          <def>
            <p>The <italic>Alu</italic> repeat family comprises short interspersed elements (SINES) present in multiple copies in the genomes of humans and other primates. The <italic>Alu</italic> sequence is approximately 300 bp in length and is found commonly in <xref ref-type="other" rid="bid.1323">intron</xref>s, 3&prime; untranslated regions of genes, and intergenic genomic regions. They are mobile elements and are present in the human genome in extremely high copy number. Almost 1 million copies of the <italic>Alu</italic> sequence are estimated to be present, making it the most abundant mobile element. The <italic>Alu</italic> sequence is so named because of the presence of a recognition site for the <italic>Alu</italic>I endonuclease in the middle of the <italic>Alu</italic> sequence. Because of the widespread occurrence of the <italic>Alu</italic> repeat in the genome, the <italic>Alu</italic> sequence is used as a universal primer for PCR in animal cell lines; it binds in both forward and reverse directions. The <italic>Alu</italic> universal primer sequence is as follows: 5&prime;-GTG GAT CAC CTG AGG TCA GGA GTT TC-3&prime; (26-mer).<!--only one source for sequence. Could not independently verify. http://www.accessexcellence.org/LC/SS/PS/PCR/PCR_technology.html--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1240">allele</term>
          <def>
            <p>One of the variant forms of a gene at a particular <xref ref-type="other" rid="bid.1332">locus</xref> on a chromosome. Different alleles produce variation in inherited characteristics such as hair color or blood type. In an individual, one form of the allele (the dominant one) may be expressed more than another form (the recessive one). When &ldquo;genes&rdquo; are considered simply as segments of a nucleotide sequence, allele refers to each of the possible alternative nucleotides at a specific position in the sequence. For example, a CT polymorphism such as CCT[C/T]CCAT would have two alleles: C and T.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1241">API</term>
          <def>
            <p>Application Programming Interface. An API is a set of routines that an application uses to request and carry out lower-level services performed by a computer's operating system. For computers running a graphical user interface, an API manages an application's windows, icons, menus, and dialog boxes.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1242">ASN.1</term>
          <def>
            <p>Abstract Syntax Notation 1 is an international standard data-representation format used to achieve interoperability between computer platforms. It allows for the reliable exchange of data in terms of structure and content by computer and software systems of all types.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1243">BAC</term>
          <def>
            <p>Bacterial Artificial Chromosome. A BAC is a large segment of DNA (100,000&ndash;200,000 bp) from another species cloned into bacteria. Once the foreign DNA has been cloned into the host bacteria, many copies of it can be made.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1244">BankIt</term>
          <def>
            <p>BankIt is a tool for the online submission of one or a few sequences into <xref ref-type="other" rid="bid.1299">GenBank</xref> and is designed to make the submission process quick and easy. (BankIt also automatically uses <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/VecScreen/VecScreen.html">VecScreen</ext-link> to identify segments of nucleic acid sequence that may be of vector, adapter, or linker origin to combat the problem of vector contamination in GenBank.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1245">bit score</term>
          <def>
            <p>The value S&prime; is derived from the raw alignment score S in which the statistical properties of the scoring system used have been taken into account. By normalizing a raw score using the formula: <inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="glossf1" mime-subtype="gif"/> a &ldquo;bit score&rdquo; S&prime; is attained, which has a standard set of units, and where K and <italic>lambda</italic> are the statistical parameters of the scoring system. Because bit scores have been normalized with respect to the scoring system, they can be used to compare alignment scores from different searches.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1246">BLAST</term>
          <def>
            <p>Basic Local Alignment Search Tool (Altschul et al., J Mol Biol 215:403-410; 1990). A sequence comparison <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/BLAST_algorithm.html">algorithm</ext-link> that is optimized for speed and used to search sequence databases for optimal local alignments to a query. See the <xref ref-type="other" rid="bid.610">BLAST chapter</xref> (Chapter 15) or the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/tut1.html">tutorial</ext-link> or the narrative <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/guide.html">guide</ext-link> to BLAST.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1247">blastn</term>
          <def>
            <p>nucleotide&ndash;nucleotide BLAST. blastn takes nucleotide sequences in <xref ref-type="other" rid="bid.1290">FASTA</xref> format, <xref ref-type="other" rid="bid.1299">GenBank</xref> Accession numbers, or <xref ref-type="other" rid="bid.1304">GI</xref> numbers and compares them against the NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/blast/html/blastcgihelp.html#nucleotide_databases">Nucleotide databases</ext-link>.<!--or BLASTn or BLASTN--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1248">blastp</term>
          <def>
            <p>protein&ndash;protein BLAST. blastp takes protein sequences in <xref ref-type="other" rid="bid.1290">FASTA</xref> format, <xref ref-type="other" rid="bid.1299">GenBank</xref> Accession numbers, or <xref ref-type="other" rid="bid.1304">GI</xref> numbers and compares them against the NCBI <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/blast/html/blastcgihelp.html#protein_databases">Protein databases</ext-link>.<!--or BLASTp--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1249">BLAT</term>
          <def>
            <p>A DNA/Protein sequence analysis program to quickly find sequences of 95%<!--%--> and greater similarity of length 40 bases or more. It may miss more divergent or shorter sequence alignments. BLAT on proteins finds sequences of 80%<!--%--> and greater similarity of length 20 amino acids or more. BLAT is not BLAST. (See the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://genome.ucsc.edu/cgi-bin/hgBlat?command=start">BLAT web page</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1250">BLink</term>
          <def>
            <p>BLAST Link. BLink displays the results of <xref ref-type="other" rid="bid.1246">BLAST</xref> searches that have been done for every protein sequence in the Entrez Protein data domain. It can be accessed by following the BLink link displayed beside any hit in the results of an Entrez Protein search. In contrast to Entrez's <named-content content-type="book-spec-emph">Related Sequences</named-content> feature, which lists the titles of similar sequences, BLink displays the graphical output of precomputed <xref ref-type="other" rid="bid.1248">blastp</xref> results against the non-redundant (nr) protein database. The output includes the positions of up to 200 BLAST hits on the query sequence, scores, and alignments. BLink offers a variety of display options, including the distribution of hits by taxonomic grouping, the best hit to each organism, the protein domains in the query sequence, similar sequences that have known 3D structures, and more. Additional options allow you to specify from which taxa you would like to exclude, increase, or decrease the BLAST cutoff score or filter the BLAST hits to show only those from a specific source database, such as <xref ref-type="other" rid="bid.1392">RefSeq</xref> or <xref ref-type="other" rid="bid.1412">SWISS-PROT</xref>. See the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/blinkhelp.html">BLink help document</ext-link> for additional information.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1251">BLOB</term>
          <def>
            <p>Binary Large Object (or binary data object). BLOB refers to a large piece of data, such as a bitmap. A BLOB is characterized by large field values, an unpredictable table size, and data that are formless from the perspective of a program. It is also a keyword designating the BLOB structure, which contains information about a block of data.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1252">BLOSUM 62</term>
          <def>
            <p>Blocks Substitution Matrix. A substitution matrix in which scores for each position are derived from observations of the frequencies of substitutions in blocks of local alignments in related proteins. Each matrix is tailored to a particular evolutionary distance. In the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/Scoring2.html">BLOSUM 62 matrix</ext-link>, for example, the alignment from which scores were derived was created using sequences sharing no more than 62%<!--62%--> identity. Sequences more identical than 62%<!--%--> are represented by a single sequence in the alignment to avoid overweighting closely related family members (Henikoff and Henikoff, Proc Natl Acad Sci U S A 89:10915-10919; 1992).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1253">Boolean</term>
          <def>
            <p>This term refers to binary algebra that uses the logical operators AND, OR, XOR, and NOT; the outcomes consist of logical values (either TRUE or FALSE). The keyword boolean indicates that the expression or constant expression associated with the identifier takes the value TRUE or FALSE. The logical-AND (&amp;&amp;)<!--(&&)--> operator produces the value 1 if both operands have nonzero values; otherwise, it produces the value 0. The logical-OR (׀׀)<!--(||)--> operator produces the value 1 if either of its operands has a nonzero value. The logical-NOT (!)<!--(!)--> operator produces the value 0 if its operand is true (nonzero) and the value 1 if its operand is FALSE (0). The exclusive OR (XOR) operator yields TRUE only if one of its operands are TRUE and the other is FALSE. If both operands are the same (either TRUE or FALSE), the operation yields FALSE.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1723">build</term>
          <def>
            <p>A run of the genome assembly and annotation process of the set of products generated by that run.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1254">CCAP</term>
          <def>
            <p>Cancer Chromosome Aberration Project. CCAP was designed to expedite the definition and detailed characterization of the distinct chromosomal alterations that are associated with malignant transformation. The project is a collaboration among the <xref ref-type="other" rid="bid.1355">NCI</xref>, the <xref ref-type="other" rid="bid.1353">NCBI</xref>, and numerous research labs. <!--ch 10--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1255">CD</term>
          <def>
            <p>Conserved Domain. CD refers to a domain (a distinct functional and/or structural unit of a protein) that has been conserved during evolution. During evolution, changes at specific positions of an amino acid sequence in the protein have occurred in a way that preserve the physico-chemical properties of the original residues, and hence the structural and/or functional properties of that region of the protein.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1256">CDART</term>
          <def>
            <p>Conserved Domain Architecture Retrieval Tool. When given a protein query sequence, CDART displays the functional domains that make up the protein and lists proteins with similar domain architectures. The functional domains for a sequence are found by comparing the protein sequence to a database of conserved domain alignments, <xref ref-type="other" rid="bid.1257">CDD</xref> using <xref ref-type="other" rid="bid.1396">RPS-BLAST</xref>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1257">CDD</term>
          <def>
            <p>Conserved Domain Database. This database is a collection of sequence alignments and profiles representing protein domains conserved during molecular evolution.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1258">cDNA</term>
          <def>
            <p>complementary DNA. A <xref ref-type="other" rid="bid.1274">DNA</xref> sequence obtained by reverse transcription of a messenger RNA (<xref ref-type="other" rid="bid.1351">mRNA</xref>) sequence.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1259">CDS</term>
          <def>
            <p>coding region, coding sequence. CDS refers to the portion of a genomic DNA sequence that is translated, from the start codon to the stop codon, inclusively, if complete.  A partial CDS lacks part of the complete CDS (it may lack either or both the start and stop codons).  Successful translation of a CDS results in the synthesis of a protein.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1724">CEPH</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.cephb.fr/">Centre d'Etude du Polymorphism Humain</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1260">CGAP</term>
          <def>
            <p>Cancer Genome Anatomy Project. CGAP is an interdisciplinary program to identify the human genes expressed in different cancerous states, based on cDNA (<xref ref-type="other" rid="bid.1283">EST</xref>) libraries, and to determine the molecular profiles of normal, precancerous, and malignant cells. The project is a collaboration among the <xref ref-type="other" rid="bid.1355">NCI</xref>, the <xref ref-type="other" rid="bid.1353">NCBI</xref>, and numerous research labs.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1261">CGH</term>
          <def>
            <p>Comparative Genomic Hybidization. CGH is a fluorescent molecular cytogenetic technique that identifies chromosomal aberrations and maps these changes to metaphase chromosomes. CGH can be used to generate a map of DNA copy number changes in tumor genomes. CGH is based on quantitative two-color fluorescence <italic>in situ</italic> hybridization (<xref ref-type="other" rid="bid.1293">FISH</xref>). DNA extracted from tumor cells is labeled in one color (e.g., green) and mixed in a 1:1 ratio with DNA from normal cells, which is labeled in a different color (e.g., red). The mixture is then applied to normal metaphase chromosomes. Portions of the genome that are equally represented in normal and tumor cells will appear orange, regions that are deleted in the tumor sample relative to the normal sample will appear red, and regions that are present in higher copy number in the tumor sample (because of amplification) will appear green. Special image analysis tools are necessary to quantitate the ratio of green-to-red fluorescence to determine whether a given region is more highly represented in the normal or in the tumor sample.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1262">CGI</term>
          <def>
            <p>Common Gateway Interface. A mechanism that allows a Web server to run a program or script on the server and send the output to a Web browser.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1263">cluster</term>
          <def>
            <p>A group that is created based on certain criteria. For example, a gene cluster may include a set of genes whose similar expression profiles are found to be similar according to certain criteria, or a cluster may refer to a group of clones that are related to each other by homology.<!--{check}--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1264">Cn3D</term>
          <def>
            <p>&ldquo;See in 3-D&rdquo; is a structure and sequence alignment viewer for NCBI databases. It allows viewing of 3-D structures and sequence&ndash;structure or structure&ndash;structure alignments. Cn3D can work as a helper application to the browser or as a client&ndash;server application that retrieves structure records from the Molecular Modeling Database (MMDB, see below) directly from the internet. The <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/CN3D/cn3d.shtml">Cn3D homepage</ext-link> provides access to information on how to install the program, a tutorial to get started, and a comprehensive help document.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1725">codon</term>
          <def>
            <p>Sequence of three nucleotides in DNA or mRNA that specifies a particular amino acid during protein synthesis; also called a triplet. Of the 64 possible codons, 3 are stop codons, which do not specify amino acids.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1265">COGs</term>
          <def>
            <p>Clusters of Orthologous Groups (of proteins) were delineated by comparing protein sequences from completely sequenced genomes. Each COG consists of individual proteins or groups of paralogs from at least three lineages and thus corresponds to an ancient conserved domain.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1266">consensus sequence</term>
          <def>
            <p>The nucleotides or amino acids found most commonly at each position in the sequences of homologous DNAs, RNAs, or proteins.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1267">contig</term>
          <def>
            <p>A contiguous segment of the genome made by joining overlapping clones or sequences. A clone contig consists of a group of cloned (copied) pieces of DNA representing overlapping regions of a particular chromosome. A sequence contig is an extended sequence created by merging primary sequences that overlap. A contig map shows the regions of a chromosome where contiguous DNA segments overlap. Contig maps provide the ability to study a complete and often large segment of the genome by examining a series of overlapping clones, which then provide an unbroken succession of information about that region.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1726">Coriell</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://locus.umdnj.edu/nia/">Coriell Institute of Aging Cell Repository</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1268">CPU</term>
          <def>
            <p>Central Processing Unit. The CPU is the computational and control unit of a computer, the device that interprets and executes instructions.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1269">CSS</term>
          <def>
            <p>Cascading Style Sheets. CSS specify the formatting details that control the presentation and layout of <xref ref-type="other" rid="bid.1312">HTML</xref> and <xref ref-type="other" rid="bid.1435">XML</xref> elements. CSS can be used for describing the formatting behavior and text decoration of simply structured XML documents but cannot display structure that varies from the structure of the source data.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1270">Cubby</term>
          <def>
            <p>A tool of <xref ref-type="other" rid="bid.1282">Entrez</xref>, the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/login.fcgi?call=so.SignOn..Login">Cubby</ext-link> stores search strategies that may be updated at any time, stores LinkOut preferences to specify which LinkOut providers have to be displayed in PubMed, and changes the default document delivery service.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1271">DCMS</term>
          <def>
            <p>Data Creation and Maintenance System</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1272">DDBJ</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ddbj.nig.ac.jp/Welcome-e.html">DNA Data Bank of Japan</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1273">definition line</term>
          <def>
            <p>A sequence in FASTA format begins with a single-line description, followed by lines of sequence data. The definition line or description line is distinguished from the sequence data by a &ldquo;greater than&rdquo; (&gt;)<!--(">")--> symbol in the first column (see <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/fasta.html">example</ext-link>); also DEFLINE, as in a flatfile.
		</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1274">DNA</term>
          <def>
            <p>Deoxyribonucleic acid is the chemical inside the nucleus of a cell that carries the genetic instructions for making living organisms. DNA is composed of two anti-parallel strands, each a linear polymer of nucleotides. Each nucleotide has a phosphate group linked by a phosphoester bond to a pentose (a five-carbon sugar molecule, deoxyribose), that in turn is linked to one of four organic bases, adenine, guanine, cytosine, or thymine, abbreviated A, G, C, and T, respectively. The bases are of two types: purines, which have two rings and are slightly larger (A and G); and pyrimidines, which have only one ring (C and T). Each nucleotide is joined to the next nucleotide in the chain by a covalent phosphodiester bond between the 5&prime; carbon of one deoxyribose group and the 3&prime; carbon of the next. DNA is a helical molecule with the sugar&ndash;phosphate backbone on the outside and the nucleotides extending toward the central axis. There is specific base-pairing between the bases on opposite strands in such a way that A always pairs with T and G always pairs with C.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1275">domain</term>
          <def>
            <p>A &ldquo;domain&rdquo; refers to a discrete portion of a protein assumed to fold independently of the rest of the protein and which possesses its own function.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1276">draft sequence</term>
          <def>
            <p>Draft sequence refers to DNA sequence that is not yet finished but is generally of high quality (i.e., an accuracy of greater than 90%<!--%-->). Draft sequence data are mostly in the form of 10,000 base pair-sized fragments, the approximate chromosomal locations of which are known. The following keywords are associated with draft sequence: phase 0, light-pass coverage of a clone, generally only 1×<!--multiplication--> coverage; phase 1, 4&ndash;10× coverage of a <xref ref-type="other" rid="bid.1243">BAC</xref> clone (order and orientation of the fragments are unknown); and phase 2, 4&ndash;10× coverage of a BAC clone (order and orientation of the fragments are known). Phase 3 refers to the completely <xref ref-type="other" rid="bid.1292">finished sequence</xref>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1277">DTD</term>
          <def>
            <p>Document Type Definition. The DTD is an optional part of the prolog of an XML document that defines the rules of the document. It sets constraints for an XML document by specifying which elements are present in the document and the relationships between elements, e.g., which tags can contain other tags, the number and sequence of the tags, and attributes of the tags. The DTD helps to validate the data when the receiving application does not have a built-in description of the incoming data.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1278">DUST</term>
          <def>
            <p>A program for filtering low-complexity regions from nucleic acid sequences.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1279">E-value</term>
          <def>
            <p>Expect value. The E-value is a parameter that describes the number of hits one can &ldquo;expect&rdquo; to see by chance when searching a database of a particular size. It decreases exponentially with the score (S) that is assigned to a match between two sequences. Essentially, the E-value describes the random background noise that exists for matches between sequences. For example, an E-value of 1 assigned to a hit can be interpreted as meaning that in a database of the current size, one might expect to see one match with a similar score simply by chance. This means that the lower the E-value, or the closer it is to &ldquo;0&rdquo;, the higher is the &ldquo;significance&rdquo; of the match. However, it is important to note that searches with short sequences can be virtually identical and have relatively high E-value. This is because the calculation of the E-value also takes into account the length of the query sequence. This is because shorter sequences have a high probability of occurring in the database purely by chance. For more information, see the following <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nih.gov/BLAST/tutorial/Altschul-1.html">tutorial</ext-link>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1280">EC number</term>
          <def>
            <p>A number assigned to a type of enzyme according to a scheme of standardized enzyme nomenclature developed by the Enzyme Commission of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (IUBMB). EC numbers may be found in <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://us.expasy.org/enzyme/">ENZYME</ext-link>, the Enzyme nomenclature database, maintained at the <xref ref-type="other" rid="bid.1289">ExPASy</xref> molecular biology server.<!--Ch 15--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1281">EMBL</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www1.embl-heidelberg.de/">European Molecular Biology Laboratory</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1282">Entrez</term>
          <def>
            <p>Entrez is a retrieval system for searching several linked databases. It provides access to the following NCBI databases: <xref ref-type="other" rid="bid.1387">PubMed</xref>, <xref ref-type="other" rid="bid.1299">GenBank</xref>, Protein, Structure, Genome, PopSet, <xref ref-type="other" rid="bid.1361">OMIM</xref>, Taxonomy, Books, ProbeSet, 3D Domains, <xref ref-type="other" rid="bid.1425">UniSTS</xref>, SNP, and <xref ref-type="other" rid="bid.1257">CDD</xref>. (See the <xref ref-type="other" rid="bid.588">Entrez chapter</xref> or the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Entrez/">Entrez web page</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1283">EST</term>
          <def>
            <p>Expressed Sequence Tag. ESTs are short (usually approximately 300&ndash;500 base pairs), single-pass sequence reads from <xref ref-type="other" rid="bid.1258">cDNA</xref>. Typically, they are produced in large batches. They represent the genes expressed in a given tissue and/or at a given developmental stage. They are tags (some coding, others not) of expression for a given cDNA library. They are useful in identifying full-length genes and in mapping.<!--ESTs--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1284">e-PCR</term>
          <def>
            <p>Electronic <xref ref-type="other" rid="bid.1366">PCR</xref> is used to compare a query sequence to mapped sequence-tagged sites (<xref ref-type="other" rid="bid.1410">STS</xref>s) to find a possible map location for the query sequence. e-PCR finds STSs in DNA sequences by searching for subsequences that closely match the PCR primers present in mapped markers. The subsequences must have the correct order, orientation, and spacing that they could plausibly prime the amplification of a PCR product of the correct molecular weight.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1285">epub citation</term>
          <def>
            <p>&ldquo;Ahead-of-print&rdquo; citation. <xref ref-type="other" rid="bid.1387">PubMed</xref> now accepts citations from publishers for articles that have been published electronically ahead of the printed issue. PubMed displays the category &ldquo;[epub ahead of print]&rdquo; in the part of the citation where the volume and pagination would ordinarily display. For example: Proc Natl Acad Sci U S A. 2000 May 2 [epub ahead of print].</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1286">ExoFish</term>
          <def>
            <p>Exon Finding by Sequence Homology. Exofish is a tool based on homology searches for the rapid and reliable identification of human genes. It relies on the sequence of another vertebrate, the pufferfish <italic>Tetraodon nigroviridis</italic> (similar to Fugu), to detect conserved sequences with a very low background. The genome of <italic>T. nigroviridis</italic> is eight times more compact than the human genome and has been used in the comparative identification of human genes from the rough draft of the human genome (Roest Crollius et al., Nat Genet 25:235-238; 2000).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1287">exon</term>
          <def>
            <p>Refers to the portion of a gene that encodes for a part of that gene's mRNA. A gene may comprise many exons, some of which may include only protein-coding sequence; however, an exon may also include 5' or 3' untranslated sequence.  Each exon codes for a specific portion of the complete protein. In some species (including humans), a gene's exons are separated by long regions of DNA (called <xref ref-type="other" rid="bid.1323">intron</xref>s or sometimes &ldquo;junk DNA&rdquo;) that often have no apparent function but have been shown to encode small untranslated RNAs or regulatory information. (See also <xref ref-type="other" rid="bid.1407">splice sites</xref>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1288">exon-trapped</term>
          <def>
            <p>Exon trapping is a technique for cloning exon sequences from genomic DNA by selecting for functional splice sites, relying on the cellular splicing machinery. The genomic DNA containing the putative exon(s) is cloned into an exon-trap vector, which has a promoter, polyadenylation signals, and splice sites, and then transfected into a cell line. If there are functional splice sites in the genomic DNA fragment, the segments of DNA between the splice sites will be removed. Total RNA is isolated and reverse-transcribed. After <xref ref-type="other" rid="bid.1258">cDNA</xref> synthesis and <xref ref-type="other" rid="bid.1366">PCR</xref> amplification, the exon of interest is cloned.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1289">ExPASy</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.expasy.org/">Expert Protein Analysis System</ext-link> is a proteomics server of the Swiss Bioinformatics Institute (SIB).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1290">FASTA</term>
          <def>
            <p>The first widely used algorithm for similarity searching of protein and DNA sequence databases. The program looks for optimal local alignments by scanning the sequence for small matches called &ldquo;words&rdquo;. Initially, the scores of segments in which there are multiple word hits are calculated (&ldquo;init1&rdquo;). Later, the scores of several segments may be summed to generate an &ldquo;initn&rdquo; score. An optimized alignment that includes gaps is shown in the output as &ldquo;opt&rdquo;. The sensitivity and speed of the search are inversely related and controlled by the &ldquo;k-tup&rdquo; variable, which specifies the size of a &ldquo;word&rdquo; (Pearson and Lipman). Also refers to a <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/BLAST/fasta.html">format</ext-link> for a nucleic acid or protein sequence.<!--This definition was taken from NCBI sitemap. Should there be a figure/formula?--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1291">fingerprint</term>
          <def>
            <p>The pattern of bands on a gel produced by a clone when restricted by a particular enzyme, such as <italic>Hin</italic>dIII.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1292">finished sequence</term>
          <def>
            <p>High-quality, low-error DNA sequence that is free of gaps. To qualify as a finished sequence, only a single error out of every 10,000 bases (i.e., an accuracy of 99.999%<!--%-->) is allowed.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1293">FISH</term>
          <def>
            <p>Fluorescence <italic>in situ</italic> hybridization. In this technique, fluorescent molecules are used to label a <xref ref-type="other" rid="bid.1274">DNA</xref> probe, which can then hybridize to a specific DNA sequence in a chromosome spread so that the site becomes visible through a microscope. FISH has been used to highlight the locations of genes, subchromosome regions, entire chromosomes, or specific DNA sequences. It has been used for mapping and the detection of genomic rearrangements, as well as studies on DNA replication.<!-- Appears in chapter 10--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1294">flatfile or flat file</term>
          <def>
            <p>A flat file is a data file that contains records (each corresponding to a row in a table); however, these records have no structured relationships. To interpret these files, the format properties of the file should be known. For example, a database management system may allow the user to export data to a comma-delimited file. Such a file is called a flat file because it has no inherent information about the data, and interpretation requires additional information. Files in a database management system have more complex storage structures. <!--Flat file, Flatfile--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1727">freeze</term>
          <def>
            <p>To copy changing data so as to preserve the dataset as it existed at a particular point in time. Also used to refer to the resulting set of frozen data.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1295">FTP</term>
          <def>
            <p>File Transfer Protocol. A method of retrieving files over a network directly to the user's computer or to his/her home directory using a set of protocols that govern how the data are to be transported.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1296">gap</term>
          <def>
            <p>A gap is a space introduced into an alignment to compensate for insertions and deletions in one sequence relative to another. To prevent the accumulation of too many gaps in an alignment, introduction of a gap causes the deduction of a fixed amount (the gap score) from the alignment score. Extension of the gap to encompass additional nucleotides or amino acid is also penalized in the scoring of an alignment. (See the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/Alignment_Scores2.html">figure</ext-link> for more information.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1297">GB</term>
          <def>
            <p>gigabytes</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1298">GBFF</term>
          <def>
            <p><xref ref-type="other" rid="bid.1299">GenBank</xref> Flat File. Refers to a format .gbff.<!--check: ch13--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1299">GenBank</term>
          <def>
            <p>GenBank is a database of nucleotide sequences from more than 100,000 organisms. Records that are annotated with coding region features also include amino acid translations. GenBank belongs to an international collaboration of sequence databases that also includes <xref ref-type="other" rid="bid.1281">EMBL</xref> and <xref ref-type="other" rid="bid.1272">DDBJ</xref>. [See the <xref ref-type="other" rid="bid.2">GenBank</xref> chapter (Chapter 1) or the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Genbank/index.html">GenBank web page</ext-link>.]</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1300">genetic code</term>
          <def>
            <p>The instructions in a gene that tell the cell how to make a specific protein. A, T, G, and C are the &ldquo;letters&rdquo; of the <xref ref-type="other" rid="bid.1274">DNA</xref> code; they stand for the chemicals adenine, thymine, guanine, and cytosine, respectively, that make up the nucleotide bases of DNA. Each gene's code combines the four chemicals in various ways to spell out three-letter &ldquo;words&rdquo; that specify which amino acid is needed at every position for making a protein.<!--ch4, ch12--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1301">GenomeScan</term>
          <def>
            <p>A gene identification algorithm that is used to identify exon&ndash;intron structures in genomic DNA sequence.<!--ch10--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1302">genotype</term>
          <def>
            <p>The genetic identity of an individual that does not show as outward characteristics. The genotype refers to the pair of alleles for a given region of the genome that an individual carries.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1303">GEO</term>
          <def>
            <p>Gene Expression Omnibus. GEO is a gene expression data repository and online resource for the retrieval of gene expression data from any organism or artificial source. Many types of gene expression data from platform types, such as spotted microarray, high-density oligonucleotide array, hybridization filter, and serial analysis of gene expression (<xref ref-type="other" rid="bid.1397">SAGE</xref>) data, are accepted, accessioned, and archived as a public dataset. [See the <xref ref-type="other" rid="bid.337">GEO chapter</xref> (Chapter 6) or the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo/">GEO web page</ext-link>.]</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1304">GI</term>
          <def>
            <p>The GenInfo Identifier is a sequence identification number for a nucleotide sequence. If a nucleotide sequence changes in any way, a new GI number will be assigned. A separate GI number is also assigned to each protein translation within a nucleotide sequence record, and a new GI is assigned if the protein translation changes in any way. GI sequence identifiers run parallel to the new accession.version system of sequence identifiers (see the description of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Sitemap/samplerecord.html#VersionB">Version</ext-link>).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1305">GSS</term>
          <def>
            <p>Genome Survey Sequences are analogous to <xref ref-type="other" rid="bid.1283">EST</xref>s except that the sequences are genomic in origin, rather than cDNA (mRNA). The GSS division of <xref ref-type="other" rid="bid.1299">GenBank</xref> contains (but is not limited to) the following types of data: random &ldquo;single-pass read&rdquo; genome survey sequences, cosmid/<xref ref-type="other" rid="bid.1243">BAC</xref>/<xref ref-type="other" rid="bid.1438">YAC</xref> end sequences, <xref ref-type="other" rid="bid.1288">exon-trapped</xref> genomic sequences, and <xref ref-type="other" rid="bid.1239"><italic>Alu</italic></xref>-PCR sequences.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1306">heterozygosity</term>
          <def>
            <p>The probability that a diploid individual will have two different alleles at a particular genome locus. These individuals are defined as heterozygous, whereas individuals who have two identical alleles at the locus are defined as homozygous. The probability can be estimated by sampling a representative number of individuals from the population and dividing the number of heterozygotes by the total number sampled.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1307">HIV</term>
          <def>
            <p>Human Immunodeficiency Virus. HIV-1 is a retrovirus that is recognized as the causative agent of AIDS (Acquired Immunodeficiency Syndrome).<!--ch4, ch18, ch19--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1308">HNPCC</term>
          <def>
            <p>Hereditary nonpolyposis colon cancer</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1309">homogeneously staining region</term>
          <def>
            <p>A region of the chromosome identified cytologically by DNA staining or the <xref ref-type="other" rid="bid.1293">FISH</xref> technique because of the presence of multiple copies of a subchromosomal region resulting from amplification.<!--ch10--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1310">homologous</term>
          <def>
            <p>The term refers to similarity attributable to descent from a common ancestor. Homologous chromosomes are members of a pair of essentially identical chromosomes, each derived from one parent. They have the same or allelic genes with genetic loci arranged in the same order. Homologous chromosomes synapse during meiosis.<!--check this definition--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1311">HTGS</term>
          <def>
            <p>High-Throughput Genomic Sequences. The source of HTGS are large-scale genome sequencing centers; <xref ref-type="other" rid="bid.1423">unfinished sequence</xref>s are in phases 0, 1, and 2, and <xref ref-type="other" rid="bid.1292">finished sequence</xref>s are in phase 3.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1728">HTGS_CANCELLED</term>
          <def>
            <p>A keyword added to GenBank entries by sequencing centers to indicate that work has stopped on a clone and that the existing sequence will not be finished. Sequencing centers may stop work because the clone is redundant or for various other reasons.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1729">HTGS_PHASE0, HTGS_PHASE1, HTGS_PHASE2, HTGS_PHASE3</term>
          <def>
            <p>Keywords added to GenBank entries by sequencing centers to indicate the status (phase) of the sequence (see phase definitions described under <xref ref-type="other" rid="bid.1276">draft sequence</xref>).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1312">HTML</term>
          <def>
            <p>Hypertext Markup Language. HTML is derived from <xref ref-type="other" rid="bid.1400">SGML</xref>. It is a text-based mark-up language and is used to primarily display information using a web browser and to link pieces of information via hyperlinks. The tags used in an HTML document provide information only on how the content is to be displayed but do not provide information about the content they encompass.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1313">HUP</term>
          <def>
            <p>Hold Until Published. HUP refers to the category for data that is electronically submitted for when it should be released to the public.<!--{check}--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1314">ICBN</term>
          <def>
            <p>International Code of Botanical Nomenclature</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1315">ICD</term>
          <def>
            <p>International Classification of Diseases<!--ICD-030. Version: ICD, third edition. Now at version tenth edition; ICD-10. Just left the abbreviation for ICD--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1316">ICD-O-3</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://training.seer.cancer.gov/module_icdo3/icdo3_home.html">International Classification of Diseases for Oncology, 3rd edition</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1317">ICNB</term>
          <def>
            <p>International Code of Nomenclature of Bacteria</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1318">ICNCP</term>
          <def>
            <p>International Code of Nomenclature for Cultivated Plants</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1319">ICTV</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/ICTV/">International Committee on Taxonomy of Viruses</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1320">ICVCN</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/ICTV/rules.html">International Code of Virus Classification and Nomenclature</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1321">ICZN</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.iczn.org/">International Code of Zoological Nomenclature</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1730">ideogram</term>
          <def>
            <p>A diagrammatic representation of the karyotype of an organism.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1322">IMAGE Consortium</term>
          <def>
            <p>Integrated Molecular Analysis of Genomes and their Expression. A consortium of academic groups that share high-quality, arrayed cDNA libraries and place sequence, map, and expression data of the clones in these arrays into the public domain. With the use of this information, unique clones can be rearrayed to form a &ldquo;master array&rdquo;, with the aim of ultimately having a representative cDNA from every gene in the genome under study. To date, human, mouse, rat, zebrafish, and <italic>Xenopus laevis</italic> genomes have been studied.<!--Chapter 21, reference to paper in text.--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1323">intron</term>
          <def>
            <p>Refers to that portion of the DNA sequence that is present in the primary transcript and that is removed by splicing during RNA processing and is not included in the mature, functional <xref ref-type="other" rid="bid.1351">mRNA</xref>, rRNA, or tRNA. Also called an intervening sequence. (See also <xref ref-type="other" rid="bid.1407">splice sites</xref>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1324">ISAM</term>
          <def>
            <p>Indexed Sequential-Access Method. ISAM is a database access method. It allows data records in a database to be accessed either sequentially (in the order in which they were entered) or randomly (using an index). In the index, each record has a unique key that enables its rapid location. The key is the field used to reference the record.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1325">ISCN</term>
          <def>
            <p>International System for Human Cytogenetic Nomenclature<!--ch10 and ch19--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1326">ISO</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.iso.ch/iso/en/ISOOnline.openerpage">International Organization for Standardization</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1327">ISSN</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.issn.org:8080/English/pub/">International Standard Serial Number</ext-link>. The ISSN is an eight-digit number that identifies periodical publications, including electronic serials.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1328">karyotype</term>
          <def>
            <p>The particular chromosome complement of an individual or a related group of individuals, as defined by both the number and morphology of the chromosomes, usually in mitotic metaphase, and arranged by pairs according to the standard classification.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1329">LANL</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.lanl.gov/worldview/">Los Alamos National Lab</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1330">LIMS</term>
          <def>
            <p>Laboratory Information Management Systems. LIMS comprise software that helps biological and chemical laboratories handle data generation, information management, and data archiving.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1331">LinkOut</term>
          <def>
            <p>A registry service to create links from specific articles, journals, or biological data in <xref ref-type="other" rid="bid.1282">Entrez</xref> to resources on external web sites. Third parties can provide a URL, resource name, brief description of their web sites, and specification of the NCBI data from which they would like to establish links. The specification can be written as a valid Boolean query to Entrez or as a list of identifiers for specific articles or sequences. Entrez PubMed users can then select which external links are visible in their searches through the NCBI Cubby service (see above). (See the <xref ref-type="other" rid="bid.650">LinkOut</xref> chapter or <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/linkout/doc/linkoutoverview.html">web page</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1332">locus</term>
          <def>
            <p>In a genomic contect, locus refers to position on a chromosome. It may, therefore, refer to a marker, a gene, or any other landmark that can be described.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1333">LocusID</term>
          <def>
            <p>Each new <xref ref-type="other" rid="bid.1334">LocusLink</xref> record is assigned a unique identifying number&mdash;a LocusID (although coding regions on genomic sequences found by gene prediction software are an exception to this).
		</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1334">LocusLink</term>
          <def>
            <p>LocusLink provides a single query interface to curated sequence and descriptive information about genetic loci. LocusLink issues a stable ID (<xref ref-type="other" rid="bid.1333">LocusID</xref>) for each locus and presents information on official nomenclature, aliases, sequence Accession numbers, phenotypes, <xref ref-type="other" rid="bid.1280">EC number</xref>s, OMIM numbers, UniGene clusters, map information, and relevant web sites. LocusLink is a collaborative effort among <xref ref-type="other" rid="bid.1353">NCBI</xref>, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.gene.ucl.ac.uk/nomenclature/">Human Gene Nomenclature Committee</ext-link>, <xref ref-type="other" rid="bid.1361">OMIM</xref>, and others. LocusLink currently contains human, mouse, rat, zebrafish, and fruit fly loci; organisms can be searched together or separately. (See the <xref ref-type="other" rid="bid.788">LocusLink</xref> chapter or <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/LocusLink/">web page</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1335">MACAW</term>
          <def>
            <p>Multiple Alignment Construction and Analysis Workbench. MACAW is a program for locating, analyzing, and editing blocks of localized sequence similarity among multiple seqences and linking them into a composite multiple alignment.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1336">Map Viewer</term>
          <def>
            <p>The Map Viewer is a software component of <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Genome">Entrez Genomes</ext-link> that provides special browsing capabilities for a subset of organisms. It allows one to view and search an organism's complete genome, display chromosome maps, and zoom into progressively greater levels of detail, down to the sequence data for a region of interest. If multiple maps are available for a chromosome, it displays them aligned to each other based on shared marker and gene names and, for the sequence maps, based on a common sequence coordinate system. The organisms currently represented in the Map Viewer are listed in the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PMGifs/Genomes/MapViewerHelp.html">Entrez Map Viewer help document</ext-link>, which provides general information on how to use that tool. The number and types of available maps vary by organism and are described in the &ldquo;data and search tips&rdquo; file provided for each organism.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1337">MB</term>
          <def>
            <p>megabytes</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1339">MEDLINE</term>
          <def>
            <p>MEDLINE is <xref ref-type="other" rid="bid.1358">NLM</xref>'s database of indexed journal citations and abstracts in the fields of biomedicine and healthcare. It encompasses nearly 4,500 journals published in the United States and more than 70 other countries. (For more information, see the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/pubs/factsheets/dif_med_pub.html">Fact Sheet</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1340">MegaBLAST</term>
          <def>
            <p>MegaBLAST is a program for aligning sequences that differ slightly as a result of sequencing or other similar &ldquo;errors&rdquo;. When larger word size is used, it is up to 10 times faster than more common sequence-similarity programs. MegaBLAST is also able to efficiently handle much longer DNA sequences than the <xref ref-type="other" rid="bid.1247">blastn</xref> program of the traditional BLAST algorithm. It uses the GREEDY algorithm for a nucleotide sequence alignment search.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1341">MeSH</term>
          <def>
            <p>Medical Subject Headings. MeSH refers to the controlled vocabulary of <xref ref-type="other" rid="bid.1358">NLM</xref> used for indexing articles in PubMed. MeSH terminology provides a consistent way to retrieve information that may use different terminology for the same concepts. (See the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/mesh/meshhome.html">MeSH homepage</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1342">Metathesaurus</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ncievs.nci.nih.gov/indexMetaphrase.html">Metathesaurus</ext-link> is a National Cancer Institute browser containing different biomedical vocabularies, including the International Classification of Diseases for Oncology <xref ref-type="other" rid="bid.1316">ICD-O-3</xref>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1343">mFASTA</term>
          <def>
            <p>Multi-FASTA format.
			<!--check--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1344">MGC</term>
          <def>
            <p>Mammalian Gene Collection. <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://mgc.nci.nih.gov/">MGC</ext-link> is a project of the <xref ref-type="other" rid="bid.1357">NIH</xref> to provide a complete set of full-length (open reading frame) sequences and cDNA clones of expressed genes for human and mouse.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1345">MGD</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.informatics.jax.org/mgihome/MGD/aboutMGD.shtml">Mouse Genome Database</ext-link>. MGD contains information on mouse genetic markers, molecular segments, phenotypes, comparative mapping data, experimental mapping data, and graphical displays for genetic, physical, and cytogenetic maps.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1346">MGI</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.informatics.jax.org/mgihome/">Mouse Genome Informatics</ext-link>. MGI houses a database that provides integrated access to data on the genetics, genomics, and biology of the laboratory mouse.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1347">microsatellite</term>
          <def>
            <p>Repetitive stretches of short sequences of DNA used as genetic markers to track inheritance in families (e.g., CC[TATATATA]CCCT). Also known as short tandem repeats (STRs).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1348">MIM</term>
          <def>
            <p>Mendelian Inheritance in Man. First published in 1966, <italic><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.press.jhu.edu/press/books/titles/f97/f97mcme.htm">Mendelian Inheritance in Man (MIM)</ext-link></italic> is a genetic knowledge base that serves clinical medicine and biomedical research, including the Human Genome Project.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1731">minimal tiling path</term>
          <def>
            <p>An ordered list or map that defines the minimal set of overlapping clones needed to provide complete coverage of a chromosome or other extended segment of DNA (compare with <xref ref-type="other" rid="bid.1419">tiling path</xref>).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1349">MMDB</term>
          <def>
            <p>Molecular Modeling Database. MMDB is a database of three-dimensional biomolecular structures derived from X-ray crystallography and nuclear magnetic resonance (NMR) spectroscopy.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1350">MMDB-ID</term>
          <def>
            <p>Molecular Modeling Database Accession number.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1351">mRNA</term>
          <def>
            <p>messenger RNA. mRNA describes the section of a genomic DNA sequence that is transcribed, and can include the 5' untranslated region (5'UTR), <xref ref-type="other" rid="bid.1259">CDS</xref>, and 3' untranslated region (3'UTR).  Successful translation of the CDS section of an mRNA results in the synthesis of a protein.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1352">mutation</term>
          <def>
            <p>A permanent structural alteration in DNA. In most cases, DNA changes have either no effect or cause harm, but occasionally a mutation can improve an organism's chance of surviving, and the beneficial change is passed on to the organism's descendants. Typically, mutations are more rare than polymorphisms in population samples because natural selection recognizes their lower fitness and removes them from the population.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1353">NCBI</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1354">NCBI Toolkit</term>
          <def>
            <p>Contains supported software tools from the Information Engineering Branch (IEB) of the NCBI. The NCBI Toolkit describes the three components of the ToolBox: data model, data encoding, and programming libraries. Provides access to documentation for the DataModel, C Toolkit, C++ Toolkit, NCBI C Toolkit Source Browser, XML Demo Program, XML DTDs, and the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/toolbox/">FTP site</ext-link>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1355">NCI</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nci.nih.gov/">National Cancer Institute</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1356">NEXUS</term>
          <def>
            <p>NEXUS refers to a file format designed to contain data for processing by computer programs. NEXUS files should end with .nxs or .nex for purposes of clarity (Maddison et al., Syst Biol 46:590-621; 1997).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1357">NIH</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nih.gov/">National Institutes of Health</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1358">NLM</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/">National Library of Medicine</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1359">NMR</term>
          <def>
            <p>Nuclear Magnetic Resonance. NMR is a spectroscopic technique used for the determination of protein structure.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1360">nr-PDB</term>
          <def>
            <p>non-redundant Protein Data Bank</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1361">OMIM</term>
          <def>
            <p>Online Mendelian Inheritance in Man. OMIM is a directory of human genes and genetic disorders, with links to literature references, sequence records, maps, and related databases.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1362">ortholog</term>
          <def>
            <p>Orthology describes genes in different species that derive from a single ancestral gene in the last common ancestor of the respective species.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1363">orthology</term>
          <def>
            <p>Orthology describes genes in different species that derive from a common ancestor, i.e., they are direct evolutionary counterparts.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1364">paralog</term>
          <def>
            <p>A paralog is one of a set of homologous genes that have diverged from each other as a consequence of gene duplication. For example, the mouse &alpha;-<italic>globin</italic> and &beta;-<italic>globin</italic> genes are paralogs. The relationship between mouse &alpha;-<italic>globin</italic> and chick &beta;-<italic>globin</italic> is also considered paralogous (see the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/Orthology.html">figure</ext-link>).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1365">paralogy</term>
          <def>
            <p>Paralogy describes the relationship of homologous genes that arose by gene duplication.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1366">PCR</term>
          <def>
            <p>Polymerase Chain Reaction. A technique for amplifying a specific DNA segment in a complex mixture. Also present in the DNA mixture are short oligonucleotide primers to the DNA segment of interest and reagents for DNA synthesis. PCR relies on the ability of DNA to separate into its two complementary strands at high temperature (a process called denaturation) and for the two strands to anneal at an optimal lower temperature (annealing). The annealing phase is followed by a DNA synthesis step at an optimal temperature for a heat-stable DNA polymerase. After multiple rounds of denaturation, annealing, and DNA synthesis, the DNA sequence specified by the oligonucleotide primers is amplified.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1367">PDB</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.rcsb.org/pdb/">Protein Data Bank</ext-link>. The PDB is a database for 3D macromolecular structure data.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1368">Pfam</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://pfam.wustl.edu/index.html">Pfam</ext-link> is a database housing a large collection of multiple sequence alignments and hidden Markov models covering many common protein domains.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1369">phenotype</term>
          <def>
            <p>The observable traits or characteristics of an organism, e.g., hair color, weight, or the presence or absence of a disease. Phenotypic traits are not necessarily genetic.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1370">PHRAP</term>
          <def>
            <p>A computer program that assembles raw sequence into sequence contigs (see above) and assigns to each position in the sequence an associated &ldquo;quality score&rdquo;, on the basis of the <xref ref-type="other" rid="bid.1371">PHRED</xref> scores of the raw sequence reads. A PHRAP quality score of <italic>X</italic> corresponds to an error probability of approximately 10<sup>-<italic>X</italic>/10</sup><!--10-X/10-->. Thus, a PHRAP quality score of 30 corresponds to 99.9%<!--%--> accuracy for a base in the assembled sequence.<!--From Nature Genome glossary--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1371">PHRED</term>
          <def>
            <p>A computer program that analyses raw sequence to produce a &ldquo;base call&rdquo; with an associated &ldquo;quality score&rdquo; for each position in the sequence. A PHRED quality score of <italic>X</italic> corresponds to an error probability of approximately 10<sup>-<italic>X</italic>/10</sup><!--10-X/10-->. Thus, a PHRED quality score of 30 corresponds to 99.9%<!--entity for % sign--> accuracy for the base call in the raw read.<!--From Nature Genome glossary--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1372">phyletic pattern</term>
          <def>
            <p>Pattern of presence&ndash;absence of a cluster of orthologs (COG) in different species.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1373">PHYLIP</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://evolution.genetics.washington.edu/phylip.html">PHYLogeny Inference Package</ext-link>. A package of programs for various computer platforms to infer phylogenies or evolutionary trees, freely available from the Web.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1374">PIR</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://pir.georgetown.edu/">Protein Information Resource</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1375">PMC</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.pubmedcentral.nih.gov/">PubMed Central</ext-link>. NLM's digital archive of life sciences journal literature.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1376">PMID</term>
          <def>
            <p>PubMed ID number</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1377">PNG</term>
          <def>
            <p>Portable Network Graphics. An extensible file format for the lossless, well-compressed storage of raster images (images that are composed of horizontal lines of pixels, such as those created by a computer screen). Compression of image, media, and application files is necessary to reduce the transmission time across the web. The technique of lossless compression reduces the size of the file without sacrificing any original data, and the image after expansion is exactly as it was before compression. PNG overcomes the patent issues of GIF (Graphic Interchange Format) and can replace many common uses of TIFF (Tagged Image File Format). Several features such as indexed color, grayscale, and truecolor are supported, as well as an optional alpha-channel. PNG is designed to work well in online viewing applications and is supported as an image standard by the <xref ref-type="other" rid="bid.1434">WWW</xref>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1378">poly A</term>
          <def>
            <p>A string of adenylic acid residues that are added to the 3&prime; end of the primary <xref ref-type="other" rid="bid.1351">mRNA</xref> transcript. Poly(A) polymerase is the enzyme that adds the poly A tail, which is between 100 and 250 bases long.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1379">polymorphism</term>
          <def>
            <p>A common variation in the sequence of <xref ref-type="other" rid="bid.1274">DNA</xref> among individuals. Genetic variations occurring in more than 1%<!--%--> of the population would be considered useful polymorphisms for genetic linkage analysis.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1732">polypeptide</term>
          <def>
            <p>Linear polymer of amino acids connected by peptide bonds. Proteins are large polypeptides, and the two terms are commonly used interchangeably.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1380">PRF</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.prf.or.jp/en/">Protein Research Foundation</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1381">private polymorphism</term>
          <def>
            <p>Variations that are only common in specific populations. Usually such populations are reproductively isolated from other, larger groups. These variations may be completely absent in other groups.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1382">ProtEST</term>
          <def>
            <p>A database of protein sequences from eight organisms: human (<italic>Homo sapiens</italic>), mouse (<italic>Mus musculus</italic>), rat (<italic>Rattus norvegicus</italic>), fruitfly (<italic>Drosophila melanogaster</italic>), worm (<italic>Caenorhabditis elegans</italic>), yeast (<italic>Saccharomyces cerevisiae</italic>), plant (<italic>Arabidopsis thaliana</italic>), and bacteria (<italic>Escherichia coli</italic>). (See the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/UniGene/ProtEST/">ProtEST web page</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1383">PROW</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/PROW/">Protein Reviews On the Web</ext-link>. An online resource that features PROW Guides&mdash;authoritative, short, structured reviews on proteins and protein families. The Guides provide approximately 20 standardized categories of information (abstract, biochemical function, ligands, references, etc.) for each protein.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1384">pseudogene</term>
          <def>
            <p>A sequence of DNA that is very similar to a normal gene but that has been altered slightly so that it is not expressed. Such genes were probably once functional but, over time, acquired one or more mutations that rendered them incapable of producing a protein product.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1385">PSI-BLAST</term>
          <def>
            <p>Position-Specific Iterated BLAST. PSI-BLAST (Altschul et al., J Mol Biol 215:403-410; 1990) is used for iterative protein&ndash;sequence similarity searches using a position-specific score matrix (<xref ref-type="other" rid="bid.1386">PSSM</xref>). It is a program for searching protein databases using protein queries to find other members of the same protein family. All statistically significant alignments found by <xref ref-type="other" rid="bid.1246">BLAST</xref> are combined into a multiple alignment, from which a PSSM is constructed. This matrix is used to search the database for additional significant alignments, and the process may be iterated until no new alignments are found.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1386">PSSM</term>
          <def>
            <p>Position-Specific Score Matrix. The PSSM gives the log-odds score for finding a particular matching amino acid in a target sequence.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1387">PubMed</term>
          <def>
            <p>A retrieval system containing citations, abstracts, and indexing terms for journal articles in the biomedical sciences. It includes literature citations supplied directly to NCBI by publishers as well as <xref ref-type="other" rid="bid.1427">URL</xref>s to full text articles on the publishers' web sites. PubMed contains the complete contents of the <xref ref-type="other" rid="bid.1339">MEDLINE</xref> and PREMEDLINE databases. It also contains some articles and journals considered out of scope for MEDLINE, based on either content or on a period of time when the journal was not indexed and, therefore, is a superset of MEDLINE.<!--Link to Pubmed: http://www.ncbi.nlm.nih.gov/Sitemap/index.html#PubMed--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1388">PXML</term>
          <def>
            <p>PubMed Central XML file</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1389">QBLAST</term>
          <def>
            <p>A queuing system to BLAST that allows users to retrieve their results at their convenience and format their results multiple times with different formatting options.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1390">QTL</term>
          <def>
            <p>Quantitative Trait Locus. A QTL is a hypothesis that a certain region of the chromosome contains genes that contribute significantly to the expression of a complex trait. QTLs are generally identified by comparing the linkage of polymorphic molecular markers and phenotypic trait measurements. The density of the linkage map is important in the accurate and precise location of QTLs; the higher the map density, the more precise the location of the putative QTL, although there is increased likelihood that false positives will be detected. Once QTLs have been mapped to a relatively small chromosomal region, other molecular methods can be used to isolate specific genes.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1391">RCSB</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.rcsb.org/index.html">Research Collaboratory for Structural Bioinformatics</ext-link>. RCSB is a nonprofit consortium that works toward the elucidation of biological, macromolecular, 3-D structures.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1733">Reciprocal best hits</term>
          <def>
            <p>Reciprocal best hits are proteins from different organisms that are each other's top BLAST hit, when the proteomes from those organisms are compared to each other. For example, proteins A&ndash;Z in organism 1 are compared against proteins AA&ndash;ZZ in organism 2. If protein A has a best hit to protein RR, and RR's best hit, when it is compared to all the proteins in organism 1, also turns out to protein A, then A and RR are reciprocal best hits. However, if RR's best hit is to B rather than to A, then A and RR are not reciprocal best hits.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1392">RefSeq</term>
          <def>
            <p>RefSeq is the NCBI database of reference sequences; a curated, non-redundant set including genomic DNA contigs, mRNAs and proteins for known genes, and entire chromosomes.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1734">RepeatMasker</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://ftp.genome.washington.edu/">Program</ext-link> that screens DNA sequences for interspersed repeats and low-complexity DNA sequences.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1393">RFLP</term>
          <def>
            <p>Restriction Fragment Length Polymorphism. Genetic variations at the site where a restriction enzyme cuts a piece of DNA. Such variations affect the size of the resulting fragments. These sequences can be used as markers on physical maps and linkage maps. RFLP is also pronounced &ldquo;rif lip&rdquo;.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1394">RH map</term>
          <def>
            <p>Radiation Hybrid map. A genome map in which <xref ref-type="other" rid="bid.1410">STS</xref>s are positioned relative to one another on the basis of the frequency with which they are separated by radiation-induced breaks. The frequency is assayed by analyzing a panel of human&ndash;hamster hybrid cell lines. These hybrids are produced by irradiating human cells, which damages the cells and fragments the DNA. The dying human cells are fused with thymidine kinase negative (TK&minus;) live hamster cells. The fused cells are grown under conditions that select against hamster cells and favor the growth of hybrid cells that have taken up the human <italic>TK</italic> gene. In the RH maps, the unit of distance is centirays (cR), denoting a 1%<!--%--> chance of a break occurring between two loci.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1395">RNA</term>
          <def>
            <p>Ribonucleic Acid. A single-stranded nucleic acid, similar to <xref ref-type="other" rid="bid.1274">DNA</xref>, but having a ribose sugar, instead of deoxyribose, and uracil instead of thymine as one of its bases.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1396">RPS-BLAST</term>
          <def>
            <p>Reverse Position-Specific BLAST. A program used to identify conserved domains in a protein query sequence. It does this by comparing a query protein sequence to position-specific score matrices (<xref ref-type="other" rid="bid.1386">PSSM</xref>)s that have been prepared from conserved domain alignments. RPS-BLAST is a &ldquo;reverse&rdquo; version of position-specific iterated BLAST (<xref ref-type="other" rid="bid.1385">PSI-BLAST</xref>); however, RPS-BLAST compares a query sequence against a database of profiles prepared from ready-made alignments, whereas PSI-BLAST builds alignments starting from a single protein sequence.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1397">SAGE</term>
          <def>
            <p>Serial Analysis of Gene Expression. An experimental technique designed to quantitatively measure gene expression.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1398">Sequin</term>
          <def>
            <p>Sequin is a stand-alone software tool developed by the <xref ref-type="other" rid="bid.1353">NCBI</xref> for submitting and updating entries to the <xref ref-type="other" rid="bid.1299">GenBank</xref>, <xref ref-type="other" rid="bid.1281">EMBL</xref>, or <xref ref-type="other" rid="bid.1272">DDBJ</xref> sequence databases. It is capable of handling simple submissions that contain a single, short mRNA sequence and complex submissions containing long sequences, multiple annotations, segmented sets of DNA, or phylogenetic and population studies.<!--from http://www.ncbi.nlm.nih.gov/Sequin/index.html}--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1399">SGD</term>
          <def>
            <p>Saccharomyces Genome Database. A database for the molecular biology and genetics of <italic>Saccharomyces cerevisceae</italic>, also known as baker's yeast.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1400">SGML</term>
          <def>
            <p>Standard Generalized Markup Language. The international standard for specifying the structure and content of electronic documents. SGML is used for the markup of data in a way that is self-describing. SGML is not a language but a way of defining languages that are developed along its general principles. A subset of SGML called <xref ref-type="other" rid="bid.1435">XML</xref> is more widely used for the markup of data. <xref ref-type="other" rid="bid.1312">HTML</xref> (Hypertext Markup Language) is based on SGML and uses some of its concepts to provide a universal markup language for the display of information and the linking of different pieces of that information.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1401">SKY</term>
          <def>
            <p>Spectral Karyotyping. SKY is a technique that allows for the visualization of all of an organism's chromosomes together, each labeled with a different color. This is achieved by using chromosome-specific, single-stranded DNA probes (each labeled with a different fluorophore) to hybridize or bind to the chromosomes of a cell; resulting in each chromosome being painted a different color. This technique is useful for identifying chromosome abnormalities because it is easy to spot instances where a chromosome painted in one color has a small piece of another chromosome, painted in a different color, attached to it. (Also see <xref ref-type="other" rid="bid.1293">FISH</xref>, <xref ref-type="other" rid="bid.1261">CGH</xref>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1402">SKYGRAM</term>
          <def>
            <p>1. A software tool to automatically convert the short-form karyotype into an image representation of a cell or clone, with each chromosome displayed in a different color, with band overlay. The program will also incorporate the number of cells for each structural abnormality, which is displayed in brackets. 2. The full ideogram or a cell or clone, with each chromosome displayed in a different color, with band overlay.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1403">SMART</term>
          <def>
            <p>Simple Modular Architecture Research Tool. A tool to allow automatic identification and annotation of domains in user-supplied protein sequences. For example, the <xref ref-type="other" rid="bid.1412">SWISS-PROT</xref> database is an extensively annotated and nonredundant collection of protein sequences. SWISS-PROT annotations have been mined for SMART-derived annotations of alignments.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1404">SMD</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://genome-www5.stanford.edu/MicroArray/SMD/index.shtml">Stanford Microarray Database</ext-link>. SMD stores raw and normalized data from microarray experiments, as well as their corresponding image files. In addition, the SMD provides interfaces for data retrieval, analysis, and visualization. Data are released to the public at the researcher's discretion or upon publication.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1405">SNP</term>
          <def>
            <p>Common, but minute, variations that occur in human DNA at a frequency of 1 every 1,000 bases. An SNP is a single base-pair site within the genome at which more than one of the four possible base pairs is commonly found in natural populations. Several hundred thousand SNP sites are being identified and mapped on the sequence of the genome, providing the densest possible map of genetic differences. SNP is pronounced &ldquo;snip&rdquo;.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1406">SOFT</term>
          <def>
            <p>Simple Omnibus Format in Text. SOFT is an ASCII text format that was designed to be a machine-readable representation of data retrieved from, or submitted to, the Gene Expression Omnibus (<xref ref-type="other" rid="bid.1303">GEO</xref>). SOFT is also a line-based format, making it easy to parse, using commonly available text processing and formatting languages. (For examples of SOFT, see the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/geo/info/soft.cgi">guide</ext-link>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1407">splice sites</term>
          <def>
            <p>Refers to the location of the exon-intron junctions in a pre-mRNA (i.e., the primary transcript that must undergo additional processing to become a mature RNA for translation into a protein). Splice sites can be determined by comparing the sequence of genomic DNA with that of the <xref ref-type="other" rid="bid.1258">cDNA</xref> sequence. In mRNA, introns (non-protein coding regions) are removed by the splicing machinery; however, exons can also be removed. Depending on which exons (or parts of exons) are removed, different proteins can be made from the same initial RNA or gene. Different proteins created in this way are &ldquo;splice variants&rdquo; or &ldquo;alternatively spliced&rdquo;.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1408">SSAHA</term>
          <def>
            <p>Sequence Search and Alignment by Hashing Algorithm. SSAHA is a software tool for very fast matching and alignment of DNA sequences and is used for searching databases containing large amounts (gigabases) of genome sequence. It achieves its fast search speed by converting sequence information into a &ldquo;hash table&rdquo; data structure, which can then be searched very rapidly for matches (Ning et al., Genome Res 11:1725-1729; 2001).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1409">SSLP</term>
          <def>
            <p>Simple Sequence Length Polymorphisms. SSLPs are markers based on the variation in the number of short tandem repeats in DNA.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1410">STS</term>
          <def>
            <p>A short DNA segment that occurs only once in the human genome, the exact location and order of bases of which are known. Because each is unique, STSs are helpful for chromosome placement of mapping and sequencing data from many different laboratories. STSs serve as landmarks on the physical map of the human genome.<!--or STSs or STS marker--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1411">substitution matrix</term>
          <def>
            <p>A substitution matrix containing values proportional to the probability that amino acid i mutates into amino acid j for all pairs of amino acids. Such matrices are constructed by assembling a large and diverse sample of verified pairwise alignments of amino acids. If the sample is large enough to be statistically significant, the resulting matrices should reflect the true probabilities of mutations occurring through a period of evolution. (See also <xref ref-type="other" rid="bid.1252">BLOSUM 62</xref>.)</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1412">SWISS-PROT</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ebi.ac.uk/swissprot/">SWISS-PROT</ext-link> is a curated protein sequence database that provides a high level of annotation (such as the description of protein function, domain structures, post-translational modifications, variants, etc.), a minimal level of redundancy, and high level of integration with other databases.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1413">Sybase</term>
          <def>
            <p>A trademarked family of products that include databases, development tools, integration middleware<!--???-->, enterprise portals, and mobile and wireless servers.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1414">synteny</term>
          <def>
            <p>On the same strand. The phrase &ldquo;conserved synteny&rdquo; refers to conserved gene order on chromosomes of different, related species.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1415">Tax BLAST</term>
          <def>
            <p>BLAST Taxonomy Reports page. Tax BLAST groups BLAST hits by source organism, according to information in <xref ref-type="other" rid="bid.1353">NCBI</xref>'s Taxonomy database. Species are listed in order of sequence similarity with the query sequence, the strongest match listed first.<!--May appear in chapter/s as TaxBlast--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1416">taxID</term>
          <def>
            <p>Taxonomy Identifier. The taxID is a stable unique identifier for each taxon (for a species, a family, an order, or any other group in the taxonomy database). The taxID is seen in the <xref ref-type="other" rid="bid.1299">GenBank</xref> records as a &ldquo;source&rdquo; feature table entry; for example, /db_xref=&ldquo;taxon:&lt;9606&gt;&rdquo; is the taxID for <italic>Homo sapiens</italic>, and the line is therefore found in all recent human sequence records.<!--/db_xref="taxon: <9606>"--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1417">taxid</term>
          <def>
            <p>See <xref ref-type="other" rid="bid.1416">taxID</xref>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1735">termination codon or stop codon</term>
          <def>
            <p>One of three codons that do not specify any amino acid and hence causes translation of mRNA into protein to be terminated. These codons mark the end of a protein coding sequence.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1418">TIGR</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.tigr.org">The Institute for Genomic Research</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1419">tiling path</term>
          <def>
            <p>An ordered list or map that defines a set of overlapping clones that covers a chromosome or other extended segment of DNA.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1420">TPA</term>
          <def>
            <p>Third-Party Annotation<!--expand--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1736">TPF</term>
          <def>
            <p>Tiling Path Format. A table format used to specify the set of clones that will provide the best possible sequence coverage for a particular chromosome, the order of the clones along the chromosome, and the location of any gaps in the clone tiling path. Also used to refer to a file (Tiling Path File) in which the <xref ref-type="other" rid="bid.1731">minimal tiling path</xref> of clones covering a chromosome is specified in Tiling Path Format or to the minimal tiling path of clones so defined.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1737">translation start site</term>
          <def>
            <p>The position within an mRNA at which synthesis of a protein begins. The translation start site is usually an AUG codon, but occasionally, GUG or CUG codons are used to initiate protein synthesis.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1421">UID</term>
          <def>
            <p>Unique Identifier<!--Delete? Can also mean User ID--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1422">UMLS</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.nlm.nih.gov/research/umls/">Unified Medical Language System</ext-link>. A project of the National Library of Medicine for the development and distribution of multipurpose, electronic &ldquo;Knowledge Sources&rdquo;, and associated lexical programs. The purpose of the UMLS is to aid the development of systems that help health professionals and researchers retrieve and integrate electronic biomedical information from a variety of sources and to make it easy for users to link disparate information systems, including computer-based patient records, bibliographic databases, factual databases, and expert systems.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1423">unfinished sequence</term>
          <def>
            <p>See <xref ref-type="other" rid="bid.1276">draft sequence</xref>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1424">UniGene cluster</term>
          <def>
            <p><xref ref-type="other" rid="bid.1283">EST</xref>s and full-length mRNA sequences organized into clusters such that each represents a unique known or putative gene within the organism from which the sequences were obtained. UniGene clusters are annotated with mapping and expression information when possible (e.g., for human) and include cross-references to other resources. Sequence data can be downloaded by cluster through the UniGene web pages, or the complete dataset can be downloaded from the <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="ftp" xlink:href="ftp://ftp.ncbi.nih.gov/repository/UniGene/">repository/UniGene directory</ext-link> of the FTP site.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1425">UniSTS</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=unists">UniSTS</ext-link> presents a unified, non-redundant view of sequence-tagged sites (<xref ref-type="other" rid="bid.1410">STS</xref>s). UniSTS integrates marker and mapping data from a variety of public resources. If two or more markers have different names but the same primer pair, a single STS record is presented for the primer pair, and all the marker names are shown.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1426">UNIX</term>
          <def>
            <p>UNIX is an operating system that was developed by Dennis Ritchie and Kenneth Thompson at Bell Labs more than 30 years ago. It allows multitasking and multiuser capabilities and offers portability with other operating systems. It comes with hundreds of programs that are of two types: integral utilites, such as the command line interpreter; and tools such as email, which are not necessary for the operation of UNIX but provide additional capabilities to the user. It is functionally organized at three levels: the kernel, which schedules tasks and manages storage; the shell, which connects and interprets user's commands, calls programs from memory, and executes them; and tools and applications, which offer additional functionality to the operating system, such as word processing and business applications. UNIX<sup>&reg;</sup> was registered by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.lucent.com/">Bell Laboratories</ext-link> as a trademark for computer operating systems. Today, this mark is owned by <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.opengroup.org/">The Open Group</ext-link>.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1427">URL</term>
          <def>
            <p>Uniform Resource Locator. The address of a resource on the Internet. URL syntax is in the form of protocol://host/localinfo, where &ldquo;protocol&rdquo; specifies the means of fetching the object (such as HTTP, used by <xref ref-type="other" rid="bid.1434">WWW</xref> browsers and servers to exchange information, or <xref ref-type="other" rid="bid.1295">FTP</xref>), &ldquo;host&rdquo; specifies the remote location where the object resides, and &ldquo;localinfo&rdquo; is a string (often a file name) passed to the protocol handler at the remote location. Also called Uniform Resource Identifier (URI).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1428">UTF-8</term>
          <def>
            <p>UCS (Universal Character Set) Transformation Format. An AscII-preserving encoding method for Unicode (a standard to provide a unique number for every character irrespective of the platform, program, or language).</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1429">UTR</term>
          <def>
            <p>Untranslated Region. The 3&prime; UTR is that portion of an <xref ref-type="other" rid="bid.1351">mRNA</xref> from the position of the last codon that is used in translation to the 3&prime; end. The 5&prime; UTR is that portion of an mRNA from the 5&prime; end to the position of the first codon used in translation.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1430">VAST</term>
          <def>
            <p><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Structure/VAST/vast.shtml">Vector Alignment Search Tool</ext-link>. A computer algorithm used to identify similar protein 3D structures.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1431">weight</term>
          <def>
            <p>An assignment of importance to a term in a search query. If a term in a search query is found to match a word in a document, that word is given a &ldquo;weight&rdquo;. The exact weight of the word will depend on the emphasis given to the word by the author or its position in the document. For example, a word that occurs in a chapter title will have a higher weight than the same word if it occurs in the body of the chapter. Similarly, words that occur in data collections are also assigned weights, depending on how frequently the terms occur in the collection. <!--Protein weight matrix: the scoring table which describes the similarity of each amino acid to each other. DNA weight matrix: the scores assigned to matches and mismatches.--></p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1432">WGS sequence</term>
          <def>
            <p>Whole Genome Shotgun sequence. In this semi-automated sequencing technique, high-molecular-weight DNA is sheared into random fragments, size selected (usually 2, 10, 50, and 150 kb), and cloned into an appropriate vector. The clones are then sequenced from both ends. The two ends of the same clone are referred to as mate pairs. The distance between two mate pairs can be inferred if the library size is known and has a narrow window of deviation. The sequences are aligned using sequence assembly software. Proponents of this approach argue that it is possible to sequence the whole genome at once using large arrays of sequencers, which makes the whole process much more efficient than the traditional approaches.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1433">WHO</term>
          <def>
            <p>
              <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.who.int/en/">World Health Organization</ext-link>
            </p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1434">WWW</term>
          <def>
            <p>World Wide Web. A <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://www.w3.org/">consortium</ext-link> (W3C) that develops technologies such specifications, guidelines, software, and tools for the internet.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1435">XML</term>
          <def>
            <p>Extensible Markup Language. XML describes a class of data objects called XML documents and partially describes the behavior of computer programs that process them. XML is a subset of SGML, and XML documents are conforming SGML documents. XML documents are made up of storage units called entities, which contain either parsed or unparsed data. Parsed data is made up of characters (a unit of text), some of which form character data, and some of which form markup. Markup includes tags that provide information about the data, i.e., a description of the structure and content of the document. Character data comprises all the text that is not markup. XML provides a mechanism to impose constraints on the storage layout and logical structure.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1436">XSL</term>
          <def>
            <p>Extensible Stylesheet Language. XSL is used for the transformation of XML-based data into HTML or other presentation formats, for display in a web browser. This is a two-part process. First, the structure of the input XML tree must be transformed into a new tree (e.g., HTML), allowing reordering of the elements, addition of text, and calculations&mdash;all without modification to the source document. This process is described by <xref ref-type="other" rid="bid.1437">XSLT</xref>. Second, XSL-FO (XSL Formatting Objects, an XML vocabulary for formatting) is used for formatting the output, defining areas of the display page and their properties. In this way, the source XML document can be maintained from the perspective of &ldquo;pure content&rdquo; and can be separated from the presentation. An XML document can be delivered in different formats to different target audiences by simply switching style sheets.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1437">XSLT</term>
          <def>
            <p>Extensible Stylesheet Language: Transformations. XSLT is a language for transforming the structure of an XML document. XSLT is designed for use as part of <xref ref-type="other" rid="bid.1436">XSL</xref>, the stylesheet language for XML. A transformation expressed in XSLT describes a sequence of template rules for transforming a source tree into a result tree; elements from the source tree can be filtered and reordered, and a different structure can be added. A template rule has two parts: a pattern that is matched against nodes in the source tree; and a template that can be instantiated to form part of the result tree. This makes XSLT a declarative language because it is possible to specify what output should be produced when specific patterns occur in the input, which distinguishes it from procedural programming languages, where it is necessary to specify what tasks have to be performed in what order. XSLT makes use of the expression language defined by XPath (a language for addressing the parts of an XML document) for selecting elements for processing, for conditional processing, and for generating text.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1438">YAC</term>
          <def>
            <p>Yeast Artificial Chromosome. Extremely large segments of DNA from another species spliced into the DNA of yeast. YACs are used to clone up to one million bases of foreign DNA into a host cell, where the DNA is propagated along with the other chromosomes of the yeast cell.</p>
          </def>
        </def-item>
        <def-item>
          <term id="bid.1439">ZFIN</term>
          <def>
            <p>Zebrafish Information Network. <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="http://zdb.wehi.edu.au:8282//">ZFIN</ext-link> is a database for the zebrafish model organism that holds information on wild-type stocks, mutants, genes, gene expression data, and map markers.</p>
          </def>
        </def-item>
      </def-list>
    </glossary>
</back><?mtl /hellip?>
</book>