<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" article-type="research-article" xml:lang="en">
  <?properties open_access?>
  <front>
    <journal-meta>
      <journal-id journal-id-type="nlm-ta">BMC Res Notes</journal-id>
      <journal-id journal-id-type="iso-abbrev">BMC Res Notes</journal-id>
      <journal-title-group>
        <journal-title>BMC Research Notes</journal-title>
      </journal-title-group>
      <issn pub-type="epub">1756-0500</issn>
      <publisher>
        <publisher-name>BioMed Central</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="pmid">23116482</article-id>
      <article-id pub-id-type="pmc">3532326</article-id>
      <article-id pub-id-type="publisher-id">1756-0500-5-615</article-id>
      <article-id pub-id-type="doi">10.1186/1756-0500-5-615</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Technical Note</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>CooVar: Co-occurring variant analyzer</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" id="A1">
          <name>
            <surname>Vergara</surname>
            <given-names>Ismael A</given-names>
          </name>
          <xref ref-type="aff" rid="I1">1</xref>
          <xref ref-type="aff" rid="I2">2</xref>
          <email>ivergara@gmail.com</email>
        </contrib>
        <contrib contrib-type="author" id="A2">
          <name>
            <surname>Frech</surname>
            <given-names>Christian</given-names>
          </name>
          <xref ref-type="aff" rid="I1">1</xref>
          <email>frech.christian@gmail.com</email>
        </contrib>
        <contrib contrib-type="author" corresp="yes" id="A3">
          <name>
            <surname>Chen</surname>
            <given-names>Nansheng</given-names>
          </name>
          <xref ref-type="aff" rid="I1">1</xref>
          <email>chenn@sfu.ca</email>
        </contrib>
      </contrib-group>
      <aff id="I1"><label>1</label>Department of Molecular Biology and Biochemistry, Simon Fraser University, 8888 University Drive, Burnaby, B.C., V5A 1S6, Canada</aff>
      <aff id="I2"><label>2</label>GenomeDx Biosciences Inc, 1595 West 3rd Avenue, Vancouver, BC, V6J 1J8, Canada</aff>
      <pub-date pub-type="collection">
        <year>2012</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>1</day>
        <month>11</month>
        <year>2012</year>
      </pub-date>
      <volume>5</volume>
      <fpage>615</fpage>
      <lpage>615</lpage>
      <history>
        <date date-type="received">
          <day>20</day>
          <month>7</month>
          <year>2012</year>
        </date>
        <date date-type="accepted">
          <day>26</day>
          <month>10</month>
          <year>2012</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>Copyright &#xA9;2012 Vergara et al.; licensee BioMed Central Ltd.</copyright-statement>
        <copyright-year>2012</copyright-year>
        <copyright-holder>Vergara et al.; licensee BioMed Central Ltd.</copyright-holder>
        <license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/2.0">
          <license-p>This is an Open Access article distributed under the terms of the Creative Commons Attribution License (<ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/2.0">http://creativecommons.org/licenses/by/2.0</ext-link>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="http://www.biomedcentral.com/1756-0500/5/615"/>
      <abstract>
        <sec>
          <title>Background</title>
          <p>Evaluating the impact of genomic variations (GV) on protein-coding transcripts is an important step in identifying variants of functional significance. Currently available programs for variant annotation depend on external databases or annotate multiple variants affecting the same transcript independently, which limits program use to organisms available in these databases or results in potentially incorrect or incomplete annotations.</p>
        </sec>
        <sec>
          <title>Findings</title>
          <p>We have developed CooVar (Co-occurring Variant Analyzer), a database-independent program for assessing the impact of GVs on protein-coding transcripts. CooVar takes GVs, reference genome sequence, and protein-coding exons as input and provides annotated GVs and transcripts as output. Other than similar programs, CooVar considers the combined impact of all GVs affecting the same transcript, generating biologically more accurate annotations. CooVar is operated from the command-line and supports standard file formats VCF, GFF/GTF, and GVF, which makes it easy to integrate into existing computational pipelines. We have extensively tested CooVar on worm and human data sets and demonstrate that it generates correct annotations in only a short amount of time.</p>
        </sec>
        <sec>
          <title>Conclusions</title>
          <p>CooVar is an easy-to-use and lightweight variant annotation tool that considers the combined impact of GVs on protein-coding transcripts. CooVar is freely available at <ext-link ext-link-type="uri" xlink:href="http://genome.sfu.ca/projects/coovar/">http://genome.sfu.ca/projects/coovar/</ext-link>.</p>
        </sec>
      </abstract>
      <kwd-group>
        <kwd>Variant effect prediction</kwd>
        <kwd>Variant annotation</kwd>
        <kwd>Genomic variation</kwd>
        <kwd>Sequence analysis</kwd>
        <kwd>Protein-coding transcript</kwd>
        <kwd>Indel</kwd>
        <kwd>SNV</kwd>
        <kwd>Insertion</kwd>
        <kwd>Deletion</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Findings</title>
      <sec>
        <title>Introduction</title>
        <p>One central goal of many genomics projects is to detect different types of genomic variations (GVs) and to understand how these GVs explain differences at the phenotypic level, for example, between healthy and diseased individuals [<xref ref-type="bibr" rid="B1">1</xref>,<xref ref-type="bibr" rid="B2">2</xref>]. Accurate and comprehensive detection of GVs, including single-nucleotide variations (SNVs), insertions and deletions, has been greatly facilitated by the development of next generation sequencing technologies [<xref ref-type="bibr" rid="B3">3</xref>] and variation detection methods [<xref ref-type="bibr" rid="B4">4</xref>]. After GVs are defined, evaluation of their functional impact on protein-coding transcripts becomes the primary focus. Many programs have been developed for this task, of which Ensembl&#x2019;s Variant Effect Predictor (VEP) [<xref ref-type="bibr" rid="B5">5</xref>], GATK&#x2019;s VariantAnnotator [<xref ref-type="bibr" rid="B6">6</xref>], Sequence Variant Analyzer (SVA) [<xref ref-type="bibr" rid="B7">7</xref>] and ANNOVAR [<xref ref-type="bibr" rid="B8">8</xref>] are among the more popular ones.</p>
        <p>Current variant annotation programs have important limitations. First, they assess the effects of multiple co-occurring GVs on the same transcript independently, which can be problematic when nearby GVs alter each other&#x2019;s effect [<xref ref-type="bibr" rid="B9">9</xref>]. For example, a small deletion can restore the open reading frame (ORF) disrupted by a small insertion co-occurring nearby on the same transcript. A second limitation is that most programs are tightly coupled to external databases, making their use inconvenient or even impractical for users who work on organisms whose genome sequence or annotation is not available in these databases.</p>
      </sec>
      <sec>
        <title>Implementation</title>
        <p>We have developed an easy-to-use Perl program named CooVar (Co-occurring Variant Analyzer) to address these limitations. CooVar takes as input (i) a list of GVs in the popular Variant Call Format (VCF) [<xref ref-type="bibr" rid="B10">10</xref>] or in a simpler tab-delimited file format, (ii) the reference genomic DNA sequence in FASTA format, and (iii) protein-coding exon coordinates in GFF or GTF format.</p>
        <p>The core output of CooVar are two files: a GVF file [<xref ref-type="bibr" rid="B11">11</xref>] reporting the functional impact of each input GV on transcripts, and a GFF file including (i) the transcript models, (ii) all GVs impacting each transcript, and (iii) a prediction of how GVs impact the function of each transcript. The functional impact of GVs on protein-coding transcripts is annotated as: ORF_INTACT, if the transcript is not impacted by any GVs; ORF_PRESERVED, if the transcript is impacted by GVs but these GVs do not introduce internal stop or splice site variants; ORF_DISRUPTED, if an internal stop or splice site variant is present; and FULLY_DELETED, if the transcript is deleted. In the case of a transcript that has its ORF disrupted by an internal stop codon, CooVar provides the percentage location of the first internal stop codon in the variant peptide compared to the reference.</p>
        <p>CooVar classifies GVs according to the GVF v1.05 specification for structural variants described in the Sequence Ontology (SO) Project [<xref ref-type="bibr" rid="B12">12</xref>]. For SNVs, these categories include <italic>silent_mutation</italic>, <italic>synonymous_codon</italic>, <italic>conservative_missense_codon</italic>, <italic>non_conservative_missense_codon</italic>, <italic>stop_gained</italic>, <italic>stop_lost</italic>, <italic>splice_acceptor_variant</italic>, and <italic>splice_donor_variant</italic>. Insertions and deletions are classified into the categories <italic>silent_mutation</italic>, <italic>frameshift_variant</italic>, <italic>inframe_variant</italic>, <italic>splice_acceptor_variant</italic>, and <italic>splice_donor_variant</italic>. The functional impact of missense SNVs causing amino acid changes is further evaluated with the Grantham score [<xref ref-type="bibr" rid="B13">13</xref>] and annotated as <italic>CONSERVATIVE</italic> or <italic>MODERATELY_CONSERVATIVE</italic> (both classified as <italic>conservative_missense_codon</italic>) and <italic>MODERATELY_RADICAL</italic> or <italic>RADICAL</italic> (both classified as <italic>non_conservative_missense_codon</italic>) [<xref ref-type="bibr" rid="B14">14</xref>]. For SNVs impacting protein-coding exons, CooVar also reports both the amino acid change and the codon change between the reference genome and the variant. This allows the user to observe immediately if a change in a codon is caused by one, two or three co-occurring substitutions at the same codon. Furthermore, CooVar lists separately all those SNVs that fall into multiple categories by impacting two or more protein-coding transcripts differently (e.g. synonymous <italic>vs</italic>. missense).</p>
        <p>In addition to the annotation of individual transcripts and GVs, CooVar outputs various summary statistics. For example, CooVar generates statistics on the codon bias for synonymous versus non-synonymous SNVs. In two other files CooVar outputs the length distribution of indels across the whole genome versus the length distribution of indels impacting only protein-coding transcripts as a way to detect biases towards non-frameshift indels in exonic regions. The file <italic>variant.stat</italic> provides information on the distribution of internal stop codons and on the total number of transcripts affected by SNVs, insertions and deletions, or by any combination of those. If the --<italic>circos</italic> flag is used, CooVar computes the genomic distribution of SNVs, insertions, deletions and coding exons in a format compatible with the Circos tool for visualization [<xref ref-type="bibr" rid="B15">15</xref>].</p>
        <p>One advantage of CooVar over other programs is that it provides full-length variant transcript and protein sequences in FASTA format as output, which can be useful for downstream analyses (for example for sequence alignments). The same information is provided at the exon level in two additional files. Since a direct comparison between the reference and variant transcript is also desirable, CooVar provides an exon-based alignment of reference and variant sequences for each transcript, with variant nucleotides marked in uppercase. This makes it easy to spot all SNVs, insertions and deletions that impact a given protein-coding transcript in a region of interest.</p>
        <p>Another commonly requested feature in variant annotation is to identify GVs that overlap with protein domains. This is because GVs affecting conserved domains are more likely to be of functional importance. With CooVar this analysis can be performed in two steps. First, the script <italic>protein2genome.pl</italic> can be used to map protein (domain) coordinates to the genome, which generates a GFF file with genomic coordinates. The script <italic>annotate-regions.pl</italic> can then be used to compute the overlap between this GFF file and the CooVar GVF output file. Overlap computation is performed efficiently using interval trees and generally finishes within a few minutes, even for very large data sets. The result of this two-step process is a new GVF file in which GVs are annotated with the protein domains they overlap with. It is worth mentioning that <italic>annotate-regions.pl</italic> script is generic and can also be used to annotate GVs that overlap with non-protein-coding regions (for example transcription factor binding sites) as long as coordinates for these regions are provided in the required input GFF format.</p>
        <p>More detailed information about program parameters and input file formats can be found in the program README file or in the Perl scripts themselves.</p>
      </sec>
      <sec>
        <title>Results and discussion</title>
        <p>We have tested CooVar on two datasets, both of which are available from the project website. The first dataset corresponds to 120,638 GVs (116,999 SNVs, 1,553 insertions ranging from 1 to 34 bp in length and 2,086 deletions ranging from 1 to 24 bp in length) detected in the Hawaiian isolate CB4856 of the model organism <italic>Caenorhabditis elegans</italic>. CB4856 GVs, the N2 reference genome (isolated in Bristol, England) and 24,256 annotated N2 protein-coding transcript models were obtained from WormBase release WS210 [<xref ref-type="bibr" rid="B16">16</xref>]. CooVar had a processing time of 10 minutes for this data set and classified 15,293 transcripts as ORF_INTACT, 8,446 transcripts as ORF_PRESERVED, and 517 transcripts as ORF_DISRUPTED. Figure&#x2009;<xref ref-type="fig" rid="F1">1</xref> shows a Circos image with the distribution of SNVs, deletions, and insertions along the six <italic>C. elegans</italic> chromosomes. This image was generated by using the CooVar output files (option --<italic>circos</italic>) as input for Circos (version 0.62-1) [<xref ref-type="bibr" rid="B15">15</xref>].</p>
        <fig id="F1" position="float">
          <label>Figure 1</label>
          <caption>
            <p><bold>Distribution of SNVs, insertions, and deletions along the </bold><bold><italic>C. elegans </italic></bold><bold>genome for Hawaiian isolate CB4856.</bold> Segmented rings on the outside represent the six <italic>C. elegans</italic> chromosomes. Going from outside to inside, the line plot shows SNV density (inward pointing peaks = higher density), histograms represent the density of deletions (also drawn inwards), and the heatmap depicts the density of insertions (dark red = higher density) detected in the Hawaiian isolate CB4856. Note the generally higher density of SNVs towards the telomeres and the presence of chromosome-internal peaks on chromosome IV and V. Data points for this image were automatically generated by CooVar using the --<italic>circos</italic> option. Circos was then used to generate the image. Circos configuration files necessary to create this type of image are provided with the <italic>C. elegans</italic> test data set at <ext-link ext-link-type="uri" xlink:href="http://genome.sfu.ca/projects/coovar/">http://genome.sfu.ca/projects/coovar/</ext-link>.</p>
          </caption>
          <graphic xlink:href="1756-0500-5-615-1"/>
        </fig>
        <p>The second dataset contains 4,044,200 human GVs detected in an anonymous individual (HG00732-200-37-ASM) sequenced by Complete Genomics. This data set was recently made publicly available for the research community as part of a larger 69 genome data set [<xref ref-type="bibr" rid="B17">17</xref>,<xref ref-type="bibr" rid="B18">18</xref>]. HG00732-200-37-ASM variants were first extracted from the 69-sample VCF file using <italic>vcf-subset</italic>[<xref ref-type="bibr" rid="B10">10</xref>]. We then discarded all but the first alternative allele and used the filtered VCF file as input to CooVar. The genomic reference sequence and the protein-coding gene set were both obtained from the Ensembl web site (release GRCh37.68, hg19). For comparison, we annotated the exact same VCF file with Ensembl&#x2019;s Variant Effect Predictor (VEP) [<xref ref-type="bibr" rid="B5">5</xref>]. VEP was run locally using the Perl script <italic>variant_effect_predictor</italic>.<italic>pl</italic> and configured to retrieve Ensembl data (release 68) over the internet (&#x2212;&#x2212;<italic>host useastdb.ensembl.org</italic>). The VEP output for the HG00732-200-37-ASM data set can be downloaded from the CooVar project homepage.</p>
        <p>Results of this comparison are summarized in Table&#x2009;<xref ref-type="table" rid="T1">1</xref>. CooVar took 36 minutes to process the complete human data set and reported 4,158,840 annotated variants. VEP outputted 4,043,939 variants and finished within 37 hours and 24 minutes. The overall increased number of variants reported by CooVar is because CooVar decomposes compound input VCF variants into sometimes multiple SNV and indel variants, which are then annotated and reported separately. For example, CooVar decomposes a VCF input variant with the reference allele &#x201C;ATG&#x201D; and the alternative allele &#x201C;AC&#x201D; into one SNV (T-&gt;C) and one deletion variant (with G being the deleted base in this case). In contrast, VEP will annotate such compound variants as a single variant, resulting in sometimes ambiguous or nonspecific classification results (see examples below).</p>
        <table-wrap position="float" id="T1">
          <label>Table 1</label>
          <caption>
            <p>Comparison of GVs annotated with CooVar and VEP for human individual HG00732-200-37-ASM</p>
          </caption>
          <table frame="hsides" rules="groups" border="1">
            <colgroup>
              <col align="left"/>
              <col align="right"/>
              <col align="right"/>
            </colgroup>
            <thead valign="top">
              <tr>
                <th align="left">&#xA0;</th>
                <th align="right">
                  <bold>CooVar</bold>
                </th>
                <th align="right">
                  <bold>VEP</bold>
                </th>
              </tr>
            </thead>
            <tbody valign="top">
              <tr>
                <td align="left" valign="bottom">Runtime<hr/></td>
                <td align="right" valign="bottom">36m<hr/></td>
                <td align="right" valign="bottom">37h 24m<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">Total reported GVs<hr/></td>
                <td align="right" valign="bottom">4,158,840<hr/></td>
                <td align="right" valign="bottom">4,043,939<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;Intronic/intergenic/UTR<hr/></td>
                <td align="right" valign="bottom">4,133,885<hr/></td>
                <td align="right" valign="bottom">4,019,490<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;Impacting protein-coding exon<hr/></td>
                <td align="right" valign="bottom">24,955<hr/></td>
                <td align="right" valign="bottom">24,449<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Synonymous/stop retained<hr/></td>
                <td align="right" valign="bottom">11,585<hr/></td>
                <td align="right" valign="bottom">11,434<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Missense<hr/></td>
                <td align="right" valign="bottom">12,011<hr/></td>
                <td align="right" valign="bottom">11,576<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;&#x2003;Conservative (%)<sup>$</sup><hr/></td>
                <td align="right" valign="bottom">9,526 (79.3)<hr/></td>
                <td align="right" valign="bottom">7,447 (64.3)<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;&#x2003;Non-conservative (%)<sup>+</sup><hr/></td>
                <td align="right" valign="bottom">2,485 (20.7)<hr/></td>
                <td align="right" valign="bottom">2,110 (18.3)<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;&#x2003;Unknown consequence (%)<hr/></td>
                <td align="right" valign="bottom">0 (0)<hr/></td>
                <td align="right" valign="bottom">2,019 (17.4)<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Splice donor/acceptor<hr/></td>
                <td align="right" valign="bottom">97<hr/></td>
                <td align="right" valign="bottom">184<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Stop lost<hr/></td>
                <td align="right" valign="bottom">47<hr/></td>
                <td align="right" valign="bottom">46<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Stop gained<hr/></td>
                <td align="right" valign="bottom">137<hr/></td>
                <td align="right" valign="bottom">134<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Frameshift<hr/></td>
                <td align="right" valign="bottom">470<hr/></td>
                <td align="right" valign="bottom">490<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Inframe<hr/></td>
                <td align="right" valign="bottom">199<hr/></td>
                <td align="right" valign="bottom">165<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Other<hr/></td>
                <td align="right" valign="bottom">0<hr/></td>
                <td align="right" valign="bottom">31<hr/></td>
              </tr>
              <tr>
                <td align="left" valign="bottom">&#x2003;&#x2003;Multiple*<hr/></td>
                <td align="right" valign="bottom">409<hr/></td>
                <td align="right" valign="bottom">389<hr/></td>
              </tr>
              <tr>
                <td align="left">ORF disrupted</td>
                <td align="right">782<sup>&#xA5;</sup></td>
                <td align="right">871<sup>&#xA7;</sup></td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <p>Each GV reported by CooVar (v0.05) and VEP (v2.6) was assigned to one (and only one) of the above categories. $ CooVar: Grantham score <italic>conservative</italic> or <italic>moderately conservative</italic>; VEP: SIFT <italic>benign</italic>; + CooVar: Grantham score <italic>moderately radical</italic> or <italic>radical</italic>; VEP: SIFT <italic>deleterious.</italic> * GVs assigned to more than one category due to differential impact on different transcripts. &#xA5; Presence of internal stop codon within the first 70% of ORF length, after applying all variants. &#xA7; Predicted frameshift or stop gain variant within the first 70% of ORF length. Abbreviations: GV &#x2026; genomic variation; ORF &#x2026; open reading frame; VEP&#x2026; Variant Effect Predictor; UTR&#x2026; untranslated region.</p>
          </table-wrap-foot>
        </table-wrap>
        <p>As expected, both programs classify the vast majority of variants as not impacting protein-coding exons (Table&#x2009;<xref ref-type="table" rid="T1">1</xref>, category <italic>intronic/intergenic/UTR</italic>). Only 0.6% of all variants (24,955 variants by CooVar and 24,449 by VEP) are predicted to impact protein-coding exons in some form. To allow for a detailed comparison of annotation results, we assigned variants impacting protein-coding exons into one (and only one) of the following categories: variants not altering protein translation (<italic>synonymous/stop retained</italic>); variants altering protein translation (<italic>missense</italic>); variants impacting AG/GT splice site di-nucleotides (<italic>splice donor/acceptor</italic>); variants leading to stop codon loss (<italic>stop lost</italic>) or gain (<italic>stop gained</italic>); and insertions or deletions that shift (<italic>frameshift</italic>) or preserve the open reading frame (<italic>inframe</italic>). Thirty-one VEP variants could not be assigned to one of these categories and were classified as <italic>other</italic>. This includes variants that VEP nonspecifically annotated as <italic>coding_sequence_variant</italic>. A number of variants (409 for CooVar, 389 for VEP) could not be unambiguously assigned to a single category because they impact multiple transcripts differently and were classified as <italic>multiple</italic>.</p>
        <p>Overall, we find that numbers of GVs in each category agree well between CooVar and VEP (Table&#x2009;<xref ref-type="table" rid="T1">1</xref>). Both programs predict ~11,500 synonymous variants and about the same number of missense variants. CooVar&#x2019;s Grantham score classifies ~20% of missense variants as <italic>moderately radical</italic> or <italic>radical</italic>, which agrees well with the VEP SIFT classification scheme that predicts 18% of missense variants to be <italic>deleterious</italic>. Both programs predict about 50 stop lost mutations and 135 stop gain mutations. Interestingly, VEP predicts 20 more frameshift variants than CooVar (490 vs. 470 variants) and 34 less inframe variants (165 vs. 199 variants). Also, the number of predicted splice site variants is markedly different between the two programs, with almost twice as many splice site variants predicted by VEP (184 variants) than CooVar (97 variants).</p>
        <p>To understand the nature of these differences, we performed a detailed manual analysis of GVs that were differently annotated between CooVar and VEP. In general, we find that the main source of discrepancy between CooVar and VEP is due to the fact that CooVar but not VEP recognized the presence of SNVs within more complex or compound VCF input variants. For example, reference and alternative allele in the VCF input variant 11:11,292,688:GGGTCAGGACGC<underline>G</underline>-&gt;GGGTCAGGACGC<underline>C</underline> differ by only a single SNV (G-&gt;C, underlined). CooVar correctly reports this variant as synonymous SNV while VEP annotates it less specifically as <italic>coding_sequence_variant,</italic> without information on codon impact. The different numbers in stop lost and stop gained variants are attributable to the same effect. For example, CooVar interprets multi-SNV variant 19:43,922,549:AGA-&gt;TGC as stop lost variant while VEP annotates it as <italic>coding_sequence_variant</italic> and <italic>3_prime_UTR_variant</italic>. Manual inspection showed that the first of the three SNVs encoded by this input variant (A-&gt;T) indeed changes the stop codon of transcript ENST00000253435 from TAA to TAT, suggesting that the CooVar prediction is correct.</p>
        <p>The much larger number of splice site variants predicted by VEP is also explained by the higher resolution with which CooVar decomposes complex VCF input variants. For example, variant 11:117,303,853:CCCAGT-&gt;CCCAGC is annotated as splice donor variant by VEP but not CooVar, which reports it as synonymous SNV. Manual inspection showed that the coordinates of this variant (117,303,853&#x2013;117,303,858) indeed overlap with a donor splice site of transcript ENST00000527706, but the actual SNV encoded by this variant (T-&gt;C) is in fact synonymous. Thus, in this case, a simple coordinate overlap analysis as seems to be performed by VEP produces an incorrect result. Other examples of this type include variant 7:101,194,424:CGTAA-&gt;TGTAA (CooVar: synonymous), 5:159,835,654:TACCA-&gt;TACCG (CooVar: missense), or 19:16,612,363:GTG-&gt;GTA (CooVar: silent).</p>
        <p>Complex input variants also account for discrepancies observed between indel annotations. We randomly picked and examined 10 of the 34 inframe variants predicted by CooVar but not VEP. For 9 out of these, we find that they are genuine inframe indels that VEP classified as missense (for example 10:126,715,151:TGCAGAGGAGC-&gt;TGCGGAGGAGCCGCAGGCTGGGGCTGCAGGGC or 12:53,045,625:CT-&gt;CCGCTGCCGCCTCCAAAGCC; note the length difference is a multiple of 3 in both cases). Why VEP classifies these variants as missense variants was not obvious to us. The remaining variant of these ten variants was actually classified as inframe variant by VEP but assigned to category <italic>multiple</italic> by our classification scheme because VEP predicts it as both inframe and stop gained variant.</p>
        <p>Another main source of discrepancy in indel classification arose from so called &#x201C;boundary indels&#x201D;. We refer to boundary indels as indels that fall right next to the start or end of coding exons, thus leaving some uncertainty about the exact impact of these variants on the protein-coding transcript. Variant 7:142,494,013 is an example of an insertion where the exact placement of the inserted sequence is ambiguous, resulting in a predicted frameshift insertion by CooVar but in a predicted <italic>coding_sequence_variant</italic> and <italic>5_prime_UTR_variant</italic> by VEP. Most of the frameshift variants predicted by VEP but not CooVar represent boundary indels. Representative examples include 11:111,853,106:G-&gt;GC (1-bp insertion right before coding exon), 16:76,311,602:G-&gt;GT (1-bp insertion right after coding exon), 16:31,770,696:GA-&gt;GAA (1-bp insertion into start codon), and 17:39,254,335:AT-&gt;ATT (1-bp insertion before start codon). We manually inspected all 20 frameshift variants predicted by VEP but not CooVar and confirm that CooVar predictions appear to be correct, <italic>i.e.</italic> these variants are likely not causing frameshift mutations in affected transcripts.</p>
        <p>We were also interested in the number of ORFs that were predicted to be disrupted by both CooVar and VEP. For this particular comparison, we defined a CooVar ORF as being disrupted if an internal stop codon occurred within the first 70% of the ORF&#x2019;s length after applying all GVs to a transcript. CooVar provides the position of the first internal stop codon as part of its output. VEP does not provide ORF status information in its output, so we defined a VEP ORF as being disrupted if VEP predicted at least one frameshift or stop gain variant within the first 70% of the ORF&#x2019;s length. Using these criteria, we find that CooVar predicts 782 ORFs to be disrupted while VEP predicts 871 ORFs as disrupted (Table&#x2009;<xref ref-type="table" rid="T1">1</xref>). We inspected about half (48) of the transcripts that had assigned a different ORF status by the two programs and found that most of them (20 ORFs, <italic>e.g.</italic> transcript ENST00000376343) carry a frame-shifting indel that does not introduce an internal stop codon albeit it changes the translated protein sequence downstream. Thus, although for these transcripts a significant portion of the ORF (&gt;30%) is changed in terms of its protein sequence, the length of the ORF remains intact. Seventeen of the 48 inspected transcripts (<italic>e.g.</italic> ENST00000222270) had a frameshift predicted by VEP but not CooVar due to boundary indels as discussed above. Five of the 48 transcripts had already internal stop codons in the reference sequence and hence were not annotated as disrupted by CooVar.</p>
        <p>Most importantly, the remaining six ORFs predicted to be disrupted by VEP but not CooVar carry neighboring indels that cancel each others effect, restoring the open reading frame. Figure&#x2009;<xref ref-type="fig" rid="F2">2A</xref> shows one such example affecting Ensembl transcript ENST00000253255. This transcript carries a 1bp insertion at position 22:46,658,224 and a nearby 1bp deletion at position 22:46,658,220. When evaluated independently, the impact of the insertion and the deletion on this transcript is a frameshift mutation, disrupting the ORF. But when evaluating the joint effect that these GVs have on the transcript it results in a preserved ORF, as reported by CooVar. A similar issue arises from co-occurring SNVs. In the example shown in Figure&#x2009;<xref ref-type="fig" rid="F2">2B</xref>, VEP classifies SNV 10:27,702,726:G-&gt;A as synonymous because it changes codon CTG on the reverse strand (coding for leucine) to TTG (also coding for leucine). However, CooVar considers the combined effect of this SNV and a neighboring SNV (10:27,702,725:A-&gt;G) that affects the same codon. When evaluated together, 10:27702726:G-&gt;A is recognized by CooVar as missense SNV that changes the codon from CTG to TCG, which codes for amino acid serine. These two examples illustrate the importance of assessing the impact of co-occurring GVs together to correctly judge their functional impact [<xref ref-type="bibr" rid="B9">9</xref>].</p>
        <fig id="F2" position="float">
          <label>Figure 2</label>
          <caption>
            <p><bold>GVs affecting the same protein-coding transcript must be assessed together to correctly predict their functional impact.</bold> Panel <bold>A</bold> shows an example where two neighboring frameshift indels (1-bp insertion and 1-bp deletion, indicated by arrows) cancel each others effect, restoring the original ORF. Panel <bold>B</bold> shows an example where an otherwise synonymous SNV (G-&gt;A, indicated by arrow) causes a missense mutation due to the effect of a neighboring SNV (A-&gt;G). In both panels, the first three rows show the reference nucleotide sequence on the forward strand, the reference nucleotide sequence from the reverse strand, and the reference protein sequence translation from the annotated ORF. Note that both ORFs are encoded on the reverse strand, so sequences must be read from right to left. The track below shows the variant sequence detected in human individual HG00732-200-37-ASM, with critical GVs highlighted by arrows. The blue horizontal bar represents the Ensembl protein-coding transcript spanning this genomic region.</p>
          </caption>
          <graphic xlink:href="1756-0500-5-615-2"/>
        </fig>
        <p>We conclude that CooVar is a fast and light-weight alternative to currently existing variant annotation tools that is particularly useful for non-model organisms. CooVar produces very similar results as other popular tools, but, under certain circumstances, generates biologically more accurate annotations by considering the combined effect of co-occurring GVs on protein-coding transcripts.</p>
      </sec>
    </sec>
    <sec>
      <title>Availability and requirements</title>
      <p><bold>Project name</bold>: CooVar: Co-occurring Variant Analyzer</p>
      <p><bold>Project home page</bold>: <ext-link ext-link-type="uri" xlink:href="http://genome.sfu.ca/projects/coovar">http://genome.sfu.ca/projects/coovar</ext-link></p>
      <p><bold>Operating system(s)</bold>: Windows, Linux, Mac OS-X</p>
      <p><bold>Programming language</bold>: Perl 5.8.8</p>
      <p><bold>Other requirements</bold>: The following Perl modules are required by CooVar and need to be installed: Cwd, Getopt::Long, POSIX, File::Basename, List::Util, Bio::DB::Fasta, Bio::Seq, Bio::SeqUtils, Bio::SeqIO, Set::IntervalTree, Set::IntSpan</p>
      <p><bold>License</bold>: GNU GPL</p>
      <p><bold>Any restrictions to use by non-academics</bold>: none</p>
      <p>The latest version of the program can be obtained from the project webpage. CooVar version 0.05 is included as online supplementary material (Additional file <xref ref-type="supplementary-material" rid="S1">1</xref>).</p>
    </sec>
    <sec>
      <title>Abbreviations</title>
      <p>GV: Genomic Variation; GFF: Generic Feature Format; GTF: Gene Transfer Format; SNV: Single Nucleotide Variant; Indel: Insertion or deletion; GVF: Genomic Variant Format; ORF: Open Reading Frame; VCF: Variant Call Format.</p>
    </sec>
    <sec>
      <title>Competing interests</title>
      <p>The authors declare that they have no competing interest.</p>
    </sec>
    <sec>
      <title>Authors&#x2019; contribution</title>
      <p>IAV and NC conceived the study. IAV developed the method and original program. CF co-developed and tested the program. IAV, CF, and NC wrote the manuscript. All authors read and approved the final manuscript.</p>
    </sec>
    <sec>
      <title>Authors&#x2019; information</title>
      <p>Ismael A Vergara and Christian Frech were joint first author.</p>
    </sec>
    <sec sec-type="supplementary-material">
      <title>Supplementary Material</title>
      <supplementary-material content-type="local-data" id="S1">
        <caption>
          <title>Additional file 1</title>
          <p>CooVar program tarball (version 0.05), including README and test scripts.</p>
        </caption>
        <media xlink:href="1756-0500-5-615-S1.gz" mimetype="application" mime-subtype="x-zip-compressed">
          <caption>
            <p>Click here for file</p>
          </caption>
        </media>
      </supplementary-material>
    </sec>
  </body>
  <back>
    <sec>
      <title>Acknowledgements</title>
      <p>The authors would like to acknowledge Tammy Wong, Jeffrey Chu, Shanshan Zou, Jiarui Li and all other lab members for insightful discussion and for testing the program. We thank Duncan Napier for IT support. This study is supported by a Discovery Grant from the Natural Sciences and Engineering Research Council of Canada (NSERC) to NC. NC is a Michael Smith Foundation for Health Research (MSFHR) Scholar and a Canadian Institutes of Health Research (CIHR) New Investigator. IAV was supported by an MBB Alumni Graduate Scholarship. CF received SFU graduate fellowships and was supported by BC Pacific Century, Weyerhaeuser, and Sulzer Pumps graduate scholarships.</p>
    </sec>
    <ref-list>
      <ref id="B1">
        <mixed-citation publication-type="journal">
          <name>
            <surname>MacArthur</surname>
            <given-names>DG</given-names>
          </name>
          <name>
            <surname>Tyler-Smith</surname>
            <given-names>C</given-names>
          </name>
          <article-title>Loss-of-function variants in the genomes of healthy humans</article-title>
          <source>Hum Mol Genet</source>
          <year>2010</year>
          <volume>19</volume>
          <issue>R2</issue>
          <fpage>R125</fpage>
          <lpage>R130</lpage>
          <pub-id pub-id-type="doi">10.1093/hmg/ddq365</pub-id>
          <pub-id pub-id-type="pmid">20805107</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B2">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Stankiewicz</surname>
            <given-names>P</given-names>
          </name>
          <name>
            <surname>Lupski</surname>
            <given-names>JR</given-names>
          </name>
          <article-title>Structural variation in the human genome and its role in disease</article-title>
          <source>Annu Rev Med</source>
          <year>2010</year>
          <volume>61</volume>
          <fpage>437</fpage>
          <lpage>455</lpage>
          <pub-id pub-id-type="doi">10.1146/annurev-med-100708-204735</pub-id>
          <pub-id pub-id-type="pmid">20059347</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B3">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Shendure</surname>
            <given-names>J</given-names>
          </name>
          <name>
            <surname>Ji</surname>
            <given-names>H</given-names>
          </name>
          <article-title>Next-generation DNA sequencing</article-title>
          <source>Nat Biotechnol</source>
          <year>2008</year>
          <volume>26</volume>
          <issue>10</issue>
          <fpage>1135</fpage>
          <lpage>1145</lpage>
          <pub-id pub-id-type="doi">10.1038/nbt1486</pub-id>
          <pub-id pub-id-type="pmid">18846087</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B4">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Medvedev</surname>
            <given-names>P</given-names>
          </name>
          <name>
            <surname>Stanciu</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Brudno</surname>
            <given-names>M</given-names>
          </name>
          <article-title>Computational methods for discovering structural variation with next-generation sequencing</article-title>
          <source>Nat Methods</source>
          <year>2009</year>
          <volume>6</volume>
          <issue>11 Suppl</issue>
          <fpage>S13</fpage>
          <lpage>S20</lpage>
          <pub-id pub-id-type="pmid">19844226</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B5">
        <mixed-citation publication-type="journal">
          <name>
            <surname>McLaren</surname>
            <given-names>W</given-names>
          </name>
          <name>
            <surname>Pritchard</surname>
            <given-names>B</given-names>
          </name>
          <name>
            <surname>Rios</surname>
            <given-names>D</given-names>
          </name>
          <name>
            <surname>Chen</surname>
            <given-names>Y</given-names>
          </name>
          <name>
            <surname>Flicek</surname>
            <given-names>P</given-names>
          </name>
          <name>
            <surname>Cunningham</surname>
            <given-names>F</given-names>
          </name>
          <article-title>Deriving the consequences of genomic variants with the Ensembl API and SNP Effect Predictor</article-title>
          <source>Bioinformatics</source>
          <year>2010</year>
          <volume>26</volume>
          <issue>16</issue>
          <fpage>2069</fpage>
          <lpage>2070</lpage>
          <pub-id pub-id-type="doi">10.1093/bioinformatics/btq330</pub-id>
          <pub-id pub-id-type="pmid">20562413</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B6">
        <mixed-citation publication-type="journal">
          <name>
            <surname>McKenna</surname>
            <given-names>A</given-names>
          </name>
          <name>
            <surname>Hanna</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Banks</surname>
            <given-names>E</given-names>
          </name>
          <name>
            <surname>Sivachenko</surname>
            <given-names>A</given-names>
          </name>
          <name>
            <surname>Cibulskis</surname>
            <given-names>K</given-names>
          </name>
          <name>
            <surname>Kernytsky</surname>
            <given-names>A</given-names>
          </name>
          <name>
            <surname>Garimella</surname>
            <given-names>K</given-names>
          </name>
          <name>
            <surname>Altshuler</surname>
            <given-names>D</given-names>
          </name>
          <name>
            <surname>Gabriel</surname>
            <given-names>S</given-names>
          </name>
          <name>
            <surname>Daly</surname>
            <given-names>M</given-names>
          </name>
          <etal/>
          <article-title>The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data</article-title>
          <source>Genome Res</source>
          <year>2010</year>
          <volume>20</volume>
          <issue>9</issue>
          <fpage>1297</fpage>
          <lpage>1303</lpage>
          <pub-id pub-id-type="doi">10.1101/gr.107524.110</pub-id>
          <pub-id pub-id-type="pmid">20644199</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B7">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Ge</surname>
            <given-names>D</given-names>
          </name>
          <name>
            <surname>Ruzzo</surname>
            <given-names>EK</given-names>
          </name>
          <name>
            <surname>Shianna</surname>
            <given-names>KV</given-names>
          </name>
          <name>
            <surname>He</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Pelak</surname>
            <given-names>K</given-names>
          </name>
          <name>
            <surname>Heinzen</surname>
            <given-names>EL</given-names>
          </name>
          <name>
            <surname>Need</surname>
            <given-names>AC</given-names>
          </name>
          <name>
            <surname>Cirulli</surname>
            <given-names>ET</given-names>
          </name>
          <name>
            <surname>Maia</surname>
            <given-names>JM</given-names>
          </name>
          <name>
            <surname>Dickson</surname>
            <given-names>SP</given-names>
          </name>
          <etal/>
          <article-title>SVA: software for annotating and visualizing sequenced human genomes</article-title>
          <source>Bioinformatics</source>
          <year>2011</year>
          <volume>27</volume>
          <issue>14</issue>
          <fpage>1998</fpage>
          <lpage>2000</lpage>
          <pub-id pub-id-type="doi">10.1093/bioinformatics/btr317</pub-id>
          <pub-id pub-id-type="pmid">21624899</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B8">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Wang</surname>
            <given-names>K</given-names>
          </name>
          <name>
            <surname>Li</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Hakonarson</surname>
            <given-names>H</given-names>
          </name>
          <article-title>ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data</article-title>
          <source>Nucleic Acids Res</source>
          <year>2010</year>
          <volume>38</volume>
          <issue>16</issue>
          <fpage>e164</fpage>
          <pub-id pub-id-type="doi">10.1093/nar/gkq603</pub-id>
          <pub-id pub-id-type="pmid">20601685</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B9">
        <mixed-citation publication-type="journal">
          <name>
            <surname>MacArthur</surname>
            <given-names>DG</given-names>
          </name>
          <name>
            <surname>Balasubramanian</surname>
            <given-names>S</given-names>
          </name>
          <name>
            <surname>Frankish</surname>
            <given-names>A</given-names>
          </name>
          <name>
            <surname>Huang</surname>
            <given-names>N</given-names>
          </name>
          <name>
            <surname>Morris</surname>
            <given-names>J</given-names>
          </name>
          <name>
            <surname>Walter</surname>
            <given-names>K</given-names>
          </name>
          <name>
            <surname>Jostins</surname>
            <given-names>L</given-names>
          </name>
          <name>
            <surname>Habegger</surname>
            <given-names>L</given-names>
          </name>
          <name>
            <surname>Pickrell</surname>
            <given-names>JK</given-names>
          </name>
          <name>
            <surname>Montgomery</surname>
            <given-names>SB</given-names>
          </name>
          <etal/>
          <article-title>A systematic survey of loss-of-function variants in human protein-coding genes</article-title>
          <source>Science</source>
          <year>2012</year>
          <volume>335</volume>
          <issue>6070</issue>
          <fpage>823</fpage>
          <lpage>828</lpage>
          <pub-id pub-id-type="doi">10.1126/science.1215040</pub-id>
          <pub-id pub-id-type="pmid">22344438</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B10">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Danecek</surname>
            <given-names>P</given-names>
          </name>
          <name>
            <surname>Auton</surname>
            <given-names>A</given-names>
          </name>
          <name>
            <surname>Abecasis</surname>
            <given-names>G</given-names>
          </name>
          <name>
            <surname>Albers</surname>
            <given-names>CA</given-names>
          </name>
          <name>
            <surname>Banks</surname>
            <given-names>E</given-names>
          </name>
          <name>
            <surname>DePristo</surname>
            <given-names>MA</given-names>
          </name>
          <name>
            <surname>Handsaker</surname>
            <given-names>RE</given-names>
          </name>
          <name>
            <surname>Lunter</surname>
            <given-names>G</given-names>
          </name>
          <name>
            <surname>Marth</surname>
            <given-names>GT</given-names>
          </name>
          <name>
            <surname>Sherry</surname>
            <given-names>ST</given-names>
          </name>
          <etal/>
          <article-title>The variant call format and VCFtools</article-title>
          <source>Bioinformatics</source>
          <year>2011</year>
          <volume>27</volume>
          <issue>15</issue>
          <fpage>2156</fpage>
          <lpage>2158</lpage>
          <pub-id pub-id-type="doi">10.1093/bioinformatics/btr330</pub-id>
          <pub-id pub-id-type="pmid">21653522</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B11">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Reese</surname>
            <given-names>MG</given-names>
          </name>
          <name>
            <surname>Moore</surname>
            <given-names>B</given-names>
          </name>
          <name>
            <surname>Batchelor</surname>
            <given-names>C</given-names>
          </name>
          <name>
            <surname>Salas</surname>
            <given-names>F</given-names>
          </name>
          <name>
            <surname>Cunningham</surname>
            <given-names>F</given-names>
          </name>
          <name>
            <surname>Marth</surname>
            <given-names>GT</given-names>
          </name>
          <name>
            <surname>Stein</surname>
            <given-names>L</given-names>
          </name>
          <name>
            <surname>Flicek</surname>
            <given-names>P</given-names>
          </name>
          <name>
            <surname>Yandell</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Eilbeck</surname>
            <given-names>K</given-names>
          </name>
          <article-title>A standard variation file format for human genome sequences</article-title>
          <source>Genome Biol</source>
          <year>2010</year>
          <volume>11</volume>
          <issue>8</issue>
          <fpage>R88</fpage>
          <pub-id pub-id-type="doi">10.1186/gb-2010-11-8-r88</pub-id>
          <pub-id pub-id-type="pmid">20796305</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B12">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Eilbeck</surname>
            <given-names>K</given-names>
          </name>
          <name>
            <surname>Lewis</surname>
            <given-names>SE</given-names>
          </name>
          <name>
            <surname>Mungall</surname>
            <given-names>CJ</given-names>
          </name>
          <name>
            <surname>Yandell</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Stein</surname>
            <given-names>L</given-names>
          </name>
          <name>
            <surname>Durbin</surname>
            <given-names>R</given-names>
          </name>
          <name>
            <surname>Ashburner</surname>
            <given-names>M</given-names>
          </name>
          <article-title>The Sequence Ontology: a tool for the unification of genome annotations</article-title>
          <source>Genome Biol</source>
          <year>2005</year>
          <volume>6</volume>
          <issue>5</issue>
          <fpage>R44</fpage>
          <pub-id pub-id-type="doi">10.1186/gb-2005-6-5-r44</pub-id>
          <pub-id pub-id-type="pmid">15892872</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B13">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Grantham</surname>
            <given-names>R</given-names>
          </name>
          <article-title>Amino acid difference formula to help explain protein evolution</article-title>
          <source>Science</source>
          <year>1974</year>
          <volume>185</volume>
          <issue>4154</issue>
          <fpage>862</fpage>
          <lpage>864</lpage>
          <pub-id pub-id-type="doi">10.1126/science.185.4154.862</pub-id>
          <pub-id pub-id-type="pmid">4843792</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B14">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Li</surname>
            <given-names>WH</given-names>
          </name>
          <name>
            <surname>Wu</surname>
            <given-names>CI</given-names>
          </name>
          <name>
            <surname>Luo</surname>
            <given-names>CC</given-names>
          </name>
          <article-title>Nonrandomness of point mutation as reflected in nucleotide substitutions in pseudogenes and its evolutionary implications</article-title>
          <source>J Mol Evol</source>
          <year>1984</year>
          <volume>21</volume>
          <issue>1</issue>
          <fpage>58</fpage>
          <lpage>71</lpage>
          <pub-id pub-id-type="doi">10.1007/BF02100628</pub-id>
          <pub-id pub-id-type="pmid">6442359</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B15">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Krzywinski</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Schein</surname>
            <given-names>J</given-names>
          </name>
          <name>
            <surname>Birol</surname>
            <given-names>I</given-names>
          </name>
          <name>
            <surname>Connors</surname>
            <given-names>J</given-names>
          </name>
          <name>
            <surname>Gascoyne</surname>
            <given-names>R</given-names>
          </name>
          <name>
            <surname>Horsman</surname>
            <given-names>D</given-names>
          </name>
          <name>
            <surname>Jones</surname>
            <given-names>SJ</given-names>
          </name>
          <name>
            <surname>Marra</surname>
            <given-names>MA</given-names>
          </name>
          <article-title>Circos: an information aesthetic for comparative genomics</article-title>
          <source>Genome Res</source>
          <year>2009</year>
          <volume>19</volume>
          <issue>9</issue>
          <fpage>1639</fpage>
          <lpage>1645</lpage>
          <pub-id pub-id-type="doi">10.1101/gr.092759.109</pub-id>
          <pub-id pub-id-type="pmid">19541911</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B16">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Harris</surname>
            <given-names>TW</given-names>
          </name>
          <name>
            <surname>Antoshechkin</surname>
            <given-names>I</given-names>
          </name>
          <name>
            <surname>Bieri</surname>
            <given-names>T</given-names>
          </name>
          <name>
            <surname>Blasiar</surname>
            <given-names>D</given-names>
          </name>
          <name>
            <surname>Chan</surname>
            <given-names>J</given-names>
          </name>
          <name>
            <surname>Chen</surname>
            <given-names>WJ</given-names>
          </name>
          <name>
            <surname>De La Cruz</surname>
            <given-names>N</given-names>
          </name>
          <name>
            <surname>Davis</surname>
            <given-names>P</given-names>
          </name>
          <name>
            <surname>Duesbury</surname>
            <given-names>M</given-names>
          </name>
          <name>
            <surname>Fang</surname>
            <given-names>R</given-names>
          </name>
          <etal/>
          <article-title>WormBase: a comprehensive resource for nematode research</article-title>
          <source>Nucleic Acids Res</source>
          <year>2010</year>
          <volume>38</volume>
          <issue>Database issue</issue>
          <fpage>D463</fpage>
          <lpage>D467</lpage>
          <pub-id pub-id-type="pmid">19910365</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B17">
        <mixed-citation publication-type="journal">
          <name>
            <surname>Drmanac</surname>
            <given-names>R</given-names>
          </name>
          <name>
            <surname>Sparks</surname>
            <given-names>AB</given-names>
          </name>
          <name>
            <surname>Callow</surname>
            <given-names>MJ</given-names>
          </name>
          <name>
            <surname>Halpern</surname>
            <given-names>AL</given-names>
          </name>
          <name>
            <surname>Burns</surname>
            <given-names>NL</given-names>
          </name>
          <name>
            <surname>Kermani</surname>
            <given-names>BG</given-names>
          </name>
          <name>
            <surname>Carnevali</surname>
            <given-names>P</given-names>
          </name>
          <name>
            <surname>Nazarenko</surname>
            <given-names>I</given-names>
          </name>
          <name>
            <surname>Nilsen</surname>
            <given-names>GB</given-names>
          </name>
          <name>
            <surname>Yeung</surname>
            <given-names>G</given-names>
          </name>
          <etal/>
          <article-title>Human genome sequencing using unchained base reads on self-assembling DNA nanoarrays</article-title>
          <source>Science</source>
          <year>2010</year>
          <volume>327</volume>
          <issue>5961</issue>
          <fpage>78</fpage>
          <lpage>81</lpage>
          <pub-id pub-id-type="doi">10.1126/science.1181498</pub-id>
          <pub-id pub-id-type="pmid">19892942</pub-id>
        </mixed-citation>
      </ref>
      <ref id="B18">
        <mixed-citation publication-type="other">
          <source>Complete Genomics 69 Genomes Data</source>
          <comment>
            <ext-link ext-link-type="ftp" xlink:href="ftp://ftp2.completegenomics.com/Multigenome_summaries/Complete_Public_Genomes_69genomes_B37_mkvcf.vcf.bz2">ftp://ftp2.completegenomics.com/Multigenome_summaries/Complete_Public_Genomes_69genomes_B37_mkvcf.vcf.bz2</ext-link>
          </comment>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>
