Subject: cdna AF160881 I was very interested by the new complete D.melanogaster cds. AF160881 that you deposited in GenBank on 8/3/99. My interest is that it includes the coding region of a complete alpha-carbonic anhydrase (bp 287-1099 encode a 270 aa isozyme), which on the basis of analyzing a P1 genomic sequence (L39622) in 1995, we named CAH1 (Hewett-Emmett and Tashian 1996). At that time, we could not from the s There is however one problem with your cDNA clone and its current annotation. The translated sequence you describe (1420-2628) is from a transposon Tn10 (cf. GenBank entry AP000342, which gives the almost identical plasmid R100 sequence) which interrupts the 3' untranslated region of This is written because of my interest in the CAH1 gene and in seeing the annotations of it in the databases are as helpful to other users as possible. REF: D. Hewett-Emmett and R.E. Tashian (1996) Functional diversity, conservation, and convergence in the evolution of the alpha-, beta-, and gamma-carbonic anhydrase gene families. Mol. Phyl. Evol. 5: 50-77. In admiration for the wealth of Drosophila genome project data that is being made available and with best wishes David Hewett-Emmett Ph.D. Human Genetics Center, P.O. Box 20334, U.Texas - Houston, Houston, TX 77225-0334. > Subject: Re: cdna AF160881 David - this CA is included in the Adh region that has been fully annotated and is discussed in a paper in press in GENETICS (Sept 99 issue). Indeed I returned the proofs just 4 hours ago. You can see the paper on the www.fruitfly.org web site as a PDF file. We call the gene BG:DS00941.1 in this paper but say that it encodes a CA. Here is the translation we used: > BG:DS00941.1 270 AAs MSHHWGYTEENGPAHWAKEYPQASGHRQSPVDITPSSAKKGSELNVAPLK WKYVPEHTKSLVNPGYCWRVDVNGADSELTGGPLGDQIFKLEQFHCHWGC TDSKGSEHTVDGVSYSGELHLVHWNTTKYKSFGEAAAAPDGLAVLGVFLK AGNHHAELDKVTSLLQFVLHKGDRVTLPQGCDPGQLLPDVHTYWTYEGSL TTPPCSESVIWIVFKTPIEVSDDQLNAMRNLNAYDVKEECPCNEFNGKVI NNFRPPLPLGKRELREIGGH I have just checked this with BLASTP at thee NCBI and it looks very convincing ! I cannot answer any other questions raised in yr email to Berkeley, but I am sure that they will be back to you. I missed your paper in MPE, indeed it seems not to be in FlyBase. Would it be possible for you to fax me a copy ? I will _try_ (but cannot promise) to add a note in proof to our paper. Michael Ashburner > Subject: cdna AF160881 Michael: Thanks for such a quick response. The 270 aa protein looks to be identical to what I see. I will FAX the paper this afternoon. Also, I will look at the website. I do need to make a relatively minor correction to one statement in my email. In the 1996 paper we named the Drosophila gene CAH as you will see. In the interim, I found another distinct Drosophila CAH -- a fragment coded by an EST (AA246259). About a year ago, I wrote a lengthy review for a book which will likely come out late this year or in 2000. In this later review, I used the nomenclature CAH1 and CAH2. I have just looked at AA246259 again and it has re-kindled my interest in CAH2! There are now several overlapping ESTs (e.g. AI109097, AI260114) giving a protein sequence that looks complete -- unlike CAH1, there appear to be no genomic sequence corresponding to them. Best Wishes David Hewett-Emmett > Subject: CAH1 Dear Dr. Hewett-Emmett, Michael passed on your correspondence about CAH1 and also your paper 'Hewett-Emmett, Tashian, 1996 Molec. Phylog. Evol. 5: 50--77' for me to curate for FlyBase. I will curate the information that you wish the gene to be called CAH1 rather than CAH as you called it originally as a personal communication from you to FlyBase. Would it be OK for me to include the following information about CAH2 from your e-mail in the personal communication: >In the interim, I found another distinct Drosophila CAH -- a fragment >coded by an EST (AA246259). About a year ago, I wrote a lengthy review >for a book which will likely come out late this year or in 2000. In >this later review, I used the nomenclature CAH1 and CAH2. >I have just looked at AA246259 again and it has re-kindled my interest >in CAH2! There are now several overlapping ESTs (e.g. AI109097, >AI260114) giving a protein sequence that looks complete -- unlike CAH1, >there appear to be no genomic sequence corresponding to them. I would create a new gene called 'CAH2' and it would have the information that it corresponds to the EST's AA246259, AI109097 and AI260114 under it. Gillian \-------------------------------------------------------------- Gillian Millburn. FlyBase (Cambridge), \-------------------------------------------------------------- > Subject: RE: CAH1/CAH2 That's fine to release it now. If you give me a day or so, I can give you the complete list of ESTs for CAH2 but I need to check them out individually. David > Subject: RE: CAH2 Gillian: I have checked the ESTs. The first batch of 10 are almost identical (they vary a bit in length and where they start and finish). The central region and 3' end have only a single representative each. 5' end (encodes N-terminal region): AA246259, AI297502, AI404891 (longest EST), AI257332, AI113987, AA263558, AI517161, AA942429, AI109423, AI388709. Central: AI109097 3' end (Encodes C-terminal region): AI26011 The encoded protein is probably 335 aa: MRRCRNTPFAIVIAPILICASLVLAQDFGYEGRHGPEHWSEDYARCSGKH 50 QSPINIDQVSAVEKKFPKLEFFNFKVVPDNLQMTNNGHTVLVKMSYNEDE 100 IPSVRGGPLAEKTPLGYQFEQFHFHWGENDTIGSEDLINNRAYPAELHVV 150 LRNLEYPDFASALDKDHGIAVMAFFFQVGDKSTGGYEGFTNLLSQIDRKG 200 KSVNMTNPLPLGEYISKSVESYFSYTGSLTTPPCSEEVTWIDFTTPIDIT 250 EKQLNAFRLLTANDDHLKNNFRPIQPLNDRTLYKNYIEIPIHNMGSIPLV 300 DAENAAGKWRAQAAAVLLPLVVLAALSRTSIFRGF* 335 NOTES: (1) Most active-site residues conserved in the active alpha- CA enzymes are conserved in this sequence. Exception is \#136 Asp (His in all active carbonic anhydrases except the prokaryotes Erwinia and Klebsiella which have Asn). (2) At residue \#233 (shown as 'P') which is also a conserved site (Pro) in active alpha-CAs: AI260114 encodes Pro (bp 229-231 of DNA sequence) AI109097 encodes Ser (late in the DNA sequence) (3) At residue \#177bp (shown as 'Q'), AI260114 encodes a 'stop' (bp 32-34 of the DNA sequence). AI109097 bp.408-410 encodes Gln. Examination of bp 1-34 of AI260114 shows no similarity with bp 377-410 of AI109097. Sequencing of additional ESTs and/or genomic sequence will eventually resolve whether the alternatives chosen in NOTES (2) and (3) are correct. David Hewett-Emmett Ph.D. Human Genetics Center, University of Texas-Houston HSC, P.O. Box 20334, Houston, TX 77225-0334, USA