Dear Drosophila database masters: I have a few general questions regarding what is being generated through the database and if there is a 'better' means of assessing some of the results. I am a protein chemist/structural biologist working on three drosophila proteins. The first protein is Roundabout-1 (Robo-1). This protein has been well characterized and demonstrates the common motifs found in cell adhesion proteins, and has thus been characterized as a cell adhesion protein in your database. This characterization is apparently wrong however because the Ig domains in robo are known to bind the slit protein and have no impact on cell adhesion. This leads to my first question about how experimental data is being incorporated into the databases that we currently have access to. I do understand that the task is difficult given the large number of sequences to be used. I feel some of this could be solved by integration with databases such as provided at the interactive fly: This would provide some experimental links and possibly help to identify homologous proteins. http://sdb.bio.purdue.edu/fly/aimain/1aahome.htm Is there some work in progress to link experimental data to the sequences? How or when can this be accessed? My second question is a bit more difficult to answer, but please provide any assistance that you think will be useful. As for the robo proteins, there are three members of the family robo-1,-2, and-3. Robo-1 has been characterized, but the other two have not. When the database is searched, robo-2 and robo-3 are found, but the identified cDNAs only include the extracellular portions of the proteins-which are composed of homologous Ig domains. The cytoplasmic tails (about 450 AA) are not found in the same entry, but are listed under the genomic scaffolds dumped into NCBI by Celera. Having worked on these domains, I know they are important for the function of the intact protein. The main problem (I believe) in identifying the protein is the absence of any homologous domains, which makes the translated cDNA appear to be a possible 'psuedo gene.' So my next question is given the large number of sequences of 'unknown' function, how does one tell if the protein produced is from a psuedo gene? Secondly, is there an entry or a compiled list of translated cDNA which are of 'unknown' function, but have been verified not to be a psuedo gene product? . Thanks for your assistance. You are doing a great job considering how much data is being generated. Sincerely, Thomas L. Selby, Ph. D. The Scripps Research Institute