Fresh NCBI Gnomon annotation files are available for D. grimshawi, this time with RNA-seq evidence so it’s a decent dataset (you skipped the original dataset on my advice because it had no RNA-seq and wasn’t very good). I also posted the BED graphs for all the RNA-seq tracks, in the same format that I provided for the earlier test files (I still need to get the other files off tape for you). Files are available as two tarballs. ftp://ftp.ncbi.nlm.nih.gov/pub/murphyte/.annotations_for_flybase/Dgri_101.20160610.tar.gz ftp://ftp.ncbi.nlm.nih.gov/pub/murphyte/.annotations_for_flybase/Dgri_101_rnaseq.20160610.tar.gz The files in the first tarball are the same as before, with the addition of four new files. These are the annotation in a form pretty-similar to how we would put them directly into RefSeq. So they are pre-filtered for the subset of models that we would retain (the ACCEPT set), and GeneIDs and RefSeq transcript and protein accessions are tracked where we would retain the same identifiers. They also integrate protein names based on best BLAST hits to SwissProt (with thresholds). The files are: final_asn_markup.asn - ASN.1 format, for me final_asn_markup.gff3 - GFF3 conversion of the first file final_asn_rna.fa - RefSeq RNA transcripts FASTA final_asn_prot.fa - RefSeq proteins FASTA For “new” RefSeqs and GeneIDs, they’re assigned dummy values starting with 99. These are placeholders only, and should never be used as permanent identifiers. If it was a full annotation, then permanent accessions and GeneIDs would have been assigned, but we have to work these a different way. I see the tRNA GeneIDs didn’t get tracked in this form, likely because we’re running the annotation somewhat differently and the code doesn’t think the existing GeneIDs are real. We haven’t yet added any short RNA predictions, and our mechanism for pulling data from miRBase doesn’t work with the way the annotation is run, so you’ll have to add those in like before. They don’t have FBgn, FBtr, FBpr Dbxrefs, but you can match to those through GeneID and RefSeq transcript/protein. So you can take a look at those files and see if they make things any easier to manage, or just process using the original files like you did before. I think all the file formats for those are unchanged, but there’s a chance a tab-delimited file or two had some columns altered, which I expect would cause a dramatic failure at some step and we can sort out what changed. So please take a look and see if these will flow smoothly into FlyBase. Best regards, Terrence Murphy NIH/NLM/NCBI