Difference between revisions of "Specifications for source files"
From WormBaseWiki
Jump to navigationJump to searchLine 6: | Line 6: | ||
#Sentence number in source file; starting with 1 and ending with whatever total number of sentences are in the file | #Sentence number in source file; starting with 1 and ending with whatever total number of sentences are in the file | ||
− | # | + | #SSC (stands for Textpresso Sentence SCore) |
#Numerical Textpresso sentence score value | #Numerical Textpresso sentence score value | ||
− | #PID (stands for | + | #PID (stands for Paper IDentifier) |
#Database code:numerical identifier | #Database code:numerical identifier | ||
##For WormBase and dictyBase, who are sending annotations to the Protein2GO tool via web services, this identifier needs to be either a PubMed ID or a doi | ##For WormBase and dictyBase, who are sending annotations to the Protein2GO tool via web services, this identifier needs to be either a PubMed ID or a doi | ||
##If neither a PubMed ID nor a doi exists, then the annotation cannot be sent to Protein2GO | ##If neither a PubMed ID nor a doi exists, then the annotation cannot be sent to Protein2GO | ||
##For TAIR, who is not yet using Protein2GO, it can still be the TAIR document ID | ##For TAIR, who is not yet using Protein2GO, it can still be the TAIR document ID | ||
+ | #SID (stands for Textpresso Sentence IDentifier) | ||
+ | #Numerical value of the sentence number in the document in Textpresso | ||
+ | #Gene product name or synonym as identified by Textpresso search | ||
+ | ##WormBase: C. elegans (protein) |
Revision as of 13:55, 26 March 2013
Specifications for Textpresso for CCC Source Files
- The source files can be simplified a bit, but will retain the key information we need for curation and search and retrieval functions for the curation form.
- The format will continue to be a tab-delimited file containing, in order:
- Sentence number in source file; starting with 1 and ending with whatever total number of sentences are in the file
- SSC (stands for Textpresso Sentence SCore)
- Numerical Textpresso sentence score value
- PID (stands for Paper IDentifier)
- Database code:numerical identifier
- For WormBase and dictyBase, who are sending annotations to the Protein2GO tool via web services, this identifier needs to be either a PubMed ID or a doi
- If neither a PubMed ID nor a doi exists, then the annotation cannot be sent to Protein2GO
- For TAIR, who is not yet using Protein2GO, it can still be the TAIR document ID
- SID (stands for Textpresso Sentence IDentifier)
- Numerical value of the sentence number in the document in Textpresso
- Gene product name or synonym as identified by Textpresso search
- WormBase: C. elegans (protein)