Difference between revisions of "WormBase-Caltech Weekly Calls"

From WormBaseWiki
Jump to navigationJump to search
 
(468 intermediate revisions by 10 users not shown)
Line 21: Line 21:
 
[[WormBase-Caltech_Weekly_Calls_2019|2019 Meetings]]
 
[[WormBase-Caltech_Weekly_Calls_2019|2019 Meetings]]
  
 +
[[WormBase-Caltech_Weekly_Calls_2020|2020 Meetings]]
  
GoToMeeting link: https://www.gotomeet.me/wormbase1
+
= 2021 Meetings =
  
= 2020 Meetings =
+
[[WormBase-Caltech_Weekly_Calls_January_2021|January]]
  
[[WormBase-Caltech_Weekly_Calls_January_2020|January]]
+
[[WormBase-Caltech_Weekly_Calls_February_2021|February]]
  
[[WormBase-Caltech_Weekly_Calls_February_2020|February]]
+
[[WormBase-Caltech_Weekly_Calls_March_2021|March]]
  
[[WormBase-Caltech_Weekly_Calls_March_2020|March]]
 
  
 +
== April 1, 2021 ==
  
== April 2, 2020 ==
+
=== Antibodies ===
 +
* Alignment of the antibody class to Alliance:
 +
** Propose to move possible_pseudonym (192) and Other_animal (37) to remarks. Those tags are not currently used for curation.
 +
*** Other animal is sometimes used for older annotations, e.g. authors say that the antibodies were raised both  in rats and rabbits. Standard practice would create 2 records, one for the rat antibody and one for the rabbit.
 +
*** Possible pseudonym was used when  a curator was not able to unambiguously assign a previous antibody to a record. (we have a Other name -synonym- tag to capture unambiguous ones). When moving to remarks we can keep a controlled vocabulary for easy future parsing, e.g. “possible_pseudonym:”
 +
** Antigen field: currently separated into Protein, peptide, and other_antigen (e.g.: homogenate of early C.elegans embryos, sperm). Propose to use just one antigen field to capture antigen info.
  
=== Community phenotype requests ===
+
All changes proposed above were approved by the group
* March 9-28
 
* 2,548 emails went out; 89 bounced; 6 resent; 13 backup; 2,478 successful emails
 
* 361 annotations overall
 
* 48 papers requested received curation (2% response rate)
 
* 53 distinct papers overall (5 papers without request)
 
* 53 distinct persons overall
 
  
=== Community curation volunteers ===
+
=== textpress-dev clean up ===
* Tracking volunteers [https://docs.google.com/spreadsheets/d/1ldECC44PXMilcDO6ctz-8AkRZntfDoV0Wtc4F-T_Zvg/edit?usp=sharing here]
+
* Michael has asked curators to assess what they have on textpresso-dev as it will not be around forever :-(
* 14 volunteers so far, all have been assigned a WBPerson ID
+
* is it okay to transfer data and files we want to keep to tazendra? and then to our own individual machines?
* Chris will set up a webinar tutorial in the coming week or two
+
* Direct access may be possible via Caltech VPN
 +
* Do we want to move content to AWS? May be complicated; it is still easy and cheap to maintain local file systems/machines
  
=== AFP pipeline ===
+
=== Braun servers ===
* Will resend email requests to authors that haven't already responded
+
* 3 servers stored in Braun server room; is there a new contact person for accessing these servers?
* May also send out for older papers
+
* Mike Miranda replacement just getting settled; Paul will find out who is managing the server room and let Raymond know
* May work with people to help
 
* Does the old AFP form still work? It should
 
* If someone has a link to the old form, they won't get one for the new form
 
* Maybe could set up an automatic redirect from the old form to the new form
 
* Received many submissions recently (>20% response rate)
 
  
=== Ontology Annotator ===
+
=== Citace upload ===
* Need to work on Genotype OA dumper
+
* Next Friday, April 9th, by end of the day
* Turns out semicolons are problematic (currently in genotypes and transgenes) for object names (ontology fields)
+
* Wen will contact Paul Davis for the frozen WS280 models file
* Ampersands (&) are also problematic for object names in the OA
 
** 20237  | Is[Pgcy-5::daf-2a::venus; Punc-122::mCherry]                          | 2014-10-08 10:32:45.874519-07
 
** 20239  | Ex[Pgcy-5::casy-1::venus; Pgcy-5::aman-2::mCherry; Punc-122::mCherry] | 2014-10-08 10:45:23.202362-07
 
** 20238  | Is[Pgcy-5::daf-2c::venus; Punc-122::mCherry]                          | 2014-10-08 10:38:19.859078-07
 
** 25249  | Ex[Prheb-1::rheb-1::GFP; unc-119(+]                                  | 2018-06-29 10:16:40.784295-07
 
** 16283  | [hlh-13::GFP;unc-119(+)]                                              | 2013-02-07 17:43:22.384819-08
 
** 26131  | Ex[pedc-3EDC-3::DsRed;pRF4]                                          | 2019-08-14 08:44:49.91063-07
 
  
=== Use Slack More ===
 
* Slack is a good tool for quick communication among team members; would be good for all curators to join Slack to enable efficient communication
 
  
 +
== April 8, 2021 ==
  
== April 9, 2020 ==
+
=== Braun server outage ===
 +
* Raymond fixed; now Spica, wobr and wobr2 are back up
  
=== Volunteer curators ===
+
=== Textpresso API ===
* Have sent out emails to schedule tutorials
+
* Was down yesterday affecting WormiCloud; Michael has fixed
* Chris had one tutorial with Michael Davies (Alyson Ashe's lab) yesterday
+
* Valerio will learn how to manage the API for the future
* One already scheduled for next Monday with Wilber and Stephanie from Paul's lab
 
* Two others already scheduled for next Tuesday with Lina Dahlberg and Colin Dolphin
 
  
===TAGC is virtual (4.22-25.2020)===
+
=== Grant opportunities ===
FYI in case you missed it
+
* Possibilities to apply for supplements
*You still have to register (it's free), if you hadn't before
+
* May 15th deadline
https://genetics-gsa.org/tagc-2020/registration/
+
* Druggable genome project
 +
** Pharos: https://pharos.nih.gov/
 +
** could we contribute?
 +
* Visualization, tools, etc.
 +
* Automated person descriptions?
 +
* Automated descriptions for proteins, ion channels, druggable targets, etc.?
  
===summer students===
+
=== New WS280 ONTOLOGY FTP directory ===
* Caltech SURF students (and other summer students worldwide) now are looking for projects
+
* Changes requested here: https://github.com/WormBase/website/issues/7900
* Maybe they could curate for WormBase
+
* Here's the FTP URL: ftp://ftp.wormbase.org/pub/wormbase/releases/WS280/ONTOLOGY/
* In addition to phenotype, they could curate:
+
* Known issues (Chris will report):
** Allele/lesion sequence curation (using Allele Sequence form); maybe Paul Davis could make a tutorial video?
+
** Ontology files are provided as ".gaf" in addition to ".obo"; we need to remove the ".gaf" OBO files
** Anatomy function, looking for novel info; opportunity to program/code
+
** Some files are duplicated and/or have inappropriate file extensions
  
=== OA semicolon issue ===
+
=== Odd characters in Postgres ===
* Juancarlos has fixed the issues on sandbox
+
* Daniela and Juancarlos discovered some errors with respect to special characters pasted into the OA
* Curators should test on Mangolassi
+
* Daniela would like to automatically pull in micropublication text (e.g. figure captions) into Postgres
 +
* We would need an automated way to convert special characters, like degree symbols ° into html unicode \&deg\;
 +
* Juancarlos and Valerio will look into possibly switching from a Perl module to a Python module to handle special characters
  
=== Textmining/automation ===
 
* Daniela will discuss with Christina Zorn from Xenbase
 
* Will discuss SVM, AFP, Textpresso, etc.
 
  
=== Retracted WBPapers ===
+
== April 15, 2021 ==
* Jae & Kimberly put in GitHub ticket to make retractions clear on WormBase site
 
* https://github.com/WormBase/website/issues/7637
 
* Can we systematically detect retractions? Yes
 
* What about finding papers that cite retractions? Maybe, but likely tricky
 
  
 +
=== Special characters in Postgres/OA ===
 +
* Juancarlos working on/proposing a plan to store UTF-8 characters in Postgres and the OA which would then get converted, at dumping, to HTML entities (e.g. α) for the ACE files
 +
* There is still a bit of cleanup needed to fix or remove special characters (not necessarily UTF-8) that apparently got munged upon copy/pasting into the OA in the past
 +
* Note: copy/paste from a PDF often works fine, but sometimes does not work as expected so manual intervention would be needed (e.g. entering Greek characters by hand in UTF-8 format)
 +
* Would copy/pasting from HTML be better than PDF?
 +
* For Person curation it would be good to be able to faithfully store and display appropriate foreign characters (e.g. Chinese characters, Danish characters, etc.)
 +
* Mangolassi script called "get_summary_characters.pl" located here: /home/postgres/work/pgpopulation/grg_generegulation/20200618_summary_characters
 +
** Juancarlos will modify script to take a data type code as an argument on the command line and return all Postgres tables (and their respective PGIDs) that have special characters, e.g.
 +
*** $ ./get_summary_characters.pl exp
 +
*** $ ./get_summary_characters.pl int
 +
*** $ ./get_summary_characters.pl grg
 +
** or could pass just the datatype + field (postgres table). e.g.
 +
*** $ ./get_summary_characters.pl pic_description
 +
** Juancarlos will email everyone once it's ready.  It's ready, email sent.  Script is at /home/postgres/work/pgpopulation/oa_general/20210411_unicode_html/get_summary_characters.pl  Symlink this to your directory and run it from there, it will create files in the directory you are at when running it.
 +
* Action items:
 +
** Juancarlos will update the "get_summary_characters.pl" script as described above
 +
** Curators should use the "get_summary_characters.pl" to look for (potentially) bad characters in their OAs/Postgres tables
 +
** Need to perform bulk (automated) replacement of existing HTML entities into corresponding UTF-8 characters
 +
** Curators will need to work with Juancarlos for each OA to modify the dumper
 +
** Juancarlos will write (or append to existing) Postgres/OA dumping scripts to:
 +
*** 1) Convert UTF-8 characters to HTML entities in ACE files
 +
*** 2) Convert special quote and hyphen characters into simple versions that don't need special handling
  
== April 16, 2020 ==
+
=== CeNGEN pictures ===
 +
* Model change went in to accommodate images from the CeNGEN project
 +
* Want gene page images for CeNGEN data; have the specifications for such images been worked out? Maybe not yet
 +
* Raymond and Daniela will work with data producers to acquire images when ready
  
=== Community Phenotype Curation Tutorials ===
+
=== Supplement opportunities ===
* Chris has run 6 tutorials, recorded 4
+
* Money available for software development to "harden" existing software
* MPG files saved on DropBox; ask Chris for access
+
* Might be possible to make Eduardo's single cell analysis tools more sustainable
* Plan to edit videos to make tutorial video to post on WB YouTube channel
+
* Could make WormiCloud adapted to Alliance?
 +
* Put Noctua on more stable production footing? (GO cannot apply as they are in final year of existing grant)
  
=== Author First Pass ===
+
=== Student project for Textpresso ===
* May run a webinar and use Zoom to record
+
* Create tool to allow user to submit text and return a list of similar papers
* May make a short tutorial video
+
* Use cases:
* Jae: Is there documentation for terminology used in the form?
+
** curator wants an alert to find papers similar to what they've curated
 +
** look for potential reviewers of a paper based on similar text content
  
=== Zoom accounts ===
 
* People can try to use Caltech Zoom account
 
  
 +
== April 22, 2021 ==
  
== April 23, 2020 ==
+
=== LinkML hackathon ===
 +
* Need to consider who works on what and how to coordinate
 +
* Need to practice good Git practice
 +
** Merge main branch into local branch before merging back into main branch to make sure everything works
 +
* How will we best handle AceDB hash structures? likely use something like Mark QT demonstrated
 +
** Do we have any/many hash-within-hash structures? #Molecular_change is used as a hash and tags within that model all reference the #Evidence hash
 +
* GO annotation extensions offer an interesting challenge
  
=== Community Phenotype Curation Tutorials ===
+
=== IWM workshop ===
* Chris has finished first round of tutorials; 8 tutorials, 6 video recordings
+
* Need to submit a workshop schedule (who speaks about what and when) by next Thursday April 29th
* There are ~8 new volunteers; will setup tutorials for them soon
+
* An initial idea was to promote data in ACEDB that may be underutilized or many users may be unaware of
 +
** An example might be transcription factor data: the ?Transcription_factor class and the modENCODE TF data
 +
** Single cell data and tools? CeNGEN, Eduardo's single cell tools
 +
** RNA-Seq FPKM values for genes and related data; Wen will write script to pull out FPKM values from SRA data and send to Magdalena
 +
* In addition to WB data types, we will cover Alliance, AFP, and community curation
 +
* Google doc for workshop here: https://docs.google.com/document/d/1H9ARhBRMKBNuOhjyxVQ_1o6cysvpppI7uA-TJrO_UZ4/edit?usp=sharing
  
=== ECO code implementation ===
+
=== WB Progress Report ===
* ?ECO_term to replace ?GO_code in ACEDB models
+
* Due April 30th
* GAF files with three-letter codes can still be generated by mapping
+
* There will be two documents: progress and plans
 +
* Place text in the appropriate places (don't write as a single integrated unit)
 +
* Paul S will put together a Google doc
 +
* We CAN include Alliance harmonization efforts
 +
* 2020 Progress report: https://docs.google.com/document/d/1f3ettnkvwoKKiaAA4TSrpSQPEF7FmVVn6u2UdflA_So/edit?usp=sharing
 +
* Last year milestone was WS276; we will compare to WS280
 +
* Google "WormBase Grants" folder: https://drive.google.com/drive/folders/1p8x9tEOfZ4DQvTcPSdNR5-JoPJu--ZAu?usp=sharing
 +
* 2021 Progress Report document here: https://docs.google.com/document/d/13E9k5JvDpUN4kWnrTm4M2iphnAJSTpk02ZiGl8O6bM4/edit?usp=sharing
  
=== Simplemine for Alliance ===
 
* Wen has presented proposal to Search group
 
* Plan is to have a link to the Alliance Simplemine prototype from the Alliance web page
 
  
=== Venn diagram tool ===
+
== April 29, 2021 ==
* Conceived by Jae, implemented by Sibyl
 
* Currently used for interactions data
 
* Could use for other data types like phenotype (e.g. comparing RNAi vs. allele phenotype)
 
* Could also use for Expression data, e.g. comparing results from different methods
 
* Could maybe use for disease data
 
  
=== AFP tutorial ===
+
=== IWM Workshop Schedule ===
* Daniela, Kimberly, Valerio will run through the AFP form with Nikita from Gupta lab tomorrow
+
* Schedule format due today (April 29th)
* May record in the future to make a tutorial video
+
* [https://docs.google.com/document/d/1H9ARhBRMKBNuOhjyxVQ_1o6cysvpppI7uA-TJrO_UZ4/edit#bookmark=id.jrjo4xhfnh7b Tentative schedule here]
* Daniela may (re-)start curating markers for relevant expression patterns
+
* Format proposal is 4, 15-minute talks followed by 30 minutes of open discussion / Q&A
* Wen noticed that many tissue markers are artificial (not necessarily endogenous sequence)
+
* Still need someone to speak (~15 minutes) about the Alliance
  
=== Expression markers ===
+
=== WB Progress Report ===
* SURF student projects: Identifying good expression markers? Maybe, but may require more curation experience
+
* 2021 documents in [https://drive.google.com/drive/folders/1p8x9tEOfZ4DQvTcPSdNR5-JoPJu--ZAu?usp=sharing this Google Drive folder]
* Wen looked at expression cluster data; hard to find good, very specific (i.e. neuron) markers
+
* Note: there is one [https://docs.google.com/document/d/13E9k5JvDpUN4kWnrTm4M2iphnAJSTpk02ZiGl8O6bM4/edit?usp=sharing 2021 "Progress" document] and a second (separate) [https://docs.google.com/document/d/1j0HkCwuimK6DD-ui1tAkYMNpLRhxR9xb1FdSDZXFXCI/edit?usp=sharing "Future Plans" document]
* Daniela may (re-)start curating markers for relevant expression patterns
+
* Existing future plans text has been moved to the "Future Plans" document
* Wen noticed that many tissue markers are artificial (not necessarily endogenous sequence)
 
* Already have an "Expression markers" widget on anatomy term pages
 
* Could combinations of genes (e.g. cGal) act as markers?
 
  
== April 30, 2020 ==
+
=== OpenBiosystems RNAi clone IDs ===
 +
* User looking to map Open Biosystems RNAi clone names to WB clone names
 +
* We may need to get a mapping file from Open Biosystems
  
=== Adding ?ECO_term class for WS278 ===
+
=== FPKM data ===
* Proposed[https://wiki.wormbase.org/index.php/Evidence_Code_Ontology#.3FECO_term_Model ?ECO_term model]
+
* Wen has produced a csv file of FPKM values; can generate as part of the SPELL pipeline
** How are the Parent/Child and Ancestor/Descendant tags used in WB for ontology classes?  Do we still need them in .ace files?
+
* May be better to generate at Hinxton
*Confirm proposed changes to class models that will use this tag:
 
** ?GO_annotation
 
** ?Phenotype
 
** ?Disease_model_annotation
 
  
=== Ontology term models in WB ===
+
=== OA Dumpers ===
* Discuss using ?RO_term values in our WB ontology term models
+
* Daniela and Juancarlos have been working on the Picture OA and Expr OA dumpers
* Currently relations between ontology terms are captured with text that is sometimes inconsistent for the same concept, e.g. is_a
+
* Inconsistencies have accumulated for all OA dumpers as each has been made separately
* Where possible, should be use ?RO_term to express the relations between ontology terms in our WB models?
+
* Juancarlos is working on a generalized, modular way to handle dumping
* Impact on web display?
+
* Should we handle historical genes in the same way across OAs?
 +
** Sure, but we need the "Historical_gene" tag in the respective ACEDB model
 +
** Decision: we will continue to only dump historical genes for specific OAs, with a plan to maybe make consistent across OAs in the future
 +
* Could we retroactively deal with paper-gene connections? We could possibly look in Postgres history tables to see which genes had been replaced previously (by Kimberly)
  
===Entries in the new Genotype OA===
+
=== Gene name ambiguities ===
*21 genotype entries created in the Genotype OA required for disease curation
+
* Jae noticed that some gene names associated with multiple WBGene IDs (e.g. one public name is the same as another gene's other name) have the same references attached
*Few more to come, and at some point need to work on the dumper, in order to submit for WS278
+
* May require updating the paper-gene connections for some of these
*The use of the Genotype class across disease related classes waiting on Paul D. for approval, will need dumper changes as well; hopefully we have enough time to get all this done for WS278
+
* One example is cep-1 gene. It associates with 3 diff WBgeneID and sharing papers in the reference widget.
 +
 
 +
=== NIH Supplement for AI readiness ===
 +
* Could we set up curation for neural circuits using a knowledge graph (e.g. GO-CAM)?
 +
** Maybe we could convert the anatomy function model to LinkML -> OWL statements?
 +
** Maybe setup a graphical curation interface?
 +
* Transcriptional regulation
 +
** Would be good to establish a common model (for the Alliance?)
 +
** CeNGEN project produced lots of predictions of TF binding sites based on single-cell expression data; Eduardo: these models should be able to be regenerated each time new data sets are published, but this requires greater integration in a central, sustainable resource
 +
* Paul S can send a link for the supplement
 +
 
 +
=== Variant First Pass Pipeline ===
 +
* Valerio: Are there any existing pipelines to make allele-paper and/or strain-paper associations?
 +
* Not sure, should ask Karen

Latest revision as of 19:13, 29 April 2021

Previous Years

2009 Meetings

2011 Meetings

2012 Meetings

2013 Meetings

2014 Meetings

2015 Meetings

2016 Meetings

2017 Meetings

2018 Meetings

2019 Meetings

2020 Meetings

2021 Meetings

January

February

March


April 1, 2021

Antibodies

  • Alignment of the antibody class to Alliance:
    • Propose to move possible_pseudonym (192) and Other_animal (37) to remarks. Those tags are not currently used for curation.
      • Other animal is sometimes used for older annotations, e.g. authors say that the antibodies were raised both in rats and rabbits. Standard practice would create 2 records, one for the rat antibody and one for the rabbit.
      • Possible pseudonym was used when a curator was not able to unambiguously assign a previous antibody to a record. (we have a Other name -synonym- tag to capture unambiguous ones). When moving to remarks we can keep a controlled vocabulary for easy future parsing, e.g. “possible_pseudonym:”
    • Antigen field: currently separated into Protein, peptide, and other_antigen (e.g.: homogenate of early C.elegans embryos, sperm). Propose to use just one antigen field to capture antigen info.

All changes proposed above were approved by the group

textpress-dev clean up

  • Michael has asked curators to assess what they have on textpresso-dev as it will not be around forever :-(
  • is it okay to transfer data and files we want to keep to tazendra? and then to our own individual machines?
  • Direct access may be possible via Caltech VPN
  • Do we want to move content to AWS? May be complicated; it is still easy and cheap to maintain local file systems/machines

Braun servers

  • 3 servers stored in Braun server room; is there a new contact person for accessing these servers?
  • Mike Miranda replacement just getting settled; Paul will find out who is managing the server room and let Raymond know

Citace upload

  • Next Friday, April 9th, by end of the day
  • Wen will contact Paul Davis for the frozen WS280 models file


April 8, 2021

Braun server outage

  • Raymond fixed; now Spica, wobr and wobr2 are back up

Textpresso API

  • Was down yesterday affecting WormiCloud; Michael has fixed
  • Valerio will learn how to manage the API for the future

Grant opportunities

  • Possibilities to apply for supplements
  • May 15th deadline
  • Druggable genome project
  • Visualization, tools, etc.
  • Automated person descriptions?
  • Automated descriptions for proteins, ion channels, druggable targets, etc.?

New WS280 ONTOLOGY FTP directory

Odd characters in Postgres

  • Daniela and Juancarlos discovered some errors with respect to special characters pasted into the OA
  • Daniela would like to automatically pull in micropublication text (e.g. figure captions) into Postgres
  • We would need an automated way to convert special characters, like degree symbols ° into html unicode \&deg\;
  • Juancarlos and Valerio will look into possibly switching from a Perl module to a Python module to handle special characters


April 15, 2021

Special characters in Postgres/OA

  • Juancarlos working on/proposing a plan to store UTF-8 characters in Postgres and the OA which would then get converted, at dumping, to HTML entities (e.g. α) for the ACE files
  • There is still a bit of cleanup needed to fix or remove special characters (not necessarily UTF-8) that apparently got munged upon copy/pasting into the OA in the past
  • Note: copy/paste from a PDF often works fine, but sometimes does not work as expected so manual intervention would be needed (e.g. entering Greek characters by hand in UTF-8 format)
  • Would copy/pasting from HTML be better than PDF?
  • For Person curation it would be good to be able to faithfully store and display appropriate foreign characters (e.g. Chinese characters, Danish characters, etc.)
  • Mangolassi script called "get_summary_characters.pl" located here: /home/postgres/work/pgpopulation/grg_generegulation/20200618_summary_characters
    • Juancarlos will modify script to take a data type code as an argument on the command line and return all Postgres tables (and their respective PGIDs) that have special characters, e.g.
      • $ ./get_summary_characters.pl exp
      • $ ./get_summary_characters.pl int
      • $ ./get_summary_characters.pl grg
    • or could pass just the datatype + field (postgres table). e.g.
      • $ ./get_summary_characters.pl pic_description
    • Juancarlos will email everyone once it's ready. It's ready, email sent. Script is at /home/postgres/work/pgpopulation/oa_general/20210411_unicode_html/get_summary_characters.pl Symlink this to your directory and run it from there, it will create files in the directory you are at when running it.
  • Action items:
    • Juancarlos will update the "get_summary_characters.pl" script as described above
    • Curators should use the "get_summary_characters.pl" to look for (potentially) bad characters in their OAs/Postgres tables
    • Need to perform bulk (automated) replacement of existing HTML entities into corresponding UTF-8 characters
    • Curators will need to work with Juancarlos for each OA to modify the dumper
    • Juancarlos will write (or append to existing) Postgres/OA dumping scripts to:
      • 1) Convert UTF-8 characters to HTML entities in ACE files
      • 2) Convert special quote and hyphen characters into simple versions that don't need special handling

CeNGEN pictures

  • Model change went in to accommodate images from the CeNGEN project
  • Want gene page images for CeNGEN data; have the specifications for such images been worked out? Maybe not yet
  • Raymond and Daniela will work with data producers to acquire images when ready

Supplement opportunities

  • Money available for software development to "harden" existing software
  • Might be possible to make Eduardo's single cell analysis tools more sustainable
  • Could make WormiCloud adapted to Alliance?
  • Put Noctua on more stable production footing? (GO cannot apply as they are in final year of existing grant)

Student project for Textpresso

  • Create tool to allow user to submit text and return a list of similar papers
  • Use cases:
    • curator wants an alert to find papers similar to what they've curated
    • look for potential reviewers of a paper based on similar text content


April 22, 2021

LinkML hackathon

  • Need to consider who works on what and how to coordinate
  • Need to practice good Git practice
    • Merge main branch into local branch before merging back into main branch to make sure everything works
  • How will we best handle AceDB hash structures? likely use something like Mark QT demonstrated
    • Do we have any/many hash-within-hash structures? #Molecular_change is used as a hash and tags within that model all reference the #Evidence hash
  • GO annotation extensions offer an interesting challenge

IWM workshop

  • Need to submit a workshop schedule (who speaks about what and when) by next Thursday April 29th
  • An initial idea was to promote data in ACEDB that may be underutilized or many users may be unaware of
    • An example might be transcription factor data: the ?Transcription_factor class and the modENCODE TF data
    • Single cell data and tools? CeNGEN, Eduardo's single cell tools
    • RNA-Seq FPKM values for genes and related data; Wen will write script to pull out FPKM values from SRA data and send to Magdalena
  • In addition to WB data types, we will cover Alliance, AFP, and community curation
  • Google doc for workshop here: https://docs.google.com/document/d/1H9ARhBRMKBNuOhjyxVQ_1o6cysvpppI7uA-TJrO_UZ4/edit?usp=sharing

WB Progress Report


April 29, 2021

IWM Workshop Schedule

  • Schedule format due today (April 29th)
  • Tentative schedule here
  • Format proposal is 4, 15-minute talks followed by 30 minutes of open discussion / Q&A
  • Still need someone to speak (~15 minutes) about the Alliance

WB Progress Report

OpenBiosystems RNAi clone IDs

  • User looking to map Open Biosystems RNAi clone names to WB clone names
  • We may need to get a mapping file from Open Biosystems

FPKM data

  • Wen has produced a csv file of FPKM values; can generate as part of the SPELL pipeline
  • May be better to generate at Hinxton

OA Dumpers

  • Daniela and Juancarlos have been working on the Picture OA and Expr OA dumpers
  • Inconsistencies have accumulated for all OA dumpers as each has been made separately
  • Juancarlos is working on a generalized, modular way to handle dumping
  • Should we handle historical genes in the same way across OAs?
    • Sure, but we need the "Historical_gene" tag in the respective ACEDB model
    • Decision: we will continue to only dump historical genes for specific OAs, with a plan to maybe make consistent across OAs in the future
  • Could we retroactively deal with paper-gene connections? We could possibly look in Postgres history tables to see which genes had been replaced previously (by Kimberly)

Gene name ambiguities

  • Jae noticed that some gene names associated with multiple WBGene IDs (e.g. one public name is the same as another gene's other name) have the same references attached
  • May require updating the paper-gene connections for some of these
  • One example is cep-1 gene. It associates with 3 diff WBgeneID and sharing papers in the reference widget.

NIH Supplement for AI readiness

  • Could we set up curation for neural circuits using a knowledge graph (e.g. GO-CAM)?
    • Maybe we could convert the anatomy function model to LinkML -> OWL statements?
    • Maybe setup a graphical curation interface?
  • Transcriptional regulation
    • Would be good to establish a common model (for the Alliance?)
    • CeNGEN project produced lots of predictions of TF binding sites based on single-cell expression data; Eduardo: these models should be able to be regenerated each time new data sets are published, but this requires greater integration in a central, sustainable resource
  • Paul S can send a link for the supplement

Variant First Pass Pipeline

  • Valerio: Are there any existing pipelines to make allele-paper and/or strain-paper associations?
  • Not sure, should ask Karen