After running variant calling and creating consensus sequences of the transcriptomic samples (STAR version 2.5; Samtools version 0.1.19), using Pst isolate 104E137A-1, use the snpb_script.py to extract the average and standard deviation SNP/b, for the housekeeping genes and PST130_P495001.
After running variant calling and creating consensus sequences of the transcriptomic samples (STAR version 2.5; Samtools version 0.1.19), using Pst isolate 104E137A-1, use the gene_variant_identify.ipynb to:
- Find the different haplotype variants of the gene of interest, given no ambiguities are present in the consensus sequence.
- Find the different isoforms from the gene of interest.
- Align the different haplotypes and save a
.pngimage of the MSA alignment.
Extract information from the Expression Browser2
To identify which isolates of the ones included in the expression browser, lack expression of PST130_P495001, the data was downloaded (directory: REB_for_figshare - downloaded from figshare) and everything apart from the Dobon et al. samples (as they are from a timeseries) were used in the analysis. The transcript counts were used to remove isolates that had a low expression of the housekeeping genes, as a way of filtering to include only high quality reads. For this, any isolates with a mean expression <15% the median were removed. Out of the 888 remaining isolates, 10 were identified to lack expression of PST130_P495001. This data analysis was carried out using check_expression.ipynb.
To obtain the abundance counts of the transcriptomic reads, pseudoalignments were performed using Kallisto version 0.51.1. The transcript counts were then analysed using Sleuth version 0.30.1, using sleuth.R.
Footnotes
-
Schwessinger, B. et al. A Near-Complete Haplotype-Phased Genome of the Dikaryotic Wheat Stripe Rust Fungus Puccinia striiformis f. sp. tritici Reveals High Interhaplotype Diversity. mBio 9 (2018). 10.1128/mBio.02275-17 ↩ ↩2