nf-core/plasmodiumdrugres
Pipeline for analyzing drug resistance markers from Plasmodium microhaplotype data. It translates variants into amino acid changes at drug resistance loci and estimates allele frequencies and prevalences at both single-locus and multi-locus levels. Microhaplotype data can be supplied in the form of an allele table or a PMO file.
Introduction
This document describes the output produced by the pipeline. All paths below are relative to the top-level results directory (--outdir).
The main results are two summary tables: sl_summary.tsv for individual drug resistance codons and ml_summary.tsv for haplotypes across groups of codons. Most users only need these two files. The other outputs explain how the summaries were reached and help with troubleshooting.
Pipeline overview
The pipeline steps and estimation methods are described in the introduction. The outputs, in the order they are produced, are:
- Translated Loci: amino acid calls for each specimen at each locus of interest
- Single Locus Allele Frequencies: prevalence and frequency of each amino acid at each locus of interest
- Multi Locus Allele Frequencies: prevalence and frequency of haplotypes across groups of loci
- SL-from-ML Summary: single-locus frequencies derived from the multi-locus estimates
- Unassigned specimens: specimens left out because they have no population
- Raw summary tables: the full method-specific tables behind the summaries
- Pipeline information: reports and software versions from the run
Key terms
- Prevalence is the proportion of specimens in which an allele was detected at all. A specimen with a mixed (polyclonal) infection can count towards the prevalence of more than one allele at the same locus, so prevalences at a locus can add up to more than 1.
- Frequency is the estimated proportion of parasites in the population that carry an allele. Frequencies at a locus add up to 1. Estimating frequency from polyclonal infections is not straightforward, which is why the pipeline offers several methods.
- Population is the group of specimens an estimate is made for (for example a country, a health facility or a year). See population grouping in the usage docs. Without population grouping, all specimens form one population named by
--population_label(defaultpop1). - Variant identifiers use variant string notation, as used by STAVE. The gene, the amino acid position(s) and the amino acid(s) are separated by
:, with the gene given by itsgene_idfrom--loci_of_interest_bed:- single locus:
gene_id:aa_position:aa, for examplePF3D7_0417200.1:51:Iis isoleucine at codon 51 of dhfr - multi-locus haplotype: positions and amino acids are each joined by
_, in the same order, for examplePF3D7_0417200.1:51_59_108:I_R_Nis I at codon 51, R at 59 and N at 108
- single locus:
Translated Loci
The first step converts the microhaplotype sequences into amino acid calls. Each distinct microhaplotype is aligned to the reference sequence of its panel target, and the codon at each locus in --loci_of_interest_bed that the target covers is extracted and translated. This is done by translate_loci_of_interest from PGEcore.
When a locus is covered by more than one panel target, the calls are combined into one call per specimen, locus and amino acid. By default the reads are summed across all covering targets. With --collapse_calls_by_summing false, only the target with the most reads for that specimen and locus is used. The collapsed calls are the input to all the estimates below.
Use these files to check which loci your panel covers, how many specimens have a call at each locus, and what was called for a given specimen.
Output files
translated_loci/collapsed_amino_acid_calls.tsv.gz: One row per specimen, locus and amino acid, with the reference amino acid, the read count and the target(s) the call came from. This is the table used for estimation.amino_acid_calls.tsv.gz: The calls before collapsing: one row per specimen, target, microhaplotype and locus, with the microhaplotype sequence, the observed and reference codon, and the observed and reference amino acid.loci_covered_by_target_samples_info.tsv: One row per locus of interest, with the target covering it (covered_by_target), the reference amino acid, and the number of specimens with a call (n_samples) out of the total (total_samples). Use it to check that every locus you are interested in is covered by your panel.loci_of_interest_for_target_for_microhap.tsv.gz: The codon and amino acid each distinct microhaplotype sequence gives at each locus it covers. Useful for checking how a particular sequence was translated.
Single Locus Allele Frequencies
For each population, the pipeline estimates the prevalence and frequency of every amino acid seen at each locus of interest, for example the proportion of parasites carrying the crt K76T mutation.
Prevalence is calculated directly from the amino acid calls. Frequency is estimated with the method chosen by --slaf_method:
naive(default): a simple estimate from the calls in each specimen, using either the within-specimen proportion of reads (--naive_slaf_method read_count_prop, default) or presence/absence (presence_absence). Implemented in PGEcore asestimate_allele_frequency_naive.IDM: the incomplete data model, which estimates frequencies while accounting for polyclonal infections. Run through a PGEcore wrapper.mhaps_freq: microhaplotype frequencies are first estimated with Dcifer, then converted to amino acid frequencies at each locus. Run through a PGEcore wrapper.
The results for all populations are combined into one table with the same columns whichever method was used.
Output files
sl_summary.tsv: Single-locus prevalence and frequency for every population.
sl_summary.tsv columns
| Column | Description |
|---|---|
population |
Population label for this row (a user-defined grouping of samples; see usage docs). |
variant |
Single-locus variant in variant string notation, gene_id:aa_position:aa. |
prev |
Estimated prevalence for the variant in the population. |
sample_count |
Number of samples with the variant (for prevalence estimate). |
sample_total |
Total number of samples considered (for prevalence estimate). |
freq |
Estimated single-locus allele frequency. |
Method-specific columns are left out of this table and kept in raw_summaries/raw_sl_summary.tsv.
Multi Locus Allele Frequencies
Resistance to some drugs depends on a combination of mutations, for example the dhfr N51I, C59R and S108N triple mutant. Multi-locus estimates give the prevalence and frequency of each combination (haplotype) of amino acids across a group of loci. The groups are defined in --loci_groups; this step only runs when that file is provided.
In a polyclonal infection it is usually not possible to tell which amino acids came from the same parasite, so haplotype frequencies have to be estimated. The method is chosen with --mlaf_method:
naive(default): uses specimens with a call at every locus in the group, and infers their haplotypes when only one locus is mixed or when one allele dominates at every locus (its share of reads is above--naive_multilocus_wsaf_cut_off, default 0.70). Prevalence and frequency are then calculated from the within-specimen allele proportions (--naive_mlaf_method wsaf_prop, default) or presence/absence (presence_absence). Implemented in PGEcore asmultilocus_prevfreq_naive.MLBM: the MultiLociBiallelicModel. Run through a PGEcore wrapper.FEM: the FreqEstimationModel, a Bayesian model that also reports credible intervals. It needs the average complexity of infection, set with--fem_coi. Run through a PGEcore wrapper.
Output files
ml_summary.tsv: Multi-locus haplotype estimates for every population and loci group. When--loci_groupsis omitted, this file is still written but contains only the header row.
ml_summary.tsv columns
| Column | Description |
|---|---|
population |
Population label for this row. |
group_id |
Group identifier from --loci_groups. |
variant |
Haplotype in variant string notation, for example PF3D7_0417200.1:51_59_108:I_R_N. |
prev |
Estimated prevalence (included when available). |
sample_count |
Number of samples with the variant (included when available). |
sample_total |
Total number of samples considered (included when available). |
freq |
Estimated multi-locus haplotype frequency. |
Not all methods report prev, sample_count and sample_total. When present, they are included in the order shown above; otherwise the table contains only population, group_id, variant and freq.
Method-specific columns are left out of this table and kept in raw_summaries/raw_ml_summary.tsv.
SL-from-ML Summary
The haplotype frequencies from the multi-locus step can be turned back into single-locus frequencies by combining the frequencies of all haplotypes that carry each amino acid. Comparing these with the frequencies in sl_summary.tsv is a useful consistency check, because the two come from different methods.
Output files
raw_summaries/raw_sl_from_ml_summary.tsv: Single-locus frequencies derived from the multi-locus estimates. The columns depend on--mlaf_method(see Raw summary tables). When--loci_groupsis omitted, this file is still written but contains only the header row.
Unassigned specimens
When populations are assigned with --population_assignment or --pmo_population_fields, any specimen that has amino acid calls but no population is left out of every estimate. So that this does not go unnoticed, the pipeline lists those specimens in this file and logs a warning with their number during the run and again when it finishes.
Output files
unassigned_specimens.txt: Specimens left out because they have no population, one per line. Only written when at least one specimen is unassigned.
Raw summary tables
To keep sl_summary.tsv and ml_summary.tsv the same whichever method is used, columns that only some methods produce are removed from them. The full tables, with every column each method produces, are kept here. Use them if you need method-specific details, such as the credible intervals from FEM.
Output files
raw_summaries/raw_sl_summary.tsv: Full single-locus prevalence and frequency table.raw_ml_summary.tsv: Full multi-locus table.raw_sl_from_ml_summary.tsv: Single-locus frequencies derived from multi-locus estimates (see SL-from-ML Summary).
Single-locus method-specific columns
Extra columns in raw_summaries/raw_sl_summary.tsv:
| SLAF method | Extra columns |
|---|---|
IDM |
None. |
naive |
allele_count, allele_total (from prevalence estimation). |
mhaps_freq (via Dcifer) |
sample_total_for_allele_freq (sample total associated with the frequency estimate output). |
Multi-locus method-specific columns
Extra columns in raw_summaries/raw_ml_summary.tsv:
| MLAF method | Extra columns |
|---|---|
MLBM |
None. |
FEM |
sequence, median_freq, CI_2.5, CI_97.5 |
naive |
allele_count, allele_total |
raw_sl_from_ml_summary.tsv columns by MLAF method
| MLAF method | Columns in raw_sl_from_ml_summary.tsv |
|---|---|
MLBM |
population, variant, freq |
FEM |
population, variant, freq |
naive |
population, group_id, variant, prev, sample_count, sample_total, allele_count, allele_total, freq |
Pipeline information
Nextflow and the pipeline write reports about the run itself. Use them to troubleshoot errors, to see run times and resource usage, and to record the parameters and software versions used, for example when publishing results.
Output files
pipeline_info/- Reports generated by Nextflow:
execution_report.html,execution_timeline.html,execution_trace.txtandpipeline_dag.html. - Reports generated by the pipeline:
pipeline_report.html,pipeline_report.txtandnf_core_plasmodiumdrugres_software_versions.yml. Thepipeline_report*files are only present if--email/--email_on_failis set. - Parameters used by the pipeline run:
params.json.
- Reports generated by Nextflow:
See the Nextflow reports documentation for more on the execution reports.