Introduction

This document describes the output produced by the pipeline.

The directories listed below will be created in the results directory after the pipeline has finished. All paths are relative to the top-level results directory.

Pipeline overview

The pipeline is built using Nextflow and processes data using the following steps:

  • References - Genomic bins and chromosome sizes used for analysis.
  • IndexFiles - Processed and indexed BAM files.
  • Counts - Binned count matrices for Histones and Methylation.
  • EpiSegMix - Trained models, segmentation BED files, and diagnostic plots.
  • Pipeline information - Report metrics generated during the workflow execution.

References

Output files
  • References/
    • [Genome]/
      • *_bins.bed: Genomic windows (e.g., 200bp) used for signal aggregation.
      • *.chrom.sizes: The chromosome sizes file fetched for the reference genome.

This directory contains the structural files generated during Genome Preparation. These files ensure that all downstream counting and modeling are performed on a consistent genomic coordinate system.

IndexFiles

Output files
  • IndexFiles/
    • [SampleID]/
      • *.nochr.bam: Filtered BAM files used for the counting process.
      • *.nochr.bam.bai: Coordinate-sorted index files for the BAMs.

Processed alignment files that have been filtered (e.g., “nochr” suffix) and indexed to allow for efficient count matrix generation.

Counts

Output files
  • Counts/
    • [SampleID]_Histone/: Histone count matrices generated from BAM input (*.tab).
    • [SampleID]_Methylation/: Methylation/coverage count files generated from BED input (*.tab and intermediate binned BED files).

EpiSegMix

This is the core results directory, containing the output of the segmentation modeling. Files are organized by sample and state number (e.g., _10).

1. Segmentation

Output files
  • EpiSegMix/[SampleID]/Segmentation/
    • *.bed.gz: Compressed BED file containing the genomic coordinates and assigned chromatin states.
    • *.tab: A tab-delimited text version of the segmentation results.

2. Models

Output files
  • EpiSegMix/[SampleID]/Model/
    • final-model-*.json: The trained HMM parameters.
    • *.yaml: The configuration used for the modeling run.
    • *.log: Log files tracking the training and decoding steps.
    • *-train-counts.txt: The data matrix used during the training phase.

3. Plots

Output files
  • EpiSegMix/[ModelID]/Plots/
    • <model-id>-correlation.png: Model-level correlation matrix of input marks (generated once per model, not once per sample).
    • <model-id>-histogram.png: Model-level signal distributions (generated once per model).
    • <model-id>-methylation-density.png: Model-level methylation density plot; LDM may also publish a per-sample copy.
    • <sample-id>-meanEmission*.png, <sample-id>-normEmission*.png, <sample-id>-transitionMatrix.png: Per-sample emission and transition plots. DM filenames include -viterbi for some plots.
    • <sample-id>-stateDistribution.png, <sample-id>-stateLength*.png, <sample-id>-stateMembership*.png, <sample-id>-state-colors.png: Per-sample state plots; exact available plots depend on the selected model.
    • <sample-id>_report.md: Per-sample Markdown report linking to that sample’s plots and the shared model-level plots. These are Markdown files, not interactive HTML reports.

Pipeline information

Output files - Reports generated by Nextflow: `execution_report.html`, `execution_timeline.html`, `execution_trace.txt` and `pipeline_dag.dot`/`pipeline_dag.svg`. - Reports generated by the pipeline: `pipeline_report.html`, `pipeline_report.txt` and `software_versions.yml`. The `pipeline_report*` files will only be present if the `--email` / `--email_on_fail` parameter's are used when running the pipeline. - Reformatted samplesheet files used as input to the pipeline: `samplesheet.valid.csv`. - Parameters used by the pipeline run: `params.json`.

Nextflow provides excellent functionality for generating various reports relevant to the running and execution of the pipeline. This will allow you to troubleshoot errors with the running of the pipeline, and also provide you with other information such as launch commands, run times and resource usage.