nf-core/epigenomesegmentation
An nf-core pipeline for epigenome segmentation using EpiSegMix/Meth — a hidden Markov model with flexible read count distributions and state duration modeling for histone, open chromatin, and methylation signals.
Introduction
This document describes the output produced by the pipeline.
The directories listed below will be created in the results directory after the pipeline has finished. All paths are relative to the top-level results directory.
Pipeline overview
The pipeline is built using Nextflow and processes data using the following steps:
- References - Genomic bins and chromosome sizes used for analysis.
- IndexFiles - Processed and indexed BAM files.
- Counts - Binned count matrices for Histones and Methylation.
- EpiSegMix - Trained models, segmentation BED files, and diagnostic plots.
- Pipeline information - Report metrics generated during the workflow execution.
References
Output files
References/[Genome]/*_bins.bed: Genomic windows (e.g., 200bp) used for signal aggregation.*.chrom.sizes: The chromosome sizes file fetched for the reference genome.
This directory contains the structural files generated during Genome Preparation. These files ensure that all downstream counting and modeling are performed on a consistent genomic coordinate system.
IndexFiles
Output files
IndexFiles/[SampleID]/*.nochr.bam: Filtered BAM files used for the counting process.*.nochr.bam.bai: Coordinate-sorted index files for the BAMs.
Processed alignment files that have been filtered (e.g., “nochr” suffix) and indexed to allow for efficient count matrix generation.
Counts
Output files
Counts/[SampleID]_Histone/: Histone count matrices generated from BAM input (*.tab).[SampleID]_Methylation/: Methylation/coverage count files generated from BED input (*.taband intermediate binned BED files).
EpiSegMix
This is the core results directory, containing the output of the segmentation modeling. Files are organized by sample and state number (e.g., _10).
1. Segmentation
Output files
EpiSegMix/[SampleID]/Segmentation/*.bed.gz: Compressed BED file containing the genomic coordinates and assigned chromatin states.*.tab: A tab-delimited text version of the segmentation results.
2. Models
Output files
EpiSegMix/[SampleID]/Model/final-model-*.json: The trained HMM parameters.*.yaml: The configuration used for the modeling run.*.log: Log files tracking the training and decoding steps.*-train-counts.txt: The data matrix used during the training phase.
3. Plots
Output files
EpiSegMix/[ModelID]/Plots/<model-id>-correlation.png: Model-level correlation matrix of input marks (generated once per model, not once per sample).<model-id>-histogram.png: Model-level signal distributions (generated once per model).<model-id>-methylation-density.png: Model-level methylation density plot; LDM may also publish a per-sample copy.<sample-id>-meanEmission*.png,<sample-id>-normEmission*.png,<sample-id>-transitionMatrix.png: Per-sample emission and transition plots. DM filenames include-viterbifor some plots.<sample-id>-stateDistribution.png,<sample-id>-stateLength*.png,<sample-id>-stateMembership*.png,<sample-id>-state-colors.png: Per-sample state plots; exact available plots depend on the selected model.<sample-id>_report.md: Per-sample Markdown report linking to that sample’s plots and the shared model-level plots. These are Markdown files, not interactive HTML reports.
Pipeline information
Output files
- Reports generated by Nextflow: `execution_report.html`, `execution_timeline.html`, `execution_trace.txt` and `pipeline_dag.dot`/`pipeline_dag.svg`. - Reports generated by the pipeline: `pipeline_report.html`, `pipeline_report.txt` and `software_versions.yml`. The `pipeline_report*` files will only be present if the `--email` / `--email_on_fail` parameter's are used when running the pipeline. - Reformatted samplesheet files used as input to the pipeline: `samplesheet.valid.csv`. - Parameters used by the pipeline run: `params.json`.Nextflow provides excellent functionality for generating various reports relevant to the running and execution of the pipeline. This will allow you to troubleshoot errors with the running of the pipeline, and also provide you with other information such as launch commands, run times and resource usage.