Reports & Summaries
- Overview
- Sample Reads, Assemblies, & Taxonomy
- Cluster Assignments
- Core Genome Alignments
- Recombination
- SNP Variants
- Quality Control
- Distances
- Summary
- Phylogeny
- Reports
This page describes results saved to --outdir. Files written to the BigBacter database (--db) are described on the BigBacter Database page.
Overview
All outputs intended for routine interpretation and reporting are organized under a run-specific subdirectory named by Unix timestamp (${timestamp}). This allows results from multiple runs to be saved to a common --outdir without overwriting previous results and also provides baked-in tracibility π₯.
Below is an overview of the standard outputs produced by BigBacter.
${outdir}/
βββ ${timestamp}
β βββ ${sample}
β β βββ asm
β β β βββ ${sample}.fa.gz
β β βββ taxa
β β β βββ ${sample}_gambit.csv
β β βββ reads
β β βββ ${sra}_1.fastq.gz
β β βββ ${sra}_2.fastq.gz
β βββ ${taxa}
β β βββ ${cluster}
β β β βββ aln
β β β β βββ ${timestamp}-${taxa}-${cluster}.aln
β β β β βββ ${timestamp}-${taxa}-${cluster}_full.aln
β β β β βββ ${timestamp}-${taxa}-${cluster}_full.csv
β β β β βββ ${timestamp}-${taxa}-${cluster}.masked.aln
β β β β βββ ${timestamp}-${taxa}-${cluster}_full.masked.aln
β β β β βββ ${timestamp}-${taxa}-${cluster}_full.masked.csv
β β β βββ dist
β β β β βββ ${timestamp}-${taxa}-${cluster}_db-dist.csv
β β β β βββ ${timestamp}-${taxa}-${cluster}_snp-dist.csv
β β β β βββ ${timestamp}-${taxa}-${cluster}_db-dist.masked.csv
β β β β βββ ${timestamp}-${taxa}-${cluster}_snp-dist.masked.csv
β β β βββ qc
β β β β βββ ${timestamp}-${taxa}-${cluster}_cg-plot.html
β β β β βββ ${timestamp}-${taxa}-${cluster}_cg-plot.masked.html
β β β βββ recomb
β β β β βββ ${timestamp}-${taxa}-${cluster}.vcf
β β β β βββ ${timestamp}-${taxa}-${cluster}.bed
β β β β βββ ${timestamp}-${taxa}-${cluster}.per_branch_statistics.csv
β β β βββ report
β β β β βββ ${timestamp}-${taxa}-${cluster}.microreact
β β β β βββ ${timestamp}-${taxa}-${cluster}.masked.microreact
β β β βββ summary
β β β β βββ ${timestamp}-${taxa}-${cluster}_summary.csv
β β β β βββ ${timestamp}-${taxa}-${cluster}_summary.masked.csv
β β β βββ tree
β β β β βββ ${timestamp}-${taxa}-${cluster}.nwk
β β β β βββ ${timestamp}-${taxa}-${cluster}.masked.nwk
β β β βββ var
β β β βββ ${sample}.tar.gz
β β βββ cluster
β β βββ clusters.csv
β β βββ global_containment.csv
β βββ multiqc
β β βββ multiqc_report.html
β βββ ${timestamp}.csv
β βββ ${timestamp}.masked.csv
βββ pipeline_info
βββ ...
Sample Reads, Assemblies, & Taxonomy
Raw reads from NCBI SRA, de novo genome assemblies from GenBank or created via Shovill, and taxonomic classifications determined via GAMBIT are published in sample-specific subdirectories.
β βββ ${sample}
β β βββ asm
β β β βββ ${sample}.fa.gz
β β βββ taxa
β β β βββ ${sample}_gambit.csv
β β βββ reads
β β βββ ${sra}_1.fastq.gz
β β βββ ${sra}_2.fastq.gz
| File | Description |
|---|---|
${sample}.fa.gz | Compressed genome assembly in FASTA format |
${sample}_gambit.csv | GAMBIT taxonomic classification results |
${sra}_1.fastq.gz | Forward reads downloaded from NCBI SRA |
${sra}_2.fastq.gz | Reverse reads downloaded from NCBI SRA |
Cluster Assignments
Samples are assigned to clusters within each taxon using MinHash-based dissimilarity. Cluster assignments and global containment scores are published at the taxon level.
β βββ cluster
β βββ clusters.csv
β βββ global_containment.csv
| File | Description |
|---|---|
clusters.csv | Per-sample cluster assignments for the current run |
global_containment.csv | Per sample MinHash global containment scores of nearest matching cluster |
Core Genome Alignments
Core genome SNP alignments produced by Polycore are published per cluster. Files with .masked in the filename are produced using the recombination-masked alignment from Gubbins.
β βββ aln
β βββ ${timestamp}-${taxa}-${cluster}.aln
β βββ ${timestamp}-${taxa}-${cluster}_full.aln
β βββ ${timestamp}-${taxa}-${cluster}_full.csv
β βββ ${timestamp}-${taxa}-${cluster}.masked.aln
β βββ ${timestamp}-${taxa}-${cluster}_full.masked.aln
β βββ ${timestamp}-${taxa}-${cluster}_full.masked.csv
| File | Description |
|---|---|
*.aln | Core genome SNP alignment in FASTA format |
*_full.aln | Full core genome alignment including invariant sites |
*_full.csv | Per-site summary of the full alignment |
Recombination
Recombinant regions identified by Gubbins are published per cluster. Only produced when recombination masking is enabled and the cluster has sufficient samples.
β βββ recomb
β βββ ${timestamp}-${taxa}-${cluster}.vcf
β βββ ${timestamp}-${taxa}-${cluster}.bed
β βββ ${timestamp}-${taxa}-${cluster}.per_branch_statistics.csv
| File | Description |
|---|---|
*.vcf | Recombinant SNPs identified by Gubbins in VCF format |
*.bed | Recombinant regions in BED format |
*.per_branch_statistics.csv | Per-branch recombination statistics |
SNP Variants
Per-sample Snippy output tarballs are published per cluster. These files are also published to the BigBacter database (--db) when using --push true.
β βββ var
β βββ ${sample}.tar.gz
| File | Description |
|---|---|
${sample}.tar.gz | Compressed Snippy output directory for each sample |
Quality Control
A per-cluster core genome plot is produced by Polycore and a MultiQC report aggregating per-sample QC metrics is produced at the run level. Files with .masked in the filename are produced using the recombination-masked alignment from Gubbins.
β βββ qc
β β βββ ${timestamp}-${taxa}-${cluster}_cg-plot.html
β β βββ ${timestamp}-${taxa}-${cluster}_cg-plot.masked.html
βββ multiqc
βββ multiqc_report.html
| File | Description |
|---|---|
*_cg-plot.html | Interactive plot of core genome size and SNP density per cluster |
multiqc_report.html | Aggregated QC report including FastQC and fastp metrics for all samples |
Distances
Pairwise MinHash and core SNP distance matrices are published per cluster. Files with .masked in the filename are produced using the recombination-masked alignment from Gubbins.
β βββ dist
β βββ ${timestamp}-${taxa}-${cluster}_db-dist.csv
β βββ ${timestamp}-${taxa}-${cluster}_snp-dist.csv
β βββ ${timestamp}-${taxa}-${cluster}_db-dist.masked.csv
β βββ ${timestamp}-${taxa}-${cluster}_snp-dist.masked.csv
| File | Description |
|---|---|
*_db-dist.csv | Pairwise MinHash distances computed by Floc during clustering |
*_snp-dist.csv | Pairwise core SNP distances |
Summary
Summary tables combining cluster assignments, QC metrics, and SNP distances are produced at two levels: per cluster and per run. Both levels share the same columns (see Summary Columns). Files with .masked in the filename are produced using the recombination-masked alignment from Gubbins.
Cluster Summary
A per-cluster summary table is published within each cluster subdirectory.
β βββ summary
β βββ ${timestamp}-${taxa}-${cluster}_summary.csv
β βββ ${timestamp}-${taxa}-${cluster}_summary.masked.csv
| File | Description |
|---|---|
*_summary.csv | Per-sample summary for all samples in a single cluster, without recombination masked |
*_summary.masked.csv | Per-sample summary for all samples in a single cluster, with recombination masked |
Run Summary
A run-level summary table is published at the top of the run directory. It combines the cluster summaries from every taxon and cluster included in the run.
β βββ ${timestamp}.csv
β βββ ${timestamp}.masked.csv
| File | Description |
|---|---|
${timestamp}.csv | Per-sample summary for all taxa and clusters included in the run, without recombination masked |
${timestamp}.masked.csv | Per-sample summary for all taxa and clusters included in the run, with recombination masked |
Summary Columns
The cluster-level and run-level summary files contain the following columns.
| Column | Description |
|---|---|
id | Sample identifier (same as supplied in samplesheet) |
run | Run timestamp (Unix time) associated with the sample |
status | Whether the sample was added in the current run (new) or already existed in the BigBacter database (old) |
included | Whether the sample was included in the cluster analysis (TRUE/FALSE). Samples are excluded if their genome fraction falls below min_genome_fraction. |
taxa | Taxon assigned to the sample (from samplesheet or GAMBIT) |
cluster | Cluster assigned to the sample within its taxon (from samplesheet or floc) |
strong_links | Samples genetically linked to this sample within the cluster, listed as colon-separated sample pairs. Linkages based on strong_link_threshold. |
inter_links | Samples linked to this sample from other clusters, listed as colon-separated sample pairs. Linkages based on inter_link_threshold |
genome_fraction | Fraction of the reference genome length with called bases, calculated as (length β missing) / length |
core_fraction | Fraction of the core genome that remains after this sample is added to the analysis. This value is used to create the βprogressive core genomeβ plot |
length | Length of the reference genome in base pairs. Will be the same for all samples in a cluster. |
masked | Number of sites masked in the sample |
missing | Number of sites with no base call (e.g., insufficient coverage) |
mixed | Number of sites with mixed or heterozygous base calls |
variants | Number of variant sites relative to the reference |
recomb_masked | Whether recombination masking with Gubbins was applied (TRUE/FALSE) |
partition | Partition assigned to the sample within the cluster. NOTE: Unlike clusters, which are stable between runs, partitions are subject to change depending on which samples are included in the analysis! |
Phylogeny
A maximum likelihood phylogenetic tree is produced per cluster for clusters with sufficient samples. Files with .masked in the filename are produced using the recombination-masked alignment from Gubbins.
β βββ tree
β βββ ${timestamp}-${taxa}-${cluster}.nwk
β βββ ${timestamp}-${taxa}-${cluster}.masked.nwk
| File | Description |
|---|---|
*.nwk | Maximum likelihood phylogenetic tree in Newick format produced by IQ-TREE |
Reports
A Microreact report is produced per cluster, combining the phylogenetic tree, Floc and SNP distance matrices, per-sample summary, and core genome plot. When recombination masking is enabled, two report are produced β one using the standard outputs and one using the masked outputs.
β βββ report
β βββ ${timestamp}-${taxa}-${cluster}.microreact
β βββ ${timestamp}-${taxa}-${cluster}.masked.microreact
| File | Description |
|---|---|
*.microreact | Microreact project file for interactive visualization at microreact.org |