qcpipe_integrated.py is an integrated quality control (QC) pipeline for targeted next-generation sequencing (tNGS) data. It performs end-to-end QC across FASTQ reads, BAM alignments, enrichment regions, insert size distributions, taxonomic classification, and optional Castanet outputs.
It generates a single per-sample QC report (*_qc.csv) and optional diagnostic plots.
- FASTQ read length QC (raw + trimmed)
- Alignment QC using
samtools - Insert size distribution analysis from BAM files
- Enrichment analysis using BED-defined regions, usually in mitochondrial regions
- Kraken2 taxonomic summarisation
- Castanet BAM + depth integration
- Insert size visualisation plots
bwa-mem2samtools($\ge$ v1.10)seqkit- ** Python 3.x Environment**: Requires
pandas,numpy, andmatplotlib.
python qcpipe_integrated.py --sample <SAMPLE_ID> --batch <BATCH_ID> --outdir <OUTPUT_DIR> [OPTIONS]| Flag | Long Flag | Requirement | Description |
|---|---|---|---|
-s |
--sample |
Required | Unique identifier name string for the processing sample. |
-b |
--batch |
Required | Running batch code identity identifier. |
-o |
--outdir |
Required | Targeted target directory route path to dump generated files. |
-r |
--raw |
Optional | Path to raw forward ( |
-t |
--trimmed |
Optional | Path to trimmed forward ( |
-k |
--kraken |
Optional | Input data file path pointing to a standard kraken2 report. |
-m |
--mttarget |
Optional | Fasta file of enrichment target used in panel (select from enrichment_targets folder). (triggers enrichment analysis) |
--bedfile |
--bedfile |
Optional | BED file corresponding to --mttarget fasta file that contains enriched regions (select from enrichment_targets folder) |
--castanetbam |
--castanetbam |
Optional | Path to castanet BAM file. |
--castanetdepth |
--castanetdepth |
Optional | Path to castanet _depth.csv file. |
--humantargets |
Optional | FASTA reference stem/file for human targets included in the capture panel. Used to run the integrated Castanet human-target analysis. | |
--humanmappref |
Optional | Castanet mapping reference table for the human targets. Required with --humantargets. |
|
--threads |
Optional | Number of threads to use for the integrated Castanet human-target analysis. Defaults to 1. |
|
--keeptmp |
Optional Flag | Prevents cleanup deletion routines targeting intermediate alignments. |
Evaluates basic spatial track sizes across processed paired data files without parsing deeper coordinate indexes.
python qcpipe_integrated.py \
--sample SRR1234567 \
--batch Run_2026_06_A \
--trimmed data/trimmed_R1.fastq data/trimmed_R2.fastq \
--outdir ./qc_resultsA complete pipeline run involving target alignments to an engine index (e.g., mitochondrial genome), sorting coordinate enrichment via a structural BED template, parsing taxonomy distributions, and extracting deduplication metadata tables.
python qcpipe_integrated.py \
--sample Patient_Sample_04 \
--batch Clinical_Batch_42 \
--raw data/raw_R1.fastq.gz data/raw_R2.fastq.gz \
--trimmed data/trimmed_R1.fastq.gz data/trimmed_R2.fastq.gz \
--mttarget enrichment_targets/HumanMt.fasta \
--bedfile enrichment_targets/HumanMt_regions.bed \
--kraken kraken2_report.txt \
--castanetbam sample.bam \
--castanetdepth sample_depth.csv \
--humantargets human_targets/reference_stem \
--humanmappref human_targets/mapping_reference.csv \
--threads 8 \
--outdir ./qc_reportsThe script generates a single comma-delimited output file ([sample_id]_qc.csv) containing headers mapped dynamically based on input flags. The full data schema is broken down below:
| CSV Header Block | Data Type | Description |
|---|---|---|
batch |
String | User-specified batch run identifier. |
sampleid |
String | Extracted or provided specimen id |
R1 / R2
|
String | path locations for forward and reverse raw reads. |
trimmedR1 / trimmedR2
|
String | path locations for filtered forward and reverse trimmed reads. |
rawreads |
Integer | Total absolute read count combined across both raw paired-end fastq inputs. |
trimmedreads |
Integer | Total absolute read count combined across both processed paired-end fastq inputs. |
avglen / medlen
|
Float | mean and median read lengths for all trimmed reads. |
r1count / r2count
|
Integer | Individual absolute processed read counts split across |
r1avglen / r1medlen
|
Float | average and median length for |
r2avglen / r2medlen
|
Float | average and median length for |
Generated when supplying --mttarget and --bedfile arguments.
| CSV Header Block | Data Type | Description |
|---|---|---|
enrichmentloci_avginsert |
Float | Mean fragment size from enrichment target mapped reads. |
enrichmentloci_stdinsert |
Float | Standard deviation fragment size from enrichment target mapped reads. |
enrichmentloci_insert25 |
Float | 25th Percentile fragment size from enrichment target mapped reads. |
enrichmentloci_insert50 |
Float | Median fragment size from enrichment target mapped reads. |
enrichmentloci_insert75 |
Float | 75th Percentile fragment size from enrichment target mapped reads. |
enrichedMedian |
Float | Computed median mapping depth across coordinates tagged as enriched. |
unenrichedMedian |
Float | Computed median mapping depth across coordinates tagged as unenriched. |
enrichmentRatio |
Float | Enrichment ratio |
enrichmentloci |
String | enrichment locus used (selected by highest read depth). |
Generated when supplying a --kraken classification file.
| CSV Header Block | Data Type | Description |
|---|---|---|
kraken:Eukaryota |
Integer | Read count for Eukaryota from krakenreport. |
kraken:Bacteria |
Integer | Read count for Bacteria from krakenreport. |
kraken:Archaea |
Integer | Read count for Archaea from krakenreport. |
kraken:Viruses |
Integer | Read count for Viruses from krakenreport. |
kraken:Fungi |
Integer | Read count for Fungi from krakenreport. |
kraken:Caudoviricetes |
Integer | Read count for Caudoviricetes (phages) from krakenreport. |
kraken:Homo sapiens |
Integer | Read count for Homo sapiens from krakenreport |
Generated when supplying --castanetbam and --castanetdepth metrics.
| CSV Header Block | Data Type | Description |
|---|---|---|
all_mapped_avginsert |
Float | Raw mean insert size from castanet bam mapped reads. |
all_mapped_stdinsert |
Float | Raw standard deviation insert size from castanet bam mapped reads. |
all_mapped_insert[25/50/75] |
Float | Raw percentile metrics ( |
filtered_mapped_avginsert |
Float | High-confidence mean insert length (Filtered to MapQ |
filtered_mapped_stdinsert |
Float | High-confidence standard deviation of insert size. |
filtered_mapped_insert[25/50/75] |
Float | Percentile metrics ( |
castanet_total_mapped_reads |
Integer | Sum of all reads mapped to all targets from castanet depths file. |
castanet_dedup_reads |
Integer | Sum of all deduplicated reads mapped to all targets from castanet depths file. |
Generated when supplying both --humantargets and --humanmappref. The script runs Castanet on the trimmed read folder, filters the Castanet depth table to rows whose probetype contains human, and emits columns dynamically from the probe names present in that file.
| CSV Header Pattern | Data Type | Description |
|---|---|---|
<probetype>_all |
Integer/Float | Total mapped reads for each human probetype. Uses n_reads_all when present in the Castanet depth CSV, otherwise reads_for_mapping. |
allhuman_all |
Integer/Float | Sum of the mapped-read counts across all human probe types. |
<probetype>_dedup |
Integer/Float | Deduplicated mapped reads (n_reads_dedup) for each human probetype. |
allhuman_dedup |
Integer/Float | Sum of deduplicated mapped reads across all human probe types. |
<probetype>_nc2 |
Float | prop_npos_cov2 value for each human probetype, representing the proportion of target positions with coverage of at least 2. |