Quality Control and Preprocessing of High-Throughput Sequencing Data
High-throughput sequencing (NGS) technologies have revolutionized biological research by enabling the massive parallel sequencing of billions of DNA or RNA fragments. However, the raw data generated by these platforms—typically in FASTQ format—is inherently "noisy." It contains various artifacts, including sequencing adapters, low-quality base calls, technical duplicates, and potential biological contamination.
If these errors are not addressed before downstream analysis, they can lead to catastrophic failures in the pipeline, such as decreased alignment rates, false-positive variant calls, and inaccurate gene expression quantification. Therefore, Quality Control (QC) and preprocessing are not merely optional steps; they are the foundational pillars that determine the reliability and reproducibility of all subsequent biological conclusions.
Initial Data Assessment: Evaluating Raw FASTQ Files
The first step in any NGS pipeline is to perform a comprehensive assessment of the raw reads. This stage aims to identify systematic errors or library preparation issues. Using industry-standard tools like FastQC for individual samples and MultiQC for aggregating reports across large cohorts, researchers should focus on several key metrics:
- Base Quality Scores (Phred Scores): The Phred quality score ($Q$) represents the probability of an incorrect base call. A score of $Q30$ indicates a 1 in 1,000 error rate (99.9% accuracy). While it is normal for quality to decline toward the 3' end of a read due to signal decay, a sudden, sharp drop in quality across the entire read length often signals a failure in the sequencing run.
- Adapter Contamination: If the DNA insert is shorter than the read length, the sequencer will read into the synthetic adapter sequences. Identifying these is crucial, as they will prevent accurate alignment to a reference genome.
- GC Content Distribution: The GC content should ideally follow a normal distribution consistent with the target organism. Significant deviations or multi-modal distributions often suggest contamination from other species or heavy PCR bias during library amplification.
- Duplication Levels: High levels of sequence duplication can indicate low input DNA amounts or over-amplification during PCR, which can skew quantitative results (especially in RNA-seq).
- N-content and Read Length: A high frequency of "N" bases (undetermined nucleotides) or unexpected variations in read length can indicate technical issues during the sequencing reaction.
Core Preprocessing Steps: Cleaning the Data
Once the quality issues are identified, the data must be "cleaned" through a series of computational trimming and filtering steps. The goal is to remove noise while preserving as much high-quality biological information as possible.
- Adapter Trimming: Identifying and removing synthetic adapter sequences and adapter dimers from the ends of the reads.
- Quality Trimming: Utilizing a sliding window approach to trim low-quality bases from the ends of reads. For example, a window might move along the read, and once the average quality falls below a threshold (e.g., $Q20$), the remainder of the read is discarded.
- Read Filtering: Entire reads should be discarded if they fail to meet specific criteria, such as having an excessive proportion of low-quality bases or too many "N" nucleotides.
- Length Filtering: After trimming, some reads may become too short to be uniquely mapped. These are typically removed if they fall below a minimum length threshold (e.g., 50 bp).
- Decontamination: Depending on the study design, it may be necessary to filter out sequences from non-target organisms (e.g., removing human DNA from a metagenomic sample).
Essential Bioinformatics Toolset
The following table summarizes the most widely used tools in the preprocessing landscape:
| Tool | Primary Function | Key Characteristics |
|---|---|---|
| FastQC | Raw data quality assessment | The industry standard for generating comprehensive QC reports. |
| MultiQC | Aggregate QC reporting | Essential for managing large-scale projects by summarizing multiple reports. |
| fastp | All-in-one preprocessing | Extremely fast; performs adapter trimming, quality filtering, and provides before/after reports. |
| Trimmomatic | Trimming and filtering | A highly flexible and classic tool for customized trimming parameters. |
| Cutadapt | Adapter removal | Specialized in identifying and removing specific adapter sequences. |
| Picard | Post-alignment processing | The standard for marking duplicates and calculating alignment statistics. |
Currently, fastp has become a preferred choice for many bioinformaticians due to its high speed and its ability to perform multiple cleaning tasks in a single pass.
Post-Alignment Quality Control
Preprocessing does not end with the FASTQ files. Once reads are aligned to a reference genome (using aligners like BWA-MEM or Bowtie2), a second layer of QC is required to validate the alignment quality:
- Alignment Rate: For Whole Genome Sequencing (WGS), an alignment rate $>95%$ is generally expected. For Whole Exome Sequencing (WES), rates between $70%–90%$ are common. Low rates may indicate contamination or poor library quality.
- Duplication Rate: High duplication rates in the BAM/SAM files suggest PCR artifacts, which should be addressed using tools like Picard to mark and remove duplicates.
- Coverage Depth and Uniformity: It is vital to ensure that the target regions are covered adequately and uniformly. Large "gaps" in coverage can lead to missed variants.
- Insert Size and Strand Bias: Analyzing the distribution of insert sizes and checking for strand bias helps detect errors introduced during library construction or alignment.
Tailoring QC Strategies to Application Scenarios
A "one-size-fits-all" approach to QC is rarely optimal. The focus of preprocessing must shift depending on the biological question:
- Whole Genome Re-sequencing (WGS): The priority is high coverage uniformity and low duplication rates to ensure accurate variant calling.
- Exome/Targeted Sequencing (WES): The focus is on capture efficiency—ensuring that the enrichment process successfully pulled down the intended genomic regions.
- Transcriptomics (RNA-seq): One must monitor rRNA contamination and 5'/3' end biases. Crucially, researchers should avoid aggressive trimming, as over-trimming can interfere with accurate gene quantification.
- Metagenomics: The emphasis is on host depletion (removing host DNA) and identifying potential cross-species contamination.
Conclusion and Best Practices
To ensure a robust and reproducible analysis, a standardized workflow is recommended:
FastQC (Initial Check) $\rightarrow$ fastp (Cleaning) $\rightarrow$ FastQC/MultiQC (Post-cleaning Check) $\rightarrow$ Alignment $\rightarrow$ Picard/samtools (Post-alignment QC).
A professional bioinformatician should always document the specific parameters used (e.g., minimum length, Phred thresholds) and maintain version control over the software used. While quality control does not directly yield biological insights, it is the bedrock of the entire analytical chain. Investing time in meticulous preprocessing is the most effective way to prevent costly errors and unreliable results in downstream biological interpretation.