Statistical Assumptions of Quantitative Genetics

In the field of quantitative genetics and plant/animal breeding, statistical models serve as the essential bridge connecting observable phenotypes to the underlying genetic architecture. However, these models are not mere mathematical abstractions; they are built upon a foundation of specific statistical assumptions. The validity of any inference—whether it concerns heritability, breeding values, or genetic correlations—depends entirely on how well the biological reality aligns with these theoretical premises.

Understanding these assumptions is not just a mathematical exercise; it is a prerequisite for avoiding erroneous breeding decisions that can lead to significant economic losses. This article explores the core statistical assumptions of quantitative genetics and discusses their implications for modern breeding practices.
The cornerstone of quantitative genetics is the decomposition of the phenotype ($P$) into its constituent parts: the genotype ($G$) and the environment ($E$), expressed by the linear model:
$$P = G + E$$

To make this model computationally tractable and biologically interpretable, several critical assumptions are made regarding the nature of $G$:

  • Dominance of Additive Effects: The model primarily assumes that the genetic value is the sum of individual additive effects. In this context, the "breeding value" represents the average effect of substituting one allele for another, independent of the other alleles present at that locus or elsewhere in the genome. This assumption is the bedrock of selection theory, as additive variance is the component that is predictably transmitted from parents to offspring.
  • Residualization of Non-Additive Effects: While biological systems are replete with dominance (interactions between alleles at the same locus) and epistasis (interactions between alleles at different loci), standard linear models often treat these non-additive components as secondary. They are frequently subsumed into the random residual error term. While this simplification ensures the solvability of the model, it carries a risk: if epistasis or dominance plays a major role in a specific trait, a purely additive model will provide a biased estimate of genetic variance and breeding values.

Assumptions of Population Genetic Structure

The accuracy of parameter estimation, such as variance components and heritability, is deeply contingent upon the genetic state of the population being studied.

  • Random Mating and Absence of Systematic Selection: Most classical models assume that the reference population is in a state of random mating (panmixia) and is not undergoing intense, systematic selection. This assumption ensures that allele frequencies remain stable enough to allow for the extrapolation of genetic parameters from the current generation to the next. If a population is under heavy directional selection, the non-random association of alleles can distort the relationship between genotype and phenotype, leading to biased estimates of additive variance.
  • Alignment with Hardy-Weinberg Equilibrium (HWE): While HWE is a concept rooted in population genetics, its implications are vital for quantitative models. For statistical derivations of variance components to remain unbiased, the genotype frequencies in the population should ideally conform to HWE. Deviations from equilibrium—caused by inbreeding, population structure, or selection—can complicate the partitioning of genetic variance.

Normality and Independence of Environmental Effects

In the framework of Analysis of Variance (ANOVA) and Linear Mixed Models (LMM), the distribution of environmental noise is a central concern for statistical inference.

  • The Normality Assumption: It is standard to assume that environmental effects and random residuals follow a normal distribution with a mean of zero ($E \sim N(0, \sigma_e^2)$). This assumption is supported by the Central Limit Theorem, which suggests that when a phenotype is influenced by a vast number of small, independent environmental factors, their cumulative effect tends toward a normal distribution. This normality is crucial for the optimality of estimation methods like Restricted Maximum Likelihood (REML), which are used to provide unbiased estimates of variance components.
  • Independence and the Absence of G×E Interaction: Models typically assume that environmental effects are independent and that there is no Genotype-by-Environment (G×E) interaction. However, in real-world breeding, a genotype may perform exceptionally well in one environment but poorly in another. If significant G×E interaction exists, a single environmental estimate cannot represent the true genetic potential of a line. In such cases, breeders must move beyond simple models toward Multi-Environment Trial (MET) models that explicitly account for these interactions.

The Infinitesimal Model and Genetic Architecture

The way we conceptualize the number and effect size of genes significantly shapes our statistical approach.

  • The Infinitesimal Model: Classical quantitative genetics relies on the assumption that traits are controlled by an infinite number of loci, each having an infinitesimally small effect. This assumption allows for the continuous distribution of phenotypes and justifies the use of normal distribution models. When a trait is actually governed by a few large-effect Quantitative Trait Loci (QTL), the phenotypic distribution may deviate from normality, requiring more complex mixed models that treat major QTLs as fixed effects to prevent them from inflating the estimates of polygenic background variance.
  • Linkage and Independence: Basic models often assume that different genetic loci are in linkage equilibrium, meaning they are inherited independently. In reality, Linkage Disequilibrium (LD)—the non-random association of alleles at different loci—is a pervasive feature of genomes. High levels of LD can make it difficult to partition variance components accurately, as the effects of closely linked genes become statistically inseparable.

Implications for Modern Breeding Practice

The transition from theoretical assumptions to practical application is where the true value of quantitative genetics lies.

  • Reliability of Heritability Estimates: A breeder's ability to predict genetic gain depends on the accuracy of heritability ($h^2$) estimates. Because $h^2$ is a ratio of additive variance to total phenotypic variance, any violation of the assumptions of additivity, normality, or population equilibrium will directly impact the reliability of this ratio.
  • Model Selection and BLUP: The Best Linear Unbiased Prediction (BLUP) method, a staple of modern breeding, is a direct application of these statistical frameworks. When breeders recognize that the "infinitesimal" or "additive-only" assumptions are insufficient, they adapt by incorporating dominance or epistatic terms into the covariance structure, or by using factor analytic models to handle complex environmental correlations.
  • The Evolution into Genomic Selection: With the advent of high-density marker data, the field has shifted from pedigree-based models to Genomic Selection (GS). While the fundamental assumptions of linearity and normality remain, the "infinitesimal model" has evolved into the Genomic Relationship Matrix (GBLUP) approach. Here, the assumption of infinite small-effect genes is replaced by the reality of a finite but very large number of marker effects. This shift allows for much higher precision in predicting breeding values by directly capturing the LD between markers and causal variants.

In conclusion, the statistical assumptions of quantitative genetics provide the logical scaffolding required to extract meaningful biological signals from noisy phenotypic data. For the professional breeder or geneticist, mastery of these assumptions is essential for selecting the appropriate analytical tools, interpreting results with caution, and ultimately making data-driven decisions that advance genetic progress.