Genetic Analysis of Repeated Measures and Longitudinal Data

In quantitative genetics and breeding programs, many economically and biologically significant traits are not adequately captured by a single measurement. Milk yield in dairy cattle is recorded across lactation months, body weight in pigs is tracked by week, tree diameter is assessed annually, and human height or blood pressure is monitored through longitudinal cohorts. These datasets are broadly categorized as repeated measures or longitudinal data. Unlike cross-sectional snapshots, they capture the dynamic trajectory of individuals over time, enabling the disentanglement of additive genetic effects, permanent environmental effects, and temporary environmental residuals. This decomposition significantly enhances the precision of genetic parameter estimation and the accuracy of breeding value predictions.

Positioned within the broader domain of population and quantitative genetics, this article focuses on the universal principles, model comparisons, and application landscapes of genetic analysis for repeated measures and longitudinal data. Foundational concepts such as Hardy-Weinberg equilibrium, quantitative trait definitions, and narrow-sense heritability are covered in dedicated discussions and are only briefly referenced here where necessary.
Repeated measures and longitudinal data possess several defining features that distinguish them from single-record datasets:

  • Within-individual correlation: Multiple observations taken on the same individual are inherently more similar than observations taken between different individuals.
  • Temporal ordering: Longitudinal data possess an inherent time sequence, and the intervals between measurements may be unequal.
  • Data imbalance: Individuals frequently differ in their total number of records and the specific time points at which they are measured.
  • Multilevel hierarchy: Observations are nested within individuals, and individuals are subsequently nested within families or broader populations.

The core objectives of genetic analysis for such data structures include:

  1. Estimating variance components: Specifically partitioning phenotypic variance into additive genetic variance, permanent environmental variance, and temporary environmental (residual) variance.
  2. Calculating repeatability and heritability: Repeatability establishes the reliability of a single record as an indicator of an individual's permanent merit, while heritability quantifies the proportion of total variance attributable to additive genetic effects.
  3. Predicting breeding values: Estimating an individual's genetic merit and tracking how this merit shifts across its lifespan.
  4. Estimating genetic correlations across time: Revealing the developmental genetic architecture by determining whether the same genes or different genes influence a trait at different life stages.

The Universal Framework: Mixed Linear Models

The genetic analysis of repeated measures and longitudinal data is fundamentally anchored in the mixed linear model (MLM) framework. The general matrix notation is expressed as:

[
\mathbf{y} = \mathbf{X}\mathbf{b} + \mathbf{Z}\mathbf{u} + \mathbf{W}\mathbf{p} + \mathbf{e}
]

Where:

  • (\mathbf{y}) is the vector of phenotypic observations;
  • (\mathbf{b}) represents fixed effects (e.g., herd, year, season, sex, parity);
  • (\mathbf{u}) is the vector of random additive genetic effects, typically assumed to follow (\mathbf{u} \sim N(\mathbf{0}, \mathbf{A}\sigma_a^2)), where (\mathbf{A}) is the numerator relationship matrix;
  • (\mathbf{p}) captures permanent environmental effects, accounting for consistent, non-genetic influences specific to the individual;
  • (\mathbf{e}) is the vector of temporary environmental effects or residuals;
  • (\mathbf{X}), (\mathbf{Z}), and (\mathbf{W}) are the corresponding incidence design matrices.

Depending on the assumptions made regarding how genetic and permanent environmental effects behave over time, this foundational model diverges into two primary approaches: the repeatability model and the random regression model.

Repeatability Model vs. Random Regression Model

The Repeatability Model

The repeatability model operates under the strict assumption that an individual's additive genetic effect and permanent environmental effect remain constant across all time points. Its typical formulation is:

[
y_{ij} = \mu + \text{fixed effects} + a_i + p_i + e_{ij}
]

Here, (a_i) is the additive genetic effect for individual (i), (p_i) is the permanent environmental effect, and (e_{ij}) is the temporary environmental effect for the (j)-th measurement.

  • Strengths: The model is highly parsimonious, computationally efficient, and straightforward to implement. It is ideal for scenarios with few time points or traits that are biologically stable over time.
  • Limitations: It cannot capture temporal changes in genetic effects, nor can it estimate genetic correlations between different time points.
  • Typical Applications: Multiple lactation milk yield records in dairy cows, repeated litter size records in sows, and individual plant yield in perennial crops.

The Random Regression Model

Random regression models (RRM) allow both genetic and permanent environmental effects to vary as a function of a continuous covariate, such as time, age, or an environmental gradient. The general form is:

[
y_{ij} = \text{fixed effects} + \sum_{k=0}^{m} \phi_k(t_{ij}) a_{ik} + \sum_{k=0}^{m} \phi_k(t_{ij}) p_{ik} + e_{ij}
]

Where (\phi_k(t_{ij})) represents a basis function evaluated at time (t_{ij}) (commonly Legendre polynomials, B-splines, or linear splines), and (a_{ik}) and (p_{ik}) are the random regression coefficients for individual (i).

  • Strengths: RRM can estimate heritability as a continuous function of time and compute genetic correlations between any two points in time. It is exceptionally well-suited for modeling dynamic traits like growth or lactation curves.
  • Limitations: The model is parameter-heavy, demanding substantial data and computational resources. Selecting the optimal order of the basis functions requires care to avoid overparameterization.
  • Typical Applications: Growth curves for body weight in livestock, lactation trajectory modeling in dairy cattle, and annual growth increments in forestry.

Key Comparative Insights

  • Assumption divergence: The repeatability model assumes genetic effects are time-invariant; RRM assumes they are dynamic.
  • Output divergence: The repeatability model yields a single, static heritability and repeatability estimate; RRM generates heritability curves and comprehensive genetic correlation matrices.
  • Data requirements: RRM typically requires a robust number of records per individual distributed across a wide temporal span to stabilize variance component estimates.
  • Computational burden: RRM is vastly more computationally intensive than the repeatability model, escalating rapidly as the dimensionality of random effects increases.

Application Landscape

The genetic analysis of repeated measures and longitudinal data plays a pivotal role across multiple biological disciplines:

  • Animal breeding: Evaluating repeated milk yield tests, weekly growth velocity in pigs, and continuous egg production in poultry. RRM enables the formulation of time-specific selection indices, improving genetic gain.
  • Plant breeding: In multi-year, multi-location trials, the same genotype is evaluated across varying seasons and sites. Repeatability models assess varietal stability, while RRM can dissect genotype-by-environment interactions as a function of time.
  • Human genetics: Longitudinal cohort studies tracking height, body mass index, or blood pressure rely on these models to quantify how genetic control dictates developmental trajectories and aging processes.
  • Aquaculture and forestry: Repeated body weight measurements in fish and annual diameter-at-breast-height records in trees help optimize the age of selection and refine breeding strategies for long-lived species.

Practical Guidelines and Considerations

When implementing these models in real-world analyses, several critical factors warrant close attention:

  1. Define the time covariate clearly: Select a biologically meaningful time scale, such as days on feed, days in milk, or an environmental index, to ensure proper model convergence and interpretability.
  2. Model selection: Employ statistical criteria such as likelihood ratio tests (LRT), AIC, or BIC to objectively compare the repeatability model against RRM with varying polynomial orders.
  3. Data hygiene: Diligently screen for outliers, handle missing data appropriately, and exclude extreme time points that fall outside the bulk of the data, as they can severely bias variance component estimates.
  4. Computational tools: Software selection should align with model complexity and data scale. Industry standards include ASReml, BLUPF90, and R packages such as sommer and MCMCglmm.
  5. Result validation: Employ cross-validation or assess predictive accuracy on independent datasets to confirm that the fitted model generalizes well beyond the training data.

Conclusion

The genetic analysis of repeated measures and longitudinal data serves as a vital conduit between quantitative genetic theory and modern breeding practice. Mastering the mixed linear model framework, understanding the comparative strengths of repeatability and random regression models, and recognizing their broad application potential are foundational steps for any quantitative geneticist. These tools not only refine the estimation of core parameters like heritability and breeding values but also illuminate the dynamic nature of gene action over time. As analytical capabilities and computational power continue to expand, the adoption of longitudinal genetic models will remain essential for driving genetic progress across animal' animal, plant, and human genetics.