JBS Haldane wrote his “Mathematical Theory of Natural and Artificial Selection,” as a series of ten papers between 1924 and 1934. These were published across four journal titles, three of which were published by Cambridge Philosophical Society. Haldane’s work frequently is placed alongside the mathematical population genetics of Sewall Wright and Ronald Aylmer Fisher (Provine 1971), but his work has received relatively little substantial analysis (an exception is Sarkar (1992)). Provine collected Wright’s papers. Bennett collected Fisher’s papers. Dronamraju has published a collected works volume for Haldane, including some of the papers in this series (numbers 1, 4, 5, and 8). Haldane’s mathematical work appears in condensed form in Haldane (1932) The Causes of Evolution.
Haldane, Fisher, and Wright were pursuing significantly different research agendas, something historians of the synthesis period need to investigate further. Some of Haldane’s archives are held at UCL Special Collections and have been digitised by Wellcome Collection.

Papers in the Haldane Series
- Haldane, J. B. S. (1924). A Mathematical Theory of Natural and Artificial Selection. Part I. Transactions of the Cambridge Philosophical Society, 23(2), 19-41. (nb. Don’t be confused by the “II” in the title of this paper. It is the second paper in the journal issue, not the second paper in Haldane’s series.)
- Haldane, J. B. S. (1924). A Mathematical Theory of Natural and Artificial Selection. Part II. The influence of partial self-fertilisation, inbreeding, assortative mating, and selective fertilisation on the composition of Mendelian populations, and on natural selection. Proceedings of the Cambridge Philosophical Society. Biological Sciences, 1(2), 158-163.
- Haldane, J. B. S. (1926). A Mathematical Theory of Natural and Artificial Selection. Part III. Proceedings of the Cambridge Philosophical Society, 23(4), 363-372.
- Haldane, J. B. S. (1927). A Mathematical Theory of Natural and Artificial Selection. Part IV. Proceedings of the Cambridge Philosophical Society, 23(5), 607-615.
- Haldane, J. B. S. (1927). A Mathematical Theory of Natural and Artificial Selection. Part V. Selection and Mutation. Proceedings of the Cambridge Philosophical Society, 23(7), 838-844.
- Haldane, J. B. S. (1930). A Mathematical Theory of Natural and Artificial Selection. Part VI. Isolation. Proceedings of the Cambridge Philosophical Society, 26(2), 220-230.
- Haldane, J. B. S. (1931). A Mathematical Theory of Natural and Artificial Selection. Part VII. Selection Intensity as a Function of Mortality Rate. Proceedings of the Cambridge Philosophical Society, 27(1), 131-136.
- Haldane, J. B. S. (1931). A Mathematical Theory of Natural Selection. Part VIII. Metastable Populations. Proceedings of the Cambridge Philosophical Society, 27(1), 137-142.
- Haldane, J. B. S. (1932). A Mathematical Theory of Natural and Artificial Selection. Part IX. Rapid Selection. Proceedings of the Cambridge Philosophical Society, 28(2), 244-248.
- Haldane, J. B. S. (1934). A Mathematical Theory of Natural and Artificial Selection. Part X. Some Theorems on Artificial Selection. Genetics, 19(5), 412-429.

Analysis
These summary points were created by Google Notebook LLM based only on all ten papers in the series. This Google Notebook resource is shared and can be used for further research.
What are the main claims in this series, considered as a whole?
Considered as a whole, Haldane’s ten-part series, “A Mathematical Theory of Natural and Artificial Selection”(1924–1934), establishes a comprehensive quantitative and mathematical framework for evolutionary biology. By systematically applying mathematical equations to Mendelian inheritance, Haldane’s overarching claim is that Darwinian natural selection is fully compatible with, and mathematically sufficient under, Mendelian genetics to drive evolutionary change at realistic rates.
Across the decade-long series, several core mathematical and biological claims emerge when the papers are considered as a cohesive whole:
1. The Mathematical Sufficiency of Selection
Haldane proves that natural selection does not require massive, unrealistic advantages to drive evolutionary change. He demonstrates that even a very small selection coefficient (where a favoured genotype has a slight viability advantage, such as \(k = 0.001\) or a 1% advantage) is mathematically sufficient to spread a gene through a population in a geologically brief period.
2. The Ineffectiveness of Selection on Rare Recessives
One of Haldane’s most famous mathematical proofs is that natural selection is highly inefficient at eliminating or promoting autosomal recessive traits when they are rare in a population. Because rare recessives are almost exclusively hidden in heterozygous carriers who do not show the trait, selection has virtually nothing to act upon. Conversely, selection is highly rapid and effective on dominant traits and sex-linked traits (since sex-linked recessives are immediately exposed in the hemizygous/heterogametic sex).
3. The Accelerating Power of Non-Random Mating
Haldane demonstrates that inbreeding, self-fertilisation, and assortative mating dramatically accelerate the rate of selection. By increasing homozygosity, these breeding systems force recessive genes out of their “hidden” heterozygous state, allowing natural or artificial selection to act on them directly and eliminate or fix them far more rapidly than under random mating.
4. The Selection-Mutation Balance
Haldane proves that selection does not operate in a vacuum but is in a constant dynamic equilibrium with mutation. He establishes that recurrent mutations can act as a counterweight to selection, preventing natural selection from ever completely weeding out rare, harmful recessive genes. Conversely, when selective advantages or disadvantages are negligible, mutation rates alone determine the course of gene frequency changes.
5. Swamping and the Necessity of Isolation
Through modeling migration, Haldane claims that local adaptation is easily “swamped” by immigrationfrom surrounding populations unless a geographical or reproductive barrier exists. Mathematically, he proves that a newly favoured gene can only successfully establish itself in a limited area if its selection coefficient (\(k\)) exceeds the rate of immigration (\(l\)) from the unimproved population; if \(k < l\), selection is rendered entirely ineffective.
6. Metastability and Multi-Gene Equilibria (Epistasis)
When multiple genes interact to determine fitness, Haldane claims that populations do not simply progress linearly toward a single optimal state. Instead, gene interactions (epistasis) can create multiple stable genetic equilibria (metastability). A population can remain stable at a suboptimal genetic equilibrium until a chance fluctuation, intense selection, or isolation pushes it past an unstable equilibrium “barrier,” triggering rapid evolution to a new, highly stable adaptive peak.
7. The Disconnection Between Competition, Mortality, and Selection Intensity
Haldane debunks the naive Darwinian assumption that intense competition or high mortality rates automatically translate to intense natural selection. He mathematically proves that selection intensity can actually diminish or even reverse at extremely high rates of elimination (mortality), meaning that catastrophic mortality can sometimes hinder, rather than help, adaptive evolution.
8. Overlapping Generations Maintain the Selection Dynamics
Haldane addresses the biological reality that generations in nature are rarely discrete. He proves that when generations overlap (requiring complex integral equations rather than simple finite difference equations), the progress of slow selection remains mathematically very similar to the dynamics of populations with separate generations.
9. Rigorous Mathematical Guidelines for Breeders
In his final paper, Haldane synthesises his population genetics models to provide practical, quantitative theorems for agriculturalists and experimentalists. He provides exact equations showing how different breeding systems (random mating, selfing, and brother-sister inbreeding) affect the speed at which multiple desired genes can be fixed and double recessives can be eliminated.
Summarise the main points of each paper, from part 1 to part 10.
J.B.S. Haldane’s decade-long, ten-part series, “A Mathematical Theory of Natural and Artificial Selection” (1924–1934), systematically built the quantitative foundation of population genetics.
Here is the chronological summary of the main points, core claims, and mathematical focuses of each paper in the series:
Part I: Selection of a Simple Mendelian Character (1924)
- Focus: Haldane establishes the mathematical baseline for natural selection of a single completely dominant autosomal factor (or its absence) under random mating, self-fertilisation, and asexual reproduction. He also introduces equations for sex-linked, bisexual, and unisexual selection.
- Main Claims:
- Selection is highly ineffective on rare autosomal recessives in random-mating populations. Because rare recessives are almost entirely hidden in heterozygous carriers, selection has virtually no phenotypic variation to act upon.
- Selection on sex-linked recessives is much faster than on autosomal recessives because sex-linked recessive genes are immediately exposed to selection in the heterogametic sex.
- Direct selection (such as in asexual reproduction or complete self-fertilisation) is vastly more efficient at altering gene frequencies than selection under random mating.
Part II: Mating Systems and Selection (1924)
- Focus: This paper moves beyond pure random mating to model how other reproductive systems—specifically partial self-fertilisation, partial inbreeding (brother-sister mating), assortative mating, and selective fertilisation (gamete selection)—influence natural selection.
- Main Claims:
- Inbreeding and self-fertilisation greatly accelerate natural selection. By systematically increasing the proportion of homozygotes, these mating systems force recessive genes out of their “hidden” heterozygous state, allowing selection to eliminate or fix them far more rapidly than under random mating.
- Assortative mating (mating based on phenotypic similarity) also increases the rate of selection, though less intensely than close inbreeding.
Part III: Incomplete Dominance and Interacting Factors (1926)
- Focus: Haldane expands his mathematical models to incorporate biological complexities, including incomplete dominance, multiple duplicate/interacting factors (epistasis), genetic linkage, and polyploidy.
- Main Claims:
- Polymorphic Equilibria: When dominance is incomplete and the heterozygote holds a selective advantage over both homozygotes, a stable genetic polymorphism is established and maintained.
- Linkage and Polyploidy: Haldane mathematically proves that genetic linkage has a negligible effect on slow selection unless the linkage is extremely tight. Conversely, polyploidy (such as tetraploidy) significantly dampens the rate of selection, making the elimination of recessives much slower.
Part IV: Overlapping Generations (1926)
- Focus: This paper addresses the ecological reality that generations in nature are rarely discrete. Haldane replaces finite difference equations with integral equations to model selection when birth, death, and age distributions overlap.
- Main Claims:
- Despite the highly complex calculus required to model overlapping generations, Haldane proves that the rate of slow selection remains mathematically very similar to the dynamics observed in populations with separate, discrete generations.
Part V: Selection and Mutation (1927)
- Focus: The dynamic relationship between newly arising mutations and the purging force of natural selection.
- Main Claims:
- Survival of Single Mutations: Haldane calculates that a single advantageous dominant mutation has a finite chance of surviving random genetic drift (approximately \(2k\), where \(k\) is its selective advantage). A rare recessive mutation, however, has a negligible chance of survival unless close inbreeding or self-fertilisation is present to immediately expose it in homozygous form.
- Selection-Mutation Balance: Recurrent mutation acts as a counterweight to natural selection, establishing a dynamic equilibrium that prevents natural selection from ever completely weeding out rare, harmful recessive genes. Mutation rates primarily dictate the path of evolution when selective advantages or disadvantages are negligible.
Part VI: Isolation (1930)
- Focus: How migration and geographical isolation interact with local natural selection. Haldane models a localized population undergoing selection of intensity \(k\) while receiving a fraction of immigrants (\(l\)) from a surrounding, unimproved population. He tests this across ten distinct genetic scenarios.
- Main Claims:
- The “Swamping” Effect: Local adaptation is easily “swamped” by immigration unless a geographic or reproductive barrier exists.
- Selection vs. Migration Threshold: A newly favoured gene can only establish itself in a local area if its selection coefficient (\(k\)) exceeds the rate of immigration (\(l\)). If \(k < l\) (or \(k < 4l\) for recessive genes), selection is rendered entirely ineffective and the new gene is wiped out.
Part VII: Selection Intensity as a Function of Mortality Rate (1930)
- Focus: Testing the classical Darwinian assumption that intense competition and high mortality rates automatically translate to highly intense natural selection.
- Main Claims:
- Diminishing Selection Intensity: By mathematically modeling normally distributed viability, Haldane proves that selection intensity does not increase proportionally with the rate of elimination (mortality).
- Under catastrophic, extreme mortality rates (“cataclasms”), the intensity of selection can actually plateau, diminish, or reverse, meaning that severe environmental crises often hinder rather than help adaptive evolution.
Part VIII: Metastable Populations (1930)
- Focus: Multi-locus systems where genes interact epistatically to determine fitness, meaning that a combination of multiple genes produces a high selective advantage, even if each gene individually is harmful.
- Main Claims:
- Multiple Stable Peaks: Epistatic gene interactions lead to multiple stable genetic equilibria (metastability). A population can remain trapped at a suboptimal genetic equilibrium where all new single mutations are selected against.
- Rupturing Equilibria: The population will remain stable until a severe fluctuation, extreme selection event, or geographic isolation pushes it past an unstable equilibrium “barrier”. This trigger causes a rapid, dramatic evolutionary shift toward a new, much higher stable peak, which Haldane proposes as a primary mechanism for speciation.
Part IX: Rapid Selection (1932)
- Focus: Developing mathematical models for intense (rapid) selection where \(k\) is large, and comparing the evolutionary efficiency of continuous selection vs. periodic, intense selection (e.g., periodic famines or plagues).
- Main Claims:
- Cataclasms vs. Steady Selection: Haldane mathematically proves that selection in favour of dominants is more efficient when concentrated in periodic environmental cataclasms.
- Conversely, selection in favour of recessives is more efficient when selection is steady, constant, and continuous across every generation.
Part X: Some Theorems on Artificial Selection (1934)
- Focus: Formulating practical, quantitative theorems for agricultural breeders and experimentalists aiming to completely eliminate undesired genotypes or fix multiple desired autosomal genes under different mating systems.
- Main Claims:
- Inbreeding for Purity: While random mating is highly ineffective at eliminating recessives or fixing multiple dominants, self-fertilisation is incredibly fast at achieving genetic purity.
- Brother-Sister Mating (Sib-Mating): Brother-sister mating initially slows down the physical appearance of the desired dominant phenotypes compared to random mating, but it substantially increases their genetic purity (homozygosity). For example, in the 10th generation of selecting for 10 dominant genes, inbreeding yields 53% homozygous dominants, compared to only 13% under random mating.
- Irreversible Elimination: Even just two generations of brother-sister mating completely eliminates the risk of producing double recessives in many isolated breeding lines.
What are the most cited publications in these papers when considered as a whole?
When considered as a whole, J.B.S. Haldane’s ten-part series relies on several foundational publications—both biological and mathematical—to construct and validate his population genetics models.
The most frequently and significantly cited works and authors across the series can be categorized as follows:
1. Ronald A. Fisher
Fisher is the most prominent external contemporary cited by Haldane to anchor his genetic models. Haldane repeatedly references two of Fisher’s works:
- “On the dominance ratio” (Proceedings of the Royal Society of Edinburgh, Vol. 42, p. 321, 1922): This paper is cited across multiple parts, including Part III and Part V. Haldane uses it to compare his models of heterozygote selection and dominance with Fisher’s formulations of genetic variance.
- The Genetical Theory of Natural Selection (Oxford University Press, 1930): Cited in Part VIII and Part IX to contrast their views on stable genetic equilibria and the mathematical properties of Malthusian parameters under intense selection.
2. Karl Pearson
To justify his mathematical derivations of selection intensity and normal distributions of phenotypic traits, Haldane heavily cites the pioneering biometrical work of Karl Pearson:
- “Mathematical contributions to the theory of evolution” (Philosophical Transactions of the Royal Society A, Vol. 203, p. 53, 1904): Cited in Part I.
- Tables for Statisticians and Biometricians (1924): This reference is crucial for Part VII and Part I, where Haldane uses Pearson’s tables to calculate the selection intensity \(q\) as a function of the mortality rate under a normal curve.
- Pearson and Lee (1903) (Biometrika, Vol. 2, p. 371): Cited in Part I alongside other biometrical papers.
3. Alfred J. Lotka
Lotka’s demographic formulations of population growth and age structure serve as the mathematical foundation for Haldane’s models of overlapping generations:
- Lotka (1921): “The stability of the normal age distribution” (Proceedings of the National Academy of Sciences, Vol. 7, p. 168): Cited in Part II.
- Dublin and Lotka (1925) (Journal of the American Statistical Association, Vol. 20, p. 306): Cited in Part IV and Part VII to incorporate biological age distributions and finite generation intervals into his equations.
4. Conway Zirkle
Zirkle’s genetic calculations represent the primary focus of extension and simplification in the latter half of Haldane’s series:
- Zirkle (1926): “Some numerical results of selection” (Genetics, Vol. 11, pp. 531–583): Discussed extensively and cited in Part X. Haldane points out that Zirkle’s multi-genic selection models under random mating can be enormously simplified and uses them as a springboard to model selection under inbreeding.
5. Albert F. Blakeslee
Haldane regularly references Blakeslee’s empirical plant genetics experiments, specifically on the jimsonweed (Datura Stramonium), to ground his theoretical equations in physical data:
- Blakeslee, Belling, and Farnham (1920): “Chromosome duplication and Mendelian inheritance in Datura” (Science, Vol. 52, p. 388): Cited in Part III to support mathematical models of tetraploid inheritance.
- Gager and Blakeslee (1927) (Proceedings of the National Academy of Sciences, Vol. 13, p. 75): Cited in Part V regarding mutation rates.
- Buchholz and Blakeslee (1929): “Pollen-tube growth…” (Genetics, Vol. 14, p. 538): Cited in Part VII to demonstrate selection intensity during pollen fertilization.
6. Haldane’s Self-Citations
Haldane frequently links the papers together by citing his own earlier parts of the same series to construct a cumulative mathematical framework:
- Part I (1924) is cited in Part II, Part III, Part IV, Part V, Part VI, Part IX, and Part X.
- Part III (1926) is cited in Part IV and Part X.
- Part IX (1932) is cited in Part X.
Which organisms are most commonly used in these papers when considered as a whole? Present this in rank order, from most frequent.
When J.B.S. Haldane’s ten-part series is considered as a whole, it is clear that while his focus is overwhelmingly mathematical, he constantly anchors his abstract models in empirical genetic data.
Based on the papers in this series, the most commonly used and cited organisms—presented in rank order from most frequent and significant—are:
1. Humans (Man / Homo sapiens)
Humans are the most frequently recurring organism throughout the series, serving as Haldane’s primary model for real-world genetic traits and demographics:
- Sex-Linked Traits: In Part I, Haldane uses human pedigrees of colour-blindness and haemophiliato model sex-linked selection.
- Overlapping Generations: In Part IV, he relies on Alfred Lotka’s demographic tables of American women to mathematically model how selection operates when generations overlap .
- Mutation and Selection: In Part V and VII, he models human mutation rates and uses survival data from civilised human societies to calculate how selection intensity varies under modern medical and social conditions .
- Assortative Mating: In Part II, Haldane reviews the mathematical impact of assortative mating (non-random pairing) based on human physical traits.
2. Mice (Mus musculus)
Mice are Haldane’s primary model for mammalian genetics and the inheritance of complex, multi-genic traits:
- Embryonic Mortality: In Part I, Haldane cites mouse breeding studies showing that selective embryonic death in utero alters Mendelian ratios. He also models the inheritance of the “G factor” (governing light bellies and yellow-tipped hair).
- Multi-Locus Selection: In Part X, the entire mathematical formulation of selecting multiple dominant genes is biologically motivated by experiments on tumor transplantation susceptibility in mice. He models the genetic trajectory of mice lines requiring between 2 and 12 independent dominant genes for successful tumor grafts.
3. Fruit Fly (Drosophila melanogaster)
Drosophila serves as Haldane’s classic invertebrate model to test the boundaries of his equations:
- Sex-Linked Selection: In Part I, Haldane tests his sex-linked models using the “eosin” eye-colourand the sterile “fused” female mutations in Drosophila.
- Mutation Rates: In Part V, he uses Drosophila as the baseline comparison for estimated natural mutation rates .
- Epistatic Peaks: In Part VIII, he models stable genetic equilibria using Gonsalez’s data on Drosophilacarrying purple-eye, arc-wing, and axillary-speck mutations, which individually reduce fitness but in combination increase viability .
4. Jimsonweed (Datura stramonium)
Datura is Haldane’s primary plant model, specifically used to explore polyploidy and gamete-level selection:
- Autopolyploidy: In Part III, Haldane uses the tetraploid chromosome structure of Datura to formulate how selection acts on organisms with duplicated genomes.
- Pollen Competition: In Part VII, he applies his normal-curve selection model to Buchholz and Blakeslee’s empirical studies on the competitive growth of pollen tubes in Datura styles .
5. Peppered Moth (Amphidasys betularia / Biston betularia)
- Natural Selection in Real Time: Although only appearing in Part I, the peppered moth is historically one of Haldane’s most famous empirical examples. He uses historical collection data from Manchester starting in 1848 to calculate the selection coefficient of the dominant melanic form (doubledayaria), mathematically proving it had a massive 33% selective advantage over the light, typical form.
6. Wheat (Triticum species)
- Allopolyploidy and Selfing: Haldane uses wheat as his prime example of allopolyploid genetics (specifically hexaploidy) in Part VI and Part X. He models how self-fertilisation in wheat dramatically accelerates the elimination of multiple recessive “rogue” phenotypes compared to random mating.
7. Evening Primrose (Oenothera species)
- Non-Amphimictic Selection: In Part I, Haldane uses Renner’s studies on Oenothera muricata to model selection in species that do not undergo normal sexual reproduction (amphimixis), showing how permanent heterozygotes survive.
8. Maize (Corn / Zea mays)
- Gamete Selection: In Part II, Haldane models the dynamics of selective fertilisation (where pollen carrying a certain gene grows faster than pollen carrying its allele) using Jones’s empirical maize breeding experiments.
9. Beans (Vicia faba)
- Continuous Variation: In Part VII, Haldane utilizes Johansen’s pure-line bean breadth data to mathematically demonstrate how natural selection acts upon a normally distributed phenotypic trait .
Does Haldane make any mistakes with the mathematics in these papers?
Within the text of Haldane’s ten-part series itself, there are no published errata or acknowledgements of mathematical “mistakes” or blunders in his derivations. However, Haldane is remarkably transparent, rigorous, and explicit about the mathematical limitations and approximation errors inherent in his models.
To make his equations tractable in an era before digital computers, Haldane frequently simplified his models by treating discrete generational steps as continuous variables (using differential and integral calculus rather than finite difference equations). He openly notes where these mathematical trade-offs introduce minor deviations from biological reality:
1. Approximation Errors under Strong Selection (Part I)
Haldane warns readers that his equations for the rate of slow selection are “not rigorously accurate”.
- He explains that the generation number (\(n\)) is actually a higher transcendental function, meaning his simplified formulations for \(kn\) carry minor inexactitudes of the order \(k^2n\) (where \(k\) is the selection coefficient).
- In calculating Table III (slow selection on a sex-linked character), he explicitly flags that the figures are “not quite accurate unless selection is very slow,” pointing out that when \(k = 0.2\), the approximation error in the number of generations reaches nearly 10%.
2. Deviations in Extreme Selection Scenarios (Part IX)
When modeling highly intense, rapid selection (where the selective disadvantage is very large), Haldane notes the limits of his series expansions. He explicitly states that under extremely intense selection (such as when \(k = -4\), representing a massive survival advantage for recessives over dominants), “the error here exceeds 10% over a certain range of values of \(u\)”.
3. Deliberate 1% Errors for Simplified Breeder Formulas (Part X)
In his tenth paper, when formulating equations to help breeders estimate how many generations of selection are needed to reduce double dominants to 1%, Haldane employs a highly approximate solution derived from his 1932 work. He openly tells the breeder that this equation “has an error of about 1 percent in the neighbourhood of \(x_n = 2\),”though it becomes much more accurate for smaller values.
List the main variables Haldane develops in these articles when considered as a whole and give some examples of their application.
Across his ten-part series, J.B.S. Haldane establishes a highly structured mathematical lexicon to model how genetic frequencies shift over time under various evolutionary forces.
Considered as a whole, the main variables Haldane develops—along with their biological applications—can be grouped into three core categories: central state variables, evolutionary force coefficients, and demographic/structural parameters.
1. Central State Variables (Genetic & Phenotypic Frequency)
These variables track the shifting genetic composition of a population across successive generations.
- \(u\) (or \(u_n\), \(v_n\)): The Gametic (or Genic) Ratio
- Definition: Haldane’s central independent variable. It represents the ratio of the frequency of the dominant allele \(A\) to the recessive allele \(a\) in the gametes forming the \(n\)-th generation. (If gamete frequencies are \(p\) and \(q\), then \(u = p/q\)).
- Application: Haldane uses \(u\) to calculate the speed of natural selection under different mating systems. In Part I, he demonstrates that if recessives are completely eliminated, the gametic ratio simply increases linearly: \(u_n = u_0 + n\).
- \(y\) (or \(y_n\)): The Proportion of Recessives (Phenotypic Frequency)
- Definition: The proportion of homozygous recessive individuals (\(aa\)) in the population in generation \(n\). Under random mating, this is mathematically tied to the gametic ratio by the Hardy-Weinberg relation: \(y_n = (1 + u_n)^{-2}\).
- Application: Haldane plots \(y_n\) to construct selection curves (such as in Part I). He famously uses this to prove that natural selection is highly inefficient at purging harmful traits when they are rare, showing that reducing \(y_n\) from \(0.01%\) to \(0.001%\) under random mating takes an immense number of generations compared to more common frequencies.
- \(Z(i, j)\) and \(g_m\): Multilocus Zygotic and Gametic Proportions
- Definition: Developed in Part X to scale up selection to multiple gene loci. \(Z(i, j)\) represents the proportion of zygotes homozygous for \(i\) dominant and \(j\) recessive genes, while \(g_m\) represents the proportion of gametes carrying \(m\) dominant genes.
- Application: Used to calculate how fast complex, multi-locus phenotypes can be fixed or eliminated by artificial breeders.
2. Evolutionary Force Coefficients
These parameters quantify the strength of the external pressures acting upon a population’s gene pool.
- \(k\) (or \(k_1, k_2, k_f, k_m\)): The Selection Coefficient
- Definition: A measure of selection intensity. If the favoured phenotype has a fitness of \(1\), the unfavoured phenotype has a relative fitness (viability/reproduction) of \(1-k\). Haldane uses positive values of \(k\) for selection favouring dominants and negative values for selection favouring recessives.
- Application: Haldane famously applies this to the peppered moth (Biston betularia) in Manchester. By analyzing historical collections from 1848 to 1898, he calculates that the dominant black form (doubledayaria) had a selection coefficient of \(k = 0.332\)—proving it possessed a massive \(33%\) survival advantage over the typical light form.
- \(l\): The Migration Rate (Immigration Coefficient)
- Definition: The fraction of a local population replaced in each generation by migrants arriving from a surrounding, unimproved population.
- Application: In Part VI (Isolation), Haldane models a population in a cave where normal-eyed aquatic animals migrate into an environment where blind individuals are selected for. He mathematically proves that selection of intensity \(k\) is entirely “swamped” and ineffective unless \(k > l\) (and \(k > 4l\) for autosomal recessive traits).
- \(p\) and \(q\) (or \(\mu\) and \(\nu\)): Mutation Rates
- Definition: The rate of forward mutation (\(p\) or \(\mu\), dominant \(A \to\) recessive \(a\)) and backward mutation (\(q\) or \(\nu\), recessive \(a \to\) dominant \(A\)).
- Application: In Part V, Haldane couples mutation with selection to calculate stable genetic equilibria. He proves that lethal recessive mutations will persist in a population at a stable equilibrium of \(y = \mu\) (or \(q = \sqrt{\mu}\)), showing why natural selection can never completely purge harmful hereditary diseases.
3. Demographic & Structural Parameters
These variables incorporate the biological realities of reproduction, age, and genomic architecture.
- \(n\): The Number of Generations
- Definition: Represents evolutionary time, treated either as a discrete step or integrated as a continuous variable.
- Application: Used to solve equations for practical breeding timelines. In Part X, Haldane solves a transcendental equation to show a breeder that it takes exactly \(12\) generations of random mating to reduce the proportion of double dominant rogues to \(1%\) (\(D_n = 0.01\)).
- \(k\) (in Part X): Number of Interacting Gene Loci
- Definition: Separate from the selection coefficient, this \(k\) represents the integer count of independent, unlinked gene pairs governing a single trait.
- Application: Biologically applied to experiments on tumor transplantation susceptibility in mice, which requires between \(2\) and \(12\) independent dominant genes. Haldane proves that selecting for multiple dominants under sibling mating yields homozygous purity at a rate of \((1 – r_n)^k\), where \(r_n\) is the recessive proportion for a single gene.
- \(K(x)\) and \(S(x)\): Age-Specific Demographics
- Definition: Utilized in Part IV for overlapping generations. \(S(x)\) is the probability of a female surviving to age \(x\), and \(K(x)\) is the probability of a female of age \(x\) producing a female offspring.
- Application: Haldane applies this to demographical tables of American women to scale continuous selection. He proves that the rate of selection under overlapping generations matches the discrete generation models when scaled by the generation interval \(\tau = \int_{0}^{\infty} x K(x) dx\).
Combine Haldane’s mathematical models into one that represents a cumulative synthesis of the whole series of papers.
To synthesize Haldane’s decade-long series into a single, cumulative mathematical model, we must construct a “Grand Unified Haldane Equation” (GUHE). This master equation must mathematically integrate the distinct evolutionary forces Haldane modeled separately across his ten papers—namelyselection, mating systems, mutation, migration, and generational overlap—into a single differential equation for the rate of change of gene frequency.
Here is a proposed blueprint to build this synthesized mathematical framework for you.
Proposed Structure of the Synthesis
I will construct the unified model around the rate of change of the gametic ratio \(u\) (the ratio of dominant allele \(A\) to recessive allele \(a\), where \(u = p/q\)), which Haldane used as his primary variable throughout the series.
We can represent the unified model through four integrated layers:
- The Selection Core (Parts I, III, VII, IX): We will model the core selection pressure using a generalized fitness function \(W(u)\) that accounts for the selection coefficient \(k\) and incomplete dominance/heterozygous advantage \(K\) .
- The Reproductive & Mating System Modifier (Parts II, X): We will introduce a mating system coefficient (derived from Haldane’s inbreeding and assortative mating parameters) that alters the proportion of homozygotes and heterozygotes, modifying how selection acts on recessives .
- The Mutational and Migrational Flux (Parts V, VI, VIII): We will superimpose terms for recurrent mutation rates (forward mutation \(\mu\) and backward mutation \(\nu\)) and the swamping effect of migration \(l\) from a surrounding population with a constant allele ratio \(u_m\) .
- The Demographic Scale (Part IV): We will incorporate Lotka’s demographic age-structure adjustment \(\tau\) (generation interval scaling) to allow the continuous differential model to accurately represent overlapping generations .
The Proposed Master Equation
I will synthesize these forces into a master differential equation of the form:
\[\frac{du}{dt} = \frac{1}{\tau} \left[ \Psi_{\text{selection}}(u, k, K, f) + \Psi_{\text{mutation}}(u, \mu, \nu) + \Psi_{\text{migration}}(u, l, u_m) \right]\]
How I Will Present This to You
I can deliver this cumulative synthesis as a highly polished Technical Report in your Studio panel containing:
- The Unified Mathematical Derivation: Showing step-by-step how these disparate equations are combined into one continuous-time master differential equation.
- The Parameter Dictionary: A reference table mapping the parameters back to the original papers in the series.
- The “Special Cases” Proofs: Demonstrating how the master equation mathematically collapses back into Haldane’s original equations (e.g., setting \(l=0, \mu=0, \tau=1\) to recover the classic Part I selection curve).
- An Interactive Python Simulation Script: A clean, executable script you can run to simulate how changing inbreeding, migration, and selection simultaneously affects a population’s genetic trajectory over generations.
Build a technical synthesis of his mathematical model. (Download available)
Notebook has created two new files in your Studio panel representing a cumulative synthesis of Haldane’s entire ten-part series:
- haldane-unified-synthesis.md (download ZIP file): A mathematically rigorous, comprehensive technical report that merges Haldane’s separate models into a single, continuous-time master differential equation—the Grand Unified Haldane Equation (GUHE). It features a complete parameter dictionary mapping variables back to their original papers, step-by-step mathematical proofs showing how the GUHE collapses back into Haldane’s individual models, and an embedded executable Python simulation script.
- haldane-simulation-trajectories.png: A publication-quality visual plot illustrating the unified dynamics. It simulates two fundamental insights from the series: how inbreeding breaks the heterozygous “shield” to accelerate recessive selection, and how migration from unimproved populations “swamps” local selection to establish a stable, suboptimal equilibrium rather than allowing fixation.
These files are fully functional and ready for you to study, export, or run.