mannafest.faith

Biological

DNA and Information Theory

Francis Crick's 1958 sequence hypothesis established that the message in DNA lies in the order of the bases, not in their chemistry — and chemistry does not dictate that order, any more than the properties of ink dictate the sentences on a page.

medium confidenceScientific

Crick and Watson's 1953 double-helix model, Science Museum, London — the backbone is chemistry; the order of the bases is not.
Crick and Watson's 1953 double-helix model, Science Museum, London — the backbone is chemistry; the order of the bases is not.User:Alkivar · Public domain

THE CLAIM

The sequence carries the meaning

Francis Crick's sequence hypothesis, stated in "On Protein Synthesis" (1958): the specificity of a piece of nucleic acid is expressed solely by the sequence of its bases, and that sequence is a code for the amino acid sequence of a protein.

The four bases attach to the backbone with essentially equal chemical affinity. No bonding force prefers one order over another — so the order is not a product of chemistry, and it is not random either, because it specifies working proteins. Michael Polanyi's comparison (Science 160:1308, 1968): the arrangement is chemically arbitrary the way type in a printing press is mechanically arbitrary.

THE RECORD

Engineers store data in it

This is not an analogy that has been stretched. In Science 337:1628 (2012) George Church, Yuan Gao and Sriram Kosuri encoded a 5.27-megabit book into synthesised DNA and read it back out. In Science 355:950 (2017) Yaniv Erlich and Dina Zielinski demonstrated a coding scheme approaching 215 petabytes per gram.

Industry uses DNA for exactly the reason biology does: it is a high-density, sequence-based digital medium. The human genome runs to 3.1 billion base pairs — about 750 megabytes of raw sequence in every nucleated cell.

THE CLAIM

The central fact

In "On Protein Synthesis" (Symposia of the Society for Experimental Biology 12:138-163, 1958) Francis Crick stated what he called the sequence hypothesis: the specificity of a piece of nucleic acid is expressed solely by the sequence of its bases, and that sequence is a code for the amino acid sequence of a protein.

That is the whole crux. The four bases — adenine, thymine, guanine, cytosine — attach to the sugar-phosphate backbone with essentially equal chemical affinity. No bonding force prefers one order over another.

The order is therefore not a product of chemical necessity, and it is not random either, because it specifies functional proteins.

Michael Polanyi made the point in Science 160:1308 (1968): the structure of DNA is chemically arbitrary in the same way the arrangement of type in a printing press is mechanically arbitrary, and in both cases the constraint that carries the meaning is imposed from outside the chemistry.

The scale

The human genome runs to roughly 3.1 billion base pairs. At two bits per base that is about 750 megabytes of raw sequence in every nucleated cell.

The floor is not much friendlier. In 2016 Clyde Hutchison and colleagues at the J. Craig Venter Institute published JCVI-syn3.0 in Science 351:aad6253 — a synthetic minimal bacterial cell reduced to 473 genes and 531,000 base pairs. It is the smallest genome of any autonomously replicating organism, and the authors reported that the function of 149 of those genes was unknown.

Dna is literally a storage medium

This is not analogy. In Science 337:1628 (2012) George Church, Yuan Gao and Sriram Kosuri encoded a 5.27-megabit book into synthesized DNA and read it back. In Science 355:950 (2017) Yaniv Erlich and Dina Zielinski demonstrated a coding scheme approaching 215 petabytes per gram.

Engineers use DNA for the same reason biology does: it is a high-density, sequence-based digital medium.

The inference

Stephen Meyer's argument in Signature in the Cell (HarperOne, 2009) is an inference to the best explanation with a narrow shape. Sequence-specified digital information has exactly one known cause in our uniform experience: intelligence. Chance and physical necessity are the alternatives, and neither has produced a demonstrated pathway to a coded sequence.

On the difficulty of finding function by chance, Douglas Axe's paper in the Journal of Molecular Biology 341:1295-1315 (2004) used site-directed mutagenesis on a 150-residue beta-lactamase domain and estimated that functional folds occupy about one sequence in 10^77 of the possible space.

Why it bears on scripture

John 1:1-3 opens with the Word as the agent of creation, and Genesis 1 has God create by speaking. A universe in which the operating instructions of every living thing are stored as a linguistic code is not a neutral finding; it is the shape the text predicted.

WHERE IT STANDS

The state of scholarship

medium confidence

The sequence hypothesis, the code, the base-pair counts and the storage results are settled science and nobody disputes them. The design inference drawn from them is not accepted in mainstream biology, which holds that once replication and selection exist, functional information accumulates without foresight.

Axe's estimate in particular is contested. Experimental work by Hayashi and colleagues (PLoS ONE 2006) on directed evolution of a phage protein recovered function far faster than his rarity figure predicts, and critics argue his method measures the rarity of one specific fold rather than of function in general.

The honest statement is that the facts are agreed and the inference is where the argument lives.

Scripture on this

Psalm 139:13-16 — The unformed substance written in God's book before it existed — a text image for embryology.

John 1:1-3 — Creation by the Word — the universe made by an act of language.

Genesis 1:11-12 — Each yielding seed after its kind — heritable specificity stated at the outset.

Colossians 1:16-17 — All things created by Him and holding together in Him.

Job 10:11 — Clothed with skin and flesh, knit with bones and sinews — assembly by design.

THE NUMBERS

What the number means

1077the share of sequence space that adopts a functional enzyme fold, by Douglas Axe's estimate — one in 10^77, from site-directed mutagenesis on a 150-residue beta-lactamase domain (Journal of Molecular Biology 341:1295-1315, 2004)
101085

Axe's figure is the contested one, and it should be read as contested: work by Hayashi and colleagues (PLoS ONE, 2006) on directed evolution of a phage protein recovered function far faster than the rarity predicts, and critics argue the method measures the rarity of one specific fold rather than of function in general. The facts underneath it are not contested at all — the code, the sequence hypothesis, and the base-pair counts are settled science. The argument lives in the inference, not the data.

  • 106base pairs in the smallest self-replicating genome (531,000)
  • 109base pairs in the human genome (3.1 billion)
  • 1080atoms in the observable universe

6 pieces of evidence

What this establishes

  1. 01The data

    Crick's sequence hypothesis (1958): the biological message is carried by base order alone, and base order is not fixed by chemical affinity.

  2. 02The data

    Polanyi (Science 160:1308, 1968) showed the constraints carrying DNA's meaning are chemically arbitrary — irreducible to the physics of the molecule.

  3. 03The record

    The simplest known self-replicating cell, JCVI-syn3.0, still requires 473 genes and 531,000 base pairs (Science 351:aad6253, 2016) — there is no simple starting point.

  4. 04The data

    DNA is used by engineers as an actual digital storage medium at up to ~215 petabytes per gram (Science 355:950, 2017); the code is not a metaphor.

  5. 05The numbers

    Axe's mutagenesis work (J. Mol. Biol. 341:1295, 2004) estimates functional protein folds at roughly one sequence in 10^77.

  6. 06The record

    In uniform experience, sequence-specified information has one known cause: a mind.

5 passages

Key verses

Berean Standard Bible

THE RECEIPTS

Sources

Contested pointscontested
  • Shannon information measures improbability, not meaning. A random string has high Shannon information; deriving 'designed' from 'informational' equivocates between two different technical senses of the word.
  • Once replication with variation exists, mutation and selection demonstrably generate new functional sequence — gene duplication and divergence, and observed novel enzymes such as the nylon-degrading enzymes in Flavobacterium, are documented cases.
  • Axe's 1-in-10^77 figure is contested at the level of method: it measures the rarity of one particular beta-lactamase fold, and directed-evolution experiments (Hayashi et al., PLoS ONE 2006) recover function from random libraries far more easily than the figure implies.
  • The design inference is disjunctive: it argues that because chance and known necessity fail, intelligence wins. That is only compelling if the space of natural explanations has been exhausted, and origin-of-life chemistry is an active field that has not been.