DNA replication is one of natures most refined and universal mechanisms for generating classical information at the nanoscale. At the center of this process operate DNA polymerasesmolecular machines that read a DNA template strand and build its complementary copy, incorporating nucleotides one by one in a strictly directional and irreversible fashion. Despite their nanometric dimensions and the strong thermal noise of the cellular environment, these enzymes achieve astonishing precision, often committing fewer than one error per million incorporations. Their fidelity arises from the interplay between molecular recognition, energetic discrimination, mechanical motion, and nonequilibrium driving. Yet, even with abundant experimental data, we still lack a quantitative physical theory capable of explaining how the diverse polymerase types operate and why specific mutation patterns emerge.
This project aims to fill this gap by developing a first-principles framework for information processing during DNA replication. A key originality of our approach is to treat replication explicitly as classical information-handling protocols executed on a one-dimensional chain with non-Markovian dynamics. Our methodology integrates heterogeneous experimental knowledgebulk biochemical assays, single-molecule measurements, and high-resolution structural datainto a coherent complex-systems perspective. Rather than viewing these datasets as isolated observations, we merge them into a unified theoretical language based on information Hamiltonians. This allows us to characterize polymerase behavior in terms of physical constraints and sequence-dependent energetics, unveiling how fidelity and robustness emerge from the underlying physics.
By developing models that capture long-range interactions, finite-memory effects, and lesion-bypass mechanisms, we aim to predict error rates and mutation spectra across diverse conditions and polymerase types. This will enable quantitative estimates without relying exclusively on costly and timeconsuming experimental assays. The theory will also clarify how mechanical motion is coupled to computation at the molecular scale, enriching our understanding of how nanosystems can process information efficiently.
A predictive physical theory of polymerase function can illuminate the origins of mutations associated with genetic disorders, cancer progression, and viral evolution, while simultaneously offering conceptual tools that extend far beyond biology. The DNA replication machinery provides a natural blueprint for engineering information technologies capable of operating robustly at ambient conditions. This represents an alternative to quantum information technologies at the nanoscale, which generally remain fragile at room temperature and condensed environments. Furthermore, bioinpired classical architectures will enable advances in molecular engineering, synthetic biology, and nanorobotics, where resilient, low-energy, context-aware information management is essential and where the robustness of lifes own molecular machinery can guide the next generation of communication systems.