How Could the Entire Internet Fit in a Sugar Cube?
The amount of data humanity produces keeps growing exponentially, but the physical space we have to store it isn't growing nearly as fast. Hard drives, tape cassettes, and magnetic archives take up room and tend to degrade within a few decades, eventually becoming unreadable. Scientists have been looking to nature's own storage system, DNA, for years. That's because DNA still carries readable genetic information from organisms that lived millions of years ago, and it does so in an astonishingly small volume. Researchers estimate that, in theory, a single gram of DNA could hold roughly a zettabyte of data — close to the entire internet. That makes DNA one of the most serious candidates for future "cold storage" needs.
Why Is This Being Discussed Now?
The appeal isn't just density. Data centers consume enormous amounts of electricity and require constant cooling, while DNA molecules can remain intact for thousands of years at room temperature, in the dark, in a dry environment. That makes DNA both a dense and a sustainable storage candidate. While disk formats become obsolete and unreadable within a few decades, the technology to read DNA will likely exist for as long as biology itself does.
How Do Numbers Turn Into Molecules?
From Binary Code to Four Letters
Computers rely on a binary system made up of just 0s and 1s. DNA, on the other hand, is built from four building blocks: adenine (A), guanine (G), cytosine (C), and thymine (T). Researchers who want to store data in DNA first convert a digital file into binary code, then use specialized algorithms to translate that binary sequence into a chain of A, C, G, and T letters. This conversion isn't arbitrary — it follows rules that prevent the same base from repeating too many times in a row, keep the sequence chemically stable, and add redundant information to guard against reading errors. The resulting letter sequence is, in effect, a molecular translation of a photo, a text, or a piece of software.
From Synthesis to Storage
The designed sequence is then turned into actual DNA molecules using devices called DNA synthesizers. This process, known as "writing," remains the slowest and most expensive step today. The synthesized DNA fragments can be stored in small tubes, either dried or in liquid form. When the data needs to be retrieved, DNA sequencing technology "reads" the molecules, converts the resulting letter sequence back into binary code, and reconstructs the original file. Teams working to increase density are also developing "random access" methods — pulling out only the molecules representing a specific file from a mixture containing millions of others, using chemical tags. It's similar to being able to pull a single book off a shelf without disturbing an entire library.
Where Does This Density Come From?
A conventional hard drive arranges data across a magnetic surface in essentially two dimensions. DNA, being a three-dimensional molecule, can pack far more information into the same volume. Each DNA base can also represent far more information than a classic computer bit, since it can take on four different states (A, C, G, T) instead of two. Combined, these two factors shrink the physical space needed to store the same amount of data by orders of magnitude. Researchers often use this example: if all the digital data produced in the world today were stored in DNA, the entire collection wouldn't fill a truckload — it could fit in the palm of a hand.
What Obstacles Remain?
The Speed Problem
The biggest drawback of DNA storage is still speed. A file that could be written to a hard drive in seconds can take hours, or even days, to "write" into DNA, because DNA synthesis is a chemical process where each base is added one step at a time. Research labs are trying to speed things up by running thousands of synthesis spots simultaneously on electrochemical chips, but write speeds still fall far short of magnetic or solid-state drives.
The Cost Wall
Producing synthetic DNA has long been expensive; writing a single megabyte of data into DNA costs many times more than writing it to a hard drive. However, DNA sequencing costs have dropped dramatically over the past two decades thanks to genome projects, and a similar drop is expected on the synthesis side. Companies and universities are trying to bring costs down by mass-producing DNA on semiconductor chips rather than one droplet at a time.
Error Resilience
DNA molecules are sensitive to chemical reactions, and errors in individual letters can occur during synthesis or reading. That's why, before writing data, redundant backup information is added — similar to the error-correcting codes used when sending files over the internet — so the original file can still be recovered intact even if a few bases contain errors.
Who Is Working on This?
Microsoft Research and the University of Washington have been collaborating for years to turn DNA storage from a lab curiosity into a real system, and their teams introduced some of the first fully automated end-to-end DNA storage demonstrations. Various universities and biotech companies are also focused on increasing synthesis speed, lowering costs, and making random access practical. Some teams are using DNA not just to store files but to preserve artworks and even entire language archives for the long term. This work shows that DNA storage is no longer science fiction, but a serious engineering goal.
What's Expected in the Coming Years?
Experts predict that DNA storage will first be adopted for "cold storage" archiving — data that's rarely accessed but needs to be preserved for a very long time. Satellite imagery, medical records, government archives, and scientific datasets are among the first likely applications. Improvements in write and read speed could eventually extend DNA storage to broader use cases, though it isn't expected to fully replace classic disks or cloud storage; instead, it's expected to become one layer among several storage tiers, chosen based on how long data needs to last and how often it's accessed. In short, the memory of our digital civilization may one day genuinely be stored in life's own code.

