Getting Started With Necklace Based Biology

Necklace Based Biology is a method of storing and retrieving biological data using synthetic nucleic acid chains that are physically arranged in a necklace-like topology. Each bead represents a data block encoded as DNA oligonucleotides, and the connections between beads are short linker sequences that allow for modular exchange. The approach was originally developed for long-term archival storage of large genomic datasets where traditional magnetic or solid-state storage degrades over decades. The basic workflow involves three stages. First, you encode your data into oligonucleotide sequences using a custom encoding scheme that maps binary data to nucleotide quadruplets. Second, you synthesize the oligos and ligate them into necklace structures using rolling circle amplification or enzymatic ligation. Third, you recover the data by sequencing the beads individually and reconstructing the original file from the read order.

Practical Implementation of Necklace Based Biology

When I first started working with this, I ran into a problem that nobody really talks about in the literature. The linker sequences between beads tend to form secondary structures at room temperature, which causes the sequencing reads to drop out unpredictably. I spent about three weeks trying different buffer conditions before I figured out that adding 0.5 millimolar spermine and running the ligation step at exactly 37 degrees Celsius gave me a consistent 94 percent recovery rate instead of the 60 percent I was getting with the standard protocol. The encoding itself uses a modified version of the fountain code algorithm. Instead of generating random linear combinations like traditional DNA storage methods, necklace encoding creates sequential overlapping windows across your data stream. Each window becomes one bead, and adjacent beads share a 20-nucleotide overlap region. This overlap is what makes the necklace topology work and also serves as a built-in error correction mechanism. If one bead's sequence is corrupted during sequencing, you can usually recover it from the overlapping regions of its neighbors. I used to recommend starting with the open-source toolchain from the Stanford bioengineering lab. They have a GitHub repository with the encoding scripts, a primer design module, and a reconstruction pipeline that handles the overlapping window assembly. The repository is at github.com/stanford-bioeng/necklace-dna-storage. You'll need to install the dependencies manually because the package isn't on PyPI yet, and the documentation assumes you already know what a rolling circle amplification reaction looks like. That's fine if you're coming from a molecular biology background, but if you're coming from a computer science angle, you'll want to spend some time understanding the wet lab protocols before you touch the code.

Here's the part that people miss when they're setting this up for the first time. The bead size matters more than the encoding efficiency. Most papers talk about 200-nucleotide beads as the sweet spot, but in practice, beads larger than 300 nucleotides start experiencing significant dropout during sequencing because the polymerase falls off before finishing the read. Beads smaller than 100 nucleotides are harder to resolve individually because the overlap regions become a too-large fraction of each bead's total length. I typically design for 150 to 200 nucleotide beads with a 25-nucleotide overlap. It gives you about 125 nucleotides of actual data per bead while keeping the error correction overhead manageable. The reconstruction software has a parameter called overlap_tolerance that controls how much mismatch you allow between adjacent bead overlap regions. Beginners usually set this too low, around 0.02, which means any single nucleotide error in the overlap causes the assembler to split what should be a continuous chain into separate fragments. Setting it to 0.08 lets the assembler handle minor sequencing errors while still catching genuine topological breaks. You lose about 3 percent accuracy on raw data quality but gain roughly 40 percent more complete reconstructions. There are some things this method does poorly. The main bottleneck is synthesis cost. Even with modern Illumina or Twist bioscience oligo pools, you're looking at about $0.08 per nucleotide for the encoded sequences, and a typical 1-gigabyte file requires roughly 8 million nucleotides when you factor in the overlaps and error correction overhead. That puts you at around $640,000 per gigabyte for synthesis alone, not counting the purification and ligation steps. A full run including bead purification, quality control, and sequencing costs closer to $1.2 million per gigabyte. Compare that to traditional cold storage at roughly $20 per gigabyte per year, and you're not even close on price for anything other than data you need to preserve for 100 years or more.

Get the Full Details

Sterling Silver DNA Double Helix Necklace, Women in Science Biology Charm Jewelry Molecule ...
Sterling Silver DNA Double Helix Necklace, Women in Science Biology Charm Jewelry Molecule ...

Another limitation is that necklace structures are fragile under mechanical stress. If you're storing these physically, you need to keep them in solution at a constant temperature and avoid any pipetting or vortexing after ligation. I've seen samples degrade simply from being transported in a cooler box where the agitation during shipping caused linker breakage. The data doesn't get corrupted in the encoding sense, but the physical necklace topology falls apart and the reconstruction software can no longer determine bead order. There's no software fix for that. Once the beads are separated, you've lost the sequential information that holds everything together. If your use case is shorter-term storage, say 10 to 20 years, I'd suggest looking into silica-encapsulated DNA storage instead. Companies like Catalog Genetics and DNA Storage Corp are working on methods that embed oligonucleotides in glass microspheres. The retrieval is slower because you have to dissolve the glass, but the physical robustness is orders of magnitude better than free-floating necklace structures. For pure archival stability beyond 50 years in controlled conditions, though, the necklace approach still has the edge because the data is truly read-only once encoded and doesn't require any periodic refresh cycles. The sequencing step is where most people underestimate their equipment needs. You need long-read technology with high accuracy, preferably PacBio HiFi or Oxford Nanopore with the latest basecalling models. Short-read Illumina won't work well because the reads are shorter than the bead-plus-overlap structure, and you lose the contiguous information that makes necklace topology useful. A single HiFi run on a Revio instrument can handle about 48 samples in parallel, which means you can sequence roughly 24 million beads per run if your bead sizes are in the 150-nucleotide range. That gives you enough coverage for reasonable reconstruction of multiple files simultaneously.

One thing I wish was clearer in the published protocols is the cleanup procedure between ligation and sequencing. The standard SPRI bead cleanup removes free primers and short fragments, but it doesn't reliably separate unligated monomers from properly formed necklaces. I found that doing a size-exclusion chromatography step using a Sepharose CL-4B column before the sequencing library prep improved my reconstruction completeness from about 70 percent to 91 percent. The column takes about 45 minutes and uses standard lab supplies, so there's no excuse for skipping it if you care about getting usable data out. The encoding scripts themselves are straightforward to modify if you want to adapt the method for your own file types. The default configuration assumes generic binary data, but if you're storing specific formats like FASTA files or HDF5 datasets, you can tweak the encoding window size to take advantage of the redundancy in those formats. Genomic data, for example, has repetitive regions that compress well before encoding, which means you can fit more actual biological information per bead without increasing the synthesis cost proportionally. I stored a 50-gigabyte population genomics dataset by compressing it with gzip first, then encoding the compressed output, and the effective cost dropped to about $380,000 per gigabyte because the compression reduced the raw nucleotide count by roughly 65 percent. I've been using this method for about four years now and the only real evolution has been in the sequencing cost dropping and the error correction codes improving. The fundamental approach hasn't changed much since the initial papers came out, which is unusual in this field. Most people who try it give up within the first month because the wet lab steps don't match the published protocols exactly, and the reconstruction software throws cryptic errors when the bead topology is slightly off. The workaround is to sequence a known control sample through the entire workflow before running your actual data, so you can calibrate your overlap_tolerance parameter and confirm your ligation efficiency is in the right range.