Abstract
Modern gene synthesis platforms enable investigations of protein function and genome biology at an unprecedented scale. Yet, the proportion of error-free constructs in diverse gene libraries decreases with length due to the propagation of oligo synthesis errors. To rescue these error-free constructs, we developed Barcode-Assisted Retrieval CRISPR-Activated Targeting (BAR-CAT), an in vitro method that uses multiplexed dCas9-single-guide RNA (sgRNA) complexes to extract barcodes corresponding to error-free constructs. After a 15-min incubation and wash regimen, three low-bundance targets in a 300,000-member test library were enriched 600-fold, greatly reducing downstream requirements. When applied to a 384-gene DropSynth gene library, BAR-CAT enriched 12 targets up to 122-fold and revealed practical limits imposed by sgRNA competition and library complexity, which now guide ongoing protocol scaling. By eliminating laborious clone-by-clone validation and working directly on plasmid libraries, BAR-CAT provides a platform for recovering perfect synthetic genes, subsetting large libraries, and ultimately lowering the cost of functional genomics at scale.
Introduction
Plasmid DNA libraries have evolved from gene discovery and shotgun genome sequencing1,2 to enabling barcode lineage tracing, 3 CRISPR screens, 4 gene regulation studies, 5 multiplex assays of variant effects (MAVEs),6,7 and directed enzyme evolution. 8 This shift was driven by massive reductions in sequencing costs,9,10 from $1 million per human genome in 2007 to $200 in 2024 with Illumina NovaSeq.
These large-scale applications have shifted DNA libraries toward being synthetic instead of natural due to increased customizability. Short synthetic inserts (≤300 bp) have encoded peptides, exons, DNA barcodes, or guide RNAs in CRISPR screens for functional genomics, 4 protein interaction mapping, 5 or targeted exon sequencing. 11 Longer synthetic libraries (>300 bp) consist of oligos assembled into full-length sequences for protein engineering, functional assays of variant effects (MAVEs), 12 synthetic genome assembly, and synthetic metagenomics. Due to challenges in DNA assembly, error-prone polymerase chain reaction (PCR) or saturation mutagenesis is often used to introduce diversity into natural templates,8,13 but the resulting variants often remain closely related to parental sequences.13,14 However, synthetic gene libraries are the only option sufficiently programmable for MAVEs such as broad mutational scanning (BMS), which excels in profiling thousands of homologs. 15 A BMS study uncovered gain-of-function mutations in dihydrofolate reductase (DHFR) homologs, previously inaccessible by other methods, with implications for understanding antimicrobial resistance. 16
DropSynth assembles up to 1,536 genes in parallel on beads within emulsion droplets, enabling the compartmentalized polymerase cycling assembly of oligos.15–18 Despite high efficiency and low cost (∼$0.70/kbp), truncations, deletions, and substitution errors from oligo synthesis10,19,20 accumulate, limiting DropSynth assemblies to ∼1,000 bp with ∼8% perfect constructs. We predict that 2,000-bp assemblies would yield <1% perfect constructs, complicating the recovery of perfect sequences. 10
Typically, error-free oligos are enriched during synthesis, after synthesis, or after gene assembly. Some approaches include sequencing-by-synthesis to eliminate truncated oligos 21 and an obsolete approach for physically picking off perfect oligos from a sequencer. 22 Dial-out PCR can precisely retrieve genes or oligos that have been sequence-verified, but scalability is limited by the cost of adaptors and the lack of multiplexing.10,23–25 In contrast, many scalable approaches are hybridization-based, although they require very carefully designed probes.11,26,27 Functional selections are more straightforward in that they involve fusion to fluorescent or antibiotic resistance markers but at the risk of false positives.28,29 Mismatch-binding proteins, such as MutS, can enrich perfect assemblies up to 25.2-fold,10,30,31 but both MutS and T7 endonuclease I require large amounts of heteroduplex DNA and do not detect all mismatches.17,32,33 These limitations render existing approaches incompatible with multiplexed methods such as DropSynth 17 and motivated us to develop a CRISPR-based strategy for enriching perfect assemblies from synthetic gene libraries.
Here, we introduce Barcode-Assisted Retrieval CRISPR-Activated Targeting (BAR-CAT), an in vitro method that can enrich perfect gene assemblies. BAR-CAT uses a deactivated version of Streptococcus pyogenes Cas9 (dCas9) due to its high targeting specificity 34 and demonstrated compatibility with multiplexed DNA enrichment in vitro. 35 In principle, dCas9 can enable precise, programmable, and scalable targeting by adjusting the number and sequences of single-guide RNAs (sgRNAs).35–37 BAR-CAT is used by tagging gene libraries with random 20-nucleotide (nt) barcodes. After sequencing, barcodes corresponding to perfect assemblies are selected as spacers that are in vitro transcribed into sgRNA libraries that target specific subsets. 38 Each sgRNA library is combined with dCas9 to form active ribonucleoproteins (RNPs), which bind to targeted barcodes. Magnetic bead pull-down washes isolate enriched DNA and eliminate unwanted barcodes prior to final amplification (Fig. 1A). Our benchmark for BAR-CAT was at least 100-fold enrichment of 60 targeted barcodes to enable the retrieval of sequences from 384- or 1,536-gene libraries, which involved exploring CRISPR-dCas9 multiplexing limits. Unlike in vivo mammalian implementations, BAR-CAT is not constrained by the need to express multiple sgRNAs,39,40 raising the question of whether CRISPR-dCas9 can be scaled to more than a few targets.39–41

Overall workflow for BAR-CAT proof-of-concept and initial optimizations.

Increasing the input supercoiled rfp library DNA leads to large improvements in observed BAR-CAT enrichment while increasing the incubation time does not.
Initial process-level optimizations to BAR-CAT, such as increased bead washes, modestly increased the median enrichment of 18 barcodes from 6.3–20-fold in a pilot library. The most substantial improvement, 600-fold median enrichment of three targets, resulted from identifying the DNA input amount as a sensitive parameter, leading to BAR-CAT version 1.0 presented here. When applied to 384- or 1,536-gene DHFR libraries assembled by DropSynth, BAR-CAT achieved up to 149-fold median enrichment for 12 targets, although enrichment decreased at larger scales. While we discuss strategies to expand multiplexing capacity and reduce off-target binding, BAR-CAT represents, to our knowledge, the first application of a CRISPR-Cas system to enrich synthetic genes. This provides a framework for advancing CRISPR-based technologies and developing next-generation BAR-CAT workflows to efficiently access and manipulate synthetic DNA libraries.
Materials and Methods
Building and barcoding gene libraries
We constructed a barcoded pilot library and DropSynth gene libraries. For the pilot library, a PAM-less plasmid (pEVBC3) was modified (Supplementary Tables S1 and S2) to barcode an mCherry cargo gene to produce the red fluorescent protein (rfp) library, with barcode diversity limited by high-efficiency electroporation into Escherichia coli. The DropSynth libraries were produced similarly using a PAM-containing plasmid (pEVBC8; Supplementary Tables S1 and S2) to barcode 1,526- and 384-gene DHFR libraries 16 according to the published protocol. 15 The detailed methods are in the Supplementary Data.
Illumina gene library sequencing
After the diversity bottleneck of transformation, PCR added adapters and indices for Illumina sequencing. Different indices and primers were used for the rfp pilot and the DropSynth DHFR libraries (detailed in the Supplementary Data and Supplementary Table S3). Sequencing reads from MiSeq runs were merged with bbmerge (bbmap 38.18) and processed with a custom Python pipeline to extract barcodes and gene regions. 15 Barcodes were collapsed using the Starcode spheres algorithm on a distance of 1, and a majority call determined the consensus sequence for each barcode. 42 Scripts are available at the lab’s GitHub repository (https://github.com/PlesaLab/BC_Mapping).
sgRNA spacer selection pipeline
We prepared a computational pipeline written in R 4.3.0 with several Python helper scripts for the purpose of converting long-amplicon sequencing data from a barcoded gene library into a set of sgRNAs for BAR-CAT. The following criteria were evaluated by the pipeline: (i) sgRNAs must cover every targeted gene, (ii) avoid recognizing any nontargeted gene, and (iii) be normalized so that every gene targeted by a sgRNA had a similar number of sequencing reads (Supplementary Fig. S1). The complete codebase, including the small Python utilities, is available at the Plesa Lab GitHub repository (https://github.com/PlesaLab/). The detailed methods for this sgRNA selection pipeline and the filtering criteria are provided in the Supplementary Data.
Preliminary exploration of buffer compositions for magnetic bead washes following BAR-CAT enrichment
Before establishing the BAR-CAT v0.1 protocol, we conducted preliminary studies to assess how the number of bead washes and the composition of wash buffers affected stringency following BAR-CAT enrichment. Various buffer conditions (listed on Supplementary Table S4) were tested using streptavidin magnetic beads. The full experimental details are available in the Supplementary Data.
Development of BAR-CAT v1.0
The first successful iteration of BAR-CAT enriched 18 barcodes from the barcoded rfp plasmid library and was designated version 0.1 (v0.1). Ribonucleoproteins (RNPs) containing biotinylated dCas9 and in vitro transcribed sgRNA libraries were added to 50 ng of the rfp library. After 15-min incubation at 37°C, targeted barcodes were captured with streptavidin-coated magnetic beads, washed to remove nonspecific DNA, and amplified by PCR. Version 0.2 (v0.2) increased wash stringency, using more washes and a larger cumulative buffer volume to further reduce off-target binding. Version 0.3 (v0.3) introduced the proteinase K treatment of bead–DNA complexes prior to final amplification, removing DNA nonspecifically bound to beads. These optimizations yielded modest gains, but not substantial improvements in enrichment. Version 1.0 (v1.0) identified library DNA input as critical, standardizing 500 ng of input DNA for both rfp and DropSynth DHFR libraries. The detailed protocols for all BAR-CAT versions are provided in the Supplementary Data.
Nanopore data analysis and enrichment scores
Nanopore sequencing was used to analyze enrichment data. BAR-CAT v0.1-v0.3 enrichments were run on an R9.4.1 flow cell with v10 chemistry and base-called using Guppy in the super-accurate mode, while subsequent enrichments used a R10.4.1 flow cell with v14 chemistry and Guppy 6.5.7 in the super-accurate mode. Raw FASTQ files were processed with custom Python script to extract barcode regions (motifs listed in the Supplementary Data). Extracted barcodes were collapsed using Starcode spheres (distance of 1) 42 and imported into R for analysis. Barcode population fractions were calculated by normalizing the total sequencing depth, and log2-fold enrichment scores were calculated as the ratio of post-enrichment to the initial population fractions. Barcodes absent from the initial population received an initial pseudocount of 0.5, and dropouts were tracked separately. Population-level enrichment was determined as the total read-level population fraction of targets after enrichment divided by the total fraction before enrichment. RNA-seq data were analyzed as described elsewhere. 38 Significance testing was performed using a Wilcoxon paired test. CRISPRscan scores were generated using crisprScore (1.4.0). 43
Results
Successful targeted retrieval of DNA barcodes (BAR-CAT version 0.1)
We tested whether BAR-CAT could selectively enrich barcodes from a low-complexity DNA library. Using a previously described approach, 15 we added unique 20-nt barcodes to an mCherry gene, generating a barcoded rfp plasmid gene library. Prior to selecting barcode targets, we evaluated the rfp library barcode distribution by calculating the Gini coefficient, where 0 indicates perfect uniformity and 1 indicates perfect inequality.44,45 Sequencing revealed >300,000 unique barcodes (Fig. 1B), with the least abundant appearing once and the most abundant ∼1,000 times, yielding a Gini coefficient of 0.56. This skewed distribution is typical of synthetic DNA libraries, arising from biases during oligonucleotide synthesis, amplification, assembly, or cloning. 46 By contrast, a single-copy 10-kb gene in diploid human DNA is ∼47-fold more abundant than the rarest barcode in the rfp library (∼1 in 14.6 million reads). To compensate for rfp library biases, we selected 18 target barcode PAM sites (5′-NGG-3′) at their 3′ end, choosing two per abundance tier from the most abundant (∼1,000 reads) to the least (1 read) to ensure coverage across the full distribution (Fig. 1A,B). We in vitro transcribed an 18-plex sgRNA library with spacers to target these barcodes, as described elsewhere. 38
Following BAR-CAT enrichment by incubating RNP complexes (dCas9 and 18-plex sgRNAs) with the rfp plasmid library, biotinylated dCas9 was captured using streptavidin-coated magnetic beads (Fig. 1A), which were then washed to remove nonspecifically bound barcodes (Fig. 1A). To determine adequate washing, we compared four buffer compositions at room temperature (∼25°C) or 30°C in sham capture containing DNA but lacking RNPs to simulate nonspecific DNA–bead interactions (Supplementary Fig. S2A). qPCR of supernatants after the final wash evaluated wash efficiency, with higher Cq values indicating reduced residual DNA. The 2× binding and wash buffer (2× B&W; 2 M NaCl in Tris-EDTA buffer) at room temperature provided sufficient stringency and was used for all BAR-CAT bead washing steps (Supplementary Fig. S2A).
Following a full BAR-CAT run designated as version 0.1 (v0.1), nanopore sequencing showed a median log2 enrichment of 2.7 (6.3-fold) for the targeted barcodes (magenta dots; pink, Replicate 1, Fig. 1C,D). Population fraction enrichment, defined as the ratio of the sum of all targeted barcode fractions after versus before enrichment, increased by 6.4-fold compared with the initial rfp library. In practical terms, this means that 6.5-fold fewer colonies would need to be screened to recover the targeted barcodes. Analysis of off-target barcodes revealed two subpopulations (blue dots; Fig. 1C): (1) an unexpected nonenriched majority (10−5 to 10−6 fraction) and (2) an enriched minority, consistent with known CRISPR-dCas9 off-target activity.47–49 Overall, most off-target barcodes were not enriched; however, only 67.4% were depleted, and 3.3% dropped out entirely (Fig. 1D), highlighting room for improvement. Further optimization was therefore necessary to increase on-target enrichment while reducing off-target barcodes.
Increasing bead washes modestly improved on-target barcode retrieval (BAR-CAT v0.2)
We reasoned that increasing the bead washes in BAR-CAT v0.1 could reduce nonenriched, off-target barcodes after enrichment. In a sham capture experiment (beads combined with DNA in the absence of RNPs), magnetic beads were washed 15 times, and DNA from the wash supernatants and the final bead pellet was amplified and analyzed by gel electrophoresis. The PCR product was undetectable after the ninth wash, while the final bead pellet produced abundant product (Supplementary Fig. S3A). ImageJ analysis 50 indicated a decrease in DNA band intensity from 18,552 a. u. (arbitrary units; third wash) to 5,708 a.u. (ninth wash, 69.2% reduction) and 4,005 a.u. (twelfth wash, 78.4% reduction; Supplementary Fig. S3B). Therefore, the following washing regimens were compared with the BAR-CAT v0.1 control (six washes with 50 μL each, 0.3 mL total; Replicate 2): six washes of 5 mL each (30 mL total) or nine washes of 2 mL each (18 mL total).
The experimental washing regimens tested doubled enrichment relative to the BAR-CAT v0.1 washing control (Replicate 2), which achieved a median log2 enrichment of 2.9 (7.4-fold) and similar barcode distributions to the first iteration, demonstrating reproducibility. The 6 × 5-mL and 9 × 2-mL bead washing regimens increased the median log2 enrichments to 3.7 and 3.6 (12.5- and 12.2-fold), with population fraction enrichments of 13.4- and 12-fold, respectively. The off-target barcode dropout remained unchanged, although the moderately abundant off-targets (10−5 to 10−6 fraction) increased in prominence with washes (Fig. 1C, Supplementary Fig. S4B,C). The 9 × 2-mL wash condition was adopted for all subsequent experiments, establishing BAR-CAT v0.2, although further optimization was needed to increase enrichment and reduce off-targets.
Enhanced depletion of off-target barcodes (BAR-CAT v0.3)
Because increasing the bead washes did not sufficiently improve BAR-CAT performance, we reasoned that off-target barcodes were nonspecifically bound to beads associated with dCas9 and co-amplified with targets. To release enriched DNA and remove bead-associated dCas9, we compared three dCas9 denaturation approaches: 8 M urea, 35 proteinase K,51,52 and heat incubation.52,53 In parallel, we hypothesized that dCas9 might access targets more effectively with a relaxed DNA topology 53 ; thus, we compared enrichment using linearized versus supercoiled rfp libraries. We enriched 18 barcodes from linearized plasmids and applied dCas9 denaturation treatments only to supercoiled plasmid enrichments.
Neither dCas9 denaturation nor DNA topology altered barcode distributions relative to the supercoiled BAR-CAT v0.2 control (9 × 2 mL; Replicate 1; Supplementary Figs. S4B, S5A–E). The median log2 enrichment was similar across denaturation conditions: urea (4.5, 23-fold), proteinase K (4.3, 20-fold), boiling (4.5, 23-fold), and the supercoiled control (4.3, 19.5-fold). Urea had low population fraction enrichment (14-fold), while proteinase K and boiling were better (Fig. 1D). Linearization increased the median log2 enrichment to 4.9 (31-fold) and the population fraction enrichment to 26-fold compared with 23-fold for the supercoiled control. However, it is unclear whether this improvement resulted from the linearization itself or from the higher molar concentration. For instance, linear rfp molecules (827 bp) were ∼3.5-fold more numerous than 2.7-kb supercoiled plasmids at equal mass, lowering the RNP-to-DNA ratio from ∼35:1 to ∼10:1. A proteinase K dCas9 denaturation step was incorporated in BAR-CAT v0.3 because it increased barcode dropouts by 57% (7.7% total, 131,254 barcodes) relative to the supercoiled control (6.2%, 106,119 barcodes). By comparison, urea and boiling were less effective, only increasing dropouts by 12% (4.4%, 74,542 barcodes) and 14% (5.6%, 95,186 barcodes; Supplementary Fig. S6, Fig. 1D). Heat and urea may have removed some nonspecifically bound off-targets, reducing separation from target barcodes. Enrichment from the linearized library showed 27% higher dropouts (6.2%, 106,119 barcodes), potentially due to experimental variation or reduced off-target binding. These results suggested that increasing the amount of target barcodes available for dCas9 binding could further improve enrichment.
Increasing input DNA enhances targeted enrichment (BAR-CAT v1.0)
Given the BAR-CAT v0.3 results, we reasoned that incubating RNPs with more supercoiled molecules could improve enrichment from the rfp library. Unlike RNA-guided endonuclease (RGEN) methods for the targeted next-generation sequencing (NGS) of genomic DNA, 51 BAR-CAT retrieves targets from nonuniform synthetic libraries. RGEN approaches typically require ∼200-fold more input DNA, 51 suggesting that higher DNA input could enhance BAR-CAT performance. We therefore compared enrichment using 50 ng (BAR-CAT v0.3) versus 500 ng input DNA.
In parallel, we hypothesized that extending RNP–DNA incubation could provide dCas9 with a greater likelihood of finding targets among hundreds of thousands of unique sequences. BAR-CAT’s incubation time was 15 min, whereas the incubation time in other CRISPR-based enrichment methods range from 20 min to 8 h.35,51,54 Given a Cas9 binding rate constant of 0.8 ± 0.2 min−1, ∼80% of targets were bound within 1 min under simplified conditions, 52 suggesting that short times may suffice for few targets. To test the effect of the incubation duration, we incubated RNP–DNA mixtures at 37°C for 15 min, 1 h, or 8 h.
We enriched 3 of the original 18 barcodes (7, 8, and 15; Fig. 1B) from the supercoiled rfp library with synthetic sgRNAs due to a temporary pause in transcribing sgRNAs. 38 These barcodes were chosen because they showed moderate to high enrichment in BAR-CAT v0.3 (proteinase K; Fig. 1D, Supplementary Fig. S7). Across all conditions, enrichment was higher with 3 targets versus 18, indicating reduced performance at higher multiplexing (Figs. 1D, 2A). Using 500 ng input DNA, 15-min incubation yielded the highest enrichment (median log2: 9.2, 600-fold; 1,094-fold population fraction) and was designated BAR-CAT v1.0. Longer incubations (1 or 8 h) reduced the median enrichment and population scores for 500 ng DNA inputs (1 h: log2 7.7, 205-fold; 8 h: log2 7.4, 163-fold), while 50 ng DNA inputs showed similar decreases (Fig. 2A).
Off-target barcode depletion and dropouts were similar between the 50 ng control (82.2% depletion, 3.9% dropouts) and BAR-CAT v1.0 (79.1% depletion, 3.4% dropouts; Fig. 2B). Most off-targets weren’t enriched (Fig. 2C, Supplementary Fig. S8) unlike earlier experiments (Fig. 1C, Supplementary Figs. S4, Figs.S5). Interestingly, the 50 ng input incubated for 8 h had fewer off-targets (Supplementary Fig. S8A–C) and increased barcode dropouts (79%) compared with the control (from 3.9% to 7.0%), likely due to random variation (Fig. 2A). Conditions with 500 ng input DNA showed consistent target and off-target barcode distributions (3.4–5.9%; Fig. 2B) across all incubation times (Supplementary Fig. S8D–F).
Together, these results indicate that BAR-CAT v1.0, with 500 ng DNA and 15-min incubation, provided adequate conditions for scaling up the enrichment of perfect gene assemblies from DropSynth libraries. Lower DNA input or longer incubation times reduced enrichment, highlighting the importance of both parameters.
Demonstration of BAR-CAT v1.0 enrichment of perfect gene assemblies from DropSynth libraries
To evaluate BAR-CAT v1.0 in a real-world setting, we enriched perfect gene assemblies from 1,536- and 384-gene DropSynth DHFR gene libraries. After barcoding and adding a PAM site to each gene, Illumina MiSeq revealed 5,300,938 unique barcodes and 149 perfect DHFR genes in the 384-gene library (S4), from which 389 barcodes were selected. For the 1,536-gene library (S2), 93,418 unique barcodes and 684 perfect genes were identified, and 1,384 barcodes were selected. These targets were selected because they were ultra-rare (1–2 reads), ensuring normalization, and because they met the spacer selection criteria (Supplementary Fig. S1).
sgRNA libraries were in vitro transcribed as described previously, 38 but the spacer representation was skewed because T7 RNA polymerase strongly favors templates with four guanines downstream of the T7 promoter (5′GGGG).38,55 Initial sgRNA libraries included only a single 5′ guanine (5′G) at this position, exacerbating bias. We therefore also constructed more uniform sgRNA libraries with a 5′GGGG sequence preceding the spacer and directly compared enrichment versus 5′G sgRNA libraries.
To test scalability, we first enriched a single target from the 384-gene library using a synthetic sgRNA and then scaled it to 12 (5′G and 5′GGGG), 60 (5′GGGG), and 384 (5′GGGG) targets. From the 1,536-gene library (S2), 1,384 barcodes were targeted (5′G). Enrichment decreased as multiplexing increased, with many targets dropping out across DHFR libraries, regardless of the sgRNAs used (5′G or 5′GGGG). Singleplex controls (5′G, n = 2) achieved log2 enrichments of 8.4 (349-fold) and 6.8 (115-fold). Scaling to 12 targets (5′G, n = 2) reduced the median log2 enrichment to 6.9 (122-fold) and 3.5 (11.5-fold), with 18.2% dropouts in one replicate (Fig. 3A,B), likely due to sgRNA degradation (Supplementary Fig. S9). Freshly transcribed 5′GGGG sgRNAs (n = 1) provided similar results (median log2: 4.8; 29-fold), indicating modest negative effects of the 5′GGGG sgRNAs (Fig. 3A,B). Low-abundance off-target barcodes were enriched from ∼10−7–10−6 to 10−4–10−3 across singleplex and 12-plex scales (Supplementary Fig. S10A–E). The 60- and 389-plex enrichments confirmed decreasing enrichment with scale. The 60-plex enrichment (5′GGGG sgRNAs) had a median log2 and population fraction enrichment of 5.8 (58-fold) and 3.7-fold, respectively, but 89.5% of targets dropped out (Fig. 3B). The 389-plex enrichment (5′GGGG sgRNAs, n = 1) performed worse, with a median log2 of 1.6 (3.0-fold), a population fraction enrichment of 1.6 (3.0-fold), and 95.8% target dropouts (Fig. 3A,B). By contrast, 389-plex 5′ G sgRNAs (n = 2) achieved higher median log2 enrichment values of 3.8 (14-fold) and 3.6 (12-fold), with 18.6% fewer target dropouts (Fig. 3A,B). Substantial enrichment of low-abundance off-targets persisted for both the 60- and 389-plex conditions (Supplementary Fig. S11A,C–E). Off-target dropouts across all the 384-gene DHFR library enrichments averaged 97.1% ± 3.64% (n = 9) versus 5.73% ± 2.52% (n = 9) for the rfp library (Figs. 1D,3B), likely due to the library’s higher diversity.

BAR-CAT v1.0 performance declines with increasing enrichment scale and is further reduced by sgRNA spacers starting with 5′ guanine tetramers (5′GGGG) compared with single 5′ guanine (5′ G) designs.
Targeting 1,384 barcodes from the 1,536-gene library (5′G sgRNAs, n = 1) yielded no enrichment (median log2: −0.11, population fraction: 1.2-fold; Fig. 3A), with 39.8% target and 73.5% off-target dropouts. This suggests that high barcode diversity in the 384-gene library contributes to excessive dropout; yet, as the number of targeted barcodes approaches 1,384 dCas9 approaches a multiplexing limit. Nevertheless, some barcodes were still enriched, implying that the spacer sequence may affect performance (Supplementary Fig. S11B).
Lessons from BAR-CAT v1.0: Guiding future improvements for multiplexing and on-target enrichment
We wondered what factors were impeding BAR-CAT v1.0 from achieving robust, multiplexed on-target retrieval of barcoded genes from DropSynth DHFR libraries. One possibility was that the sgRNA spacers selected for this study limited BAR-CAT performance. We used CRISPRscan 43 to predict performance for the 389-plex and 1,384-plex sgRNA libraries. CRISPRscan predictions showed no correlation with experimental enrichment for either the 389-plex (5′G sgRNAs, Replicate 2, R2 = 0.009; Supplementary Fig. S12A) or 1,384-plex enrichments (5′G sgRNAs, R2 = 0.015; Supplementary Fig. S12B), indicating that sgRNA efficiency did not influence performance. Similarly, uneven spacer distributions in the sgRNA libraries 38 did not correlate with the observed log2 enrichment for either the 389-plex (Replicate 2, Supplementary Fig. S13A) or the 1,384-plex (Supplementary Fig. S13B) experiments.
Next, we asked whether excessive off-target enrichment was driven by shared sequence similarity with target protospacers. However, off-target sequences containing 7–10-bp PAM-adjacent seed matches were not significantly enriched compared with sequences without matches (data not shown), contrary to previous studies showing seed sequence-driven CRISPR off-target effects.36,37,48,57 This suggests that redesigning spacers alone may not reduce off-target enrichment in BAR-CAT.
Instead, we tested whether simple process improvements could enhance BAR-CAT enrichment and reduce off-targets. As the number of unique sgRNAs increases, free dCas9 becomes limiting until fully saturated, reducing overall sgRNA activity.56,58,59 We varied the amounts of 12-plex sgRNAs, dCas9, or the 384-gene DHFR library in BAR-CAT v1.0. Across all conditions, we observed that substantial target dropouts persisted (25–66.7%; Supplementary Fig. S14), possibly due to dCas9 reconstitution issues. Increasing inputs of sgRNAs, DNA, or dCas9 generally did not improve enrichment and often reduced the median log2 and population fraction enrichment scores, likely due to excess active RNPs increasing off-target enrichment (Supplementary Fig. S14). Doubling the total reaction volume and all the corresponding inputs had no effect on enrichment or off-target dropouts (mean: ∼43.7% ± 2.61% vs. 41.0% for the 5′G control or barcode distributions; Supplementary Fig. S15A–F).
We concluded that CRISPR prediction tools and re-designing selected spacers would not be sufficient to increase multiplexing while decreasing off-target enrichment. Instead, we postulate that increasing the amount of free dCas9 while keeping the amount of sgRNAs constant may improve enrichment without increasing off-targets, as opposed to our ineffective approach of increasing both dCas9 and sgRNAs. Our findings, as described here, will inform versions of BAR-CAT beyond v1.0 that will be better positioned for use in multiplexed gene enrichment.
Discussion
We developed BAR-CAT v1.0 as a framework for selectively enriching perfect gene assemblies, building on progressive improvements from earlier versions. The initial enrichment of 18 barcodes from a single-gene (rfp) library demonstrated feasibility, but the excessive off-target barcodes remained a challenge (Fig. 1C, Supplementary Fig. S4A). Increasing the bead washes reduced these off-targets, median enrichment and reduced nonenriched off-targets, while proteinase K treatment further separated targets from nonspecific barcodes (v0.3; Supplementary Fig. S6). Ultimately, increasing the DNA input had the most dramatic effect, yielding up to 600-fold median enrichment of three barcodes in BAR-CAT v1.0 (Fig. 2A,C). These optimizations highlight BAR-CAT’s sensitivity to reaction parameters, particularly those affecting the RNP-to-target ratio.
Interestingly, enrichment also depended on the incubation time. A 15-min enrichment incubation step was ideal, whereas longer incubations unexpectedly decreased performance (Fig. 2A). Although Cas9 is often reported to achieve near-complete target occupancy within an hour,52,57 these estimates may not fully reflect dCas9 binding dynamics across longer time scales. One possibility is that synthesis errors in the sgRNA reversibility determining region accelerate dCas9 dissociation, 49 leading to the gradual loss of complexes during extended incubations.
While BAR-CAT v1.0 represents the most effective version so far, challenges remain when enriching perfect genes from synthetic gene libraries. Enrichment decreased sharply with increasing scale (Fig. 3A), targeted barcode dropouts sometimes exceeded 90%, and off-target enrichment was substantial. Overcoming these challenges will be key for increasing the multiplexing capacity to at least 100-fold for 60 targets, which could allow the retrieval of all perfect genes from a 1,536-gene library with ∼25 BAR-CAT reactions. Developing BAR-CAT v2.0 will require additional resources, but our findings provide insights into future improvements.
We attribute BAR-CAT multiplexing limits to competition among increasing numbers of sgRNAs for a limited pool of free dCas9. Previous studies show that the effectiveness of each sgRNA decreases when >12 sgRNAs are present, supporting our observations (Fig. 3A).56,58 To overcome this, free dCas9 should be titrated as the sgRNA library size increases while keeping the sgRNA input constant. This may be sufficient to raise the multiplexing ceiling. Additionally, sgRNA spacers with a single 5′G leader are associated with higher enrichment and fewer target dropouts than those with the 5′GGGG leader. This is because the 5′GGGG sgRNAs produced lower targeted enrichment values with higher dropout rates compared with the 5′G sgRNAs (Fig. 3). This is likely due to excessive high-molecular-weight RNA diluting functional sgRNAs. 38
Excessive barcode diversity in the 384-gene DHFR library increased potential off-targets (Supplementary Figs. S10, Figs.S11A, Figs.C–E), and reducing diversity alone is insufficient to fully solve this issue. The persistent enrichment of low-abundance off-target barcodes was also present in the less diverse single-gene (rfp) library and the 1,536-gene plasmid libraries (Fig. 1C, Supplementary Figs. S4, Figs.S5, Figs.S8, Figs.S11B). The predicted efficiency of our spacers did not correlate with the observed enrichment (Supplementary Fig. S12), consistent with the limits of CRISPRscan, which predicts in vivo rather than in vitro performance.60–63 Furthermore, no seed sequence similarity was observed between the target and off-target barcodes, suggesting that spacer redesign may not reduce off-targets. Instead, promiscuous dCas9 binding to the seed region in off-target spacers, especially in large, diverse libraries, likely drives persistent off-target enrichment, provided that the reversibility-determining region is intact.49,64
Given these constraints, targeting barcodes based on their initial abundance may mitigate enrichment challenges. In BAR-CAT v1.0, the strongest enrichment came from low-to-medium-abundance barcodes (BCs 7, 14, and 15), while the most and least abundant barcodes consistently performed poorly (Supplementary Fig. S7). High-abundance barcodes are constrained by dCas9 binding and amplification limits, whereas ultra-low-abundance barcodes risk off-target capture due to excess RNPs. Low-to-medium-abundance barcodes may strike a “sweet spot” with sufficient molecules for robust enrichment. Coupling this strategy with reduced library diversity can further reduce dropouts and off-targets. Finally, we note that the improvements observed in BAR-CAT v1.0 may simply reflect higher DNA input. Future optimizations could use tRNA-assisted ethanol precipitation 65 to recover larger amounts of DNA in small reaction volumes compared with standard column-based methods.
In addition to process-level improvements, replacing dCas9 with Cas9 in BAR-CAT could enhance the stringency and scalability of enrichment. Cas9 can cleave even ultra-low-abundance barcodes, whereas dCas9 relies on binding to more abundant targets. Cleavage and linearization of supercoiled plasmids would allow the addition of unique adapters, enabling selective amplification. This strategy parallels FLASH and Cas12a-Capture, which use catalytically active Cas proteins for the targeted sequencing of low-abundance DNA while avoiding dCas9 limitations.62,66 However, off-target cleavage by Cas9 may remain comparable to that observed with dCas9. 34
Length-independent DNA enrichment is an advantage of dCas9 that may be lost with Cas9. For example, BAR-CAT v1.0 can retrieve entire plasmids (∼2.7 kbp, rfp library) and potentially DNA up to ∼20 kbp. Longer fragments may be limited by slower diffusion and steric hindrance, although slow-pipetting techniques can reduce shear stress for DNA up to 100 kbp. 67 Validated enrichment of long fragments could allow dCas9-powered BAR-CAT to capture low-abundance ancient DNA from environmental samples by binding conserved genes without cleaving fragile fragments. This could preserve fragile DNA and outperform existing RNA hybridization-based methods. 68
Conclusions
BAR-CAT v1.0 establishes the first practical framework for CRISPR-based DNA enrichment and provides insights into dCas9 behavior with synthetic gene libraries. Addressing multiplexing limits and off-target enrichment in BAR-CAT v2.0 could enable broader applications, including the subsetting of large libraries without reassembly and large-scale DNA manipulation for DNA data storage, 69 serving as a “cut and paste” mechanism. We urge the scientific community to build upon these insights to develop improved CRISPR-based tools, advancing synthetic biology and facilitating the generation of large biological datasets for various applications.
Authors’ Contributions
N.K.V.: Methodology, investigation, validation, writing—original draft, writing—review and editing, visualization. M.H.T.: Software. A.K.: Investigation, validation. C.P.: Conceptualization, methodology, software, formal analysis, writing—review and editing, supervision, visualization.
Footnotes
Acknowledgment
The authors provide special thanks to Yuki R. Gaudreault for assisting with some of the experimental work.
Author Disclosure Statement
N.K.V. and C.P. are named inventors on a patent based on this method (US20250043276). C.P. is a co-founder and holds equity in SynPlexity.
Funding Information
This work was supported by the National Science Foundation MCB-2032259 grant. The research reported in this publication was supported by the National Institute of General Medical Sciences of the National Institutes of Health under award number T32GM149387 (to N.K.V.). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Data and Material Availability
Raw Illumina MiSeq reads of the target gene libraries and nanopore sequence reads for enriched libraries were submitted to the NCBI Sequence Read Archive under the BioProject accession number PRJNA1273454 (https://www.ncbi.nlm.nih.gov/bioproject/1273454). Processed enrichment data for each spacer for all experiments are available on FigShare (
).
Supplemental Material
Supplemental Material
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
