// NATURE NEWS — SPAZIO & SCIENZA
An expanded codebook of human transcription factor DNA-binding specificity
Nature
(2026) Cite this article
Gene expression is regulated by transcription factors (TFs), which recognize specific DNA sequence motifs. Several hundred putative human TFs, identified mainly by an apparent DNA-binding domain, lack known binding motifs1. Furthermore, even for well-characterized TFs, it remains controversial the degree to which motifs accurately reflect binding sites in living cells2. Here we describe a systematic effort (‘Codebook’) to determine the sequence specificity of 332 putative and poorly characterized human TFs. More than 4,000 independent experiments, encompassing multiple in vitro and in vivo assays, produced motifs for just over half (177; 53%) of the TFs, of which most are associated with only a single protein. These results extend the vocabulary of sequence recognition encoded by human TFs by around 130 distinct motifs. Moreover, binding motifs identified in vitro are strongly enriched in cellular binding sites. Collectively, the data reveal tens of thousands of previously unknown, conserved and direct TF-binding sites across the human genome. These sites are concentrated in promoter regions and are predictive of gene expression. In summary, this new codebook provides an important step forward in decoding the human genome.
A 2018 survey of putative human TFs concluded that over one-quarter of the estimated 1,600 TFs lacked established binding motifs1. This proportion represents a striking deficit given that most of the conserved DNA in the human genome is noncoding3 and that a primary hypothesis for the function of noncoding DNA is gene regulation4. Moreover, most of the uncharacterized and putative TFs are not close paralogues of any established TF. Thus, despite possessing either literature evidence of DNA binding or protein domains that would typically bind DNA (DNA-binding domains (DBDs)), their binding motifs cannot be readily inferred. However, it cannot be excluded that these putative TFs lack DNA sequence specificity entirely.
TF DNA-binding motifs are commonly modelled as position weight matrices (PWMs), which describe the relative preference of a TF for each nucleotide base pair in the binding site5,6. Different methods for measuring TF binding, and for deriving PWMs from the resultant data, have different inherent limitations and biases5. There is also long-standing controversy regarding the contribution of the inherent sequence specificity of a TF to its cellular binding relative to the influence of chromatin and cofactors2. These uncertainties represent fundamental hurdles for the analysis of gene regulation and for a myriad of related tasks in genome analyses, including the interpretation of conserved genomic elements and sequence variants.
To address these issues, we analyse a large majority of the poorly characterized human TFs1 (defined here as having no confidently known binding motif) and several dozen previously studied control TFs7,8 using a panel of assays that provide different perspectives on DNA sequence specificity. We refer to this international collaborative project as the ‘Codebook–GRECO-BIT Collaboration’. The reagent set and laboratory experiments were initiated as the ‘Codebook project’, which alludes to the fact that TFs decode individual ‘words’ in the genome. Meanwhile, the Benchmarking Initiative by the Gene Regulation Consortium, GRECO-BIT rooted in GRECO9, was engaged for much of the data analyses.
In this paper, we present an overview of the Codebook data, major outcomes of the study and examples of prevalent phenomena and applications. We discuss the following findings:
Just over half of the 332 putative (that is, Codebook) TFs (177) display DNA sequence specificity, most with the same motifs observed in vitro and in living cells.
The resulting 177 motifs are largely different from one another and from motifs of previously studied human TFs.
Tens of thousands of genomic binding sites for these TFs display evidence of pu