Foundational models pretrained on millions of single cells (Geneformer, scGPT, and scBERT) now provide transferable embeddings that can help improve and automate cellular annotation, integration, cross-species mapping, and zero-shot predictions.49–51 While there has been considerable contribution to pretraining with immune and tumor data sets, other cell type–specific data for vascular resident parenchymal cells remain sparse, and the applications in atherosclerosis are still emerging. Therefore, overreliance on early foundational models may lead to mislabeling atherosclerosis-specific subtypes and rare cell states, and this area is still being actively investigated with improved consensus in cell types, such as vascular cells (SMC), forthcoming. While promising, we anticipate that these tools will improve with further fine-tuning and robust vascular tissue validation, and interpreted with pathway and genetics-based constraints. Currently, with more modest-sized vascular disease single-cell data sets available, probabilistic variational autoencoder–based methods, such as scVI, MultiVI, and GLUE,44,45,52 offer a good tradeoff between robustness, interpretability, and analytical efficiency. Another exciting area of research is expanding through the development and application of in silico perturbations. While TF perturbation using tools such as CellOracle has provided a first step, improved regulatory network predictions have long been pursued but are still in early stages of implementation.53,54
Deep learning approaches to study TF-DNA interactions are now being utilized to advance our understanding of the DNA regulatory grammar.55 Beyond simple chromatin syntax predictions, deep learning models can provide functional insights, affinity predictions for TF cooperativity, link allelic variation, and chromatin accessibility to cellular epigenetic and transcriptional functions (see below).56 Implementation of a deep learning approach has dramatically extended the capability of scATAC-seq to identify at single basepair resolution the TF motifs that are functional in a cell-specific context to modulate chromatin accessibility, TF binding, and gene expression (Figure 2). Furthermore, allelic variants that are identified with this method are highly enriched among those associated with the complex human traits and diseases that are being investigated. ChromBPNet is a fully convolutional neural network that uncovers the genomic grammar at dynamic enhancers in loci of interest.