Toggle light / dark theme

DNA repair enzymes favor specific sequences, shaping mutation patterns in the human genome

When a wound does not heal properly, it leaves a scar. Similarly, mutations—which are permanent changes to genetic code—are often the result of damaged DNA that has not been properly repaired. Mutations can impede the function of genes and lead to disease and aging, but they are also the source of genetic variation, which allows new traits to emerge and facilitates the evolutionary process. Scientists still do not fully understand why some damaged DNA segments are successfully repaired while others are not.

In a new study published in Nature Communications, researchers from the Weizmann Institute of Science succeeded in identifying which DNA sequences and structures are the preferred targets for several of the most important DNA repair enzymes. The findings from the laboratory of Dr. Ariel Afek suggest that these preferences shaped the human genome and could even help explain how cells become cancerous.

Every day, thousands of chemical reactions take place in every living cell, damaging the genome. “When DNA repair systems work properly, they repair most of the damage, but not all of it,” Afek explains. “Therefore, the rate at which mutations accumulate is a balance between the rate of damage and the rate of repair.

PlasmidGPT: A generative framework for plasmid analysis and generation

By training an AI model using 153,208 plasmids from Addgene, Shao et al.’s PlasmidGPT can annotate and classify existing plasmids as well as generate new functional plasmid sequences from DNA “prompts”. While existing plasmid design tools currently surpass PlasmidGPT in sophistication, future architecture and training data augmentations may allow us to automate much of the plasmid design process. I could certainly see this playing a role in high-throughput biological screening methods.


Assembly standards facilitate the construction of functional plasmids (13). Collections such SEVA (14, 15) and CIDAR MoClo (16) include ready-to-use constructs and genetic parts that can be easily assembled, making them valuable tools for microorganism bioengineering and genome editing. Tools such as Cello (17, 18) can design genetic constructs with a high success rate for specific functions, such as computation, across diverse organisms. iBioSim 3 enables the design and modeling of genetic circuits that extend beyond logic circuits (19). However, there is still no computational method capable of harnessing the existing collection of plasmid sequences for designing the full spectrum of plasmids, such as those for mammalian expression, bacterial expression, and gateway vectors. Consequently, for many applications, plasmid DNA design remains a labor-intense process that requires manual inspection, annotation, and the combination of functional sequences.

Recently, generative models such generative pretrained transformers (GPTs) (20) have demonstrated remarkable success in modeling human language. Given the similarity of human language and biological sequences such as protein and DNA, researchers have adapted these frameworks to design proteins (21) and, more recently, to generate genomic sequences that contain potentially functional regulatory elements and genes (2224). Despite these advances, it remains an open question whether language models can be leveraged to efficiently design and analyze complex engineered DNA.

Here, we introduce PlasmidGPT, a generative framework for designing and annotating plasmid DNA sequences (Fig. 1A). Our framework is built on a decoder-only transformer model that is pretrained on 153,208 plasmid sequences from Addgene (25), a public repository for engineered DNA sequences. We demonstrate that sequence embeddings generated by PlasmidGPT encode plasmid sequences into a continuous numerical space. These sequence representations facilitate the visualization of research topics across laboratories by capturing sequence-level similarities and variations. Leveraging simple machine learning models trained on these embeddings, PlasmidGPT enables the fast identification of a wide range of high-level plasmid features (vector type, selectable marker, growth strain, and lab of origin) directly from sequence, facilitating plasmid analysis tasks such as functional annotation and provenance tracking. Moreover, PlasmidGPT generates plasmids that have genetic part distributions similar to those of the training sequences. Conditional plasmid generation can be achieved either by providing a user-specified starting sequence or by fine-tuning the model using special tokens that represent specific vector types. Furthermore, we experimentally validated the functionality of two model-generated plasmids in bacterial cells.

Sixteen AI-designed viruses offer a new route against drug-resistant bacteria

In a world first, scientists led by a team from Stanford University have created 16 viable viruses that do not exist in nature and were designed by AI. Their experiment, which is published in Science, could help in the fight against superbugs by allowing researchers to design customized viruses to kill drug-resistant bacteria. Thomas Inglesby and Moritz S. Hanke have published a Perspective piece on the work and its implications in the same edition of the journal.

Artificial intelligence is already helping to speed up drug discovery by analyzing massive genetic data sets, predicting protein shapes and identifying potential medicines. But in this research, the team wanted to see if AI could go a step further by creating an entire functioning genome based on a natural virus template.

SingleCell Studies Advance Understanding of the Genetic and Molecular Basis of Atherosclerosis

Foundational models pretrained on millions of single cells (Geneformer, scGPT, and scBERT) now provide transferable embeddings that can help improve and automate cellular annotation, integration, cross-species mapping, and zero-shot predictions.49–51 While there has been considerable contribution to pretraining with immune and tumor data sets, other cell type–specific data for vascular resident parenchymal cells remain sparse, and the applications in atherosclerosis are still emerging. Therefore, overreliance on early foundational models may lead to mislabeling atherosclerosis-specific subtypes and rare cell states, and this area is still being actively investigated with improved consensus in cell types, such as vascular cells (SMC), forthcoming. While promising, we anticipate that these tools will improve with further fine-tuning and robust vascular tissue validation, and interpreted with pathway and genetics-based constraints. Currently, with more modest-sized vascular disease single-cell data sets available, probabilistic variational autoencoder–based methods, such as scVI, MultiVI, and GLUE,44,45,52 offer a good tradeoff between robustness, interpretability, and analytical efficiency. Another exciting area of research is expanding through the development and application of in silico perturbations. While TF perturbation using tools such as CellOracle has provided a first step, improved regulatory network predictions have long been pursued but are still in early stages of implementation.53,54

Deep learning approaches to study TF-DNA interactions are now being utilized to advance our understanding of the DNA regulatory grammar.55 Beyond simple chromatin syntax predictions, deep learning models can provide functional insights, affinity predictions for TF cooperativity, link allelic variation, and chromatin accessibility to cellular epigenetic and transcriptional functions (see below).56 Implementation of a deep learning approach has dramatically extended the capability of scATAC-seq to identify at single basepair resolution the TF motifs that are functional in a cell-specific context to modulate chromatin accessibility, TF binding, and gene expression (Figure 2). Furthermore, allelic variants that are identified with this method are highly enriched among those associated with the complex human traits and diseases that are being investigated. ChromBPNet is a fully convolutional neural network that uncovers the genomic grammar at dynamic enhancers in loci of interest.

Telomere-to-telomere brown rat genome could sharpen disease research models

Researchers have created the most complete genetic profile of the brown rat to date, according to a UTHealth Houston-led team, paving the way for scientists to more accurately investigate genetic links to conditions like heart disease, kidney disease, high blood pressure and stroke.

The research, published in Cell Genomics, was led by corresponding author Peter Doris, Ph.D., director of the Center for Human Genetics at The Brown Foundation Institute of Molecular Medicine within McGovern Medical School at UTHealth Houston.

The assembly of the brown rat’s genome provides a complete genetic fingerprint and reveals that the brown rat’s DNA is more complex than scientists previously understood. In addition to uncovering more than 60 new genes, many of which were previously difficult to sequence and are thought to play a role in immunity and other biological processes, the team discovered that brown rat sex chromosomes differ significantly from those in humans.

Digit regeneration in mice is stimulated by sequential treatment with FGF2 and BMP2

Basically whole body regeneration is definitely possible we just need to right genetic code to push the regeneration button in the human body much like how these mice had their digits regenerated so too we can regenerate just like Deadpool or even the axolotl.


Wound fibrosis after amputation in mammals is replaced with regeneration of amputated structural elements by sequential FGF2/BMP2 treatment. Regenerated tissues include phalangeal/sesamoid bones, tendon/ligament, synovial joint, articular cartilage.

From fragment to form: wholebody regeneration in a model urochordate Medicine

Whole body regeneration is possible just would need to find it in human beings similar genetics.


Rinkevich, Y., Rinkevich, B. From fragment to form: whole-body regeneration in a model urochordate. npj Regen Med 10, 36 (2025). https://doi.org/10.1038/s41536-025-00423-0

Download citation.

What Can We Learn From the Most Genetically Modified Human Alive? | Liz Parrish

In 2015, Liz Parrish flew to Colombia and let someone inject an untested gene therapy into her body 150 times. She texted her kids that she loved them before it started. When it was over, she went for nachos.

Ten years later, her telomeres are longer than when she started, and she has taken 12 gene therapies in total.

Her decade of self-experimentation has produced peer-reviewed data, a growing protocol of gene therapies, and a company training its sights on making biological aging optional. Our conversation goes over what it took to become patient zero and what gene therapy for aging looks like in practice today, among other things.

/* */