Monthly briefing

August 2026

August 1, 2026–August 31, 2026

AI & Machine Learning

CellTypeAI – cell annotation for scRNA-seq using local generative-AI

Single-cell RNA sequencing, or scRNA-seq, allows researchers to examine gene activity in individual cells rather than averaging signals across thousands or millions of cells. This level of detail can reveal important differences between cell populations, but it also creates a major analytical challenge, determining exactly what type of cell each sequenced cell represents. Researchers commonly identify cell types by looking for marker genes, genes known to be associated with particular types of cells. For example, a specific combination of genes may indicate that a cell is a certain type of immune cell. However, marker gene expression is not always consistent. It can change between individuals, tissues, diseases, and experimental conditions, potentially making traditional cell annotation methods less reliable. Researchers at the Lydia Becker Institute of Immunology and Inflammation at The University of Manchester developed a new computational tool called CellTypeAI to address this problem. CellTypeAI uses generative artificial intelligence to help identify cell types within single-cell RNA sequencing datasets.

Machine learning approaches for biomarker discovery using single-cell RNA sequencing

The application of single-cell RNA sequencing (scRNA-seq) for biomarker discovery promises unprecedented resolution in identifying potential biomarkers by capturing and analysing cellular heterogeneity. Traditionally, biomarker discovery efforts within single-cell transcriptomics have primarily relied on conventional statistical approaches, particularly through the application of differential gene expression analysis, to identify candidate biomarkers. However, in recent years, with the rapid advancement and growing popularity of artificial intelligence and machine learning, their application in scRNA-seq biomarker discovery has become increasingly prominent. Currently, machine learning-based approaches for scRNA-seq biomarker discovery exhibit considerable methodological diversity, which can be distinguished by factors such as the level of discovery, choice of supervised learning algorithm, feature selection methods, classification metrics, and downstream biological analyses. This review provides a comprehensive overview of the current landscape of machine learning methods for scRNA-seq biomarker discovery, offering researchers a complete and detailed understanding of the field.

Novel AI model trained on RNA-Seq data accurately detects key gene mutations and predicts biomarkers across 32 cancer types

Accurate molecular profiling from routine histopathology slides could transform clinical oncology. A Vision Transformer (ViT)–based model was developed to jointly predict the TP53 biomarker, detect 32 solid tumor types, and predict survival directly from whole slide images (WSIs). Over 11,000 primary tumor data were retrieved from the Pan-Cancer Atlas, along with corresponding somatic mutation, RNA-sequencing, and clinical outcome data. WSIs underwent tissue masking, quality control, stain normalization, patch extraction, and feature embedding using a ViT encoder. Seven task heads were developed to generate predictions for cancer type, TP53 mutation status, TP53 RNA expression levels, overall survival, progression-free interval, and their corresponding event times. Model training proceeded in two stages: initial training on tumor-only patches at multiple magnifications, followed by fine-tuning on WSIs using a content-aware strategy. Model performance was evaluated on an independent validation set of 1729 slides using evaluation metrics, including the area under the receiver operating characteristic curve, regression metrics, and the concordance index. An area under the receiver operating characteristic curve of 0.766 was found for TP53 mutation detection on an independent validation set across 32 human solid tumors. In conclusion, the ViT-based model could simultaneously infer TP53 mutation status, TP53 RNA expression levels, and tumor taxonomy directly from WSIs, supporting the existence of reproducible morphologic correlates of TP53 alterations across human cancers, whereas prognostic risk prediction remained limited.

AI Agents in Science: What Are AI Agents, and How Are They Being Used in Scientific Research?

With Artificial Intelligence (AI) rapidly expanding, it seems as if there is a constant flow of news reports, advertisements and social media posts promoting the development of a new tool that promises to accelerate or automate some aspect of our lives. The same is true in healthcare and the sciences, where AI technologies are changing the way that researchers and clinicians think about the future of biomedical research.

agentsAI

Agentic benchmarks on messy, real-world genomics tasks

Each problem includes a snapshot of real experimental data taken immediately prior to a target decision or analysis step, a description of the task through a high-level scientific lens, and a deterministic grader (e.g., Jaccard similarity of sets) that evaluates recovery of the key biological result in a verifiable manner. The benchmark is designed to test durable biological reasoning rather than method-specific implementation details and require empirical interaction with the data.

aibenchmarksbioinformatics

Anthropic courses

Claude 101 Learn how to use Claude for everyday work tasks, understand core features, and explore resources for more advanced learning on other topics. Claude Code 101 Learn how to use Claude Code effectively in your daily development workflow. Claude Platform 101 This course teaches developers to build on the Claude Platform from the ground up, whether you've made a few API calls or have only used Claude through a chat window. Introduction to Claude Cowork Learn to work alongside Claude on your real files and projects. This hands-on course covers the Cowork task loop, plugins and skills, file and research workflows, and how to steer multi-step work responsibly — so you're productive in your first week. Claude Code in Action Run long, hands-off Claude Code sessions you can trust: steer, configure, automate, and verify AI Fluency: Framework & Foundations Learn to collaborate with AI systems effectively, efficiently, ethically, and safely Building with the Claude API This comprehensive course covers the full spectrum of working with Anthropic models using the Claude API Introduction to Model Context Protocol Learn to build Model Context Protocol servers and clients from scratch using Python. Master MCP's three core primitives—tools, resources, and prompts—to connect Claude with external services AI Fluency for educators This course empowers faculty, instructional designers, and educational leaders to apply AI Fluency into their own teaching practice and institutional strategy. AI Fluency for students This course empowers students to develop AI Fluency skills that enhance learning, career planning, and academic success through responsible AI collaboration. Model Context Protocol: Advanced Topics Discover advanced Model Context Protocol implementation patterns including sampling, notifications, file system access, and transport mechanisms for production MCP server development. Claude with Amazon Bedrock As part of an accreditation program created for AWS, Anthropic launched a first-of-its-kind training for AWS employees. Here's the full course so you can follow along. Claude on Google Cloud This comprehensive course covers the full spectrum of working with Anthropic models on Google Cloud.

learning courses

Anthropic’s Complete Guide to Claude Skills Building

This guide covers the complete picture: what skills are technically, how to plan and design them, the exact file structure and naming rules, how to write instructions that Claude follows reliably, a complete working skill built from scratch, how to test and distribute, and what to do when things go wrong.

aiclaudeskills

How Good is Opus 5 at Biology?

Opus 5 results across our biosecurity, therapeutics, and -omics benchmarks are now live on benchmarks.bio, to generally positive results. I’ll share high level numbers first, and then a deeper dive into model and harness quirks we noticed across some of the 4,674 trajectories we generated.

aibenchmarksclaude

Deep learning improves cell cycle prediction from single-cell RNA sequencing

Single-cell RNA sequencing has given researchers an unprecedented view of how individual cells behave. One important piece of information scientists often want to know is where each cell is in the cell cycle, the series of stages that cells pass through as they grow and divide. Accurately identifying these stages is important because the cell cycle strongly influences gene expression. If researchers do not account for these differences, cell cycle activity can interfere with the interpretation of RNA sequencing data and make it more difficult to identify meaningful biological changes.

cell cycledeep learning

Software & Tools

Benchmarking long-read RNA sequencing tools for single-cell and spatial transcriptomics

Alternative splicing plays a crucial role in transcriptomic complexity, yet remains difficult to resolve at the single-cell level due to the limitations of short-read technologies. Coupling single-cell with long-read sequencing offers full-length transcript coverage, enabling more accurate isoform detection. Diverse computational tools tailored for single-cell and spatial long-read transcriptomics have been developed. To compare the effectiveness of these approaches, we generated paired short-read and Nanopore long-read single-cell datasets, tailored for benchmarking bioinformatics tools. We evaluated ten state-of-the-art methods, spanning four analytical dimensions: barcodes and unique molecular identifiers (UMI) detection, demultiplexing and UMI clustering, gene-level expression profiling, and isoform detection and quantification. Using real and simulated datasets across different protocols, sequencing depths and chemistries, we assessed the accuracy, robustness, and scalability of each tool. Our results revealed method-specific trade-offs, and highlight the importance of sequencing quality and UMI correction strategies. This benchmark provides a practical resource for optimizing isoform analysis and accurate gene expression profiling in single-cell and spatial transcriptomics using long-read sequencing. The workflow employed for benchmarking is designed to be reusable, thereby enabling method developers to compare their own approaches against the set of reference methods evaluated in this work.

MIRACLE keeps single-cell atlases up to date with continual learning

Single-cell sequencing has transformed our understanding of cellular heterogeneity, enabling the construction of multi-omics atlases through data integration. However, conventional atlas updates require full reintegration of all datasets, creating scalability challenges that limit the timeliness and adaptability of biomedical research. Here we present multimodal integration with continual learning (MIRACLE), an online learning framework for scalable multimodal integration. Using dynamic architecture adaptation and data rehearsal, MIRACLE continually integrates diverse datasets while preserving biological fidelity. Across evaluations, MIRACLE achieves accurate online integration with substantially improved efficiency, refining and expanding atlases with new cross-modal, cross-tissue and cross-disease data. Applied to respiratory infections, it reveals both shared and pathogen-specific immune mechanisms in coronavirus disease 2019, influenza A and tuberculosis. Overall, MIRACLE provides an efficient and collaborative solution for the continual integration, sharing and exploration of biological knowledge.

Tidesurf improves RNA velocity analysis for modern single-cell RNA sequencing

RNA velocity enables predicting future cellular states from single-cell RNA-sequencing (scRNA-seq) data by inferring the derivative of gene expression from separately quantified spliced and unspliced transcripts. Although the original implementation, velocyto, established the foundation for such analyses and remains widely used, it has not been updated to accommodate major advances in scRNA-seq technologies. Notably, some modern 10x Genomics protocols capture transcripts from the 5’ end rather than the 3’ end, introducing reversed transcript orientations and other protocol-specific features not implemented in velocyto. We demonstrate that velocyto systematically assumes opposite transcript direction compared to 10x Genomics Cell Ranger in 5’-sequencing data, leading to different count assignments than for other tools, substantial deviations in inferred velocities, and ultimately divergent biological interpretations. To address these shortcomings, we present tidesurf, a command-line tool designed for accurate quantification of spliced and unspliced molecules across both 3’- and 5’-based scRNA-seq protocols. By evaluating tidesurf on four publicly available 10x Genomics Chromium datasets and comparing it to state-of-the-art quantification approaches, we show that it reliably recovers correct transcript counts in settings where velocyto performs well (3’ chemistry) and where it produces unexpected results (5’ chemistry). These results underscore that the validity of RNA velocity analyses critically depends on reliable splicing-state quantification and that continued use of velocyto with 5’-sequencing data is inadvisable. Tidesurf provides a robust, up-to-date alternative that preserves the interpretability and reliability of RNA velocity across diverse experimental designs.

Ultrafast and reference-free sequence discovery in single-cell data

Knowledge of RNA sequences, expression, splicing, isoforms, structure and modifications is central for understanding and targeting cellular processes. Revolutionary single-cell and spatial transcriptomics technologies—for example, as deployed by consortia such as the Human Cell Atlas—partially capture this diversity and generate cellular profiles that expand at petabyte scale each year. Yet researchers cannot search sequences across these datasets: standard pipelines do not scale or rely on references, retaining only gene or isoform counts, whereas accessing raw sequences requires collecting, downloading and processing millions of large files. Here we present Malva, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts. The continuously expanding Malva Index currently comprises around 74 million cells from thousands of experiments in health and disease. Malva enables reference-free discovery—researchers can, for example, identify cell types and predict cell–cell similarity directly from sequence composition. Building on Malva’s speed and accuracy, we demonstrate how Malva can be flexibly connected to state-of-the-art neural networks and how to execute complex searches and enable automated analyses. Malva transforms single-cell atlases from static gene count tables into dynamic, sequence-resolved resources that may help to bridge human–machine reasoning about biology.

Python Polars: The Definitive Cheatsheet

Polars is a library for transforming, analyzing, and visualizing data with a fast and expressive DataFrame API. It was first released by Ritchie Vink in 2020. Consider it as a potentially faster and different philosophy to pandas.

polarspython

A systematic benchmark of bioinformatics methods for single-cell and spatial RNA-seq nanopore long reads data

Alternative splicing plays a crucial role in transcriptomic complexity, yet remains difficult to resolve at the single-cell level due to the limitations of short-read technologies. Coupling single-cell with long-read sequencing offers full-length transcript coverage, enabling more accurate isoform detection. Diverse computational tools tailored for single-cell and spatial long-read transcriptomics have been developed. To compare the effectiveness of these approaches, we generated paired short-read and Nanopore long-read single-cell datasets, tailored for benchmarking bioinformatics tools. We evaluated ten state-of-the-art methods, spanning four analytical dimensions: barcodes and unique molecular identifiers (UMI) detection, demultiplexing and UMI clustering, gene-level expression profiling, and isoform detection and quantification. Using real and simulated datasets across different protocols, sequencing depths and chemistries, we assessed the accuracy, robustness, and scalability of each tool. Our results revealed method-specific trade-offs, and highlight the importance of sequencing quality and UMI correction strategies. This benchmark provides a practical resource for optimizing isoform analysis and accurate gene expression profiling in single-cell and spatial transcriptomics using long-read sequencing. The workflow employed for benchmarking is designed to be reusable, thereby enabling method developers to compare their own approaches against the set of reference methods evaluated in this work.

Introducing the Warp Agent CLI: a CLI coding agent that does what others can't

Today we are excited to launch the Warp Agent CLI, a new standalone CLI that lets you use the Warp Agent anywhere. It’s the same multi-model agent that’s built into Warp Terminal, now available in Ghostty, iTerm 2, VSCode, the built-in Windows terminal, or whatever terminal you prefer. The Warp Agent CLI has everything you’d expect from a modern CLI coding agent: it’s a multi-model, cost-optimizing harness built for pro developers. Out of the box, you get access to frontier and open-weight models and auto-routing based on task complexity.

aicli

Posit Public Package Manager R and Python packages, built for speed

Pre-built Linux binaries, date-based snapshots, and Bioconductor support. No account required. No cost. Pre-built R binaries Install R packages in seconds. Posit Public Package Manager provides binaries for Linux, macOS, and Windows, so you skip compilation errors. Linux support spans 10 distributions with ARM64 and portable manylinux for additional distros. Snapshot URLs Pin your environment to a specific date. Snapshot URLs freeze packages to the versions available on that date, so you can reproduce your work months or years later. Vulnerability visibility Package detail pages show known CVEs from the Open Source Vulnerabilities (OSV) database. See what vulnerabilities affect a package before you install it. Bioconductor support Life sciences and genomics packages with date-based snapshots for reproducibility, hosted alongside CRAN and PyPI. Pin Bioconductor to the same snapshot date as your CRAN packages.

R

Scientific computing in the age of agentic AI: an exploratory field report

Scientific computing has become a central component of modern scientific discovery. Yet many computational tools are developed by small, specialized teams under incentives that encourage the release of rapidly prototyped tooling without commensurate attention to engineering concerns, including performance and maintainability. These gaps are particularly visible in the life sciences, where the advent of high-throughput sequencing and molecular profiling has made the production and processing of datasets routine at scales that strain reliability and cost. Recently, LLM-based agents have become increasingly capable, with publicly available systems possessing both significant domain knowledge in many scientific fields and the ability to autonomously operate over complex and specialized codebases in pursuit of well-defined goals. Together, these developments create a practical opportunity for scientific computing. Many of the persistent weaknesses of the scientific computing ecosystem stem from technical debt and a shortage of sustained engineering labor and expertise. Here, we examine coding agents as a potential way to address these weaknesses: we present an exploratory field report of eight early case studies in the application of LLM agents to scientific computing across a range of project scopes, from lightweight maintenance tasks to full performance-oriented rewrites of scientific libraries, with a focus on the life sciences. Each of these case studies is accompanied by reflections from the individual or group responsible for the work, including lessons from the process. Overall, we find that the use of coding agents in scientific computing holds great promise for accelerating scientific research and increasing the reliability of critical systems, but that outstanding concerns remain, including responsibility and ownership for such projects, and we suggest collaboration and stewardship with existing maintainers when feasible.

aibioinformaticsllm

renv 1.2.4 released

renv 1.2.4 focuses on improving reliability, reproducibility, and usability across the package management workflow. The release introduces more robust package installation and dependency resolution, significantly better diagnostics and error messages, safer shared cache behavior, expanded lockfile capabilities, and tighter integration with pak, while fixing a number of edge cases affecting package restoration, Bioconductor, and development workflows. Its getting faster too! Linux binaries hosted by Posit.

Rrenvreproducibility

Cancer Research

Genetic background sets the trajectory of experimental cancer evolution

Human cancers are heterogeneous1. Dissecting how germline genetic variation and environmental factors shape tumour evolution using human datasets is limited by inherent diversity in genetic backgrounds2 and environmental exposures3,4,5. Here, to overcome these limitations, we re-ran early tumour evolution hundreds of times in diverged inbred mouse strains, generating matched histology and whole-genome and transcriptome sequences. The sex, environment and carcinogenic exposures were all controlled, and the study design allowed us to capture genetic variation comparable with that observed across human populations while exploiting the nested hierarchical structure of strain–litter–animal–tumour relationships. Our analyses reveal that epistatic interactions between genetic background and acquired somatic mutations result in population-specific disease progression, including choice of driver mutations, occurrence of whole-genome duplication and subclonal selection dynamics that mirror both cancer susceptibility and tumour growth rate. Even modest genetic divergence, comparable with that found across human ancestry groups, can strikingly alter selection pressures during cancer development to shape both cancer risk and the trajectory of tumour evolution.

Single-nucleus multimodal spatial transcriptomics reveals spatial colocalization of neoantigen-expressing tumor cells and cognate T cells

Improved methods to identify therapeutically relevant tumor neoantigens and their cognate T cells would aid the development of precision medicines for cancer. Here, we developed Slide-GoTags, a droplet-based single-nucleus spatial transcriptomics approach that characterizes neoantigen-specific immunity by integrating targeted transcript genotyping and T cell receptor (TCR) sequencing with single-nucleus RNA sequencing from the same slice of frozen tissue. Application of Slide-GoTags to mouse and human tumors revealed colocalization of clonally expanded, neoantigen-specific T cells with tumor cells expressing their cognate neoantigen. We also identified distinct spatial immune landscapes shaped by anti-PD1 or anti-CTLA4 blockade in mouse colorectal tumors. Across human tumor types, Slide-GoTags detected TCR–neoantigen interactions through spatial proximity and identified an enrichment of interferon-driven immunogenicity niches in immunologically ‘hot’ tumors compared to ‘cold’ tumors. These niches harbored three T cell clonotypes that colocalized with genotyped neoantigens, highlighting a spatially organized antitumor immune response. Collectively, Slide-GoTags establishes a framework for in situ mapping of T cell–tumor interactions directly from individual tissue.

Technology Platforms

DRAGEN - Illumina Cloud Analysis Platform

DRAGEN supports an extensive range of applications, providing comprehensive coverage for many experiment types in a single solution. Key applications include: Germline (whole-genome and enrichment) Somatic (whole-genome and enrichment, tumor-only and tumor-normal) RNA Methylation DRAGEN includes a versatile set of pipelines that can accept input data files and create output files at different stages of the pipelines.

Illumina 5-Base DNA Prep

Illumina 5-Base DNA Prep is a whole-genome library preparation kit that simultaneously detects standard DNA bases (A, T, G, C) and modified 5-methylcytosine (5mC) in a single workflow. Using a one-step enzymatic process rather than bisulfite treatment to convert 5mC to thymine, it preserves sequence complexity and alignment efficiency across low-input samples like cell-free DNA (1–20 ng) and genomic DNA (50–100 ng). The library prep takes under 8 hours, and when paired with NovaSeq sequencers and DRAGEN secondary analysis pipelines, it enables integrated calling of single nucleotide variants, indels, copy number variations, and differential methylation within 3 days.

Illumina Single Cell 3' RNA Prep

llumina Single Cell 3' RNA Prep is a microfluidics-free single-cell RNA sequencing (scRNA-Seq) technology based on PIPseq chemistry, which utilizes particle-templated instant partitions to capture and barcode single-cell mRNA via simple vortex mixing. The workflow enables cell capture and library preparation from inputs ranging between 100 and 200,000 cells (or single nuclei) without requiring dedicated microfluidic instruments. It captures the coding transcriptome via 3' mRNA sequencing, supports multiplexing up to 96 samples per run, and is compatible with various fixed or fresh sample types, producing stranded cDNA libraries for downstream sequencing on Illumina platforms and analysis through DRAGEN pipelines.

Illumina SomaSeq Discovery

An automated proteomics solution using SOMAmer technology and Illumina sequencing for accurate and reproducible quantification of 9.5K human proteins in 2.5 days

Illumina TruPath Genome

The Illumina TruPath Genome kit is a human whole-genome sequencing assay designed for germline variant detection using on-flow cell tagmentation, bypassing traditional manual library preparation. Compatible exclusively with the NovaSeq X Series (running software v1.4 or higher), the workflow requires 350 ng of input DNA and takes approximately 10 minutes of hands-on preparation time, with a total assay and sequencing run time of roughly 29 hours. It utilizes proximity mapped read technology—combining standard short reads with physical proximity data from clusters across nearby flow cell nanowells—to preserve spatial linkage from large DNA templates. Secondary analysis via DRAGEN Germline uses these linked spatial signals to improve read alignment, assist with ultralong phasing, resolve difficult-to-map region sequences, and identify single nucleotide variants (SNVs), indels, copy number variants (CNVs), and structural variants.

Illumina StrataMap Spatial Transcriptome solution

The Illumina StrataMap Spatial Transcriptome solution delivers unbiased whole‑transcriptome spatial profiling at true single-cell resolution—combining high sensitivity, large tissue coverage, and seamless integration with Illumina sequencing and analysis to reveal tissue biology in its native context. StrataMap Spatial Transcriptome reveals biology that has not been seen before, with sensitivity, scale, and real-time cellular insights.

illuminaspatial