Back to Research papers
Research paper index

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

William Poulett, Alice Waterhouse, Ben Wallace, Scarlett Kynoch, Amaia Imaz Blanco, Michael Spence, Jonathan Pearson

arXiv:2606.26879Published June 25, 2026Updated June 26, 20260 citations
  • cs.AI

Abstract

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.