CompDiff: demographically conditioned medical image generation

CompDiff is a Stable Diffusion 2.1 model with a compositional conditioner that takes sex, race and age as typed attributes, so it can generate images for demographic intersections that are rare or missing in the training data. The text prompt carries only the clinical findings.

Paper · Project page · Code · Chest X-ray model · Fundus model

Research use only. Not for clinical use. Every image here is synthetic, produced by a generative model. It does not show a real patient and must not be used for diagnosis, screening, treatment or any clinical decision. Images can contain artefacts and may not represent pathology faithfully. Race categories are the datasets' self-reported labels, used as a conditioning signal, not a biological ground truth.

Trained on MIMIC-CXR posteroanterior radiographs of adults. Sex, race and age all go through the conditioner; age is continuous. Write findings only, as in a radiology impression. An empty prompt is replaced by "Normal chest radiograph", as in training.

Example findings
Sex
Race / ethnicity
18 100
1 4
10 100
1 15

Model weights: CreativeML OpenRAIL++-M (inherited from Stable Diffusion 2.1-base). If you use CompDiff, please cite the paper.