CompDiff: demographically conditioned medical image generation
CompDiff is a Stable Diffusion 2.1 model with a compositional conditioner that takes sex, race and age as typed attributes, so it can generate images for demographic intersections that are rare or missing in the training data. The text prompt carries only the clinical findings.
Paper · Project page · Code · Chest X-ray model · Fundus model
Research use only. Not for clinical use. Every image here is synthetic, produced by a generative model. It does not show a real patient and must not be used for diagnosis, screening, treatment or any clinical decision. Images can contain artefacts and may not represent pathology faithfully. Race categories are the datasets' self-reported labels, used as a conditioning signal, not a biological ground truth.
Trained on MIMIC-CXR posteroanterior radiographs of adults. Sex, race and age all go through the conditioner; age is continuous. Write findings only, as in a radiology impression. An empty prompt is replaced by "Normal chest radiograph", as in training.
Trained on FairGenMed scanning laser ophthalmoscopy fundus images. This is the model as currently published on the Hub (first release): sex and race go through the conditioner and age goes through the text prompt as "<age> years old.". It has three race classes. The findings are a fixed list of glaucoma-related labels; free text is out of distribution.
Model weights: CreativeML OpenRAIL++-M (inherited from Stable Diffusion 2.1-base). If you use CompDiff, please cite the paper.