Peer-reviewed work from K01 on differential privacy, synthetic clinical data, multi-omics generation and healthcare reporting standards. Open by default. Methods peer-reviewed and published.
TIM3 is a cross-domain benchmark that evaluates time series foundation models on nuclear fusion, ICU and tropical cyclone data at native resolution and modalities. Regridding to a uniform grid costs up to a factor of 5.6 in error, the cost is architecture-specific, and it changes which model wins.
→A 10.1M-parameter conditional latent diffusion model for 12-lead ECGs that takes age, sex, height, weight and heart rate as numeric conditioning alongside 71 SCP codes. On PTB-XL it tracks the real age distribution to within one standard deviation, follows a specified heart rate with a per-sample MAE of 2.4 bpm and reaches TSTR macro-AUROC 0.885. Removing the demographic axes drops it to 0.561.
→Thirteen sampling policies for training an ECG classifier on synthetic data only, all drawn from one latent diffusion generator over PTB-XL. None beats a naive bootstrap of the real training distribution, and the strongest matches it only within noise. The failures collapse onto three mechanisms: class-distribution distortion, within-class diversity collapse, and label-by-demographic decoupling.
→Across six tabular datasets, DP-SGD synthesis with MLP variational autoencoders has a sharp viability boundary at N/d ≈ 50–300. On Adult, the cost of stricter privacy is sublinear: ε = 1 needs about 2.5× more data than ε = 10. Marginal-based DP methods can be viable two orders of magnitude lower.
→On a CITE-seq PBMC dataset, grouping single-cell features into biological modules substantially beats flat baselines at matched parameter budgets. A Tensor-Train coupling adds a modest gain over dense modular coupling. Preliminary results from one dataset, three seeds.
→Three autoregressive conditional mean models on PhysioNet 2019 ICU data. Statistical similarity improves with complexity. Cross-feature clinical rules like fever co-occurring with tachycardia are over-produced relative to real rates, and the gap widens with complexity. Reordering features does not move the cross-feature rule.
→A position paper. Healthcare GenAI lacks a standardised reporting framework. We propose a 16-item checklist extending STROBE from epidemiology, covering architecture, data, privacy, bias, clinical relevance, regulatory compliance and supplementary material. First iteration, not yet validated.
→A position paper. Whether a GenAI system is a medical device under the EU MDR comes down to intended use, and Rule 11 sets the class by the consequence of the decisions it informs. We walk eight use cases through it, from synthetic training data to digital twins, and score each on decision-making relevance and therapeutic risk.
→