K01
Research

Publications.

Peer-reviewed work from K01 on differential privacy, synthetic clinical data, multi-omics generation and healthcare reporting standards. Open by default. Methods peer-reviewed and published.

ICML 2026 · FMSD Workshop · July 2026

Towards Benchmarking Time Series Foundation Models on Native Scientific Data

TIM3 is a cross-domain benchmark that evaluates time series foundation models on nuclear fusion, ICU and tropical cyclone data at native resolution and modalities. Regridding to a uniform grid costs up to a factor of 5.6 in error, the cost is architecture-specific, and it changes which model wins.

ICML 2026 · SD4H Workshop · July 2026

DCDM-ECG: Demographic-Conditional Diffusion Model for 12-Lead ECG Generation

A 10.1M-parameter conditional latent diffusion model for 12-lead ECGs that takes age, sex, height, weight and heart rate as numeric conditioning alongside 71 SCP codes. On PTB-XL it tracks the real age distribution to within one standard deviation, follows a specified heart rate with a per-sample MAE of 2.4 bpm and reaches TSTR macro-AUROC 0.885. Removing the demographic axes drops it to 0.561.

ICML 2026 · SD4H Workshop · July 2026

Empirical-Distribution Matching for Synthetic ECG Classification

Thirteen sampling policies for training an ECG classifier on synthetic data only, all drawn from one latent diffusion generator over PTB-XL. None beats a naive bootstrap of the real training distribution, and the strongest matches it only within noise. The failures collapse onto three mechanisms: class-distribution distortion, within-class diversity collapse, and label-by-demographic decoupling.

ICLR 2026 · DATA-FM Workshop · March 2026

The Viability Boundary of Differential Privacy

Across six tabular datasets, DP-SGD synthesis with MLP variational autoencoders has a sharp viability boundary at N/d ≈ 50–300. On Adult, the cost of stricter privacy is sublinear: ε = 1 needs about 2.5× more data than ε = 10. Marginal-based DP methods can be viable two orders of magnitude lower.

ICLR 2026 · Gen² Workshop · March 2026

Tensorised Modular Architectures for Multi-Omics Generation

On a CITE-seq PBMC dataset, grouping single-cell features into biological modules substantially beats flat baselines at matched parameter budgets. A Tensor-Train coupling adds a modest gain over dense modular coupling. Preliminary results from one dataset, three seeds.

EurIPS 2025 · ML4H Workshop · December 2025

Multimodal Alignment for Synthetic Clinical Time Series

Three autoregressive conditional mean models on PhysioNet 2019 ICU data. Statistical similarity improves with complexity. Cross-feature clinical rules like fever co-occurring with tachycardia are over-produced relative to real rates, and the gap widens with complexity. Reordering features does not move the cross-feature rule.

NeurIPS 2024 · GenAI4H Workshop · December 2024

Transparent Reporting for Healthcare GenAI

A position paper. Healthcare GenAI lacks a standardised reporting framework. We propose a 16-item checklist extending STROBE from epidemiology, covering architecture, data, privacy, bias, clinical relevance, regulatory compliance and supplementary material. First iteration, not yet validated.

NeurIPS 2024 · GenAI4H Workshop · December 2024

Classifying GenAI under the European Union’s Medical Device Regulation

A position paper. Whether a GenAI system is a medical device under the EU MDR comes down to intended use, and Rule 11 sets the class by the consequence of the decisions it informs. We walk eight use cases through it, from synthetic training data to digital twins, and score each on decision-making relevance and therapeutic risk.