Scale-up Unlearnable Examples Learning with High-Performance Computing

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

The rapid evolution of artificial intelligence (AI) systems, exemplified by architectures like ChatGPT, has raised significant concerns about data privacy in healthcare applications. Modern AI models frequently retain user interactions, creating potential vulnerabilities where sensitive medical imaging data - such as those processed by AI-driven diagnostic tools in radiology - could be inadvertently stored and repurposed for model training without explicit consent. This paradigm poses critical challenges to patient confidentiality, institutional intellectual property rights, and regulatory compliance in clinical environments.

To combat unauthorized data exploitation in machine learning, Unlearnable Examples (UEs) have emerged as a promising defense mechanism by systematically degrading model training through optimized perturbations. Among these approaches, Unlearnable Clustering (UC) has demonstrated particular promise through cluster-wise perturbation strategies that enhance protection efficacy with larger batch sizes. However, previous implementations have been constrained by computational limitations, typically operating on single workstations that restrict batch size scalability and dataset diversity. In this study, we present a groundbreaking implementation of UC leveraging High-Performance Computing (HPC) resources through Distributed Data Parallel (DDP) training on the Oak Ridge National Laboratory's Summit supercomputer. Our scaled implementation enables unprecedented batch sizes and comprehensive experiments across multiple medical and natural image datasets, including Oxford-IIIT Pets, MedMNIST, Flowers102, etc.

Through systematic evaluation, we demonstrate that UE effectiveness exhibits complex relationships with batch size configuration. Experiments conducted on datasets such as Pets, MedMNist, Flowers, and Flowers102 reveal that both overly large and overly small batch sizes can lead to performance instability. Furthermore, the optimal batch size is dataset-specific, highlighting the need for tailored strategies to maximize data protection. These findings establish practical guidelines for deploying UEs at scale and provide empirical evidence that HPC-enabled perturbation strategies can significantly enhance healthcare data security in AI applications. The complete implementation is openly available at https://github.com/hrlblab/UE_HPC to support reproducibility and clinical adoption.

Description

Keywords

Unlearnable Examples, Unlearnable Clusters, High Performance Computing

Citation

Endorsement

Review

Supplemented By

Referenced By