Skip to content

LDG-33: Add LLM membership inference attack (ez-mia) - #454

Draft
Muhaddisabarat wants to merge 8 commits into
mainfrom
llm-attack-ez-mia
Draft

Muhaddisabarat wants to merge 8 commits into
mainfrom
llm-attack-ez-mia

Conversation

@Muhaddisabarat

@Muhaddisabarat Muhaddisabarat commented Sep 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Ports ez-mia (an error-zone membership inference attack for causal LLMs) into the LeakPro repo as leakpro/llm_attacks/mia, a new top-level attack category for text/LLM-targeted attacks (sibling to attacks/ and synthetic_data_attacks/).
  • Stays standalone rather than plugging into AbstractMIA/AttackFactoryMIA: this attack trains its own target LLM and a single reference model (base/distillation/SFT) and scores membership from per-token log-probabilities, which doesn't fit LeakPro's shadow-model/logit-caching classifier machinery.
  • Swaps ez-mia's own HPU/CUDA/CPU device detection for this branch's leakpro.utils.device.get_device()/mark_step(), so it follows the same device-selection convention as the rest of the codebase (hence targeting this branch rather than main).
  • Adds examples/mia/llm_mia/main.ipynb + train_config.yaml, giving it the same open-and-run experience as the other examples/mia/* notebooks.
  • Adds an llm-mia optional-dependency group in pyproject.toml for its HF/PEFT/OmegaConf stack, and excludes the module from ruff's strict house-style lint pass (same treatment as leakpro/webapp) since it keeps ez-mia's own code style rather than being rewritten to LeakPro's ANN/docstring conventions.

Test plan

  • Imports resolve and python -m leakpro.llm_attacks.mia --help works
  • ruff check leakpro --exclude examples,leakpro/tests,leakpro/webapp,leakpro/llm_attacks passes
  • End-to-end run verified on physical Habana HPU hardware (small-scale config): dataset download, target model fine-tuning, reference model construction, EZ-score computation, and AUC/TPR reporting all work
  • examples/mia/llm_mia/main.ipynb verified end-to-end on CPU and HPU
  • CUDA not independently verified — no NVIDIA GPU available in this environment. No CUDA-specific code changed (device selection is entirely leakpro.utils.device.get_device(); the pin_memory/empty_device_cache CUDA branches are unchanged from ez-mia's original code), and leakpro's own test_device.py exercises the CUDA path via mocking, but this has not been run on real GPU hardware
  • Instrumented HPU memory-stats run confirms no leak across training/eval phases
  • 49 targeted existing tests pass (including the new HPU device tests); full-suite collection succeeds except one pre-existing, unrelated HPU-contention failure in syn_text_pii_scanner

🤖 Generated with Claude Code

Muhaddisabarat and others added 7 commits September 1, 2026 09:57
…ument ts2vec HPU gap

Continues the HPU integration work: the notebook now picks up the shared
device utility instead of hardcoding CUDA/CPU, and notes that TS2Vec should
be dropped from audit.yaml's signal list on Habana.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds the ez-mia error-zone membership inference attack for causal LLMs as
leakpro/llm_attacks/mia, a new top-level attack category (sibling to
attacks/ and synthetic_data_attacks/) for text/LLM-targeted attacks. It
stays standalone rather than plugging into AbstractMIA/AttackFactoryMIA:
LeakPro's classic MIA attacks audit a pre-trained classifier via shadow
models, while this attack trains its own target LLM and a single reference
model (base/distillation/SFT) and scores membership from per-token
log-probabilities, which doesn't fit that shadow-model/logit-caching
machinery.

Device selection is swapped from ez-mia's own HPU/CUDA/CPU detection to
leakpro.utils.device.get_device()/mark_step(), and RNG seeding is now
based on the selected device (respecting LEAKPRO_DEVICE overrides) rather
than independently probing CUDA/HPU availability. Everything else (CLI,
YAML configs, results.csv output) is unchanged from ez-mia.

Adds an `llm-mia` optional-dependency group for its HF/PEFT/OmegaConf
stack, and excludes the module from ruff's strict house-style lint pass
(same treatment as leakpro/webapp) since it keeps ez-mia's own style
rather than being rewritten to LeakPro's ANN/docstring conventions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
# Conflicts:
#	examples/mia/time_series_mia/main.ipynb
#	pyproject.toml
Gives leakpro.llm_attacks.mia the same open-and-run experience as the
other examples/mia/* notebooks, even though it doesn't go through
Leakpro(...)/audit.yaml: it calls run_attack(cfg) directly, with a small
smoke-test config and a full run loaded from one of the shipped YAML
configs. Verified end to end on this machine (CPU override via
LEAKPRO_DEVICE=cpu, since the physical HPU was held by another process).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Each Habana card is exclusive to one process (confirmed via hl-smi and a
memory_stats() repro: a process reserves ~101GB, the whole card, the
instant it touches the HPU, independent of actual workload size). On this
shared machine, if the notebook's kernel and another job end up on the
same physical card, Habana's runtime can hard-crash the loser mid-run
with no Python traceback ("kernel crashed"). Confirmed this is contention,
not a leak in the port: a controlled repro at comparable batch/sequence
shapes showed flat HPU memory usage (~1-2GB) across training and eval.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Matches the examples/mia/*/train_config.yaml convention used by the other
attacks, though the schema here is leakpro.llm_attacks.mia's own
AttackConfig (dataset/model/training hyperparameters consumed directly
by run_attack), not the target-model-training config those examples use
-- llm_mia doesn't go through that pipeline. The notebook's YAML-config
cell now loads this local file instead of one of the module's bundled
configs, so editing dataset/model/hyperparameters means editing the file
next to the notebook, as with the other examples.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Every field the smoke test set (dataset, ref_variant, model_name, seed,
train_total, eval_total, val_total, epochs, batch_size, sequence_length)
is a real AttackConfig field, so the same fast sanity check is just
train_config.yaml with smaller values -- no need for a second,
hardcoded config path in the notebook. Down to one config-driven flow:
imports -> train_config.yaml -> run_attack. Verified end to end with the
values temporarily lowered, matching the new instructions in the
notebook's markdown cell.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@review-notebook-app

Copy link
Copy Markdown

Check out this pull request on  ReviewNB

See visual diffs & provide feedback on Jupyter Notebooks.


Powered by ReviewNB

@Muhaddisabarat Muhaddisabarat changed the title Add standalone LLM membership inference attack (ez-mia port) Add LLM membership inference attack (ez-mia) Sep 1, 2026
The first cell was unconditionally setting PT_HPU_LAZY_MODE and
HABANA_VISIBLE_MODULES regardless of what hardware is actually present.
Harmless on GPU/CPU-only machines (Habana's runtime just never reads
them), but it read as HPU-first rather than device-agnostic. Now it only
touches Habana-specific setup when habana_frameworks is actually
installed (checked via importlib.util.find_spec, not an import, since
the vars must be set before any torch/habana import to take effect).
The real HPU -> CUDA -> CPU dispatch is unchanged: it still happens via
leakpro.utils.device.get_device() regardless of what this cell does.

Verified both branches directly (habana_frameworks present vs. a
nonexistent package name standing in for "absent", with a cleared
environment to rule out this machine's own exported Habana env vars),
and re-ran the notebook's full pipeline end to end afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@Muhaddisabarat Muhaddisabarat changed the title Add LLM membership inference attack (ez-mia) LDG-33: Add LLM membership inference attack (ez-mia) Sep 1, 2026
Base automatically changed from LDG-13-leakpro-habana to main September 8, 2026 14:16
@fazelehh
fazelehh removed the request for review from TheColdIce September 17, 2026 06:55
@Muhaddisabarat
Muhaddisabarat marked this pull request as draft September 17, 2026 07:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant