Skip to content

About

Leakage-free NIDS benchmarking and adversarial attacks constrained to be physically sendable. Measures what a permissive split protocol is worth, and what feature-space attacks overstate.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

Repository files navigation

evade-nids

Leakage-free intrusion-detection benchmarking, and adversarial attacks that could actually be sent.

CI Python License: MIT

Two claims about published NIDS results, each measured rather than asserted.

The split protocol is doing a lot of the work. Flows from one capture session share an environment: the same background traffic, the same clock, often the same host. Split those flows at random and a model can identify the session, and the session all but gives away the label. It scores brilliantly without having learned anything that transfers.

Adversarial evaluations attack a feature vector, not a network. An attacker does not have a dial marked flow_iat_std. They have packets. Every feature is a statistic of a packet sequence, so a perturbation is only real if some sequence produces it. Most published attacks produce vectors no sequence could: negative packet counts, a mean packet length below the minimum, 3.7 SYN flags.

This measures both gaps on the same data with the same models.

pip install -e ".[dev]"
python -m evade_nids.benchmark --out results/

featureguard is used for schema validation and is not on PyPI, so it is optional:

pip install "git+https://github.com/SohanBag/featureguard.git"

Without it everything runs except FlowDataset.validate(), which says so rather than silently skipping.

What the split protocol is worth, on real traffic

CTU-13 botnet captures. Real traffic, real malware, real labels: 33,013 flows, 56 features, 1,911 attack (5.8%), grouped by source host across 497 groups. One dataset, one set of seeds; only the splitter changes, so any difference is attributable to the protocol and nothing else.

Model Random Grouped by host Temporal Inflation Fold-spread ratio Significant
logistic 0.4221 ± 0.0174 0.2510 ± 0.2873 0.3159 ± 0.3044 +17.1 pts 16.5× no
random forest 0.9650 ± 0.0078 0.7379 ± 0.2739 0.4145 ± 0.4311 +22.7 pts 35.3× no
xgboost 0.9456 ± 0.0095 0.8226 ± 0.0652 0.5624 ± 0.4145 +12.3 pts 6.9× yes

Macro F1, 5 folds × 3 seeds. Full study in reports/cic_ids2017_real_world_audit.md, reproduce with python scripts/run_ctu13.py.

Quote the XGBoost row, not the largest one. Random forest shows the biggest inflation at +22.7 points, and it does not survive its own significance test: fold variance under the correct protocol is so large (± 0.2739) that the difference cannot be distinguished from noise on this sample. Only XGBoost's +12.3 points is significant.

That distinction is the reason inflation_is_significant() exists. An audit tool that reports the biggest number it can find is a worse instrument than one that reports the number it can defend.

The fold-spread ratio is above 1 in every row. The permissive split always looked more stable, whether or not its inflation was significant.

Validating the instrument on a generator

The synthetic generator exists to check the tool against known ground truth, which real data cannot provide: on CTU-13 nobody can say how much leakage there truly is, so a diagnostic cannot be scored there. On the generator, leakage is a parameter.

Model Random Grouped Temporal Inflation Fold-spread ratio
logistic 0.9045 ± 0.0075 0.2127 ± 0.2700 0.2639 ± 0.2508 +69.2 pts 35.9×
random forest 0.9966 ± 0.0008 0.0608 ± 0.1306 0.1850 ± 0.2034 +93.6 pts 155.9×
xgboost 0.9991 ± 0.0009 0.1692 ± 0.2392 0.2481 ± 0.2453 +83.0 pts 275.3×
torch MLP 0.8880 ± 0.0631 0.2180 ± 0.2471 0.3225 ± 0.2153 +67.0 pts 3.9×

Macro F1, 5 folds × 3 seeds, synthetic. Reproduce with python -m evade_nids.benchmark.

These numbers are much larger than the CTU-13 ones because the generator plants leakage deliberately and heavily. They measure the instrument, not any real dataset, and should never be quoted as a finding about network traffic. Their value is the pair of checks they support: with leakage switched on the diagnostic finds it, and with it switched off the diagnostic reports none.

The last column is worth as much as the inflation column. Under the leaky protocol the fold-to-fold spread collapses, by up to 35× on CTU-13 and 275× on the generator. A leaky protocol does not only look more accurate, it looks more reproducible, because every fold recycles the same sessions. So an unusually tight confidence interval is a reason to check the splitter, not a reason to trust the result.

Note that the strongest model on the random split is the most inflated. Capacity is what lets a model memorise sessions, so the ranking under a leaky protocol can inverted relative to what would actually deploy.

The audit knows its own failure mode

Grouped cross-validation is high-variance when groups are few, and that variance can look like inflation when there is none. Measured on the generator with leakage switched off entirely:

Groups Random Grouped Apparent inflation
8 0.7753 ± 0.0243 0.4398 ± 0.5079 +33.6 pts
16 0.8121 ± 0.0209 0.6253 ± 0.4183 +18.7 pts
40 0.8182 ± 0.0177 0.7723 ± 0.1239 +4.6 pts
80 0.8259 ± 0.0247 0.8310 ± 0.0613 −0.5 pts

There is no leakage in any row. With eight groups each test fold holds two, and a score from two sessions is mostly noise. Only by 80 groups does the audit correctly report nothing.

So SplitAudit.inflation_is_significant() compares the gap against the grouped protocol's own fold spread and the report prints NOT SIGNIFICANT when the gap is smaller, and FlowDataset.evaluation_caveats() warns below 50 groups. An audit tool that did not know when it was unreliable would be no better than the protocols it criticises.

What the constraints are worth

fgsm perturbs the feature vector freely. constrained_fgsm projects onto the feasible set after every step. Same model, same budget.

Budget FGSM Constrained FGSM Overstatement
0.05 40.3% 45.2% −4.9 pts
0.10 72.6% 45.2% +27.4 pts
0.25 100.0% 45.2% +54.8 pts
0.50 100.0% 37.1% +62.9 pts

Evasion rate against the PyTorch MLP, 62 correctly-detected attack flows. Mean overstatement across the sweep: +36.3 points.

Every FGSM row is marked infeasible by the constraint checker. Those flows cannot be sent. The constrained attack plateaus around 45% and then declines as the budget grows, because a larger step gets projected back harder: the feasible set, not the budget, is what binds.

At the smallest budget the constrained attack does slightly better, which is not a bug. It runs iteratively with ten projection steps while fgsm takes one; at a small budget the extra steps are worth more than the freedom.

The constraints

Applied in an order that took some debugging to get right:

  1. Non-negativity. Every flow feature is a count, length, duration or rate.
  2. Integrality on 15 counts. 3.7 SYN flags is not a rounding artefact.
  3. Monotonicity on 11 features. An attacker can pad a packet, send more packets, or wait longer. They cannot un-send a packet or shorten a flow that already elapsed.
  4. Ordering. packet_len_min <= packet_len_mean <= packet_len_max, and the same for the five other min/mean/max triples.
  5. Derived consistency. flow_bytes_per_s is recomputed from the byte counts and the duration. Perturbing a rate independently of its own numerator is the clearest case of an impossible vector, and the easiest to miss, because the rate alone looks fine.

Order matters, and step 2 has to precede step 3. Rounding to nearest can round down, and rounding an integer feature below where it started breaks monotonicity. Doing integrality first lets monotonicity have the last word, and the maximum of two integers is still an integer. This was found by a test, not by inspection.

Monotonicity is also measured from a feasible base, via make_feasible. "May not fall below where it started" is meaningless when where it started was itself impossible.

When a model is too bad to attack

Attacks only target attack flows the model already classifies correctly. Under the honest protocol, the random forest detects so few attacks that there is nothing left to evade, and the benchmark says so:

random_forest:
  boundary  eps=0.050  not measurable: only 0 correctly-detected attack flows to target
    nothing to attack: this model detects too few attacks under the honest protocol
    for evasion to be measurable

Reporting 0% evasion there would read as robustness. It is the opposite.

Data

The generator in evade_nids.data produces synthetic flows, clearly labelled as such in every report. It exists so the test suite and CI run offline and deterministically, and because validating the audit needs known ground truth: to show the audit reports zero inflation on a clean dataset, the dataset has to be provably clean.

Three structures, each controlling one failure mode:

  • signal_strength: the genuine, transferable attack signal. What a model should learn.
  • group_signature: a per-session offset. What leaks under a random split.
  • temporal_drift: a slow rotation of the class boundary. What a random split hides.

Set the last two to zero and a random split becomes legitimate, which is how the false-positive test is written.

Real captures

import flowlens
from evade_nids import from_flowlens, audit_protocols, make_xgboost

table = flowlens.extract("capture.pcap")
data = from_flowlens(table, labels, group_by="src_ip").validate()

audit = audit_protocols(data, make_xgboost, model_name="xgboost")
print(audit.report())

Or from the flowlens CLI's CSV:

python -m evade_nids.benchmark --csv flows.csv \
    --label-column label --group-column src_ip --time-column first_seen_us

validate() runs the batch through featureguard's schema lock, so a reordered column fails loudly rather than silently retraining on the wrong features.

How it fits together

Project Role here
flowlens Extracts the 56 flow features from a capture
featureguard Schema lock on the feature table before any model sees it
leakhunt The general form of the split audit, for any grouped dataset
evade-nids The NIDS-specific benchmark and the constrained attacks

Output

results/
├── protocol_folds.csv     one row per fold: every metric, plus group overlap
├── protocol_summary.csv   one row per model: the headline comparison
├── attacks.csv            evasion rates and constraint violations
└── benchmark.json         everything, including caveats

Limitations

Read these before citing anything from here.

  • One real dataset, two scenarios. The CTU-13 study covers two capture scenarios and two malware families, so it measures leakage from host and session structure, not generalisation across attack types. CIC-IDS2017 was the intended target and proved unobtainable; every published mirror now redirects to a landing page.
  • The generator numbers are synthetic and much larger. They measure the instrument against planted ground truth, not any real dataset, and should not be quoted as a finding about network traffic.
  • Constrained attacks are still an upper bound. The output satisfies the feature constraints, but nothing here constructs a packet sequence producing those exact statistics, nor verifies the attack still works after perturbation. That is the remaining gap between this and a true problem-space attack.
  • The constraint set is incomplete. It captures non-negativity, integrality, ordering, monotonicity and rate consistency. It does not capture TCP state-machine validity, MTU limits, or the relationship between flag counts and packet counts.
  • Groups are a proxy. Grouping by source IP approximates a capture session. On real traffic a session may span addresses, or one address may span sessions.
  • Temporal folds are expanding-window. The first fold trains on very little, so its score is noisy and drags the temporal mean down.
  • No defences. Adversarial training, feature squeezing and ensembles are not implemented, so this measures attack success and not robustness after mitigation.

Development

pytest              # 51 tests, 87% coverage
ruff check .
mypy

If pytest fails to start with a PluginValidationError mentioning nengo, an unrelated package in your environment ships an incompatible pytest plugin. Run pytest -p no:nengo, or use a clean virtual environment.

Related reading

  • Arp et al., Dos and Don'ts of Machine Learning in Computer Security, USENIX Security 2022.
  • Pierazzi et al., Intriguing Properties of Adversarial ML Attacks in the Problem Space, IEEE S&P 2020.
  • Engelen et al., Troubleshooting an Intrusion Detection Dataset, IEEE S&P Workshops 2021.
  • Apruzzese et al., The Cross-Evaluation of Machine Learning-Based Network Intrusion Detection Systems, IEEE TNSM 2022.

License

MIT. See LICENSE.

About

Leakage-free NIDS benchmarking and adversarial attacks constrained to be physically sendable. Measures what a permissive split protocol is worth, and what feature-space attacks overstate.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages