Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection

2026-07-15T07:23:54Z0d9eb8e41b719853fae1a0858a8c9c3614c0342ed41c36d7dee5899ba4a630bc
adversarial-testinganomaly-evasionbenchmarkingevaluation-metricsmachine-learningmetric-selectionmonitoringopen-sourcepip-packageresearchrobust-evaluationtime-series-anomaly-detection

What happened

This feed aggregates new ML/statistics arXiv papers (15 Jul 2026). Most security-relevant is an independent adversarial stress-test of post-point-adjustment metrics for time-series anomaly detection (TSAD): the authors evaluate 12 adopted metrics across 250 UCR Anomaly Archive series and five additional benchmarks using trivial and adversarial no-skill score generators. Key findings: affiliation-F1 and ROC-based metrics (including VUS-ROC) are widely gameable by no-skill detectors (high gameability rates across benchmarks), while PR-based metrics and PA%K are substantially more robust. The kit

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_stat_ml
Record identifier
0d9eb8e41b719853fae1a0858a8c9c3614c0342ed41c36d7dee5899ba4a630bc
Enrichment time
2026-07-15T07:23:54Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.