Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection
2026-07-15T07:23:54Z•0d9eb8e41b719853fae1a0858a8c9c3614c0342ed41c36d7dee5899ba4a630bc
adversarial-testinganomaly-evasionbenchmarkingevaluation-metricsmachine-learningmetric-selectionmonitoringopen-sourcepip-packageresearchrobust-evaluationtime-series-anomaly-detection
What happened
This feed aggregates new ML/statistics arXiv papers (15 Jul 2026). Most security-relevant is an independent adversarial stress-test of post-point-adjustment metrics for time-series anomaly detection (TSAD): the authors evaluate 12 adopted metrics across 250 UCR Anomaly Archive series and five additional benchmarks using trivial and adversarial no-skill score generators. Key findings: affiliation-F1 and ROC-based metrics (including VUS-ROC) are widely gameable by no-skill detectors (high gameability rates across benchmarks), while PR-based metrics and PA%K are substantially more robust. The kit
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_stat_ml
- Record identifier
- 0d9eb8e41b719853fae1a0858a8c9c3614c0342ed41c36d7dee5899ba4a630bc
- Enrichment time
- 2026-07-15T07:23:54Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.