
Anomaly Detection
anomaly-detection — SKILL.md
Description (triggering): Detects unusual patterns, deviations, and irregularities in datasets, time series, industrial sensor measurements, and performance indicators. Triggers for requests like "find the outliers," "is this sensor drifting," "monitor equipment health," even without the exact phrase "anomaly detection." Covers classical statistical methods (z-score, IQR, control charts, EWMA) and ML methods (Isolation Forest, LOF, One-Class SVM, autoencoders/LSTM). Key strength: distinguishing a genuine anomaly from normal variation, seasonality, or a regime change.
What the skill is for: the core idea is that anomaly detection isn't a single algorithm but a decision problem — defining what counts as "normal," choosing the right type of anomaly, and setting a threshold that can actually be justified.
Anomaly taxonomy:
- Point — a single extreme value
- Contextual — abnormal only given its context (load, time of day, operating regime)
- Collective — a whole sequence is abnormal even though no individual point is (slow drift, an irregular vibration signature)
6-step workflow: understand the data → clarify ambiguity → pick the method family → set the threshold deliberately → validate before reporting → present results usefully (graded severity, explanation of cause).
Bundled resources:
references/statistical-methods.md— z-score, modified z-score (MAD), IQR, Grubbs' test, Shewhart/EWMA/CUSUM control chartsreferences/ml-methods.md— Isolation Forest, LOF, One-Class SVM, autoencoders, LSTMreferences/industrial-diagnostics.md— multiple operating regimes, sensor fault vs. equipment fault, severity tiers, alarm fatiguescripts/detect_anomalies.py— ready-to-run (tested) CLI script applying z-score/IQR/EWMA/Isolation Forest to a CSV
Common pitfalls to avoid: a single global threshold applied across multiple regimes, confusing missing data with anomalies, assuming normality on skewed data, reporting a raw score without translating it into an action, overfitting the threshold to a demo dataset.


