Fix Predictive Alerting False Positives in Prometheus Python Models
Predictive alerting on Prometheus data pages on-call for pod restarts, not real anomalies? Here's the root cause and three fixes that work.
AWS Cost Anomaly Detection Script: A Bash Alert Checklist
A bash-based AWS cost anomaly detection script checklist to catch billing spikes in Slack before finance does — with real gotchas.
Build a Self-Hosted Log Anomaly Detector to Cut Alert Noise
How we replaced regex log alerts with a self-hosted log anomaly detection pipeline using IsolationForest, Vector, and FastAPI.
ML Anomaly Detection in Loki Logs: Per-Entity Isolation Forest
A practical look at ML anomaly detection for Loki logs — feature extraction, per-entity models, and why global thresholds fail.
Terraform Drift Detection in CI: Alert-Only vs Auto-Remediate
Terraform drift detection in CI can alert humans or auto-fix infra. I've run both — here's which one saved us and which one bit us.
Building an Online Anomaly Scoring Sidecar for Grafana Alerts
AI anomaly detection in Grafana is mostly statistics, not magic — here's how to wire online scoring into alerts without flapping.
Grafana Alerting Checklist: Wiring AI Anomaly Scores Correctly
A practical Grafana AI anomaly detection checklist for wiring model scores into unified alerting without alert storms or false negatives.
Cursor AI Guardrails Checklist for Terraform and Ansible Repos
A practical Cursor AI infrastructure code review checklist to catch hallucinated IAM wildcards and provider bumps before they hit prod.
PostgreSQL Logical Replication Across Kubernetes StatefulSets
A deep dive into PostgreSQL logical replication on Kubernetes StatefulSets — the wire mechanics, common misconfigurations, and production-safe wiring.
Ansible Dynamic Inventory for EC2 Fleets: 7 Tuning Tips
Ansible dynamic inventory for EC2 can throttle your API and break patch runs if left untuned — here's what we fixed after it bit us.
☕ Support us · 💳 Monobank