feat(alerts): add container health alerts with log excerpt on notifications (#2225)

Add a new "ContainerHealth" alert type that fires when a Docker container's
health check reports unhealthy, and resolves when it recovers. This mirrors
the existing Status (up/down) alert pattern: an alert can be armed per system
and honors the "min minutes" delay before firing.

When the alert fires, the notification (email and any configured webhook,
including Discord via shoutrrr) includes a log excerpt fetched live from the
agent for up to 2 of the unhealthy containers, prioritizing lines containing
"error" or "fatal" (falling back to the log tail if none match), capped to
keep the message well under Discord's size limit.

---------

Co-authored-by: hank <hank@henrygd.me>
This commit is contained in:
Michał Mleczko
2026-09-02 18:46:23 +02:00
committed by GitHub
parent b1895247ba
commit 5969d36856
44 changed files with 1245 additions and 16 deletions

View File

@@ -424,6 +424,14 @@ msgstr "冲突"
msgid "Connection is down"
msgstr "连接已断开"
#: src/lib/alerts.ts
msgid "Container"
msgstr "容器"
#: src/lib/alerts.ts
msgid "Container Health"
msgstr "容器健康"
#: src/components/routes/system.tsx
msgid "Containers"
msgstr "容器"
@@ -1749,6 +1757,10 @@ msgstr "当 15 分钟负载平均值超过阈值时触发"
msgid "Triggers when 5 minute load average exceeds a threshold"
msgstr "当 5 分钟内的平均负载超过阈值时触发"
#: src/lib/alerts.ts
msgid "Triggers when a Docker container's health check reports unhealthy"
msgstr "当 Docker 容器的健康检查报告不健康状态时触发"
#: src/lib/alerts.ts
msgid "Triggers when any sensor exceeds a threshold"
msgstr "当任何传感器超过阈值时触发"
@@ -1795,6 +1807,10 @@ msgstr "当任何磁盘的使用率超过阈值时触发"
msgid "Type"
msgstr "类型"
#: src/lib/alerts.ts
msgid "Unhealthy"
msgstr "不健康"
#: src/components/systemd-table/systemd-table.tsx
msgid "Unit file"
msgstr "单元文件"