feat(alerts): add container health alerts with log excerpt on notifications (#2225)

Add a new "ContainerHealth" alert type that fires when a Docker container's
health check reports unhealthy, and resolves when it recovers. This mirrors
the existing Status (up/down) alert pattern: an alert can be armed per system
and honors the "min minutes" delay before firing.

When the alert fires, the notification (email and any configured webhook,
including Discord via shoutrrr) includes a log excerpt fetched live from the
agent for up to 2 of the unhealthy containers, prioritizing lines containing
"error" or "fatal" (falling back to the log tail if none match), capped to
keep the message well under Discord's size limit.

---------

Co-authored-by: hank <hank@henrygd.me>
This commit is contained in:
Michał Mleczko
2026-09-02 18:46:23 +02:00
committed by GitHub
parent b1895247ba
commit 5969d36856
44 changed files with 1245 additions and 16 deletions

View File

@@ -424,6 +424,14 @@ msgstr "تعارض‌ها"
msgid "Connection is down"
msgstr "اتصال قطع است"
#: src/lib/alerts.ts
msgid "Container"
msgstr "کانتینر"
#: src/lib/alerts.ts
msgid "Container Health"
msgstr "سلامت کانتینر"
#: src/components/routes/system.tsx
msgid "Containers"
msgstr "کانتینرها"
@@ -1749,6 +1757,10 @@ msgstr "هنگامی که میانگین بار ۱۵ دقیقه‌ای از یک
msgid "Triggers when 5 minute load average exceeds a threshold"
msgstr "هنگامی که میانگین بار ۵ دقیقه‌ای از یک آستانه فراتر رود، فعال می‌شود"
#: src/lib/alerts.ts
msgid "Triggers when a Docker container's health check reports unhealthy"
msgstr "هنگامی که بررسی سلامت یک کانتینر Docker وضعیت ناسالم را گزارش می‌کند، فعال می‌شود"
#: src/lib/alerts.ts
msgid "Triggers when any sensor exceeds a threshold"
msgstr "هنگامی که هر حسگری از یک آستانه فراتر رود، فعال می‌شود"
@@ -1795,6 +1807,10 @@ msgstr "هنگامی که استفاده از هر دیسکی از یک آستا
msgid "Type"
msgstr "نوع"
#: src/lib/alerts.ts
msgid "Unhealthy"
msgstr "ناسالم"
#: src/components/systemd-table/systemd-table.tsx
msgid "Unit file"
msgstr "فایل واحد"