fix(agent): report LXC guest CPU usage from cgroup accounting (#2341)

Inside an LXC guest, lxcfs serves /proc/stat with the counters of the
host cores in the guest's cpuset, so an idle guest sharing a core with a
busy neighbor reported near-100% CPU.

When the agent detects it is running in an LXC guest (lxcfs mounted on
/proc/stat, container=lxc, or /run/systemd/container), derive CPU usage
from the guest's own cgroup instead: cpu.stat on cgroup v2, cpuacct on
v1, read at the cgroup mount root so it covers every process in the
guest. Usage is normalized by the usable cores (the smallest of the
affinity mask, cpuset and CPU quota).

Per-core usage is omitted in this mode, since the per-core /proc/stat
counters describe shared host cores and would contradict the total.
Iowait and steal are not available from cgroup accounting and report as
zero.

Hosts and Docker/Podman agents are unaffected and keep reading
/proc/stat, as does an LXC guest whose cgroup accounting is unreadable.

---------

Co-authored-by: hank <hank@henrygd.me>
This commit is contained in:
Santhi Prakash
2026-10-02 05:03:42 +05:30
committed by GitHub
parent c394a6b5b1
commit 01d91728f0
5 changed files with 747 additions and 3 deletions

View File

@@ -170,9 +170,13 @@ func (a *Agent) getSystemStats(cacheTimeMs uint16) system.Stats {
slog.Error("Error getting cpu metrics", "err", err)
}
// per-core cpu usage
if perCoreUsage, err := getPerCoreCpuUsage(cacheTimeMs); err == nil {
systemStats.CpuCoresUsage = perCoreUsage
// per-core cpu usage. Skipped when the total comes from cgroup accounting:
// per-core /proc/stat counters there describe shared host cores, not the
// guest, and would contradict the total.
if !cpuMetrics.fromCgroup {
if perCoreUsage, err := getPerCoreCpuUsage(cacheTimeMs); err == nil {
systemStats.CpuCoresUsage = perCoreUsage
}
}
// load average