Check /sys/class/ieee80211 before opening netlink sockets. Looking up
the
nl80211 family while cfg80211 is not loaded makes the kernel run
modprobe,
which loads the module on hosts without wireless hardware, or fails
again
on every poll when the module is unavailable.
Inside an LXC guest, lxcfs serves /proc/stat with the counters of the
host cores in the guest's cpuset, so an idle guest sharing a core with a
busy neighbor reported near-100% CPU.
When the agent detects it is running in an LXC guest (lxcfs mounted on
/proc/stat, container=lxc, or /run/systemd/container), derive CPU usage
from the guest's own cgroup instead: cpu.stat on cgroup v2, cpuacct on
v1, read at the cgroup mount root so it covers every process in the
guest. Usage is normalized by the usable cores (the smallest of the
affinity mask, cpuset and CPU quota).
Per-core usage is omitted in this mode, since the per-core /proc/stat
counters describe shared host cores and would contradict the total.
Iowait and steal are not available from cgroup accounting and report as
zero.
Hosts and Docker/Podman agents are unaffected and keep reading
/proc/stat, as does an LXC guest whose cgroup accounting is unreadable.
---------
Co-authored-by: hank <hank@henrygd.me>
Make WebSocket reconnect and SSH fallback transitions reliable across asynchronous disconnects, stale callbacks, and overlapping connections. Keep the SSH listener available while disconnected and allow a verified WebSocket connection to take precedence when it recovers.
Co-authored-by: henrygd <hank@henrygd.me>
Raises the bandwidth threshold max to 5 GB/s without making the MB/s
slider less precise. Values are still stored in MB/s, and the unit is
inferred from the stored value when the alert sheet opens.
Right after agent start the hub requests stats immediately, and the first
sample of the interval was measured from the baseline taken at startup, only
a second or so earlier. A burst of startup I/O was then stored as the rate
for the whole minute, causing large spikes in disk I/O charts.
The first sample of an interval now must span at least half the interval.
Otherwise it only stores the snapshot, and the next sample is measured from it.
The updater and on-demand requests kept separate copies of the SSH client
(sys.client and the transport's client) and synced them after each request.
That allowed a closed client to be reinstalled over a newer one and leaked
connections that were replaced without being closed.
The updater now dials, opens sessions and tears down timed-out connections
through the transport. The dial keeps the TCP keepalive and handshake
deadline, and an OnConnect callback handles the per-connection resets.
The transport is created under a lock and agentVersion is now atomic.
ssh.ClientConfig.Timeout only covers the TCP connect. A peer that accepts
the connection but never sends an SSH banner blocked ssh.NewClientConn
forever, wedging the system's updater goroutine and leaking the socket.
Set a deadline on the conn for the handshake and clear it afterward.