Port the CGroups CPU usage fix to AF5 - #182
Merged
Merged
Conversation
Ports the fix from the Axon Framework 4 line (1.10.5) to AF5. OperatingSystemMXBean.getProcessCpuLoad and getCpuLoad are container-aware since JDK 14, and under a cgroup CPU quota they divide the CPU consumed by quota * nr_periods. nr_periods counts only the scheduling periods in which the cgroup had a runnable task, not the periods that elapsed, so an idle application is measured against the slice it was allowed while it happened to be awake. The quieter it gets the higher it reads, saturating at the 1.0 the JDK clamps to. Measured on production before the fix: containers at 10 periods/s reported accurately, while ones at 3-7 periods/s reported up to 10x their real usage, or exactly zero. Two live services were reporting 0.0% CPU, and a customer with a 200m limit at 0.0017 cores was reporting 100% on both process and host CPU - the two readings share the calculation, which is why they saturate together. Both values now come from monotonic counters divided by elapsed wall time: the JVM's process CPU time, and the host's jiffies in /proc/stat. The process reading keeps its meaning - a fraction of what the container may use - by dividing by the cgroup quota where one is enforced, falling back to the visible processors. Off Linux there is no /proc/stat and equally no cgroup to distort the bean, so the host reading asks it directly. The setup payload gains an optional runtime block describing the JVM, the operating system, the garbage collectors, the heap ceiling and the control group's CPU quota, shares and memory limit. Without it a CPU or heap percentage has no denominator: a reading of 100% could mean eight cores or a tenth of one, and neither the UI nor an alert could tell the difference. The server does not read it yet.
CodeDrivenMitch
requested review from
a team,
Andrew-deVillier and
stefanmirkovic
and removed request for
a team
September 22, 2026 12:53
|
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Ports #180 to the AF5 line. The defect and the fix are identical; only the package differs.
The defect
OperatingSystemMXBean.getProcessCpuLoad/getCpuLoadare container-aware since JDK 14: under a cgroup quota the JDK computesdelta(cpu) / delta(quota * nr_periods), andnr_periodscounts only the scheduling periods in which the cgroup had a runnable task, not the periods that elapsed. An idle application is measured against the slice it was allowed while it happened to be awake, so its reading rises as it gets quieter and saturates at the 1.0 the JDK clamps to.Measured on our own production: containers at 10 periods/s reported accurately; ones at 3-7 periods/s reported up to 10x their real usage, or exactly zero. Two live services were reporting 0.0% CPU while serving traffic. A customer with a 200m limit using 0.0017 cores was reported at 100% on both process and host CPU — the two readings share the calculation, which is why they saturate together.
The fix
Both values now come from monotonic counters over elapsed wall time:
getProcessCpuTimefor the process,/proc/statjiffy deltas for the host. The process reading keeps its meaning — a fraction of what the container may use — by dividing by the cgroup quota where one is enforced (0.2 for a200mlimit), falling back to the visible processors. Off Linux there is no/proc/stat, and equally no cgroup to distort the bean, so the host reading asks it directly.Cgroupsresolves the process's own cgroup via/proc/self/cgroupand walks up to the mount root taking the tightest limit, so a systemd unit withCPUQuota=or a container started with--cgroupns=hostis not mistaken for unlimited.Setup payload
Adds an optional
runtimeblock: JVM, OS, garbage collectors, heap ceiling, and the cgroup's CPU quota, shares and memory limit. Without it a CPU or heap percentage has no denominator — 100% could mean eight cores or a tenth of one, and neither the UI nor an alert can tell. Nothing that identifies the machine or its owner is gathered, and the whole block is wrapped so that describing the runtime can never fail a connection. The server does not read it yet.Differences from #180
@JvmOverloadsonSetupPayloadCreator—AxoniqPlatformConfigurerEnhanceruses a constructor reference, which does not see Kotlin default parametersTests
19 new unit tests; 166 of 170 pass. The four failures are
AxoniqConsoleRSocketClientToxiproxyIntegrationTest, which fail identically on unmodifiedmain— verified by running that class on a clean checkout. Pre-existing, and worth a separate look.The new tests are mutation-checked: reverting
take(8)(the/proc/statguest double-count),minOrNull(ancestor limits), the/proc/self/cgroupwalk or theUNAVAILABLEinitial value each turns the suite red.