On Mon, Aug 10, 2026 at 04:10:48PM +0100, Will Deacon wrote:
On Mon, Aug 10, 2026 at 03:44:42PM +0100, Leo Yan wrote:
Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by default") made the AUX allocator use order-0 pages by default unless a PMU explicitly asks for contiguous allocations.
But that commit specifically calls out SPE as benefitting from non-contiguous pages:
"For instance, ARM SPE and TRBE operate with virtual pages, and Coresight ETR allocates a separate buffer. For these PMUs, allocating contiguous AUX pages unnecessarily exacerbates memory fragmentation. This fragmentation can prevent their use on long-running devices."
so why doesn't passing PERF_PMU_CAP_AUX_PREFER_LARGE reintroduce the problems that 18049c8cff9c was trying to solve?
The question is how "allocating contiguous AUX pages unnecessarily exacerbates memory fragmentation." The relevant information I could find is [1]:
"On Android, we collect ETM data periodically on internal user devices for AutoFDO optimization (for both userspace libraries and the kernel). Allocating a large chunk of contiguous AUX pages (4M for each CPU) periodically is almost unbearable. The kernel may need to kill many processes to fulfill the request. It affects user experience even after using PMU."
We might have missed chance to clarify how the fragmentation issue occurs in the first place. Let's say, a phone with 8 CPUs, allocating 4MB per CPU requires 32MB in total, which is a relatively small portion of 4GiB or 8GiB of RAM commonly found in phones. Moreover, once contiguous pages are freed, the buddy allocator can coalesce them again into buddy list. It is not obvious to me that PREFER_LARGE directly causes fragmentation.
One case where AUX allocation could exacerbate fragmentation is when the system is already fragmented. If a high-order allocation fails and the allocator falls back to smaller-order blocks, those allocations may consume free blocks scattered across different buddy regions and make subsequent high-order allocations more difficult.
If this is the main concern, I'd suggest using a smaller AUX buffer (e.g. 1MB or even 512KB) for TRBE/SPE to reduce memory pressure. Snapshot mode '-S' could also be considered, as it allows the buffer to be allocated once and reused for subsequent recordings by signals.
OTOH, using only order-0 pages can significantly increase TTW overhead on the trace path and lead to overflows, we observe this causes huge trace discontinuity. In the end, we need to trace-off the fragmentation concern against the trace discontinuity.
Thanks, Leo
[1] https://lore.kernel.org/lkml/CALJ9ZPNLgEBxOmDim-vztUknEETwdL-Z2gJ8K9s44TiPgK...
On 10/08/2026 18:41, Leo Yan wrote:
On Mon, Aug 10, 2026 at 04:10:48PM +0100, Will Deacon wrote:
On Mon, Aug 10, 2026 at 03:44:42PM +0100, Leo Yan wrote:
Commit 18049c8cff9c ("perf/aux: Allocate non-contiguous AUX pages by default") made the AUX allocator use order-0 pages by default unless a PMU explicitly asks for contiguous allocations.
But that commit specifically calls out SPE as benefitting from non-contiguous pages:
"For instance, ARM SPE and TRBE operate with virtual pages, and Coresight ETR allocates a separate buffer. For these PMUs, allocating contiguous AUX pages unnecessarily exacerbates memory fragmentation. This fragmentation can prevent their use on long-running devices."
so why doesn't passing PERF_PMU_CAP_AUX_PREFER_LARGE reintroduce the problems that 18049c8cff9c was trying to solve?
The question is how "allocating contiguous AUX pages unnecessarily exacerbates memory fragmentation." The relevant information I could find is [1]:
"On Android, we collect ETM data periodically on internal user devices for AutoFDO optimization (for both userspace libraries and the kernel). Allocating a large chunk of contiguous AUX pages (4M for each CPU) periodically is almost unbearable. The kernel may need to kill many processes to fulfill the request. It affects user experience even after using PMU."
This sounds like it could be an attribute to perf_event_open. We can do PREFER_LARGE by default for performance and fewer discontinuities, but on Android or small systems users can enable an option to revert back to single pages.
Or can this bit be determined at allocation time: "The kernel may need to kill many processes to fulfill the request"? If this memory pressure exists on allocation then do it one way, if not do it the other way.
We might have missed chance to clarify how the fragmentation issue occurs in the first place. Let's say, a phone with 8 CPUs, allocating 4MB per CPU requires 32MB in total, which is a relatively small portion of 4GiB or 8GiB of RAM commonly found in phones. Moreover, once contiguous pages are freed, the buddy allocator can coalesce them again into buddy list. It is not obvious to me that PREFER_LARGE directly causes fragmentation.
One case where AUX allocation could exacerbate fragmentation is when the system is already fragmented. If a high-order allocation fails and the allocator falls back to smaller-order blocks, those allocations may consume free blocks scattered across different buddy regions and make subsequent high-order allocations more difficult.
If this is the main concern, I'd suggest using a smaller AUX buffer (e.g. 1MB or even 512KB) for TRBE/SPE to reduce memory pressure. Snapshot mode '-S' could also be considered, as it allows the buffer to be allocated once and reused for subsequent recordings by signals.
OTOH, using only order-0 pages can significantly increase TTW overhead on the trace path and lead to overflows, we observe this causes huge trace discontinuity. In the end, we need to trace-off the fragmentation concern against the trace discontinuity.
Thanks, Leo
[1] https://lore.kernel.org/lkml/CALJ9ZPNLgEBxOmDim-vztUknEETwdL-Z2gJ8K9s44TiPgK...
On Tue, Aug 11, 2026 at 10:02:44AM +0100, James Clark wrote:
[...]
The question is how "allocating contiguous AUX pages unnecessarily exacerbates memory fragmentation." The relevant information I could find is [1]:
"On Android, we collect ETM data periodically on internal user devices for AutoFDO optimization (for both userspace libraries and the kernel). Allocating a large chunk of contiguous AUX pages (4M for each CPU) periodically is almost unbearable. The kernel may need to kill many processes to fulfill the request. It affects user experience even after using PMU."
This sounds like it could be an attribute to perf_event_open. We can do PREFER_LARGE by default for performance and fewer discontinuities, but on Android or small systems users can enable an option to revert back to single pages.
Or can this bit be determined at allocation time: "The kernel may need to kill many processes to fulfill the request"? If this memory pressure exists on allocation then do it one way, if not do it the other way.
Be careful with "The kernel may need to kill many processes ...".
The AUX allocator uses:
#define PERF_AUX_GFP (GFP_KERNEL | __GFP_ZERO | __GFP_NOWARN | __GFP_NORETRY)
Documentation/core-api/memory-allocation.rst says that __GFP_NORETRY causes the allocator to "fail early rather than cause disruptive reclaim" and that "the OOM killer is not invoked".
Therefore, it is unlikely to me that AUX allocation itself would cause the kernel OOM killer to kill processes in order to satisfy the request.
I am a bit suspect the original observation was related to Android's lmkd [1], where allocating the AUX buffers increased overall memory pressure and indirectly caused processes to be killed. If so, I think it is important to understand what actually triggered those kills.
Thanks, Leo
[1] https://source.android.com/docs/core/perf/lmkd?hl=en&utm_source=chatgpt....