On 20/08/2026 13:32, Marc Zyngier wrote:
> On Thu, 20 Aug 2026 11:50:24 +0100,
> Steven Price <steven.price(a)arm.com> wrote:
>>
>> its_alloc_pages_node() passes __GFP_ZERO to the page allocator before
>> calling set_memory_decrypted(). This assumes that converting a page from
>> private to shared preserves its contents.
>>
>> For Arm CCA with MEC (Memory Encryption Contexts) the key used to access
>> the page will change, and so by default the visible data will change.
>> The host could ensure that it zeros the page, but rather than relying on
>> the host's behaviour it's best if the guest simply zeros after the
>> decryption rather than before. Specifically in this case the ITS tables
>> are required to be zeroed.
>
> What are the guarantees that we want to enforce post decryption? My
> recollection is that the RME firmware cleans the caches to the PoPA,
> making the data immediately visible to the hypervisor. Obviously, this
> isn't the case anymore, since the zeroing comes after that, and I
> don't see any CMO enforcing this.
The firmware should be ensuring that things are cleaned sufficiently
that the original data is inaccessible - that's required as part of the
wiping when converting from private. However the wipe doesn't have to be
writing zeros, indeed the RMM spec suggests that two "possible
implementations" are:
* The RMM (or other platform firmware) writing either random data or
zeroes to the memory location
* The MEC of the memory location being changed
My assumption (I have to admit I haven't checked) is that the GIC code
is doing sufficient CMO to ensure that the zeros that are being written
after the conversion are visible to the hypervisor - but that's no
different to the non-CCA case.
> I'm concerned that this relies on undocumented behaviours that may
> hold today on some undisclosed combinations of HW and hypervisors, but
> that are not guaranteed at all. set_memory_decrypted() doesn't really
> say anything, and I have the feeling that we may want some hypervisor
> specific hook to perform the correct CMO magic. I don't think this is
> required right now, but I'm not excluding anything!
This is reducing how much Linux relies on undocumented behaviour - at
the moment Linux is relying on either the zeros it has written still
being visible or the firmware/hypervisor writing zeros after any
private->shared transition. This patch makes the guest do it rather than
relying on anything else.
You have a point that a spec clarification about CMOs might be worth
having - the RMM spec doesn't make clear what is required of the
firmware. Clearly for security it should be doing something to ensure
that the old data from the realm doesn't become visible.
>>
>> Mask out __GFP_ZERO from the allocation request, and do the zeroing as a
>> separate step after decryption.
>>
>> Fixes: b08e2f42e86b ("irqchip/gic-v3-its: Share ITS tables with a non-trusted hypervisor")
>> Signed-off-by: Steven Price <steven.price(a)arm.com>
>> ---
>> drivers/irqchip/irq-gic-v3-its.c | 7 ++++++-
>> 1 file changed, 6 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/irqchip/irq-gic-v3-its.c b/drivers/irqchip/irq-gic-v3-its.c
>> index 6f5811aae59c..c954bbe9f4db 100644
>> --- a/drivers/irqchip/irq-gic-v3-its.c
>> +++ b/drivers/irqchip/irq-gic-v3-its.c
>> @@ -213,10 +213,12 @@ static gfp_t gfp_flags_quirk;
>> static struct page *its_alloc_pages_node(int node, gfp_t gfp,
>> unsigned int order)
>> {
>> + bool want_zero = gfp & __GFP_ZERO;
>> struct page *page;
>> int ret = 0;
>>
>> - page = alloc_pages_node(node, gfp | gfp_flags_quirk, order);
>> + page = alloc_pages_node(node, (gfp & ~__GFP_ZERO) | gfp_flags_quirk,
>> + order);
>>
>> if (!page)
>> return NULL;
>> @@ -231,6 +233,9 @@ static struct page *its_alloc_pages_node(int node, gfp_t gfp,
>> if (ret)
>> return NULL;
>>
>> + if (want_zero)
>> + clear_pages(page_address(page), 1 << order);
>> +
>
> nit: please use BIT(order), which matches the type required for
> clear_pages().
Sure, this was matching the use in set_memory_decrypted(), but I can
update that too.
> But I'd really like some discussion about the CMO side of things.
I'm not sure what more to say about CMO - if you want changes in the
commit message(s) then please suggest something. AFAICT this patch
doesn't change anything about cache maintenance. You're welcome to raise
spec clarifications if you want to.
Thanks,
Steve
PS. I'll take a look at the Sashiko comments - but they are both
pre-existing issues not issues with this series.
Arm CCA includes "Memory Encryption Contexts" (MEC) which allows the
private and shared data accessible to a guest to have different memory
encryption keys. Consequently when converting memory to shared, the
memory encryption key used to access the physical page will change.
Both the GICv3 ITS driver and the system_cc_shared dma-buf heap
currently allocate memory with __GFP_ZERO and then decrypt it. With MEC
the zeroing is done with the wrong encryption key and the data visible
after decryption may be ciphertext. The RMM is required to scrub the
data, but may perform this scrub with a different encryption key to the
eventual key that will be used for shared access.
Fix these two sites by avoiding the __GFP_ZERO during the allocation and
performing a clear_pages() call after the decryption.
Steven Price (2):
irqchip/gic-v3-its: Zero shared pages after conversion
dma-buf: heaps: Zero system shared heap pages after conversion
drivers/dma-buf/heaps/system_heap.c | 14 +++++++++++---
drivers/irqchip/irq-gic-v3-its.c | 7 ++++++-
2 files changed, 17 insertions(+), 4 deletions(-)
--
2.43.0
On 20/08/2026 10:52, Dmitry Baryshkov wrote:
> On Wed, Aug 19, 2026 at 04:18:51PM +0200, Krzysztof Kozlowski wrote:
>> On 19/08/2026 15:17, Ekansh Gupta wrote:
>>> On 19-08-2026 00:40, Krzysztof Kozlowski wrote:
>>>> On 17/08/2026 06:47, Ekansh Gupta wrote:
>>>>> Add the skeleton of the Qualcomm DSP Accelerator (QDA) driver, a DRM
>>>>> accel driver for the Hexagon DSPs found on Qualcomm SoCs.
>>>>>
>>>>> This patch registers a DRM accel device, exposing a /dev/accel/accelN
>>>>> character device node, and binds it to the RPMsg channel used to reach
>>>>> the DSP. Buffer management, IOMMU context banks and the FastRPC
>>>>> protocol are added by later patches in this series.
>>>>>
>>>>> qda_drv.c / qda_drv.h define the drm_driver ops table, the per-file
>>>>> private state (qda_file_priv) and the main device structure (qda_dev),
>>>>> which embeds drm_device so that it can be recovered with container_of().
>>>>>
>>>>> qda_rpmsg.c binds to the "qcom,fastrpc" compatible via
>>>>> module_rpmsg_driver(), reads the DSP domain name from the "label"
>>>>> device-tree property, and registers the DRM device.
>>>>>
>>>>> Assisted-by: Claude:claude-sonnet-5
>>>>> Signed-off-by: Ekansh Gupta <ekansh.gupta(a)oss.qualcomm.com>
>>>>> ---
>>>>> Changes in v2:
>>>>> - Use module_rpmsg_driver() and drop the qda_rpmsg_register()/
>>>>> _unregister() wrappers, module_init()/module_exit() and
>>>>> qda_rpmsg.h entirely (Dmitry Baryshkov)
>>>>> - Read the "label" property directly into qdev->dsp_name (Dmitry Baryshkov)
>>>>> - Drop the probe/remove/init log messages (Dmitry Baryshkov)
>>>>> - Return the result of qda_register_device() directly (Dmitry Baryshkov)
>>>>> - Clarify the Kconfig help text (Dmitry Baryshkov)
>>>>> ---
>>>>> drivers/accel/Kconfig | 1 +
>>>>> drivers/accel/Makefile | 1 +
>>>>> drivers/accel/qda/Kconfig | 30 ++++++++++++++++
>>>>> drivers/accel/qda/Makefile | 10 ++++++
>>>>> drivers/accel/qda/qda_drv.c | 71 ++++++++++++++++++++++++++++++++++++++
>>>>> drivers/accel/qda/qda_drv.h | 61 +++++++++++++++++++++++++++++++++
>>>>> drivers/accel/qda/qda_rpmsg.c | 79 +++++++++++++++++++++++++++++++++++++++++++
>>>>> 7 files changed, 253 insertions(+)
>>>>>
>>>>> diff --git a/drivers/accel/Kconfig b/drivers/accel/Kconfig
>>>>> index bdf48ccafcf2..74ac0f71bc9d 100644
>>>>> --- a/drivers/accel/Kconfig
>>>>> +++ b/drivers/accel/Kconfig
>>>>> @@ -29,6 +29,7 @@ source "drivers/accel/ethosu/Kconfig"
>>>>> source "drivers/accel/habanalabs/Kconfig"
>>>>> source "drivers/accel/ivpu/Kconfig"
>>>>> source "drivers/accel/qaic/Kconfig"
>>>>> +source "drivers/accel/qda/Kconfig"
>>>>> source "drivers/accel/rocket/Kconfig"
>>>>>
>>>>> endif
>>>>> diff --git a/drivers/accel/Makefile b/drivers/accel/Makefile
>>>>> index 1d3a7251b950..58c08dd5f389 100644
>>>>> --- a/drivers/accel/Makefile
>>>>> +++ b/drivers/accel/Makefile
>>>>> @@ -5,4 +5,5 @@ obj-$(CONFIG_DRM_ACCEL_ARM_ETHOSU) += ethosu/
>>>>> obj-$(CONFIG_DRM_ACCEL_HABANALABS) += habanalabs/
>>>>> obj-$(CONFIG_DRM_ACCEL_IVPU) += ivpu/
>>>>> obj-$(CONFIG_DRM_ACCEL_QAIC) += qaic/
>>>>> +obj-$(CONFIG_DRM_ACCEL_QDA) += qda/
>>>>> obj-$(CONFIG_DRM_ACCEL_ROCKET) += rocket/
>>>>> \ No newline at end of file
>>>>
>>>> You have trivial patch errors.
>>> newline problem was already there, wasn't introduced as part of this
>>> patch series, so I wasn't sure to fix it here. I can fix this in v3.>
>>>> ...
>>>>
>>>>> +}
>>>>> +
>>>>> +static const struct of_device_id qda_rpmsg_id_table[] = {
>>>>> + { .compatible = "qcom,fastrpc" },
>>>>> + {},
>>>>
>>>> Device node with this compatible is already populated, so this looks
>>>> simply wrong or you are adding a duplicated driver.
>>>>
>>>> That's a no-go, you are supposed to work with existing drivers and grow
>>>> them.
>>> I'll bring the discussion again here, there was a discussion to move the
>>> driver to accel subsystem if we want to support new features/uAPI
>>> changes. Please read [1],[2] threads. The intention is to replace
>>> fastrpc driver with QDA eventually.
>>
>> None of them address the problem. You want to grow fastrpc into user of
>> dmabuf? So you move it from misc to here.
>
> It's not as easy and nice, so I think in this case it's better to repeat
I disagree. The existing fastrpc driver is not that complicated. It's
actually moderate amount of code, much less than Venus was (~7 times less).
It easily can grow to support two interfaces and the only difficulty is
how to manage these two interfaces simultaneously or exclusively, e.g.
opening first one disables the second.
> what we did for the venus/iris migration or what happened already
> several times in the kernel history (for example the AIC7xxx SCSI host
> drivers were migrated by introducing the second driver and then removing
> the first one after the grace period). The backwards compatibility is a
> separate topic, but it will be addressed before the driver can be
Backwards compatibility should be one of the first things explained in
cover letter in one of the first paragraphs.
And if you make it backwards compatible, then just remove old driver,
because there is no point to keep it there. Again, this should be one of
the first things explained in cover letter.
Best regards,
Krzysztof
On 19/08/2026 17:48, Rob Clark wrote:
> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk(a)kernel.org> wrote:
>> The rule of usptream development is that we do not accept duplicated
>> code, just because a vendor wants to write something new. This is
>> basically the concept applied all over the drivers tree, where we pushed
>> back against all sorts of duplications all over the vendors.
>>
>> What I miss in this thread is why would there be any exception here. We
>> do not grant exceptions from standard practices on "I want" reasons.
>
> I agree that we should not have duplicated drivers just for vendor
> lolz. But when it comes to adopting common frameworks and integrating
> better into the ecosystem, this doesn't seem like something we should
> actively discourage. I don't think this is a case of vendor lolz, but
No one discourages it. Following standard Linux kernel practices and
requirements is not discouraging, do not twist the narrative here.
Again, it is standard upstream review telling that we do not duplicate
drivers. Ever, unless there is serious exception needed.
I asked why there should be an exception granted? Is the reason for
exception following:
"We want to adopt common framework"
?
> rather reacting to drm/accel emerging as the standard framework for
> this sort of driver.
>
> So how do we get from here to there?
What is wrong with my proposal?
Best regards,
Krzysztof
On 8/17/26 16:07, Taimuraz Kaitmazov wrote:
> amdxdna_drm_sync_bo_ioctl() calls amdxdna_hwctx_sync_debug_bo() for every
> FROM_DEVICE sync, which answers -EINVAL when the BO has no assigned hwctx.
> Only a BO attached with ATTACH_DEBUG_BO ever gets one, so an ordinary
> read-back sync reports failure after its flush has already run.
>
> There is no debug buffer to sync in that case, so answer success.
>
> Signed-off-by: Taimuraz Kaitmazov <taimuraz(a)kaitmazov.com>
> ---
> drivers/accel/amdxdna/amdxdna_ctx.c | 3 ++-
> 1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/accel/amdxdna/amdxdna_ctx.c b/drivers/accel/amdxdna/amdxdna_ctx.c
> index 855da8c79a1c..c0d0aa53c596 100644
> --- a/drivers/accel/amdxdna/amdxdna_ctx.c
> +++ b/drivers/accel/amdxdna/amdxdna_ctx.c
> @@ -416,7 +416,8 @@ int amdxdna_hwctx_sync_debug_bo(struct amdxdna_client *client, u32 debug_bo_hdl)
> guard(mutex)(&xdna->dev_lock);
> hwctx = xa_load(&client->hwctx_xa, abo->assigned_hwctx);
> if (!hwctx) {
> - ret = -EINVAL;
> + /* Not attached as a debug BO, so there is nothing to sync. */
> + ret = 0;
It should check assigned_hwctx before entering this function:
  if (abo->assigned_hwctx != AMDXDNA_INVALID_CTX_HANDLE &&
args->direction == SYNC_DIRECT_FROM_DEVICE)
      ret = amdxdna_hwctx_sync_debug_bo(client, args->handle);
Thanks,
Lizhi
> goto put_obj;
> }
>
On 8/17/26 16:07, Taimuraz Kaitmazov wrote:
> SYNC_BO clflushes an imported BO's scatterlist. An importer may not do
> that: the memory belongs to the exporter, and dma-buf gives the importer
> no interface to ask for maintenance on it. Refuse the request instead.
>
> is_import_bo() is (obj)->attach, which covers more than foreign buffers.
> A userptr BO arrives through a ubuf, and on a carveout device every share
> BO and the device heap arrive through a cbuf, so SYNC_BO answers
> -EOPNOTSUPP for those too, including the AMDXDNA_BO_DEV path that flushes
> through its heap.
>
> Only the ubuf case gives up maintenance it was getting: on a 64 MiB
> userptr BO a 4 KiB sync and a full sync both cost 659 us, this arm having
> ignored the range. amdxdna_cbuf_map() fills in only the DMA address and
> length, so drm_clflush_sg() already walks zero pages on carveout memory.
> Userspace maintains these through the mapping it already holds, as XRT's
> buffer::sync() does unless it is told to sync through the driver.
>
> Suggested-by: Lizhi Hou <lizhi.hou(a)amd.com>
> Signed-off-by: Taimuraz Kaitmazov <taimuraz(a)kaitmazov.com>
> ---
> drivers/accel/amdxdna/amdxdna_gem.c | 7 ++++---
> 1 file changed, 4 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/accel/amdxdna/amdxdna_gem.c b/drivers/accel/amdxdna/amdxdna_gem.c
> index 4f38f985c74e..a713a9982d34 100644
> --- a/drivers/accel/amdxdna/amdxdna_gem.c
> +++ b/drivers/accel/amdxdna/amdxdna_gem.c
> @@ -1224,6 +1224,9 @@ static int amdxdna_flush_bo(struct amdxdna_gem_obj *abo, u64 offset, u64 size)
> {
> u64 end;
>
> + if (is_import_bo(abo))
> + return -EOPNOTSUPP;
> +
> if (offset >= abo->mem.size)
> return -EINVAL;
>
> @@ -1234,9 +1237,7 @@ static int amdxdna_flush_bo(struct amdxdna_gem_obj *abo, u64 offset, u64 size)
> if (!size)
> return 0;
>
> - if (is_import_bo(abo))
> - drm_clflush_sg(abo->base.sgt);
> - else if (amdxdna_gem_vmap(abo))
> + if (amdxdna_gem_vmap(abo))
Reviewed-by: Lizhi Hou <lizhi.hou(a)amd.com>
> drm_clflush_virt_range(amdxdna_gem_vmap(abo) + offset, size);
> else if (abo->base.pages)
> drm_clflush_pages(abo->base.pages, abo->mem.size >> PAGE_SHIFT);
On 8/17/26 16:07, Taimuraz Kaitmazov wrote:
> amdxdna_drm_sync_bo_ioctl() answers a failed amdxdna_flush_bo() with
> drm_WARN(). Both of that function's error returns are decided by the
> ioctl's arguments, so SYNC_BO with an offset past the end of the BO
> splats and taints the kernel from an unprivileged caller.
>
> Log it like the pin failure above it.
>
> Signed-off-by: Taimuraz Kaitmazov <taimuraz(a)kaitmazov.com>
> ---
> drivers/accel/amdxdna/amdxdna_gem.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/accel/amdxdna/amdxdna_gem.c b/drivers/accel/amdxdna/amdxdna_gem.c
> index 77a9493cd7ba..4f38f985c74e 100644
> --- a/drivers/accel/amdxdna/amdxdna_gem.c
> +++ b/drivers/accel/amdxdna/amdxdna_gem.c
> @@ -1310,7 +1310,7 @@ int amdxdna_drm_sync_bo_ioctl(struct drm_device *dev,
> amdxdna_gem_unpin(abo);
>
> if (ret) {
> - drm_WARN(&xdna->ddev, 1, "Can not get flush memory");
> + XDNA_ERR(xdna, "Flush BO %d failed, ret %d", args->handle, ret);
Should it be XDNA_DBG?
Lizhi
> goto put_obj;
> }
> }
On 8/17/26 16:07, Taimuraz Kaitmazov wrote:
> amdxdna_drm_sync_bo_ioctl() forms the range for a device BO by adding the
> caller's offset and size to the BO address without checking either, while
> amdxdna_flush_bo() one call down guards the same arithmetic with
> check_add_overflow().
>
> A size that wraps flush_end leaves it below the heap it is clamped
> against, so every heap fails the start >= end test, and a sync that asked
> for more than the address space holds reports success having flushed
> nothing. Reject it instead.
>
> Signed-off-by: Taimuraz Kaitmazov <taimuraz(a)kaitmazov.com>
> ---
> drivers/accel/amdxdna/amdxdna_gem.c | 9 +++++++--
> 1 file changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/accel/amdxdna/amdxdna_gem.c b/drivers/accel/amdxdna/amdxdna_gem.c
> index f88b5349cd4b..77a9493cd7ba 100644
> --- a/drivers/accel/amdxdna/amdxdna_gem.c
> +++ b/drivers/accel/amdxdna/amdxdna_gem.c
> @@ -1274,8 +1274,13 @@ int amdxdna_drm_sync_bo_ioctl(struct drm_device *dev,
> struct amdxdna_gem_obj *heap;
> unsigned long heap_id;
> u64 bo_start = amdxdna_gem_dev_addr(abo);
> - u64 flush_start = bo_start + args->offset;
> - u64 flush_end = flush_start + args->size;
> + u64 flush_start, flush_end;
> +
> + if (check_add_overflow(bo_start, args->offset, &flush_start) ||
> + check_add_overflow(flush_start, args->size, &flush_end)) {
> + ret = -EINVAL;
> + goto put_obj;
> + }
Reviewed-by: Lizhi Hou <lizhi.hou(a)amd.com>
>
> xa_for_each_range(&client->dev_heap_xa, heap_id, heap,
> abo->heap_start_id, abo->heap_end_id) {
On 8/17/26 16:07, Taimuraz Kaitmazov wrote:
> amdxdna_gem_vmap() flattens the iosys_map drm_gem_vmap() fills in down to
> the void * in abo->mem.kva, and iosys_map is discriminated by is_iomem, so
> an exporter answering with an I/O mapping leaves a void __iomem pointer
> there, which amdxdna_cmd_set_error() memsets and memcpys through.
>
> amdxdna_drm_va_tbl takes a dmabuf_fd, so such a BO can be any exporter's
> buffer. amdgpu cannot reach this: its .pin forces GTT for a non peer to
> peer attachment like ours. An exporter on drm_gem_prime_dmabuf_ops has
> no .pin, and drm_gem_ttm_vmap() answers iomem for a VRAM resident
> object, so an NPU paired with nouveau or radeon does.
>
> Drop such a mapping and answer NULL. Checking here rather than in the
> .vmap callback leaves that callback's iosys_map contract intact for a
> caller equipped to read I/O memory, and covers everything that takes a
> plain kernel address through this helper. vmw_gem_vmap() refuses the
> same case; unlike that one this path is reachable from an unprivileged
> ioctl, so it neither warns nor logs at error level.
>
> Signed-off-by: Taimuraz Kaitmazov <taimuraz(a)kaitmazov.com>
> ---
> drivers/accel/amdxdna/amdxdna_gem.c | 9 +++++++--
> 1 file changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/accel/amdxdna/amdxdna_gem.c b/drivers/accel/amdxdna/amdxdna_gem.c
> index cca84fa07e9d..f88b5349cd4b 100644
> --- a/drivers/accel/amdxdna/amdxdna_gem.c
> +++ b/drivers/accel/amdxdna/amdxdna_gem.c
> @@ -209,10 +209,15 @@ void *amdxdna_gem_vmap(struct amdxdna_gem_obj *abo)
>
> if (!abo->mem.kva) {
> ret = drm_gem_vmap(to_gobj(abo), &map);
> - if (ret)
> + if (ret) {
> XDNA_ERR(abo->client->xdna, "Vmap bo failed, ret %d", ret);
> - else
> + } else if (map.is_iomem) {
> + /* Callers use the result as an ordinary kernel address. */
> + XDNA_DBG(abo->client->xdna, "Vmap bo returned I/O memory");
> + drm_gem_vunmap(to_gobj(abo), &map);
> + } else {
> abo->mem.kva = map.vaddr;
> + }
Reviewed-by: Lizhi Hou <lizhi.hou(a)amd.com>
> }
> return abo->mem.kva;
> }
On 8/18/26 14:40, Taimuraz Kaitmazov wrote:
> Patch 1 needs a prerequisite. It makes amdxdna_gem_vmap() answer NULL
> on an iomem exporter, and the eight amdxdna_cmd_get_payload() callers
> in aie2_message.c check neither the pointer nor the length it leaves
> unwritten there. A command BO can be an import, so patch 1 alone turns
> a silent __iomem write into a NULL deref with an uninitialised length.
> https://lore.kernel.org/all/20260818002459.377641-1-taimuraz@kaitmazov.com/
>
amdxdna_cmd_get_op() is always called before amdxdna_cmd_get_payload()
for the same BO, and since amdxdna_gem_vmap() caches its results, the
vmap inside get_payload() is currently guaranteed to succeed. The patch
is therefore defensive rather than fixing a currently reachable crash.
Lizhi
>
> fixes the callers. Happy to respin on top if you prefer them together.
>
> Taimuraz
>
> On 8/18/26 02:07, Taimuraz Kaitmazov wrote:
>> Five fixes in and around amdxdna_drm_sync_bo_ioctl().
>>
>> Patch 1 refuses an I/O memory mapping the driver would otherwise
>> store as
>> if it were an ordinary kernel address. Patch 2 checks the device-BO
>> range
>> for overflow. Patch 4 refuses a flush of an imported BO, which is why
>> patch 3 comes first: the ioctl answers a rejected flush with drm_WARN(),
>> so without it an ordinary sync on an imported BO splats. Patch 5 stops
>> the ioctl reporting failure for a flush that succeeded.
>>
>> v3's zero-length patch has left this series. On hardware it turns out to
>> be a page fault in drm_clflush_virt_range() rather than the tidy-up its
>> commit message described, so it is a -fixes patch now, sent
>> separately as
>> "accel/amdxdna: return early from a zero-length flush" with Fixes: and
>> Cc: stable. Patch 4 here needs its hunk, so this series wants that one
>> first.
>>
>> Changes in v4:
>> Â Â - patch 1: the is_iomem check moved from the .vmap callback into
>> Â Â Â Â amdxdna_gem_vmap(), per Lizhi, and it logs at debug level rather
>> than
>> Â Â Â Â error, since an unprivileged caller can repeat it.
>> Â Â - new patch 3: an unprivileged SYNC_BO with an out-of-range offset
>> Â Â Â Â already reaches that drm_WARN() today. Sashiko's review of v3 3/3
>> Â Â Â Â flagged the same thing.
>> Â Â - new patch 4: refuses the flush for every imported BO, as asked.
>> Â Â - new patch 5: the debug-BO sync I mentioned on the v2 thread.
>> Only a BO
>> Â Â Â Â attached with ATTACH_DEBUG_BO has an assigned hwctx, so every other
>> Â Â Â Â FROM_DEVICE sync ends in -EINVAL with its flush already done. The
>> Â Â Â Â -EINVAL reproduces on a Strix Point NPU.
>>
>> On patch 4, one consequence worth deciding before it lands. A heap BO is
>> created through amdxdna_drm_create_share_bo(), so a device running on
>> carveout memory reaches its heap through a cbuf, is_import_bo() is
>> true of
>> it, and SYNC_BO on every AMDXDNA_BO_DEV now answers -EOPNOTSUPP. Today
>> that path flushes nothing anyway -- amdxdna_cbuf_map() fills in only the
>> DMA address and length, so drm_clflush_sg() walks zero pages -- so the
>> change is silent no-op to hard error, and XRT's dbg_buffer::sync()
>> reaches
>> it without Debug.force_driver_sync. If you would rather keep our own
>> exporters working, I have the variant keyed on the exporter's ops, which
>> confines the refusal to foreign buffers. Say which you prefer.
>>
>> v3:
>> https://lore.kernel.org/all/20260813164700.43960-1-taimuraz@kaitmazov.com/
>>
>> Built on drm-misc-next, each commit on its own: x86_64 with
>> DRM_ACCEL_AMDXDNA=m, clang 22.1.8, no warnings. Patch 5's reproducer was
>> run against 7.1.8's in-tree driver, where that call is unchanged.
>>
>> Taimuraz Kaitmazov (5):
>> Â Â accel/amdxdna: refuse an I/O memory mapping of an imported BO
>> Â Â accel/amdxdna: check the sync range for overflow on a device BO
>> Â Â accel/amdxdna: do not warn when a sync request is rejected
>> Â Â accel/amdxdna: refuse to flush an imported BO
>> Â Â accel/amdxdna: do not fail a sync for a BO with no debug context
>>
>> Â drivers/accel/amdxdna/amdxdna_ctx.c |Â 3 ++-
>> Â drivers/accel/amdxdna/amdxdna_gem.c | 27 +++++++++++++++++++--------
>> Â 2 files changed, 21 insertions(+), 9 deletions(-)
>>
>
>