On 9/26/26 09:12, Karl Mehltretter wrote:
> get_sg_table() merges physically contiguous pages without accounting
> for the mapping device's maximum segment size. This affects both
> importer mappings and the udmabuf misc device mapping used for CPU
> access.
>
> With DMA_API_DEBUG enabled, DMA_BUF_IOCTL_SYNC on a 64 MiB udmabuf
> reports:
>
> DMA-API: misc udmabuf: mapping sg segment longer than device claims to support [len=65884160] [max=65536]
>
> Use sg_alloc_table_from_pages_segment() with the mapping device's
> maximum segment size. Keep a PAGE_SIZE minimum because the allocator
> warns and returns -EINVAL for smaller limits.
>
> Before commit 5bf888673e0d ("udmabuf: Do not create malformed
> scatterlists"), each entry covered one page.
>
> Fixes: 5bf888673e0d ("udmabuf: Do not create malformed scatterlists")
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter(a)gmail.com>
> ---
>
> Notes:
> Tested on v7.3-rc4-70-gfe2ec83746e5 in QEMU (x86_64, TCG) with
> DMA_API_DEBUG (all_errors=1) and DMABUF_DEBUG, A/B against the same
> base:
>
> before after
> DMA_BUF_IOCTL_SYNC, 64 MiB udmabuf 1 report 0
> vivid import, 4 MiB udmabuf 2 reports 0
> vivid import, 2 MiB hugetlb udmabuf 2 reports 0
> frames captured 5/5 5/5
>
> vb2-dma-contig rejected the non-contiguous import in both runs.
>
> drivers/dma-buf/udmabuf.c | 10 +++++++---
> 1 file changed, 7 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/dma-buf/udmabuf.c b/drivers/dma-buf/udmabuf.c
> index df6dd00462423..09f1eb8432f19 100644
> --- a/drivers/dma-buf/udmabuf.c
> +++ b/drivers/dma-buf/udmabuf.c
> @@ -139,9 +139,13 @@ static struct sg_table *get_sg_table(struct device *dev, struct dma_buf *buf,
> if (!sg)
> return ERR_PTR(-ENOMEM);
>
> - ret = sg_alloc_table_from_pages(sg, ubuf->pages, ubuf->pagecount, 0,
> - ubuf->pagecount << PAGE_SHIFT,
> - GFP_KERNEL);
> + /* The SG allocator requires a segment limit of at least PAGE_SIZE. */
> + ret = sg_alloc_table_from_pages_segment(sg, ubuf->pages, ubuf->pagecount,
> + 0, ubuf->pagecount << PAGE_SHIFT,
> + max_t(unsigned int,
> + dma_get_max_seg_size(dev),
> + PAGE_SIZE),
Please return -EINVAL instead when dma_get_max_seg_size() returns that the segment size is smaller than a page.
In general I think that the sg_alloc_table_from_pages_segment() approach is because of the broken design of the old DMA API. Stuff like that should be handled by the iterator going over the DMA segments instead. But yeah that is not something you can fix in one patch.
So apart from the error handling the patch looks good to me.
Regards,
Christian.
> + GFP_KERNEL);
> if (ret < 0)
> goto err_alloc;
>
On Tue, Sep 29, 2026 at 08:56:01AM +0200, Karl Mehltretter wrote:
> On Tue, Sep 29, 2026 at 06:00:27AM +0000, sashiko-bot(a)kernel.org wrote:
> > On architectures with a 256KB PAGE_SIZE (such as Hexagon or PowerPC), this
> > new validation evaluates to 65536 < 262144, unconditionally returning
> > -EINVAL and breaking all CPU mappings.
>
> Right, the misc device has no dma_parms, so it gets the 64K default and
> sync fails with 256K pages
I would ignore this, it seems fictional to me, and if someone really
needs to support this rediculous combination it should be done inside
sg_alloc_table_from_pages_segment(:)
> The misc device has no real segment limit, so for v3 I'll give it
> dma_parms with an unlimited max segment size and keep the -EINVAL for
> importers that report a limit below PAGE_SIZE.
The params have to come from the *importing* device, not the udmabuf
misc device. You don't get to control what they are, and there is no
reason to put a dma_parms on a misc deivce.
Jason
On 9/28/26 12:36, Rob Clark wrote:
> On Mon, Sep 28, 2026 at 1:19 AM Christian König
> <christian.koenig(a)amd.com> wrote:
>>
>> On 9/26/26 04:20, Jianfeng Liu wrote:
>>> That commit fixed a dangling reference in the DMABUF_DEBUG default
>>> and thereby enabled the option - and with it the page-stripping
>>> sg_table wrapper that dma_buf_map_attachment() hands to importers -
>>> on every kernel with DEBUG_KERNEL=y, i.e. virtually every distro
>>> kernel.
>>>
>>> drm/msm is broken by the wrapper. Both of msm's map paths consume
>>> sg->length and sg_phys() of the attachment sg_table:
>>> msm_iommu_pagetable_map() for the per-process GPU pagetables, and
>>> iommu_map_sg() (via iommu_map_sgtable()) for scanout. The wrapper
>>> zeroes sg->length and strips the page pointers, so mappings of
>>> imported dma-bufs silently map nothing, and userspace observes
>>> arm-smmu translation faults from UCHE, e.g. during hardware video
>>> decode (clapper, chromium) on Adreno systems:
>>>
>>> gpu fault: ttbr0=000000088a889000 iova=000000010741c000 dir=READ
>>> type=TRANSLATION source=UCHE
>>>
>>> Bisected on a Snapdragon X1E78100 laptop as v7.3-rc3 good,
>>> v7.3-rc4 bad, culprit 143755bdabaa9.
>>>
>>> Switching msm to sg_dma_address()/sg_dma_len() is not a trivial fix
>>> either: those fields are only valid for sg_tables that msm has
>>> dma-mapped itself, which native non-MSM_BO_WC objects' sg_tables
>>> are not, so the conversion needs more work. The msm maintainer has
>>> therefore requested restoring the previous default for v7.3, to be
>>> revisited once msm no longer consumes struct page and sg->length of
>>> imported sg_tables.
>>
>> Yeah as I said before the problematic part is MSM here. We have enforced correct driver behavior for over 5 years now when that option is enabled.
>
> The problem is bigger than MSM here
Well, so far I have only heard about MSM.
But yes I mean the config option is doing exactly what it is supposed to do, pointing out when driver need some work to get this fixed.
I also agree that we shouldn't have allowed driver to touch that stuff in the first place and better document how to do things but yeah I can't change the past I can only try to fix it now.
>> What we can do is to mark MSM as broken and/or give a warning in MSM when DMABUF_DEBUG is enabled and you try to import a DMA-buf.
>
> sorry, no, we can't mark MSM as broken.. we can mark DMABUF_DEBUG as BROKEN
As I wrote the debug functionality to enforce not using struct pages has been around for over 5 years now, it was just not enabled by default.
What I can offer is to set it to default N for another few month to give you more time to fix things.
Regards,
Christian.
>
> BR,
> -R
>
>> But making the check not default to enable on debug kernels is not an option. This check here is exactly to point out broken drivers and you can manually disable it.
>>
>> Regards,
>> Christian.
>>
>>>
>>> Link: https://lore.kernel.org/linux-arm-msm/20260923074256.9357-1-liujianfeng1994…
>>> Suggested-by: Rob Clark <robin.clark(a)oss.qualcomm.com>
>>> Cc: Christian König <christian.koenig(a)amd.com>
>>> Cc: Sumit Semwal <sumit.semwal(a)linaro.org>
>>> Cc: Karl Mehltretter <kmehltretter(a)gmail.com>
>>>
>>> Signed-off-by: Jianfeng Liu <liujianfeng1994(a)gmail.com>
>>> ---
>>>
>>> drivers/dma-buf/Kconfig | 2 +-
>>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>>
>>> diff --git a/drivers/dma-buf/Kconfig b/drivers/dma-buf/Kconfig
>>> index e4f078a326a41..7efc0f0d07126 100644
>>> --- a/drivers/dma-buf/Kconfig
>>> +++ b/drivers/dma-buf/Kconfig
>>> @@ -43,7 +43,7 @@ config UDMABUF
>>> config DMABUF_DEBUG
>>> bool "DMA-BUF debug checks"
>>> depends on DMA_SHARED_BUFFER
>>> - default y if DEBUG_KERNEL
>>> + default y if DEBUG
>>> help
>>> This option enables additional checks for DMA-BUF importers and
>>> exporters. Specifically it validates that importers do not peek at the
>>> ---
>>> base-commit: 93f51579e7df248780214094418f205253383cc5
>>> branch: revert-dmabuf-debug-for-7.3
>>>
>>
On 9/26/26 20:30, Rob Clark wrote:
> This was always the way it was supposed to work, and when we drop the
> page array for imported dma-bufs our fault handling path will no longer
> work.
>
> Signed-off-by: Rob Clark <robin.clark(a)oss.qualcomm.com>
Nice to see that finally happening.
Reviewed-by: Christian König <christian.koenig(a)amd.com> for this patch here, Acked-by: Christian König <christian.koenig(a)amd.com> for the rest of the series.
Regards,
Christian.
> ---
> drivers/gpu/drm/msm/msm_gem.c | 24 ++++++++++++++++++++++++
> 1 file changed, 24 insertions(+)
>
> diff --git a/drivers/gpu/drm/msm/msm_gem.c b/drivers/gpu/drm/msm/msm_gem.c
> index e8390ebd5dd5..c90336b3b231 100644
> --- a/drivers/gpu/drm/msm/msm_gem.c
> +++ b/drivers/gpu/drm/msm/msm_gem.c
> @@ -23,6 +23,8 @@
> #include "msm_gpu.h"
> #include "msm_kms.h"
>
> +MODULE_IMPORT_NS("DMA_BUF");
> +
> static void update_device_mem(struct msm_drm_private *priv, ssize_t size)
> {
> uint64_t total_mem = atomic64_add_return(size, &priv->total_mem);
> @@ -338,6 +340,9 @@ static vm_fault_t msm_gem_fault(struct vm_fault *vmf)
> int err;
> vm_fault_t ret;
>
> + if (drm_WARN_ON_ONCE(obj->dev, drm_gem_is_imported(obj)))
> + return VM_FAULT_SIGBUS;
> +
> /*
> * vm_ops.open/drm_gem_mmap_obj and close get and put
> * a reference on obj. So, we dont need to hold one here.
> @@ -1126,6 +1131,25 @@ static int msm_gem_object_mmap(struct drm_gem_object *obj, struct vm_area_struct
> {
> struct msm_gem_object *msm_obj = to_msm_bo(obj);
>
> + if (drm_gem_is_imported(obj)) {
> + int ret;
> +
> + /* Reset both vm_ops and vm_private_data, so we don't end up with
> + * vm_ops pointing to our implementation if the dma-buf backend
> + * doesn't set those fields.
> + */
> + vma->vm_private_data = NULL;
> + vma->vm_ops = NULL;
> +
> + ret = dma_buf_mmap(obj->dma_buf, vma, 0);
> +
> + /* Drop the reference drm_gem_mmap_obj() acquired.*/
> + if (!ret)
> + drm_gem_object_put(obj);
> +
> + return ret;
> + }
> +
> vm_flags_set(vma, VM_PFNMAP | VM_DONTEXPAND | VM_DONTDUMP);
> vma->vm_page_prot = msm_gem_pgprot(msm_obj, vma_get_page_prot(vma));
>
On 9/26/26 04:20, Jianfeng Liu wrote:
> That commit fixed a dangling reference in the DMABUF_DEBUG default
> and thereby enabled the option - and with it the page-stripping
> sg_table wrapper that dma_buf_map_attachment() hands to importers -
> on every kernel with DEBUG_KERNEL=y, i.e. virtually every distro
> kernel.
>
> drm/msm is broken by the wrapper. Both of msm's map paths consume
> sg->length and sg_phys() of the attachment sg_table:
> msm_iommu_pagetable_map() for the per-process GPU pagetables, and
> iommu_map_sg() (via iommu_map_sgtable()) for scanout. The wrapper
> zeroes sg->length and strips the page pointers, so mappings of
> imported dma-bufs silently map nothing, and userspace observes
> arm-smmu translation faults from UCHE, e.g. during hardware video
> decode (clapper, chromium) on Adreno systems:
>
> gpu fault: ttbr0=000000088a889000 iova=000000010741c000 dir=READ
> type=TRANSLATION source=UCHE
>
> Bisected on a Snapdragon X1E78100 laptop as v7.3-rc3 good,
> v7.3-rc4 bad, culprit 143755bdabaa9.
>
> Switching msm to sg_dma_address()/sg_dma_len() is not a trivial fix
> either: those fields are only valid for sg_tables that msm has
> dma-mapped itself, which native non-MSM_BO_WC objects' sg_tables
> are not, so the conversion needs more work. The msm maintainer has
> therefore requested restoring the previous default for v7.3, to be
> revisited once msm no longer consumes struct page and sg->length of
> imported sg_tables.
Yeah as I said before the problematic part is MSM here. We have enforced correct driver behavior for over 5 years now when that option is enabled.
What we can do is to mark MSM as broken and/or give a warning in MSM when DMABUF_DEBUG is enabled and you try to import a DMA-buf.
But making the check not default to enable on debug kernels is not an option. This check here is exactly to point out broken drivers and you can manually disable it.
Regards,
Christian.
>
> Link: https://lore.kernel.org/linux-arm-msm/20260923074256.9357-1-liujianfeng1994…
> Suggested-by: Rob Clark <robin.clark(a)oss.qualcomm.com>
> Cc: Christian König <christian.koenig(a)amd.com>
> Cc: Sumit Semwal <sumit.semwal(a)linaro.org>
> Cc: Karl Mehltretter <kmehltretter(a)gmail.com>
>
> Signed-off-by: Jianfeng Liu <liujianfeng1994(a)gmail.com>
> ---
>
> drivers/dma-buf/Kconfig | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/dma-buf/Kconfig b/drivers/dma-buf/Kconfig
> index e4f078a326a41..7efc0f0d07126 100644
> --- a/drivers/dma-buf/Kconfig
> +++ b/drivers/dma-buf/Kconfig
> @@ -43,7 +43,7 @@ config UDMABUF
> config DMABUF_DEBUG
> bool "DMA-BUF debug checks"
> depends on DMA_SHARED_BUFFER
> - default y if DEBUG_KERNEL
> + default y if DEBUG
> help
> This option enables additional checks for DMA-BUF importers and
> exporters. Specifically it validates that importers do not peek at the
> ---
> base-commit: 93f51579e7df248780214094418f205253383cc5
> branch: revert-dmabuf-debug-for-7.3
>
On Thu, Sep 24, 2026 at 02:55:08PM +0800, Yili Zhang wrote:
> The dma_buf_phys_vec_to_sgt() routine links every physical range into
> the allocated IOVA space at offset zero instead of advancing the offset
> by the already mapped length. This breaks the THRU_HOST_BRIDGE flow
> whenever phys_vec contains more than one entry: the second and
> subsequent dma_iova_link() calls try to install mappings on top of
> PTEs that already exist, which fails and makes the whole conversion
> return an error.
>
> Fix it by passing the accumulated mapped_len as the IOVA offset. This
> matches the pattern used by all other dma_iova_link() callers and is
> consistent with the rest of the function: dma_iova_sync() and the
> error path in dma_iova_destroy() both operate on the [0, mapped_len)
> prefix of the IOVA space, which assumes that the ranges were linked
> contiguously starting at offset 0.
>
> Fixes: 3aa31a8bb11e ("dma-buf: provide phys_vec to scatter-gather mapping routine")
> Cc: stable(a)vger.kernel.org
> Signed-off-by: Yili Zhang <zhangyili01(a)baidu.com>
> ---
> drivers/dma-buf/dma-buf-mapping.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
Thanks,
Reviewed-by: Leon Romanovsky <leon(a)kernel.org>
Hi Longfang,
On 22/09/2026 10:16, liulongfang wrote:
> On 2026/9/21 21:24, Matt Evans wrote:
>> Hi Longfang,
>>
>> On 15/09/2026 13:16, liulongfang wrote:
>>> On 2026/9/12 5:41, Matt Evans wrote:
>>>> [snip]
>>>> +
>>>> + /*
>>>> + * The DMABUF begins from the mmap()'s BAR offset, i.e. the
>>>> + * start of the VMA corresponds to byte 0 of the DMABUF and
>>>> + * byte (vma_pgoff << PAGE_SHIFT) of the BAR.
>>>> + *
>>>> + * vfio_pci_dma_buf_find_pfn() reverses this offset using
>>>> + * vma_pgoff_adjust, so that ultimately a fault's offset from
>>>> + * the start of the _VMA_ has a consistent usage whether the
>>>> + * VMA originates from an mmap() of the VFIO device here or a
>>>> + * direct DMABUF mmap(). Note vma_pgoff_adjust also includes
>>>> + * the encoded VFIO region index, which cancels out the index
>>>> + * encoded in vm_pgoff.
>>>> + */
>>>> + priv->vdev = vdev;
>>>> + priv->size = req_len;
>>>> + priv->nr_ranges = 1;
>>>> + priv->vma_pgoff_adjust = vma->vm_pgoff;
>>>> +
>>>> + priv->provider = pcim_p2pdma_provider(vdev->pdev, res_index);
>>>> + if (!priv->provider) {
>>>> + ret = -EINVAL;
>>>> + goto err_free_name;
>>>> + }
>>>> +
>>>> + priv->phys_vec[0].paddr = phys_start + ((u64)vma_pgoff << PAGE_SHIFT);
>>>> + priv->phys_vec[0].len = priv->size;
>>>> +
>>>> + ret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);
>>>> + if (ret)
>>>> + goto err_free_name;
>>>> +
>>>
>>> In the current patch, the PCIe device's BAR2 configuration space can be mapped as a DMABUF.
>>> However, on an OS with a 64K page size, if a VF device's BAR2 is smaller than 64K,
>>> a problem arises where the space is forced to page-align to 64K, it will causing the VM to
>>> access memory beyond the actual size of the VF device's BAR2 space.
>>>
>>> How does your solution handle these cases where the BAR2 space is smaller than the Host OS's page size?
>>
>> Even on a 4K host, there can be BARs < PAGE_SIZE so 64K (or 16K) hosts
>> aren't a new case. These small BARs cannot be mmap()ed and DMABUFs
>> cannot be exported from them. (vfio_pci_core_mmap() errors out when
>> !bar_mmap_supported[index]. And, a DMABUF needs to be an aligned
>> multiple of PAGE_SIZE, plus vfio_pci_core_fill_phys_vec() won't allow a
>> DMABUF to be created off the end of a BAR.)
>>
>> So, although this series allows a DMABUF to be mmap()ed, the preexisting
>> checks prevent a sub-page DMABUF from existing and so there is no new
>> route to mapping a sub-page BAR.
>>
>> What's the concern on BAR2 specifically, out of interest? This logic is
>> applied to all resources equally, and tests pci_resource_len(...) so
>> there shouldn't be a PF/VF distinction either.
>>
>
> However, the typical boundary check found in VFIO, such as:
>
> if (req_start + req_len > phys_len)
> return -EINVAL;
>
> seems to be missing here.
Isn't it covered by that statement in vfio_pci_core_mmap() just before
this function is called? This helper is intended to do as it's told by
a caller that has validated the range is correct (it doesn't duplicate
the checks already done by vfio_pci_core_mmap()).
Thanks,
Matt
This series tightens the alignment requirements for buffers that are shared
between confidential-computing guests and the host, and adds a common
allocator for host-shared memory.
When a guest runs with private memory, buffers shared with the hypervisor
are not only accessed by the guest. They are also accessed by the host
kernel, and the host may manage the corresponding shared/private state at a
granularity larger than the guest page size.
This matters for CCA systems where the Realm stage-2 mappings managed by
the RMM can still operate at 4K granularity, while the non-secure host may
manage the IPA state change at a larger page size, for example 64K. In that
case, allowing a guest to convert and share only a 4K subrange of a
host-managed granule is unsafe.
Architectures such as Arm can detect incorrect accesses to Realm physical
address space PFNs through GPC faults. However, relying on that as the only
line of defence is fragile and can still lead to kernel crashes. The risk
is especially visible for shared buffers that are later mmapped into
userspace, such as guest_memfd or dma-buf backed allocations. Once
userspace can access the mapping, the kernel cannot guarantee that
applications will only touch the intended 4K region rather than the whole
host page mapped into their address space. Those userspace addresses may
also be passed back into the kernel and accessed through the linear map,
resulting in a GPC fault.
To avoid this, host-shared buffers must satisfy two constraints:
- the address must be aligned to the CoCo shared-granule size
- the size must be a multiple of that granule size
The series adds a common CoCo shared-memory layer for enforcing these
constraints. It provides shared-granule geometry and range-validation
helpers, byte-oriented private/shared transition helpers, and
alloc_cc_shared_pages() with a node-aware variant. The allocator rounds a
request to the architecture shared granule, allocates suitably aligned
contiguous pages, transitions the complete allocation to shared state, and
returns the transitioned size alongside the page.
The corresponding free helper restores the complete allocation to private
state before returning it to the buddy allocator. If private state cannot
be restored safely, the allocation is deliberately leaked rather than
returning potentially shared memory for unrelated use. Since a
private-to-shared transition may modify memory contents, __GFP_ZERO is
applied after the transition.
The generic shared-granule size defaults to PAGE_SIZE. For arm64 CCA, the
series queries the host IPA state change alignment through the Realm Host
Interface, caches it during Realm initialization, and exposes it through
the arm64 memory-encryption operations.
The common allocator is used for host-shared allocations whose backing is
owned by an individual caller:
- GIC ITS command queues and tables
- dma-direct allocations backed by CMA or the page allocator
- backing allocations for the CoCo atomic DMA pools
- dma-buf system_cc_shared heap allocations
Hyper-V users of set_memory_encrypted() and set_memory_decrypted() are not
changed by this series. Those paths are not currently used by the arm64 CCA
code path, and therefore are not part of the arm64 CCA IPA state change
alignment problem addressed here.
NOTE: I have not added explicit MAINTAINERS entries for mm/cc_shared.c and
include/linux/cc_shared.h, as I am unsure whether we need a separate section
for common CoCo-related files. I will add the entries based on feedback.
Changes from v7:
https://lore.kernel.org/all/20260921144847.501151-1-aneesh.kumar@kernel.org
* Add the following new patches:
* "irqchip/gic-v3-its: Preallocate VPE L1 tables"
* "mm: Assert CoCo shared allocations may sleep"
* "mm: Zero memory during shared memory transitions"
* Drop the arm64 RHI and shared granule size patches so that the series can
be rebased on top of upstream to enable Shashiko review.
Changes from v6:
https://lore.kernel.org/all/20260904103452.1197239-1-aneesh.kumar@kernel.org
* Add a common allocator and geometry/transition helpers for CoCo host-shared
memory.
* Convert GIC ITS, dma-direct, atomic DMA pools, and the dma-buf
system_cc_shared heap to the common allocator.
* Limit dma-buf scatterlist entries to the requested buffer size so rounded
backing is not exposed to importers.
Changes from v5:
https://lore.kernel.org/all/20260706060432.1375570-1-aneesh.kumar@kernel.org
* Rebased to latest kernel
* Drop patch arm64: realm: Move Realm memory encryption ops to RSI code
Changes from v4:
https://lore.kernel.org/all/20260427063108.909019-1-aneesh.kumar@kernel.org
* Rename the helpers to use CoCo terminology
(mem_cc_shared_granule_size() / mem_cc_align_to_shared_granule() instead of
mem_decrypt_granule_size() / mem_decrypt_align()).
* Use __DMA_ATTR_ALLOC_CC_SHARED to pass CoCo shared allocation requirements
down to CMA-based allocation helpers.
* Add validation for restricted DMA pools to reject pools that are not aligned
to the shared granule size.
* Add dma-buf system heap handling for cc-shared buffers.
* Split the previous combined DMA/SWIOTLB/ITS change into smaller subsystem
patches covering ITS, DMA direct, SWIOTLB, restricted DMA pools, dma-buf
system heap, and arm64 Realm support.
* Rework arm64 Realm support by moving Realm memory encryption ops into RSI
code and exposing the CCA shared granule size through arm64_mem_crypt_ops.
Changes from v3:
https://lore.kernel.org/all/20260309102625.2315725-1-aneesh.kumar@kernel.org
* Fix build error reported by kernel test robot <lkp(a)intel.com>
Changes from v2:
https://lore.kernel.org/all/20251221160920.297689-1-aneesh.kumar@kernel.org
* Rebase to latest kernel
* Consider swiotlb always decrypted and don't align when allocating from swiotlb.
Changes from v1:
* Rename the helper to mem_encrypt_align
* Improve the commit message
* Handle DMA allocations from contiguous memory
* Handle DMA allocations from the pool
* swiotlb is still considered unencrypted. Support for an encrypted swiotlb pool
is left as TODO and is independent of this series.
Cc: Andrew Morton <akpm(a)linux-foundation.org>
Cc: Baoquan He <baoquan.he(a)linux.dev>
Cc: Mike Rapoport <rppt(a)kernel.org>
Cc: Pasha Tatashin <pasha.tatashin(a)soleen.com>
Cc: Pratyush Yadav <pratyush(a)kernel.org>
Cc: Catalin Marinas <catalin.marinas(a)arm.com>
Cc: "Christian König" <christian.koenig(a)amd.com>
Cc: Jason Gunthorpe <jgg(a)ziepe.ca>
Cc: Joerg Roedel (AMD) <joro(a)8bytes.org>
Cc: Marc Zyngier <maz(a)kernel.org>
Cc: Marek Szyprowski <m.szyprowski(a)samsung.com>
Cc: Robin Murphy <robin.murphy(a)arm.com>
Cc: Steven Price <steven.price(a)arm.com>
Cc: Sumit Semwal <sumit.semwal(a)linaro.org>
Cc: Suzuki K Poulose <suzuki.poulose(a)arm.com>
Cc: Thomas Gleixner <tglx(a)kernel.org>
Cc: Will Deacon <will(a)kernel.org>
Cc: Russell King <linux(a)armlinux.org.uk>
Cc: Benjamin Gaignard <benjamin.gaignard(a)collabora.com>
Cc: Brian Starkey <Brian.Starkey(a)arm.com>
Cc: John Stultz <jstultz(a)google.com>
Cc: Mark Rutland <mark.rutland(a)arm.com>
Cc: Radu Rendec <radu(a)rendec.net>
Cc: "T.J. Mercier" <tjmercier(a)google.com>
Cc: Madhavan Srinivasan <maddy(a)linux.ibm.com>
Cc: Michael Ellerman <mpe(a)ellerman.id.au>
Cc: Nicholas Piggin <npiggin(a)gmail.com>
Cc: Christophe Leroy (CS GROUP) <chleroy(a)kernel.org>
Cc: Ritesh Harjani (IBM) <ritesh.list(a)gmail.com>
Cc: Shrikanth Hegde <sshegde(a)linux.ibm.com>
Cc: Alexander Gordeev <agordeev(a)linux.ibm.com>
Cc: Gerald Schaefer <gerald.schaefer(a)linux.ibm.com>
Cc: Heiko Carstens <hca(a)linux.ibm.com>
Cc: Vasily Gorbik <gor(a)linux.ibm.com>
Cc: Christian Borntraeger <borntraeger(a)linux.ibm.com>
Cc: Sven Schnelle <svens(a)linux.ibm.com>
Cc: Ingo Molnar <mingo(a)redhat.com>
Cc: Borislav Petkov <bp(a)alien8.de>
Cc: Dave Hansen <dave.hansen(a)linux.intel.com>
Cc: x86(a)kernel.org
Cc: H. Peter Anvin <hpa(a)zytor.com>
Cc: Kiryl Shutsemau <kas(a)kernel.org>
Cc: Rick Edgecombe <rick.p.edgecombe(a)intel.com>
Cc: K. Y. Srinivasan <kys(a)microsoft.com>
Cc: Haiyang Zhang <haiyangz(a)microsoft.com>
Cc: Wei Liu <wei.liu(a)kernel.org>
Cc: Dexuan Cui <decui(a)microsoft.com>
Cc: Long Li <longli(a)microsoft.com>
Cc: Paolo Bonzini <pbonzini(a)redhat.com>
Cc: Vitaly Kuznetsov <vkuznets(a)redhat.com>
Cc: Andy Lutomirski <luto(a)kernel.org>
Cc: Peter Zijlstra <peterz(a)infradead.org>
Cc: dri-devel(a)lists.freedesktop.org
Cc: iommu(a)lists.linux.dev
Cc: linaro-mm-sig(a)lists.linaro.org
Cc: linux-arm-kernel(a)lists.infradead.org
Cc: linux-kernel(a)vger.kernel.org
Cc: linux-media(a)vger.kernel.org
Cc: linux-mm(a)kvack.org
Aneesh Kumar K.V (Arm) (14):
mm: Add an allocator for CoCo shared memory
mm: Zero memory during shared memory transitions
irqchip/gic-v3-its: Resolve the default NUMA node explicitly
irqchip/gic-v3-its: Allocate shared tables using CoCo shared memory
allocator
dma-contiguous: Derive shared alignment from DMA attributes
dma-pool: Allocate CoCo atomic pools using CoCo shared memory
allocator
dma-direct: Align CoCo shared DMA allocations to the shared granule
size
swiotlb: Align shared IO TLB pools to the shared granule size
swiotlb: Reject misaligned restricted DMA pools for CoCo guests
dma-buf: system_heap: Limit scatterlist entries to the buffer size
dma-buf: system_heap: Allocate shared buffers using CoCo shared memory
allocator
swiotlb: Make rounded shared pool capacity allocatable
mm: Assert CoCo shared allocations may sleep
irqchip/gic-v3-its: Preallocate VPE L1 tables
arch/arm/mm/dma-mapping.c | 5 +-
arch/arm64/mm/pageattr.c | 3 +
arch/powerpc/platforms/pseries/svm.c | 2 +
arch/s390/mm/init.c | 3 +
arch/x86/coco/tdx/tdx.c | 3 +
arch/x86/hyperv/hv_init.c | 6 +-
arch/x86/hyperv/ivm.c | 4 +
arch/x86/kernel/kvmclock.c | 6 +-
arch/x86/mm/mem_encrypt_amd.c | 4 +
drivers/dma-buf/heaps/system_heap.c | 128 +++++-----
drivers/hv/connection.c | 41 ++--
drivers/hv/hv.c | 11 +-
drivers/hv/hv_common.c | 2 -
drivers/iommu/dma-iommu.c | 2 +-
drivers/irqchip/irq-gic-v3-its.c | 153 +++++++++---
drivers/irqchip/irq-gic-v3.c | 4 +-
drivers/virt/coco/pkvm-guest/arm-pkvm-guest.c | 3 +
include/linux/cc_shared.h | 39 +++
include/linux/dma-map-ops.h | 9 +-
include/linux/irqchip/arm-gic-v3.h | 3 +-
kernel/dma/contiguous.c | 41 +++-
kernel/dma/direct.c | 73 ++++--
kernel/dma/ops_helpers.c | 2 +-
kernel/dma/pool.c | 25 +-
kernel/dma/swiotlb.c | 80 ++++--
kernel/kexec_file.c | 3 +-
mm/Makefile | 1 +
mm/cc_shared.c | 232 ++++++++++++++++++
28 files changed, 682 insertions(+), 206 deletions(-)
create mode 100644 include/linux/cc_shared.h
create mode 100644 mm/cc_shared.c
base-commit: 704340f1cd0dcef829eb62f5b48ae95a2ce17bdf
--
2.43.0
On Wed, Sep 23, 2026 at 09:28:54AM -0700, Kameron Carr wrote:
> On 9/23/2026 8:21 AM, Jason Gunthorpe wrote:
> > On Wed, Sep 23, 2026 at 08:28:57PM +0530, Aneesh Kumar K.V wrote:
> >
> >> I'm also considering requiring the address passed to cc_make_shared() to
> >> be in the linear map. This is currently required by both TDX and CCA,
> >> while AMD SNP appears to support vmalloc addresses. The only user of
> >> that vmalloc support is Hyper-V VMBus GPADL setup
> >> (vmbus_establish_gpadl()). How should the generic CoCo shared-memory
> >> allocator handle this?
> >
> > vmbus_establish_gpadl() is the the same wrong abstraction as the arch
> > code. Get rid of it and use vmbus_establish_gpadl_caller_decrypted():
> >
> > pdata->recv_buf = vzalloc(RECV_BUFFER_SIZE);
> > if (!pdata->recv_buf) {
> > ret = -ENOMEM;
> > goto fail_free_ring;
> > }
> >
> > ret = vmbus_establish_gpadl(channel, pdata->recv_buf,
> > RECV_BUFFER_SIZE, &pdata->recv_gpadl);
> >
> >
> > So you made an allocator that returns folios, now you just need to
> > use that allocator to implement a kvzalloc wrapper. ARM can call vmap
> > on decrypted memory with pgprot_decrypted(), right?
> >
> > Or maybe this can use vmbus_alloc_buffer(), it already does it.
>
> Michael Kelley recently proposed [1] moving all ring buffer allocations
> to use vmbus_alloc_buffer(). If we move forward with this, it should
> remove the dependency on vmalloc support.
Even better vmbus_alloc_buffer() can use the new allocator API
directly so it doesn't need to open code the set_memory_decrypted arch
call.
Lets do it!
Jaon