PCI P2PDMA applies Request and Completion Redirect throughout both paths.
This misclassifies asymmetric and nested switches, and reports one answer
for every kind of TLP.
Three ACS controls act on TLP attributes the client chooses rather than
on the topology: Translation Blocking and Direct Translated P2P act on
a Request's Address Type, and Completion Redirect skips Completions carrying
Relaxed Ordering.
Evaluate each direction at the path divergence, decide every class from the
one walk, expose the provider to dma-buf importers, and let mlx5 ask rather
than assume.
Signed-off-by: Leon Romanovsky <leonro(a)nvidia.com>
---
Changes in v8:
- Added extra Reviewed-by from Logan.
- Added patches to take care of ATS per-device vs. per-mapping option
- Link to v7: https://patch.msgid.link/20260920-fix-p2p-acs-v4-0-v7-0-ca0828ab697c@nvidia…
Changes in v7:
- Removed "Document pdev->p2pdma lifetime rules" patch, it gives nothing
after p2pmem fix.
- Split dmabuf patch.
- Tushar retested the series, so added his Tested-by.
- Added Logan's ROB tags and fixed minor documentation issues pointed
by him.
- Link to v6: https://lore.kernel.org/all/20260914-fix-p2p-acs-v4-0-v6-0-5ef07ec9ef06@nvi…
Changes in v6:
- Changed DMABUF to use callbacks and not directly stored pointer.
- Link to v5: https://patch.msgid.link/20260910-fix-p2p-acs-v4-0-v5-0-856087f63c0d@nvidia…
Changes in v5:
- Rebase on the posted fixes.
- Dropped tags from changed patches.
- Remove Egress Control Vector interpretation and coverage.
- Keep enabled Egress Control conservative as a Request redirect.
- Use pci_dbg()/dev_dbg() for diagnostics and drop the "debug" prefix.
- Removed code comments from "Document the pdev->p2pdma lifetime and RCU
rules" patch and reduced description to actual lifetime explanation.
- Added note that Linux assumes that TLPs are in strict-ordering and
untranslated.
- Added code to calculate p2p paths per-TLP type.
- Converted mlx5 to use that new proposed API.
- Link to https://patch.msgid.link/20260821-fix-p2p-acs-v4-0-v4-0-94426b96de73@nvidia…
Changes in v4:
- Reject ACS Violations and unreadable routing state instead of treating
them as host-bridge redirects
- Added Tested-by tags from Tushar Dave
- Added support to asymmetric ACS routing
- Limited redirect checks to the two ports at the path divergence
- Added standalone ACS routing diagnostics for hardware retesting
- Dropped " PCI: Account for Direct Translated P2P in ACS isolation checks" patch
- Link to v3: https://patch.msgid.link/20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com
Changes in v3:
- Fixed pci_p2pdma_add_resource() error unwinding
- Made pdev->p2pdma teardown wait unconditionally for RCU readers
- Restricted pci_p2pmem_find_many() to pool-backed providers
- Documented the pdev->p2pdma lifetime and RCU rules
- Fixed calc_map_type_and_dist() handling of the verbose argument
- Required the ACS port and target to share a bus before indexing the
Egress Control Vector
- Gave pci_acs_enabled() and pci_acs_path_enabled() a scope, so the ACS
Direct Translated P2P rule no longer stops pci_enable_pasid() from
enabling PASID
- Dropped "Report ACS ports when the paths share no upstream bridge":
the mapping type cannot change without a shared upstream bridge, so
the pci=disable_acs_redir= hint was not actionable there and the ACS
walk only cost config space reads
- Folded the Request Redirect rule into pci_acs_rr_ineffective(), so
pci_acs_flags_enabled() and the Intel SPT PCH quirk share one copy
- Renamed pci_acs_egress_ctrl_set() to pci_acs_egress_ctrl_is_set(), it
reads the bit rather than setting it
- Reworded the blocked-path warning: ACS may also leave the direct route
indeterminate rather than blocked
- Added KUnit coverage for the shared-bus guard, a device with no ACS
capability and an unreadable ACS Control register
- Added the missing Fixes: tags, a second one on the
pci_p2pdma_add_resource() unwinding fix (the dangling devres action
dates to f58ef9d1d135) and one on the Egress Control isolation change
- Link to v2: https://patch.msgid.link/20260806-fix-p2p-acs-v2-0-0cec14812965@nvidia.com
Changes in v2:
- Added Logan's ROB tags
- Added commas in Documentation patch
- Link to v1: https://patch.msgid.link/20260802-fix-p2p-acs-v1-0-a7c5eb64fff6@nvidia.com
---
Leon Romanovsky (23):
PCI/P2PDMA: Document the TLP attribute assumptions
PCI/P2PDMA: Derive routing from directional ACS controls
PCI: Reject unreadable ACS controls in isolation checks
PCI/P2PDMA: Evaluate ACS controls at the path divergence
PCI/P2PDMA: Document directional ACS routing
PCI/P2PDMA: Collect the path's ACS controls before deciding
PCI/P2PDMA: Answer routing per TLP class
PCI/P2PDMA: Route Relaxed Ordering Completions directly
PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
PCI/P2PDMA: Log detailed ACS routing diagnostics
PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
PCI/P2PDMA: Test the ACS P2P routing walk
PCI: Add KUnit coverage for ACS isolation checks
PCI/P2PDMA: Document TLP-class routing
dma-buf: Let importers ask how peer-to-peer traffic is routed
vfio/pci: Hand out the P2PDMA provider behind a dma-buf
RDMA/uverbs: Hand out the P2PDMA provider behind a dma-buf
RDMA/mlx5: Ask P2PDMA whether ATS takes a direct peer-to-peer route
PCI/P2PDMA: Let a client declare that it selects ATS per mapping
RDMA/mlx5: Declare to P2PDMA that ATS is selected per memory key
PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
PCI/P2PDMA: Test the routing of clients with ATS enabled
Documentation/admin-guide/kernel-parameters.txt | 15 +-
Documentation/driver-api/pci/p2pdma.rst | 80 ++
drivers/dma-buf/dma-buf-mapping.c | 39 +
drivers/infiniband/core/uverbs_std_types_dmabuf.c | 12 +
drivers/infiniband/hw/mlx5/main.c | 2 +
drivers/infiniband/hw/mlx5/mlx5_ib.h | 36 +-
drivers/infiniband/hw/mlx5/mr.c | 47 ++
drivers/pci/Kconfig | 15 +
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 704 ++++++++++++++++--
drivers/pci/pci.c | 7 +-
drivers/pci/pci.h | 31 +
drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++
drivers/pci/quirks.c | 6 +-
drivers/vfio/pci/vfio_pci_dmabuf.c | 12 +
include/linux/dma-buf-mapping.h | 3 +
include/linux/dma-buf.h | 19 +
include/linux/pci-p2pdma.h | 62 +-
18 files changed, 1828 insertions(+), 123 deletions(-)
---
base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e(a)nvidia.com>
prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
Best regards,
--
Leon Romanovsky <leonro(a)nvidia.com>
Le 28/09/2026 à 15:32, Pavel Begunkov a écrit :
> The goal is to be able to natively use dma-buf in the read-write / IO
> path. This patch adds basic building blocks serving as a glue and API
> between drivers and upper layer subsystems providing the uAPI. Later
> patches implement it for NVMe raw block devices and expose it to the
> user space via io_uring.
>
> There are two main objects. struct dma_buf_io_ctx and struct
> dma_buf_io_map. The ctx is used during initial registration and serves
> as an interface between the upper layer user like io_uring and to the
> importer subsystem / driver. The map represents the actual dma map
> established for the target device[s] with dma_buf_map_attachment() and
> stored in a device specific format. The context is created via a new
> file operation ->init_dma_buf_io_ctx.
>
> The ctx-map separation exists to support map invalidation (see
> dma_buf_io_invalidate_mappings()). A ctx can create
> multiple maps during its lifetime, but there can only be no more than
> one (active) map attached to it. Invalidation drops the active map
> if present, and the next map will only be attempted to be created
> once there is a new request that wants to use the dma-buf IO ctx.
>
> The primary task of the dma_buf_io_map object is to count requests
> using it and to wait for their completion when we want to destroy the
> DMA map.
>
> [un]mapping and any work with dma addresses is delegated to the
> importer driver via an ops table stored in the ctx, see struct
> dma_buf_io_ops. Only the target driver / subsystem knows about devices
> it wants to use the dma-buf with, especially in case of multi-device
> filesystems or stacking in the future.
>
> Signed-off-by: Pavel Begunkov <asml.silence(a)gmail.com>
> ---
Hi,
a few nitpicks below, should it help
> +static void dma_buf_io_kill_maps(struct dma_buf_io_ctx *ctx, bool final)
> +{
> + struct dma_buf_io_map *map;
> +
> + scoped_guard(mutex, &ctx->map_mutex) {
guard() is enough. This saves indentation.
> + if (final)
> + ctx->maps_killed = true;
> +
> + map = rcu_dereference_protected(ctx->map,
> + lockdep_is_held(&ctx->map_mutex));
> + if (!map)
> + return;
> + rcu_assign_pointer(ctx->map, NULL);
> + percpu_ref_kill(&map->refs);
> + }
> +}
...
> +int dma_buf_io_ctx_create(struct file *file,
> + struct dma_buf *dmabuf,
> + enum dma_data_direction dir,
> + struct dma_buf_io_ctx **out_ctx)
> +{
> + struct dma_buf_io_ctx *ctx;
> + int ret;
> +
> + if (!file->f_op->init_dma_buf_io_ctx)
> + return -EOPNOTSUPP;
> +
> + ctx = kmalloc_obj(*ctx);
kzalloc_obj() and save the memset() below
> + if (!ctx)
> + return -ENOMEM;
> +
> + memset(ctx, 0, sizeof(*ctx));
> + ctx->dir = dir;
> + ctx->dmabuf = dmabuf;
> + get_dma_buf(dmabuf);
> + mutex_init(&ctx->map_mutex);
> + mutex_init(&ctx->map_create_mutex);
> + atomic_set(&ctx->active_maps, 0);
> + atomic_set(&ctx->all_maps, 0);
> + refcount_set(&ctx->refs, 1);
> + init_waitqueue_head(&ctx->maps_wq);
> +
> + ret = file->f_op->init_dma_buf_io_ctx(file, ctx);
> + if (ret) {
> + kfree(ctx);
> + dma_buf_put(dmabuf);
> + return ret;
> + }
> +
> + if (WARN_ON_ONCE(!ctx->dev_ops ||
> + !ctx->dev_ops->map ||
> + !ctx->dev_ops->unmap ||
> + !ctx->dev_ops->release))
> + return -EINVAL;
> +
> + *out_ctx = ctx;
> + return 0;
> +}
...
CJ
On 9/26/26 09:12, Karl Mehltretter wrote:
> get_sg_table() merges physically contiguous pages without accounting
> for the mapping device's maximum segment size. This affects both
> importer mappings and the udmabuf misc device mapping used for CPU
> access.
>
> With DMA_API_DEBUG enabled, DMA_BUF_IOCTL_SYNC on a 64 MiB udmabuf
> reports:
>
> DMA-API: misc udmabuf: mapping sg segment longer than device claims to support [len=65884160] [max=65536]
>
> Use sg_alloc_table_from_pages_segment() with the mapping device's
> maximum segment size. Keep a PAGE_SIZE minimum because the allocator
> warns and returns -EINVAL for smaller limits.
>
> Before commit 5bf888673e0d ("udmabuf: Do not create malformed
> scatterlists"), each entry covered one page.
>
> Fixes: 5bf888673e0d ("udmabuf: Do not create malformed scatterlists")
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter(a)gmail.com>
> ---
>
> Notes:
> Tested on v7.3-rc4-70-gfe2ec83746e5 in QEMU (x86_64, TCG) with
> DMA_API_DEBUG (all_errors=1) and DMABUF_DEBUG, A/B against the same
> base:
>
> before after
> DMA_BUF_IOCTL_SYNC, 64 MiB udmabuf 1 report 0
> vivid import, 4 MiB udmabuf 2 reports 0
> vivid import, 2 MiB hugetlb udmabuf 2 reports 0
> frames captured 5/5 5/5
>
> vb2-dma-contig rejected the non-contiguous import in both runs.
>
> drivers/dma-buf/udmabuf.c | 10 +++++++---
> 1 file changed, 7 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/dma-buf/udmabuf.c b/drivers/dma-buf/udmabuf.c
> index df6dd00462423..09f1eb8432f19 100644
> --- a/drivers/dma-buf/udmabuf.c
> +++ b/drivers/dma-buf/udmabuf.c
> @@ -139,9 +139,13 @@ static struct sg_table *get_sg_table(struct device *dev, struct dma_buf *buf,
> if (!sg)
> return ERR_PTR(-ENOMEM);
>
> - ret = sg_alloc_table_from_pages(sg, ubuf->pages, ubuf->pagecount, 0,
> - ubuf->pagecount << PAGE_SHIFT,
> - GFP_KERNEL);
> + /* The SG allocator requires a segment limit of at least PAGE_SIZE. */
> + ret = sg_alloc_table_from_pages_segment(sg, ubuf->pages, ubuf->pagecount,
> + 0, ubuf->pagecount << PAGE_SHIFT,
> + max_t(unsigned int,
> + dma_get_max_seg_size(dev),
> + PAGE_SIZE),
Please return -EINVAL instead when dma_get_max_seg_size() returns that the segment size is smaller than a page.
In general I think that the sg_alloc_table_from_pages_segment() approach is because of the broken design of the old DMA API. Stuff like that should be handled by the iterator going over the DMA segments instead. But yeah that is not something you can fix in one patch.
So apart from the error handling the patch looks good to me.
Regards,
Christian.
> + GFP_KERNEL);
> if (ret < 0)
> goto err_alloc;
>
On Tue, Sep 29, 2026 at 08:56:01AM +0200, Karl Mehltretter wrote:
> On Tue, Sep 29, 2026 at 06:00:27AM +0000, sashiko-bot(a)kernel.org wrote:
> > On architectures with a 256KB PAGE_SIZE (such as Hexagon or PowerPC), this
> > new validation evaluates to 65536 < 262144, unconditionally returning
> > -EINVAL and breaking all CPU mappings.
>
> Right, the misc device has no dma_parms, so it gets the 64K default and
> sync fails with 256K pages
I would ignore this, it seems fictional to me, and if someone really
needs to support this rediculous combination it should be done inside
sg_alloc_table_from_pages_segment(:)
> The misc device has no real segment limit, so for v3 I'll give it
> dma_parms with an unlimited max segment size and keep the -EINVAL for
> importers that report a limit below PAGE_SIZE.
The params have to come from the *importing* device, not the udmabuf
misc device. You don't get to control what they are, and there is no
reason to put a dma_parms on a misc deivce.
Jason
Hi all,
The goal of this series is to enable userspace driver designs that use
VFIO to export DMABUFs representing subsets of PCI device BARs, and
"vend" those buffers from a primary process to other subordinate
processes by fd. This is achieved by allowing the processes to mmap()
the DMABUFs; their access to the device is isolated to the exported
ranges. The primary userspace driver process can forcibly revoke
access to previously-shared buffers upon cleanup (without requiring
cooperation from the subordinate processes). This is an improvement
on sharing the VFIO device fd to subordinate processes, which would
allow global access. See the RFCs for background.
The existing VFIO PCI BAR mmap() becomes backed by a DMABUF too,
keeping common vm_ops and fault handler for VMAs from the VFIO device
and explicitly-exported DMABUFs. This will help future iommufd
emulation of VFIO Type1 peer-to-peer, making it easier to get a DMABUF
for a VFIO BAR as a DMA target.
Below are per-patch notes & background info, and at the bottom are
several related questions that reviewers may like to consider (worth
at least skipping to #1, a possible bug).
Notes on patches
================
mmap() conversion to use DMABUF underneath has been done for vfio-pci,
but not sub-drivers: nvgrace-gpu's mmap() override path is unchanged;
I kept this out of scope for now not least because I don't have a
thorough test setup for this system. I would prefer to help the
nvgrace-gpu maintainers enable BAR mmap() DMABUFs themselves.
vfio/pci: Remove DMABUF export dependency on vdev->memory_lock
In v5 of this series we discover that doing an export from the VFIO
mmap() path (fundamental!) adds a dependency between mmap_lock and
vdev->memory_lock(W) because export was using memory_lock(W) to
protect the vdev->dmabufs list and state within, and a deadlock
scenario leapt out to bite:
https://lore.kernel.org/kvm/9a615f22-c0d4-46ae-9654-db11e94e5fec@ozlabs.org/
The suggestion was to add a dedicated mutex/rwsem specifically for
the DMABUFs/list, which is cleaner than overloading memory_lock(W).
But whilst export could now downgrade to holding memory_lock(R) to
test __vfio_pci_memory_enabled(), a very similar deadlock can still
arise due to a memory_lock(W) elsewhere depending on a prior
memory_lock(R) to be released, and attempting to take
memory_lock(R) will queue behind the (W) for fairness (effectively
an R->R dependency). I'd overlooked that rwsem cannot guarantee
multiple readers.
To be able to export while holding mmap_lock, export cannot hold
memory_lock at all. Instead of testing __vfio_pci_memory_enabled()
(which tracks PCI_COMMAND.MSE and PM state), this patch tracks
device-global DMABUF revocation state in a new vdev->bars_revoked
flag updated by vfio_pci_dma_buf_move(). If an mmap() is somehow
performed during a period when DMABUFs are all revoked, then the
DMABUF is created revoked. Move() already bookends reset, PM
transitions etc., so subsequent revocation state changes work
as-is. This flag is also protected by the dmabuf_lock and thus
memory_lock is not required to export.
vfio/pci: Un-revoke DMABUFs in LOW_POWER_ENTRY_WITH_WAKEUP resume
On LOW_POWER entry, DMABUFs have a move(revoke=true), but the
runtime resume path didn't un-revoke. This adds a corresponding
move(revoke=false), which will later turn into
vfio_pci_unrevoke_bars().
NOTE: the unrevoke is reordered _before_ the eventfd_signal() in
vfio_pci_core_runtime_resume() to remove a window in which a waking
waiter could have observed the DMABUF state as still revoked (or
pm_runtime_engaged = true). (The UAPI docs state the event means
the resume's complete, so observing otherwise seemed unintended.)
Also, when DMABUFs are later mmap()ed, a waiter waking and taking a
fault could have seen the unrevoked state and SIGBUS just before
taking the memory_lock. (If the handler gets to acquire the lock,
though, the resume sequence is complete, and the handler observes
pm_runtime_engaged = false.)
This fix is in this series because the issue will impact CPU access
to the VMA as well (once they use DMABUFs), and so it's a strict
dependency of later commits.
dma-buf: Provide dma_buf_set_name()
Makes dma_buf_set_name() available for use (by helper patch),
taking a kernel-allocated string. The pre-existing local helper
becomes a wrapper copying a __user string for the set-name ioctl.
vfio/pci: Add a helper to look up PFNs for DMABUFs
Adds a DMABUF VMA fault handler helper to determine arbitrary-sized
PFNs from ranges in DMABUF.
vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA
Refactors DMABUF export for use by the existing export feature, and
adds a helper that creates a DMABUF corresponding to a VFIO BAR
mmap() request.
There was a request for decent debug naming in /proc/<pid>/maps
etc. comparable to the existing VFIO names: since the VMAs are
DMABUFs, they have a "dmabuf:" prefix and can't be 100% identical
to before. This is a user-visible change, but this patch at least
now gives us extra info on the BDF & BAR being mmap()ed. The name
is installed with the new dma_buf_set_name() above.
vfio/pci: Convert BAR mmap() to use a DMABUF
The vfio-pci core mmap() creates a DMABUF with the helper above,
and the vm_ops fault handler uses the other helper to resolve the
fault. Because this depends on DMABUF structs/code,
CONFIG_VFIO_PCI_CORE needs to depend on CONFIG_DMA_SHARED_BUFFER.
The CONFIG_VFIO_PCI_DMABUF still conditionally enables the export
support code.
NOTE: The user mmap()s a device fd, but the resulting VMA's vm_file
becomes that of the DMABUF. The DMABUF takes ownership of the
device file and put()s it on release, which maintains the existing
behaviour of a VMA keeping the VFIO device open.
BAR zapping then happens via the existing vfio_pci_dma_buf_move()
path, which now needs to unmap PTEs in the DMABUF's address_space.
NOTE: As local LLM reviews did, Sashiko might falsely worry about
the DMABUF fd being obtained through /proc/pid/map_files and
remapped, but the file is an anon inode and doesn't support the
open op (-ENXIO) so AFAICT this is currently impossible.
NOTE: A side-effect of this is that is_mergeable_vma() will be
false between adjacent mappings of VFIO BARs; merging would require
rebuilding a larger DMABUF representing the union, not just
plugging the VMAs together.
The DMABUF backing the BAR is an implicit/internal export, even
when the CONFIG_VFIO_PCI_DMABUF feature is not included because
CONFIG_PCI_P2PDMA is not available. Without P2PDMA, it's
acceptable for the DMABUF to not have a P2PDMA provider. In this
configuration, the VFIO DMABUF .attach prohibits any import, which
avoids getting as far as a WARN in dma_buf_map_attachment(), which
would fail without P2PDMA anyway.
vfio/pci: Clean up BAR zap and revocation
In general (see NOTE!) the vfio_pci_zap_bars() is now obsolete,
since it unmaps PTEs in the VFIO device address_space which is now
unused. This consolidates all calls (e.g. around reset) with the
neighbouring vfio_pci_dma_buf_move()s into new functions, to
revoke/unrevoke (making the steps clearer).
NOTE: Because drivers can use their own vm_ops and override .mmap,
the core must conservatively assume an overridden .mmap might still
add PTEs to the VFIO device address_space and therefore still does
the zap. A new flag, zap_bars_on_revoke, enables the zap when
.mmap is overridden. A driver that does not need the zap can clear
this to opt-out, e.g. if the driver calls down to the common mmap
(and so uses DMABUFs). hisi-acc-vfio-pci does just this, and thus
sets the opt-out flag.
vfio/pci: Support mmap() of a VFIO DMABUF
Adds mmap() for a DMABUF fd exported from vfio-pci.
It was a goal to keep the VFIO device fd lifetime behaviour
unchanged with respect to the DMABUFs. An application can close
all device fds, and this will revoke/clean up all DMABUFs; then, no
mappings or other access can be performed. When enabling mmap() of
the DMABUFs, this means access through the VMA is also revoked.
This complicates the fault handler because whilst the DMABUF
exists, it has no guarantee that the corresponding VFIO device is
still alive. Adds synchronisation ensuring the vdev is available
before the locks in vdev are touched; this holds the device
registration so that even if the buffer has been cleaned up, vdev
hasn't been freed and so the locks can be safely taken.
vfio/pci: Revoke a DMABUF on request from userspace
This is mostly a rename of `revoked` to an enum, `status`, and
adding a third state for a buffer: usable, revoked temporary,
revoked permanent. A new VFIO feature is added,
VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE, which takes a DMABUF (exported
from the same device) and permanently revokes it. Thus a userspace
driver can guarantee any downstream consumers of a shared fd are
prevented from accessing a BAR range, and that range can be reused.
NOTE: This might block userspace, waiting on importers to detach.
The code doing revocation in vfio_pci_dma_buf_move() is moved, to a
common function used by ..._move() and this new feature.
Testing
=======
(The [RFC ONLY] userspace test program, which drives a QEMU
bochs-display function, can be found in the GitHub branch below. It
at least illustrates how the export, map, revoke, and close semantics
interoperate. WIP on a follow-up with a proper vfio-selftests style
test based on this -- this won't be part of this series.)
This code has been tested in mapping DMABUFs of single/multiple ranges
from multiple BARs, aliasing mmap()s, aliasing ranges across DMABUFs,
vm_pgoff > 0, revocation, shutdown/cleanup scenarios, and hugepage
mappings. No regressions observed on the VFIO selftests, or on our
internal vfio-pci applications. VFIO on i386 has been build-tested.
Thanks to Alex Mastro for building a (WIP) testcase for the
prior mmap_lock->memory_lock issue.
Dear Reviewers,
===============
Along the way several related issues came up that warrant more
eyes, and I'd be grateful for your input:
1. If VFIO fd is opened O_RDONLY, it currently can't be mmap()ed
(because PROT_WRITE is rejected in do_mmap(), and PROT_READ alone
drops the VM_SHARED so VFIO's mmap rejects it). BUT it seems we
can export a DMABUF from it, and then pass the resulting fd around
for P2P writes.
I don't know if this is intentional/relied on/a known limitation,
or a bug?
a) We could reject export w/ -EPERM unless the device fd's f_mode
has O_RDWR, to reflect the RW abilities of P2P
If we agree it's a bug, I want to do this fix (a), as we can now
export a DMABUF RW from an O_RDONLY device fd and then succeed to
mmap() the DMABUF with RW. (That said, even with an O_RDONLY
device fd, the device state can still be changed/reset. But it
feels cleaner to prevent export for a O_RDONLY device fd, and match
the device fd mmap() behaviour.)
In future, we could consider finer-grained RD/WR if there's a
future goal to tie DMABUF permissions to, say, iommufd
IOMMU_READ/IOMMU_WRITE permissions:
b) Instead of just failing if !O_RDWR, we could limit the
get_dma_buf.open_flags to the VFIO device fd's f_mode, such as:
VFIO device fd perms: Export flags: Result:
O_RDWR O_RDWR, O_RDONLY OK
O_RDONLY O_RDONLY OK
O_RDONLY O_RDWR -EPERM
O_WRONLY * -EPERM
* O_WRONLY -EPERM
(Skipping WRONLY because a PROT_WRITE-only mmap() won't work,
though it probably should be included for P2P.)
2. The mmap fault handler takes a bunch of locks non-interruptibly,
and potentially depends on a lot of DMABUF-related activities
completing. I'd had a go at converting them to
interruptible/killable forms, but that revealed there seems to be a
wider issue if move/revoke doesn't complete in a timely fashion
(due to buggy importers). Where I got to was that just updating
the fault handler won't fix the user experience of an unkillable
task, and move()/revocation will need thought too. I don't intend
to fix this here but wanted to start discussion so we can address
it in a follow up. There's now a dependency between mmap_lock in
the fault handler and the DMABUF resv (which might take a while to
resolve), though revocation will be rare in practice.
3. vfio_basic_config_write() has an error path if
vfio_default_config_write() fails that releases memory_lock but
doesn't un-revoke BARs in the case of PCI_COMMAND.MSE being
cleared. When can the write fail, in practice, perhaps surprise
removal?
The effect on this series would be: a write of MSE=0 revokes BARs,
vconfig[PCI_COMMAND]'s MSE becomes 0, but if the physical write
fails then the physical MSE remains 1 and BAR VMAs stay revoked.
This seemed a mess; fixing isn't as simple as un-revoking on the
error path since vfio_default_config_write() has already trampled
vconfig so that'd need unwinding. It felt like a catastrophic
scenario where BARs staying revoked isn't a bad outcome, but want
to hear your experience of the likelihood of this issue.
4. The exchange of a VFIO fd mmap()'s vma->vm_file with an implicitly
created DMABUF's file has implications on LSM. For example, an
mmap will be checked against the policy for a VFIO fd, but a
subsequent mprotect() relates to the policy of the DMABUF file
(which is anon/unique to the mapping). This is pretty confusing.
END
===
This is based on v7.3-rc4
These commits are on GitHub for easier browsing, along with
"[RFC ONLY] selftests: vfio: Add standalone vfio_dmabuf_mmap_test":
https://github.com/metamev/linux/compare/v7.3-rc4...dev/mev/vfio-dmabuf-mma…
Thanks for reading,
Matt
================================================================================
Changelog:
v7:
- Rebased, v7.3-rc4
- "dma-buf: Export dma_buf_set_name()" is now "dma-buf: Provide
dma_buf_set_name()": reworded commit message with rationale, and
indicate it is only expected to be used by exporters. Remove static
helper for user ioctl & copy from user inline in the ioctl.
- "vfio/pci: Permanently revoke a DMABUF on request" is now
"vfio/pci: Revoke a DMABUF on request from userspace", with a
clarified commit message, and removal of "temporary/permanent"
language. This still replaces the priv->revoked flag with an
enum/third state, named OK/REVOKED for the existing two. The new
state, DEAD, is "sticky" from the VFIO-internal perspective, meaning
it prevents any of the move-style transitions from REVOKED back to
OK, guaranteeing that a DEAD buffer cannot ever be attached/mapped
by a new or existing importer. Clarified language around attach vs
map to not imply a racing import will be detached; it might attach,
but cannot map (and has no existing maps) after the ioctl returns.
- Reworded several commit messages for more clarity/brevity.
- "vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA":
As of the recent e8efdf02d3a97 ("vfio: Enable cdev noiommu mode
under iommufd") dev_name() can be much larger
(e.g. noiommu_vfio1048575) and the discovery that
drivers/pci/controller/vmd.c (others?) could create a domain up to
MAX_INT means there isn't a nice way to guarantee a debug name is
available across all cdev names and maximum sizes of all properties
(e.g. 1M cdevs, 4G domains). So, removed the cdev name from the
debug name construction, which now looks like 'vfio:0001:02:03.4/5'.
The cdev can be fished out of sysfs given the domain+BDF, and it's
still useful debug. Adjusted comments & commit message.
v6: https://lore.kernel.org/all/20260911214200.33793-1-matt@ozlabs.org/
v5: https://lore.kernel.org/all/20260715174737.15287-1-matt@ozlabs.org/
v4: https://lore.kernel.org/all/20260701171245.90111-1-matt@ozlabs.org/
v3: https://lore.kernel.org/all/20260610154327.37758-1-matt@ozlabs.org/
v2: https://lore.kernel.org/all/20260527102319.100128-1-mattev@meta.com/
v1: https://lore.kernel.org/kvm/20260416131815.2729131-1-mattev@meta.com/
RFCv2: https://lore.kernel.org/kvm/20260312184613.3710705-1-mattev@meta.com/
RFCv1: https://lore.kernel.org/all/20260226202211.929005-1-mattev@meta.com/
Tech topic: https://lore.kernel.org/linux-iommu/20250918214425.2677057-1-amastro@fb.com/
Matt Evans (9):
vfio/pci: Remove DMABUF export dependency on vdev->memory_lock
vfio/pci: Un-revoke DMABUFs in LOW_POWER_ENTRY_WITH_WAKEUP resume
dma-buf: Provide dma_buf_set_name()
vfio/pci: Add a helper to look up PFNs for DMABUFs
vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA
vfio/pci: Convert BAR mmap() to use a DMABUF
vfio/pci: Clean up BAR zap and revocation
vfio/pci: Support mmap() of a VFIO DMABUF
vfio/pci: Revoke a DMABUF on request from userspace
drivers/dma-buf/dma-buf.c | 78 ++-
drivers/vfio/pci/Kconfig | 4 +-
drivers/vfio/pci/Makefile | 3 +-
.../vfio/pci/hisilicon/hisi_acc_vfio_pci.c | 14 +-
drivers/vfio/pci/vfio_pci_config.c | 30 +-
drivers/vfio/pci/vfio_pci_core.c | 222 +++++--
drivers/vfio/pci/vfio_pci_dmabuf.c | 609 +++++++++++++++---
drivers/vfio/pci/vfio_pci_priv.h | 54 +-
include/linux/dma-buf.h | 2 +
include/linux/vfio_pci_core.h | 3 +
include/uapi/linux/vfio.h | 24 +
11 files changed, 856 insertions(+), 187 deletions(-)
--
2.50.1 (Apple Git-155)
On 9/28/26 12:36, Rob Clark wrote:
> On Mon, Sep 28, 2026 at 1:19 AM Christian König
> <christian.koenig(a)amd.com> wrote:
>>
>> On 9/26/26 04:20, Jianfeng Liu wrote:
>>> That commit fixed a dangling reference in the DMABUF_DEBUG default
>>> and thereby enabled the option - and with it the page-stripping
>>> sg_table wrapper that dma_buf_map_attachment() hands to importers -
>>> on every kernel with DEBUG_KERNEL=y, i.e. virtually every distro
>>> kernel.
>>>
>>> drm/msm is broken by the wrapper. Both of msm's map paths consume
>>> sg->length and sg_phys() of the attachment sg_table:
>>> msm_iommu_pagetable_map() for the per-process GPU pagetables, and
>>> iommu_map_sg() (via iommu_map_sgtable()) for scanout. The wrapper
>>> zeroes sg->length and strips the page pointers, so mappings of
>>> imported dma-bufs silently map nothing, and userspace observes
>>> arm-smmu translation faults from UCHE, e.g. during hardware video
>>> decode (clapper, chromium) on Adreno systems:
>>>
>>> gpu fault: ttbr0=000000088a889000 iova=000000010741c000 dir=READ
>>> type=TRANSLATION source=UCHE
>>>
>>> Bisected on a Snapdragon X1E78100 laptop as v7.3-rc3 good,
>>> v7.3-rc4 bad, culprit 143755bdabaa9.
>>>
>>> Switching msm to sg_dma_address()/sg_dma_len() is not a trivial fix
>>> either: those fields are only valid for sg_tables that msm has
>>> dma-mapped itself, which native non-MSM_BO_WC objects' sg_tables
>>> are not, so the conversion needs more work. The msm maintainer has
>>> therefore requested restoring the previous default for v7.3, to be
>>> revisited once msm no longer consumes struct page and sg->length of
>>> imported sg_tables.
>>
>> Yeah as I said before the problematic part is MSM here. We have enforced correct driver behavior for over 5 years now when that option is enabled.
>
> The problem is bigger than MSM here
Well, so far I have only heard about MSM.
But yes I mean the config option is doing exactly what it is supposed to do, pointing out when driver need some work to get this fixed.
I also agree that we shouldn't have allowed driver to touch that stuff in the first place and better document how to do things but yeah I can't change the past I can only try to fix it now.
>> What we can do is to mark MSM as broken and/or give a warning in MSM when DMABUF_DEBUG is enabled and you try to import a DMA-buf.
>
> sorry, no, we can't mark MSM as broken.. we can mark DMABUF_DEBUG as BROKEN
As I wrote the debug functionality to enforce not using struct pages has been around for over 5 years now, it was just not enabled by default.
What I can offer is to set it to default N for another few month to give you more time to fix things.
Regards,
Christian.
>
> BR,
> -R
>
>> But making the check not default to enable on debug kernels is not an option. This check here is exactly to point out broken drivers and you can manually disable it.
>>
>> Regards,
>> Christian.
>>
>>>
>>> Link: https://lore.kernel.org/linux-arm-msm/20260923074256.9357-1-liujianfeng1994…
>>> Suggested-by: Rob Clark <robin.clark(a)oss.qualcomm.com>
>>> Cc: Christian König <christian.koenig(a)amd.com>
>>> Cc: Sumit Semwal <sumit.semwal(a)linaro.org>
>>> Cc: Karl Mehltretter <kmehltretter(a)gmail.com>
>>>
>>> Signed-off-by: Jianfeng Liu <liujianfeng1994(a)gmail.com>
>>> ---
>>>
>>> drivers/dma-buf/Kconfig | 2 +-
>>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>>
>>> diff --git a/drivers/dma-buf/Kconfig b/drivers/dma-buf/Kconfig
>>> index e4f078a326a41..7efc0f0d07126 100644
>>> --- a/drivers/dma-buf/Kconfig
>>> +++ b/drivers/dma-buf/Kconfig
>>> @@ -43,7 +43,7 @@ config UDMABUF
>>> config DMABUF_DEBUG
>>> bool "DMA-BUF debug checks"
>>> depends on DMA_SHARED_BUFFER
>>> - default y if DEBUG_KERNEL
>>> + default y if DEBUG
>>> help
>>> This option enables additional checks for DMA-BUF importers and
>>> exporters. Specifically it validates that importers do not peek at the
>>> ---
>>> base-commit: 93f51579e7df248780214094418f205253383cc5
>>> branch: revert-dmabuf-debug-for-7.3
>>>
>>
On 9/26/26 20:30, Rob Clark wrote:
> This was always the way it was supposed to work, and when we drop the
> page array for imported dma-bufs our fault handling path will no longer
> work.
>
> Signed-off-by: Rob Clark <robin.clark(a)oss.qualcomm.com>
Nice to see that finally happening.
Reviewed-by: Christian König <christian.koenig(a)amd.com> for this patch here, Acked-by: Christian König <christian.koenig(a)amd.com> for the rest of the series.
Regards,
Christian.
> ---
> drivers/gpu/drm/msm/msm_gem.c | 24 ++++++++++++++++++++++++
> 1 file changed, 24 insertions(+)
>
> diff --git a/drivers/gpu/drm/msm/msm_gem.c b/drivers/gpu/drm/msm/msm_gem.c
> index e8390ebd5dd5..c90336b3b231 100644
> --- a/drivers/gpu/drm/msm/msm_gem.c
> +++ b/drivers/gpu/drm/msm/msm_gem.c
> @@ -23,6 +23,8 @@
> #include "msm_gpu.h"
> #include "msm_kms.h"
>
> +MODULE_IMPORT_NS("DMA_BUF");
> +
> static void update_device_mem(struct msm_drm_private *priv, ssize_t size)
> {
> uint64_t total_mem = atomic64_add_return(size, &priv->total_mem);
> @@ -338,6 +340,9 @@ static vm_fault_t msm_gem_fault(struct vm_fault *vmf)
> int err;
> vm_fault_t ret;
>
> + if (drm_WARN_ON_ONCE(obj->dev, drm_gem_is_imported(obj)))
> + return VM_FAULT_SIGBUS;
> +
> /*
> * vm_ops.open/drm_gem_mmap_obj and close get and put
> * a reference on obj. So, we dont need to hold one here.
> @@ -1126,6 +1131,25 @@ static int msm_gem_object_mmap(struct drm_gem_object *obj, struct vm_area_struct
> {
> struct msm_gem_object *msm_obj = to_msm_bo(obj);
>
> + if (drm_gem_is_imported(obj)) {
> + int ret;
> +
> + /* Reset both vm_ops and vm_private_data, so we don't end up with
> + * vm_ops pointing to our implementation if the dma-buf backend
> + * doesn't set those fields.
> + */
> + vma->vm_private_data = NULL;
> + vma->vm_ops = NULL;
> +
> + ret = dma_buf_mmap(obj->dma_buf, vma, 0);
> +
> + /* Drop the reference drm_gem_mmap_obj() acquired.*/
> + if (!ret)
> + drm_gem_object_put(obj);
> +
> + return ret;
> + }
> +
> vm_flags_set(vma, VM_PFNMAP | VM_DONTEXPAND | VM_DONTDUMP);
> vma->vm_page_prot = msm_gem_pgprot(msm_obj, vma_get_page_prot(vma));
>
On 9/26/26 04:20, Jianfeng Liu wrote:
> That commit fixed a dangling reference in the DMABUF_DEBUG default
> and thereby enabled the option - and with it the page-stripping
> sg_table wrapper that dma_buf_map_attachment() hands to importers -
> on every kernel with DEBUG_KERNEL=y, i.e. virtually every distro
> kernel.
>
> drm/msm is broken by the wrapper. Both of msm's map paths consume
> sg->length and sg_phys() of the attachment sg_table:
> msm_iommu_pagetable_map() for the per-process GPU pagetables, and
> iommu_map_sg() (via iommu_map_sgtable()) for scanout. The wrapper
> zeroes sg->length and strips the page pointers, so mappings of
> imported dma-bufs silently map nothing, and userspace observes
> arm-smmu translation faults from UCHE, e.g. during hardware video
> decode (clapper, chromium) on Adreno systems:
>
> gpu fault: ttbr0=000000088a889000 iova=000000010741c000 dir=READ
> type=TRANSLATION source=UCHE
>
> Bisected on a Snapdragon X1E78100 laptop as v7.3-rc3 good,
> v7.3-rc4 bad, culprit 143755bdabaa9.
>
> Switching msm to sg_dma_address()/sg_dma_len() is not a trivial fix
> either: those fields are only valid for sg_tables that msm has
> dma-mapped itself, which native non-MSM_BO_WC objects' sg_tables
> are not, so the conversion needs more work. The msm maintainer has
> therefore requested restoring the previous default for v7.3, to be
> revisited once msm no longer consumes struct page and sg->length of
> imported sg_tables.
Yeah as I said before the problematic part is MSM here. We have enforced correct driver behavior for over 5 years now when that option is enabled.
What we can do is to mark MSM as broken and/or give a warning in MSM when DMABUF_DEBUG is enabled and you try to import a DMA-buf.
But making the check not default to enable on debug kernels is not an option. This check here is exactly to point out broken drivers and you can manually disable it.
Regards,
Christian.
>
> Link: https://lore.kernel.org/linux-arm-msm/20260923074256.9357-1-liujianfeng1994…
> Suggested-by: Rob Clark <robin.clark(a)oss.qualcomm.com>
> Cc: Christian König <christian.koenig(a)amd.com>
> Cc: Sumit Semwal <sumit.semwal(a)linaro.org>
> Cc: Karl Mehltretter <kmehltretter(a)gmail.com>
>
> Signed-off-by: Jianfeng Liu <liujianfeng1994(a)gmail.com>
> ---
>
> drivers/dma-buf/Kconfig | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/dma-buf/Kconfig b/drivers/dma-buf/Kconfig
> index e4f078a326a41..7efc0f0d07126 100644
> --- a/drivers/dma-buf/Kconfig
> +++ b/drivers/dma-buf/Kconfig
> @@ -43,7 +43,7 @@ config UDMABUF
> config DMABUF_DEBUG
> bool "DMA-BUF debug checks"
> depends on DMA_SHARED_BUFFER
> - default y if DEBUG_KERNEL
> + default y if DEBUG
> help
> This option enables additional checks for DMA-BUF importers and
> exporters. Specifically it validates that importers do not peek at the
> ---
> base-commit: 93f51579e7df248780214094418f205253383cc5
> branch: revert-dmabuf-debug-for-7.3
>