PCI P2PDMA applies Request and Completion Redirect throughout both paths.
This misclassifies asymmetric and nested switches, and reports one answer
for every kind of TLP.
Three ACS controls act on TLP attributes the client chooses rather than
on the topology: Translation Blocking and Direct Translated P2P act on
a Request's Address Type, and Completion Redirect skips Completions carrying
Relaxed Ordering.
Evaluate each direction at the path divergence, decide every class from the
one walk, and treat a client with ATS enabled as translating unless its
driver declares per-mapping ATS.
This completes the P2PDMA side; dma-buf and mlx5 follow separately.
Signed-off-by: Leon Romanovsky <leonro(a)nvidia.com>
---
Changes in v9:
- Removed dmabuf patches for now.
- Removed mlx5 per-mapping ATS declaration for now; it comes back with
the dmabuf patches.
- Link to v8: https://patch.msgid.link/20260928-fix-p2p-acs-v4-0-v8-0-404453b9c435@nvidia…
Changes in v8:
- Added extra Reviewed-by from Logan.
- Added patches to take care of ATS per-device vs. per-mapping option
- Link to v7: https://patch.msgid.link/20260920-fix-p2p-acs-v4-0-v7-0-ca0828ab697c@nvidia…
Changes in v7:
- Removed "Document pdev->p2pdma lifetime rules" patch, it gives nothing
after p2pmem fix.
- Split dmabuf patch.
- Tushar retested the series, so added his Tested-by.
- Added Logan's ROB tags and fixed minor documentation issues pointed
by him.
- Link to v6: https://lore.kernel.org/all/20260914-fix-p2p-acs-v4-0-v6-0-5ef07ec9ef06@nvi…
Changes in v6:
- Changed DMABUF to use callbacks and not directly stored pointer.
- Link to v5: https://patch.msgid.link/20260910-fix-p2p-acs-v4-0-v5-0-856087f63c0d@nvidia…
Changes in v5:
- Rebase on the posted fixes.
- Dropped tags from changed patches.
- Remove Egress Control Vector interpretation and coverage.
- Keep enabled Egress Control conservative as a Request redirect.
- Use pci_dbg()/dev_dbg() for diagnostics and drop the "debug" prefix.
- Removed code comments from "Document the pdev->p2pdma lifetime and RCU
rules" patch and reduced description to actual lifetime explanation.
- Added note that Linux assumes that TLPs are in strict-ordering and
untranslated.
- Added code to calculate p2p paths per-TLP type.
- Converted mlx5 to use that new proposed API.
- Link to https://patch.msgid.link/20260821-fix-p2p-acs-v4-0-v4-0-94426b96de73@nvidia…
Changes in v4:
- Reject ACS Violations and unreadable routing state instead of treating
them as host-bridge redirects
- Added Tested-by tags from Tushar Dave
- Added support to asymmetric ACS routing
- Limited redirect checks to the two ports at the path divergence
- Added standalone ACS routing diagnostics for hardware retesting
- Dropped " PCI: Account for Direct Translated P2P in ACS isolation checks" patch
- Link to v3: https://patch.msgid.link/20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com
Changes in v3:
- Fixed pci_p2pdma_add_resource() error unwinding
- Made pdev->p2pdma teardown wait unconditionally for RCU readers
- Restricted pci_p2pmem_find_many() to pool-backed providers
- Documented the pdev->p2pdma lifetime and RCU rules
- Fixed calc_map_type_and_dist() handling of the verbose argument
- Required the ACS port and target to share a bus before indexing the
Egress Control Vector
- Gave pci_acs_enabled() and pci_acs_path_enabled() a scope, so the ACS
Direct Translated P2P rule no longer stops pci_enable_pasid() from
enabling PASID
- Dropped "Report ACS ports when the paths share no upstream bridge":
the mapping type cannot change without a shared upstream bridge, so
the pci=disable_acs_redir= hint was not actionable there and the ACS
walk only cost config space reads
- Folded the Request Redirect rule into pci_acs_rr_ineffective(), so
pci_acs_flags_enabled() and the Intel SPT PCH quirk share one copy
- Renamed pci_acs_egress_ctrl_set() to pci_acs_egress_ctrl_is_set(), it
reads the bit rather than setting it
- Reworded the blocked-path warning: ACS may also leave the direct route
indeterminate rather than blocked
- Added KUnit coverage for the shared-bus guard, a device with no ACS
capability and an unreadable ACS Control register
- Added the missing Fixes: tags, a second one on the
pci_p2pdma_add_resource() unwinding fix (the dangling devres action
dates to f58ef9d1d135) and one on the Egress Control isolation change
- Link to v2: https://patch.msgid.link/20260806-fix-p2p-acs-v2-0-0cec14812965@nvidia.com
Changes in v2:
- Added Logan's ROB tags
- Added commas in Documentation patch
- Link to v1: https://patch.msgid.link/20260802-fix-p2p-acs-v1-0-a7c5eb64fff6@nvidia.com
---
Leon Romanovsky (18):
PCI/P2PDMA: Document the TLP attribute assumptions
PCI/P2PDMA: Derive routing from directional ACS controls
PCI: Reject unreadable ACS controls in isolation checks
PCI/P2PDMA: Evaluate ACS controls at the path divergence
PCI/P2PDMA: Document directional ACS routing
PCI/P2PDMA: Collect the path's ACS controls before deciding
PCI/P2PDMA: Answer routing per TLP class
PCI/P2PDMA: Route Relaxed Ordering Completions directly
PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
PCI/P2PDMA: Log detailed ACS routing diagnostics
PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
PCI/P2PDMA: Test the ACS P2P routing walk
PCI: Add KUnit coverage for ACS isolation checks
PCI/P2PDMA: Document TLP-class routing
PCI/P2PDMA: Let a client declare that it selects ATS per mapping
PCI/P2PDMA: Evaluate the ATS path for clients with ATS enabled
PCI/P2PDMA: Test the routing of clients with ATS enabled
Documentation/admin-guide/kernel-parameters.txt | 15 +-
Documentation/driver-api/pci/p2pdma.rst | 80 +++
drivers/pci/Kconfig | 15 +
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 719 ++++++++++++++++++--
drivers/pci/pci.c | 7 +-
drivers/pci/pci.h | 58 ++
drivers/pci/pci_acs_test.c | 860 ++++++++++++++++++++++++
drivers/pci/quirks.c | 6 +-
include/linux/pci-p2pdma.h | 12 +-
10 files changed, 1687 insertions(+), 86 deletions(-)
---
base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e(a)nvidia.com>
prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
Best regards,
--
Leon Romanovsky <leonro(a)nvidia.com>
Hi all,
The goal of this series is to enable userspace driver designs that use
VFIO to export DMABUFs representing subsets of PCI device BARs, and
"vend" those buffers from a primary process to other subordinate
processes by fd. This is achieved by allowing the processes to mmap()
the DMABUFs; their access to the device is isolated to the exported
ranges. The primary userspace driver process can forcibly revoke
access to previously-shared buffers upon cleanup (without requiring
cooperation from the subordinate processes). This is an improvement
on sharing the VFIO device fd to subordinate processes, which would
allow unfettered access. See the RFCs for background.
The existing VFIO PCI BAR mmap() becomes backed by a DMABUF too,
keeping common vm_ops and fault handler for VMAs from the VFIO device
and explicitly-exported DMABUFs. This will help future iommufd
emulation of VFIO Type1 peer-to-peer, making it easier to get a DMABUF
for a VFIO BAR as a DMA target.
Below are per-patch notes & background info, and at the bottom are
several related questions that reviewers may like to consider (worth
at least skipping to #1, a possible bug).
Notes on patches
================
mmap() conversion to use DMABUF underneath has been done for vfio-pci,
but not sub-drivers:
nvgrace-gpu's mmap() override path is unchanged; I kept this out of
scope for now not least because I don't have a thorough test setup for
this system. I would prefer to help the nvgrace-gpu maintainers
enable BAR mmap() DMABUFs themselves.
It's also worth noting that one reason for nvgrace-gpu's special BAR
handling is to map RESMEM with a WC memory type. This is unchanged
for existing mmap() but WC is not currently supported for mapping
DMABUFs exported for that BAR; all DMABUF VMAs currently have the same
memory type and control of UC/WC for a mapping is future work.
vfio/pci: Remove DMABUF export dependency on vdev->memory_lock
In v5 of this series we discover that doing an export from the VFIO
mmap() path (fundamental!) adds a dependency between mmap_lock and
vdev->memory_lock(W) because export was using memory_lock(W) to
protect the vdev->dmabufs list and state within, and a deadlock
scenario leapt out to bite:
https://lore.kernel.org/kvm/9a615f22-c0d4-46ae-9654-db11e94e5fec@ozlabs.org/
The suggestion was to add a dedicated mutex/rwsem specifically for
the DMABUFs/list, which is cleaner than overloading memory_lock(W).
But whilst export could now downgrade to holding memory_lock(R) to
test __vfio_pci_memory_enabled(), a very similar deadlock can still
arise due to a memory_lock(W) elsewhere depending on a prior
memory_lock(R) to be released, and attempting to take
memory_lock(R) will queue behind the (W) for fairness (effectively
an R->R dependency). I'd overlooked that rwsem cannot guarantee
multiple readers.
To be able to export while holding mmap_lock, export cannot hold
memory_lock at all. Instead of testing __vfio_pci_memory_enabled()
(which tracks PCI_COMMAND.MSE and PM state), this patch tracks
device-global DMABUF revocation state in a new vdev->bars_revoked
flag updated by vfio_pci_dma_buf_move(). If an mmap() is somehow
performed during a period when DMABUFs are all revoked, then the
DMABUF is created revoked. Move() already bookends reset, PM
transitions etc., so subsequent revocation state changes work
as-is. This flag is also protected by the dmabuf_lock and thus
memory_lock is not required to export.
vfio/pci: Un-revoke DMABUFs in LOW_POWER_ENTRY_WITH_WAKEUP resume
On LOW_POWER entry, DMABUFs have a move(revoke=true), but the
runtime resume path didn't un-revoke. This adds a corresponding
move(revoke=false), which will later turn into
vfio_pci_unrevoke_bars().
NOTE: the unrevoke is reordered _before_ the eventfd_signal() in
vfio_pci_core_runtime_resume() to remove a window in which a waking
waiter could have observed the DMABUF state as still revoked (or
pm_runtime_engaged = true). (The UAPI docs state the event means
the resume's complete, so observing otherwise seemed unintended.)
Also, when DMABUFs are later mmap()ed, a waiter waking and taking a
fault could have seen the unrevoked state and SIGBUS just before
taking the memory_lock. (If the handler gets to acquire the lock,
though, the resume sequence is complete, and the handler observes
pm_runtime_engaged = false.)
This fix is in this series because the issue will impact CPU access
to the VMA as well (once they use DMABUFs), and so it's a strict
dependency of later commits.
dma-buf: Provide dma_buf_set_name()
Makes dma_buf_set_name() available for use (by helper patch),
taking a kernel-allocated string. The pre-existing local helper
becomes a wrapper copying a __user string for the set-name ioctl.
vfio/pci: Add a helper to look up PFNs for DMABUFs
Adds a DMABUF VMA fault handler helper to determine arbitrary-sized
PFNs from ranges in DMABUF.
vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA
Refactors DMABUF export for use by the existing export feature, and
adds a helper that creates a DMABUF corresponding to a VFIO BAR
mmap() request.
There was a request for decent debug naming in /proc/<pid>/maps
etc. comparable to the existing VFIO names: since the VMAs are
DMABUFs, they have a "dmabuf:" prefix and can't be 100% identical
to before. This is a user-visible change, but this patch at least
now gives us extra info on the BDF & BAR being mmap()ed. The name
is installed with the new dma_buf_set_name() above.
vfio/pci: Convert BAR mmap() to use a DMABUF
The vfio-pci core mmap() creates a DMABUF with the helper above,
and the vm_ops fault handler uses the other helper to resolve the
fault. Because this depends on DMABUF structs/code,
CONFIG_VFIO_PCI_CORE needs to depend on CONFIG_DMA_SHARED_BUFFER.
The CONFIG_VFIO_PCI_DMABUF still conditionally enables the export
support code.
NOTE: The user mmap()s a device fd, but the resulting VMA's vm_file
becomes that of the DMABUF. The DMABUF takes ownership of the
device file and put()s it on release, which maintains the existing
behaviour of a VMA keeping the VFIO device open.
BAR zapping then happens via the existing vfio_pci_dma_buf_move()
path, which now needs to unmap PTEs in the DMABUF's address_space.
NOTE: As local LLM reviews did, Sashiko might falsely worry about
the DMABUF fd being obtained through /proc/pid/map_files and
remapped, but the file is an anon inode and doesn't support the
open op (-ENXIO) so AFAICT this is currently impossible.
NOTE: A side-effect of this is that is_mergeable_vma() will be
false between adjacent mappings of VFIO BARs; merging would require
rebuilding a larger DMABUF representing the union, not just
plugging the VMAs together.
The DMABUF backing the BAR is an implicit/internal export, even
when the CONFIG_VFIO_PCI_DMABUF feature is not included because
CONFIG_PCI_P2PDMA is not available. Without P2PDMA, it's
acceptable for the DMABUF to not have a P2PDMA provider. In this
configuration, the VFIO DMABUF .attach prohibits any import, which
avoids getting as far as a WARN in dma_buf_map_attachment(), which
would fail without P2PDMA anyway.
vfio/pci: Clean up BAR zap and revocation
In general (see NOTE!) the vfio_pci_zap_bars() is now obsolete,
since it unmaps PTEs in the VFIO device address_space which is now
unused. This consolidates all calls (e.g. around reset) with the
neighbouring vfio_pci_dma_buf_move()s into new functions, to
revoke/unrevoke (making the steps clearer).
NOTE: Because drivers can use their own vm_ops and override .mmap,
the core must conservatively assume an overridden .mmap might still
add PTEs to the VFIO device address_space and therefore still does
the zap. A new flag, zap_bars_on_revoke, enables the zap when
.mmap is overridden. A driver that does not need the zap can clear
this to opt-out, e.g. if the driver calls down to the common mmap
(and so uses DMABUFs). hisi-acc-vfio-pci does just this, and thus
sets the opt-out flag.
vfio/pci: Support mmap() of a VFIO DMABUF
Adds mmap() for a DMABUF fd exported from vfio-pci.
It was a goal to keep the VFIO device fd lifetime behaviour
unchanged with respect to the DMABUFs. An application can close
all device fds, and this will revoke/clean up all DMABUFs; then, no
mappings or other access can be performed. When enabling mmap() of
the DMABUFs, this means access through the VMA is also revoked.
This complicates the fault handler because whilst the DMABUF
exists, it has no guarantee that the corresponding VFIO device is
still alive. Adds synchronisation ensuring the vdev is available
before the locks in vdev are touched; this holds the device
registration so that even if the buffer has been cleaned up, vdev
hasn't been freed and so the locks can be safely taken.
vfio/pci: Revoke a DMABUF on request from userspace
This is mostly a rename of `revoked` to an enum, `status`, and
adding a third state for a buffer: usable, revoked temporary,
revoked permanent. A new VFIO feature is added,
VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE, which takes a DMABUF (exported
from the same device) and permanently revokes it. Thus a userspace
driver can guarantee any downstream consumers of a shared fd are
prevented from accessing a BAR range, and that range can be reused.
NOTE: This might block userspace, waiting on importers to detach.
The code doing revocation in vfio_pci_dma_buf_move() is moved, to a
common function used by ..._move() and this new feature.
Testing
=======
(The [RFC ONLY] userspace test program, which drives a QEMU
bochs-display function, can be found in the GitHub branch below. It
at least illustrates how the export, map, revoke, and close semantics
interoperate. WIP on a follow-up with a proper vfio-selftests style
test based on this -- this won't be part of this series.)
This code has been tested in mapping DMABUFs of single/multiple ranges
from multiple BARs, aliasing mmap()s, aliasing ranges across DMABUFs,
vm_pgoff > 0, revocation, shutdown/cleanup scenarios, and hugepage
mappings. No regressions observed on the VFIO selftests, or on our
internal vfio-pci applications. VFIO on i386 has been build-tested.
Thanks to Alex Mastro for building a (WIP) testcase for the
prior mmap_lock->memory_lock issue.
END
===
This is based on v7.3-rc6
These commits are on GitHub for easier browsing, along with
"[RFC ONLY] selftests: vfio: Add standalone vfio_dmabuf_mmap_test":
https://github.com/evansm7/linux/compare/v7.3-rc6...dev/mev/vfio-dmabuf-mma…
Also please see...
https://lore.kernel.org/all/f8efaabd-c06e-4f04-8cfa-489c148e37ba@ozlabs.org/
("[PATCH] dma-buf: Annul dmabuf->file on file release")
...which fixes a stale file pointer issue that could affect VFIO's
DMABUF export. Although the issue exists prior to this series, the
addition of another dmabuf->file dereference adds additional
unpleasant failure modes (for example, unmap_mapping_range() the wrong
file, or good old random data dereference, etc.)
Thanks for reading,
Matt
================================================================================
Changelog:
v8:
- Rebased, on v7.3-rc6
- "dma-buf: Provide dma_buf_set_name()": Comment and "DOC: locking
convention" updates. dma_buf_set_name() is is listed in the "do
not hold resv" category
v7: https://lore.kernel.org/all/20260924152159.49702-1-matt@ozlabs.org/
v6: https://lore.kernel.org/all/20260911214200.33793-1-matt@ozlabs.org/
v5: https://lore.kernel.org/all/20260715174737.15287-1-matt@ozlabs.org/
v4: https://lore.kernel.org/all/20260701171245.90111-1-matt@ozlabs.org/
v3: https://lore.kernel.org/all/20260610154327.37758-1-matt@ozlabs.org/
v2: https://lore.kernel.org/all/20260527102319.100128-1-mattev@meta.com/
v1: https://lore.kernel.org/kvm/20260416131815.2729131-1-mattev@meta.com/
RFCv2: https://lore.kernel.org/kvm/20260312184613.3710705-1-mattev@meta.com/
RFCv1: https://lore.kernel.org/all/20260226202211.929005-1-mattev@meta.com/
Tech topic: https://lore.kernel.org/linux-iommu/20250918214425.2677057-1-amastro@fb.com/
Matt Evans (9):
vfio/pci: Remove DMABUF export dependency on vdev->memory_lock
vfio/pci: Un-revoke DMABUFs in LOW_POWER_ENTRY_WITH_WAKEUP resume
dma-buf: Provide dma_buf_set_name()
vfio/pci: Add a helper to look up PFNs for DMABUFs
vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA
vfio/pci: Convert BAR mmap() to use a DMABUF
vfio/pci: Clean up BAR zap and revocation
vfio/pci: Support mmap() of a VFIO DMABUF
vfio/pci: Revoke a DMABUF on request from userspace
drivers/dma-buf/dma-buf.c | 84 ++-
drivers/vfio/pci/Kconfig | 4 +-
drivers/vfio/pci/Makefile | 3 +-
.../vfio/pci/hisilicon/hisi_acc_vfio_pci.c | 14 +-
drivers/vfio/pci/vfio_pci_config.c | 30 +-
drivers/vfio/pci/vfio_pci_core.c | 222 +++++--
drivers/vfio/pci/vfio_pci_dmabuf.c | 609 +++++++++++++++---
drivers/vfio/pci/vfio_pci_priv.h | 54 +-
include/linux/dma-buf.h | 16 +-
include/linux/vfio_pci_core.h | 3 +
include/uapi/linux/vfio.h | 24 +
11 files changed, 870 insertions(+), 193 deletions(-)
--
2.47.3
The dmabuf release path is split between file and dentry release:
dma_buf_file_release() is called shortly before the file is freed, and
dma_buf_release() calls an exporter's dmabuf->ops->release when the
dentry is freed. However, the dentry can outlive the file (for
example, if opened with O_PATH), meaning .release might be called some
time after the file is freed.
This presents a window, when closing a DMABUF file, in which
dmabuf->file points to freed memory yet .release op has not yet been
called.
For VFIO, if a buffer's .release has not been called it's considered
still active and subject to move/cleanup. If so, it attempts to
get_file_active() on dmabuf->file: a file close with dentry held open
will call this function with a stale pointer, a UAF.
To make this pattern safe, set dmabuf->file to NULL in
dma_buf_file_release() to reflect that the associated file is now dead
even if the DMABUF is not yet gone. A get_file_active() ... fput()
sequence concurrent with a file close will only execute as one of:
- Gets the file before file_ref_put() (dmabuf->file valid)
- Observes dmabuf->file pointing to a file, but it's DEAD (no file)
- Observes dmabuf->file = NULL (no file)
Originally, drivers could assume dmabuf->file was valid until .release
was called from fops->release. This assumption was no longer valid
after 4ab59c3c638c6 ("dma-buf: Move dma_buf_release() from fops to
dentry_ops"), which moved the callback to the dentry release (by which
point the file might have been freed). With this commit, drivers must
still consider that dmabuf->file could be NULL before .release.
Fixes: 4ab59c3c638c6 ("dma-buf: Move dma_buf_release() from fops to dentry_ops")
Signed-off-by: Matt Evans <matt(a)ozlabs.org>
---
Hi,
This issue was found (by Claude Opus 5.5) in the context of VFIO's
DMABUF export path. VFIO iterates live DMABUFs with a
get_file_active()/fput() block, which now becomes safe if the file is
closed (and memory freed!) yet DMABUF .release hasn't yet occurred.
However, there are a couple of other places that directly use
dmabuf->file and seem able to race a closing file (i.e. without holding
the file reference)? If this is so, they'd be a UAF today; with this
patch that goes away, but instead of a stale pointer dmabuf->file could
be NULL:
1. drivers/gpu/drm/vmwgfx/ttm_object.c:get_dma_buf_unless_doomed()
file_ref_get(&dmabuf->file->f_ref); on a non-refcounted DMABUF
The commit message of 90ee6ed776c0 ("fs: port files to file_ref") hints
this might be more subtle than replacing it with a get_file() variant
(so as to accept a NULL file *).
2. drivers/gpu/drm/i915/gvt/dmabuf.c:intel_vgpu_get_dmabuf()
gvt_dbg_dpy(... file_count(dmabuf->file) ...);
Respective vmwgfx & i915 maintainers, what is your view?
Or, indeed, if anyone sees any other questionable uses of dmabuf->file.
Thanks,
Matt
drivers/dma-buf/dma-buf.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c
index 4c9add51f9ef..726639130477 100644
--- a/drivers/dma-buf/dma-buf.c
+++ b/drivers/dma-buf/dma-buf.c
@@ -193,11 +193,15 @@ static void dma_buf_release(struct dentry *dentry)
static int dma_buf_file_release(struct inode *inode, struct file *file)
{
+ struct dma_buf *dmabuf = file->private_data;
+
if (!is_dma_buf_file(file))
return -EINVAL;
- __dma_buf_list_del(file->private_data);
+ __dma_buf_list_del(dmabuf);
+ /* Must be observed by __get_file_rcu() before file_free() */
+ smp_store_mb(dmabuf->file, NULL);
return 0;
}
--
2.47.3
On 10/6/26 02:48, Val Packett wrote:
> Shared memory buffers for software-rendered GUI applications are often
> large and highly fragmented in memory. To get a shareable handle to such
> a buffer from a guest VM, a virtio-gpu device backend can use the
> UDMABUF_CREATE_LIST ioctl with the memfds backing guest RAM and the list
> of guest physical pages sent by the guest driver. That list reaches 80K
> entries for a maximized window on a 4K display!
>
> Like 44e9eb5a7621 ("dma-buf/udmabuf: Disable the size limit by default")
> did for the size limit, change the default to INT_MAX since the amount
> of protection provided by this limit seems dubious.
>
> Signed-off-by: Val Packett <val(a)invisiblethingslab.com>
> ---
> drivers/dma-buf/udmabuf.c | 4 ++--
> 1 file changed, 2 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/dma-buf/udmabuf.c b/drivers/dma-buf/udmabuf.c
> index df6dd0046242..589031133588 100644
> --- a/drivers/dma-buf/udmabuf.c
> +++ b/drivers/dma-buf/udmabuf.c
> @@ -16,9 +16,9 @@
> #include <linux/vmalloc.h>
> #include <linux/iosys-map.h>
>
> -static int list_limit = 1024;
> +static int list_limit = INT_MAX;
> module_param(list_limit, int, 0644);
> -MODULE_PARM_DESC(list_limit, "udmabuf_create_list->count limit. Default is 1024.");
> +MODULE_PARM_DESC(list_limit, "udmabuf_create_list->count limit. Default is INT_MAX.");
This will completely overflow the size calculation in udmabuf_ioctl_create_list().
Please fix udmabuf_ioctl_create_list() to use memdup_array_user() and use a reasonable limit here. Something like 128k should probably do.
Apart from that I think we need to start using huge and giant pages for virtio-gpu on the client side, we already had it multiple times that we hit limits with that in multiple places.
Sharing 80k individual 4k pointers is really not very efficient.
Regards,
Christian.
>
> static int size_limit_mb = INT_MAX;
> module_param(size_limit_mb, int, 0644);
On 10/5/26 15:20, Fred Griffoul wrote:
> On 10/5/26 Christian König wrote:
>> Well filling page tables by the importer is an absolutely clear NO-GO
>> for the DMA-buf design, we have gotten down that path already and it
>> took us years to remove this functionality again.
>
> Understood, thanks for the quick and clear answers on this patch and on
> patch 3.
>
>> The problem you are facing here is that dma_buf_mmap() doesn't work
>> because you don't have a VMA.
>
> Right: the memory has no struct page and is deliberately not mapped in
> the host, so there is no VMA to hand to KVM.
>
> I'll post a new version of this series. It drops get_phys(), the ranged
> invalidation and the guest_memfd dma-buf import (patches 2-5), and it
> does not touch drivers/dma-buf, as you suggested in the PAL discussion:
>
> https://lore.kernel.org/all/c413710b-4c28-4ed8-88ec-aeb8c4482011@amd.com/
>
> KVM gets its frames through the KVM-private guest_memfd provider
> interface from David's series. iommufd gets them through a private
> interface with the exporter, as it already does for vfio-pci. The new
> version adds a small registry in iommufd for that, which replaces the
> symbol_get() of the sample in David's series. The dma-buf is then only
> the handle and the existing revoke.
Yeah that approach sounds totally sane to me.
We have drivers which stuff all kind of resources, not just memory, into a DMA-buf. That ranges from MMIO BARs to trigger FW operations (doorbells) over full HW blocks (ordered appends, global wave sync etc...) where two or more applications need to share which one is used.
That you then have a exporter private interface is for accessing this is perfectly ok.
The problem you are facing here is more that this is not limited to one driver/module but multiple ones and you don't want module inter dependencies because of the symbols.
So symbol_get() is probably the right thing to do as long as we don't have something like weak symbols or similar.
Regards,
Christian.
>
> Thanks,
> Fred
On 10/5/26 11:55, Fred Griffoul wrote:
> From: Fred Griffoul <fgriffo(a)amazon.co.uk>
>
> dma_buf_invalidate_mappings() tells every importer that the whole
> buffer changed. An exporter that changes one part of its memory cannot
> say which bytes changed, so importers throw away mappings that are
> still valid.
>
> Add an exporter helper that invalidates a byte range, and an importer
> callback that receives it. The callback means that the address, the
> attributes or the backing of the range changed. If part of the range is
> no longer backed, get_phys() returns -ENOENT for it. Importers must stop
> using their old answer before the callback returns. Importers that do
> not implement the callback still receive a whole-buffer invalidation.
Yeah that is exactly one of the reasons why we don't allow that.
Clear NAK to the whole approach. See the reply to patch #2 for a detailed description.
Regards,
Christian.
>
> Signed-off-by: Fred Griffoul <fgriffo(a)amazon.co.uk>
> ---
> drivers/dma-buf/dma-buf.c | 30 ++++++++++++++++++++++++++++++
> include/linux/dma-buf.h | 16 ++++++++++++++++
> 2 files changed, 46 insertions(+)
>
> diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c
> index 66b85d53ed22..e7010163eb2f 100644
> --- a/drivers/dma-buf/dma-buf.c
> +++ b/drivers/dma-buf/dma-buf.c
> @@ -1389,6 +1389,36 @@ void dma_buf_invalidate_mappings(struct dma_buf *dmabuf)
> }
> EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings, "DMA_BUF");
>
> +/**
> + * dma_buf_invalidate_mappings_range - notify attachments that a range changed
> + * @dmabuf: buffer whose layout changed
> + * @offset: first changed byte
> + * @length: number of changed bytes
> + *
> + * Importers with a ranged callback stop using their old mappings of the range
> + * before returning. Other importers receive the existing whole-buffer
> + * callback, which is correct but coarser. The reservation lock must be held.
> + */
> +void dma_buf_invalidate_mappings_range(struct dma_buf *dmabuf,
> + unsigned long offset,
> + unsigned long length)
> +{
> + struct dma_buf_attachment *attach;
> +
> + dma_resv_assert_held(dmabuf->resv);
> + list_for_each_entry(attach, &dmabuf->attachments, node) {
> + const struct dma_buf_attach_ops *ops = attach->importer_ops;
> +
> + if (!ops)
> + continue;
> + if (ops->invalidate_mappings_range)
> + ops->invalidate_mappings_range(attach, offset, length);
> + else if (ops->invalidate_mappings)
> + ops->invalidate_mappings(attach);
> + }
> +}
> +EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings_range, "DMA_BUF");
> +
> /**
> * dma_buf_get_phys - describe the run that starts at an offset
> * @attach: attachment to query
> diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h
> index b223962e20c2..55c3fe60a0ba 100644
> --- a/include/linux/dma-buf.h
> +++ b/include/linux/dma-buf.h
> @@ -485,6 +485,19 @@ struct dma_buf_attach_ops {
> * required behavior.
> */
> void (*invalidate_mappings)(struct dma_buf_attachment *attach);
> +
> + /**
> + * @invalidate_mappings_range: [optional] a byte range changed
> + *
> + * The exporter changed the address, attributes or backing of
> + * [@offset, @offset + @length). The importer must stop using its old
> + * answer for that range before returning.
> + * Importers without this callback receive @invalidate_mappings for
> + * the whole buffer instead.
> + */
> + void (*invalidate_mappings_range)(struct dma_buf_attachment *attach,
> + unsigned long offset,
> + unsigned long length);
> };
>
> /**
> @@ -600,6 +613,9 @@ struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *,
> void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,
> enum dma_data_direction);
> void dma_buf_invalidate_mappings(struct dma_buf *dma_buf);
> +void dma_buf_invalidate_mappings_range(struct dma_buf *dma_buf,
> + unsigned long offset,
> + unsigned long length);
> bool dma_buf_attach_revocable(struct dma_buf_attachment *attach);
> /* bits 0-7: memory type (a value, not flags) */
> #define DMA_BUF_PHYS_ATTR_TYPE_MASK GENMASK(7, 0)
> --
> 2.47.3
>
On Mon, Sep 28, 2026 at 02:32:22PM +0100, Pavel Begunkov wrote:
> +static void nvme_dmabuf_map_sync_for_cpu(struct nvme_dev *nvme_dev,
> + struct request *req)
> +{
> + struct device *dev = nvme_dev->dev;
> + enum dma_data_direction dma_dir;
> + struct bio *bio = req->bio;
> + struct nvme_dmabuf_map *map = to_nvme_dmabuf_map(bio->bi_dmabuf_map);
> +
> + dma_dir = rq_data_dir(req) == READ ? DMA_FROM_DEVICE : DMA_TO_DEVICE;
> + dma_sync_sgtable_for_cpu(dev, map->sgt, dma_dir);
This can use rq_dma_dir in the argument. Same for the other DMA API
calls.
Btw, given how larger the dmabuf code is now, have you considered
splitting it into a new nvme-pci-dmabuf.c file?
On Mon, Sep 28, 2026 at 02:32:21PM +0100, Pavel Begunkov wrote:
> Add a simple proxy implementation of init_dma_buf_io_ctx() forwarding
> the call to a new struct block_device_operations operation. Also reject
> dma-buf backed iterators for buffered IO.
>
> Reviewed-by: Christoph Hellwig <hch(a)lst.de>
> [pavel: reject dma-buf without O_DIRECT]
> Signed-off-by: Pavel Begunkov <asml.silence(a)gmail.com>
The Sashiko report about the errno on fallback here looks real,
and is probably easily fixed by setting a proper errno.