Hi all,
The goal of this series is to enable userspace driver designs that use
VFIO to export DMABUFs representing subsets of PCI device BARs, and
"vend" those buffers from a primary process to other subordinate
processes by fd. These processes then mmap() the buffers and their
access to the device is isolated to the exported ranges. This is an
improvement on sharing the VFIO device fd to subordinate processes,
which would allow unfettered access.
This is achieved by enabling mmap() of vfio-pci DMABUFs, passed by fd
to subordinate processes. Second, a new revocation mechanism is added
to allow the primary process to forcibly revoke access to
previously-shared BAR spans, even if the subordinate processes haven't
cleanly exited.
(The related topic of safe delegation of iommufd control to the
subordinate processes is not addressed here, and is follow-up work.)
The background/rationale is covered in more detail in the RFC cover
letters.
Reviews requested that we migrate the existing VFIO PCI BAR mmap() to
be backed by a DMABUF too, resulting in a common vm_ops and fault
handler for mmap()s of both the VFIO device and explicitly-exported
DMABUFs. This will help future iommufd emulation of VFIO Type1
peer-to-peer, making it easier to get a DMABUF for a VFIO BAR as a DMA
target.
mmap() conversion to use DMABUF underneath has been done for vfio-pci,
but not sub-drivers:
nvgrace-gpu's mmap() override path is unchanged; I kept this out of
scope for now not least because I don't have a thorough test setup
for this system. I would prefer to help the nvgrace-gpu maintainers
enable BAR mmap() DMABUFs themselves.
Notes on patches
================
vfio/pci: Remove DMABUF export dependency on vdev->memory_lock
In v5 of this series [1] we discover that doing an export from the
VFIO mmap() path (fundamental!) adds a dependency between mmap_lock
and vdev->memory_lock(W) because export was using memory_lock(W) to
protect the vdev->dmabufs list and state within, and a deadlock
scenario leapt out to bite. (Details in [1].)
The suggestion was to add a dedicated mutex/rwsem specifically for
the DMABUFs/list, which is cleaner than overloading memory_lock(W).
But whilst export could now downgrade to holding memory_lock just
for _read_ to test __vfio_pci_memory_enabled(), a very similar
deadlock can still arise due to a memory_lock(W) elsewhere
depending on a prior memory_lock(R) to be released, and attempting
to take memory_lock(R) will queue behind the (W) for fairness
(effectively an R->R dependency). I'd overlooked that rwsem cannot
guarantee multiple readers.
To be able to export while holding mmap_lock, export cannot hold
memory_lock at all. Instead of testing __vfio_pci_memory_enabled()
(which tracks PCI_COMMAND.MSE and PM state), this patch tracks
device-global DMABUF revocation state in a new vdev->bars_revoked
flag updated by vfio_pci_dma_buf_move(). If an mmap() is somehow
performed during a period when DMABUFs are all revoked, then the
DMABUF is created revoked. Move() already bookends reset, PM
transitions etc., so subsequent revocation state changes work
as-is. This flag is also protected by the dmabuf_lock and thus
memory_lock is not required to export.
vfio/pci: Un-revoke DMABUFs in LOW_POWER_ENTRY_WITH_WAKEUP resume
On LOW_POWER entry, DMABUFs have a move(revoke=true), but the
runtime resume path didn't un-revoke. This adds a corresponding
move(revoke=false), which will later turn into
vfio_pci_unrevoke_bars().
NOTE: the unrevoke is reordered _before_ the eventfd_signal() in
vfio_pci_core_runtime_resume(). This removes a window in which a
waking waiter could have observed the DMABUF state as still revoked
(or pm_runtime_engaged = true). (The UAPI docs state the event
means the resume's complete, so observing otherwise seemed
unintended.)
Also, when DMABUFs are later mmap()ed, a waiter waking and taking a
fault could have seen the unrevoked state and SIGBUS just before
taking the memory_lock. (If the handler gets to acquire the lock,
though, the resume sequence is complete, and the handler observes
pm_runtime_engaged = false.)
This fix is in this series because the issue will impact CPU access
to the VMA as well (once they use DMABUFs), and so it's a strict
dependency of later commits.
dma-buf: Export dma_buf_set_name()
Makes dma_buf_set_name() available for use (by helper patch),
taking a kernel-allocated string. The pre-existing local helper
becomes a wrapper copying a __user string for the set-name ioctl.
vfio/pci: Add a helper to look up PFNs for DMABUFs
Adds a DMABUF VMA fault handler helper to determine arbitrary-sized
PFNs from ranges in DMABUF.
vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA
Refactors DMABUF export for use by the existing export feature, and
adds a helper that creates a DMABUF corresponding to a VFIO BAR
mmap() request.
There was a request for decent debug naming in /proc/<pid>/maps
etc. comparable to the existing VFIO names: since the VMAs are
DMABUFs, they have a "dmabuf:" prefix and can't be 100% identical
to before. This is a user-visible change, but this patch at least
now gives us extra info on the BDF & BAR being mmap()ed. The name
is installed with the new dma_buf_set_name() above. An alternative
would be to add a creation-time name field to struct
dma_buf_export_info, but a setter function is more useful: the name
can be changed at will after init.
vfio/pci: Convert BAR mmap() to use a DMABUF
The vfio-pci core mmap() creates a DMABUF with the helper above,
and the vm_ops fault handler uses the other helper to resolve the
fault. Because this depends on DMABUF structs/code,
CONFIG_VFIO_PCI_CORE needs to depend on CONFIG_DMA_SHARED_BUFFER.
The CONFIG_VFIO_PCI_DMABUF still conditionally enables the export
support code.
NOTE: The user mmap()s a device fd, but the resulting VMA's vm_file
becomes that of the DMABUF. The DMABUF takes ownership of the
device file and put()s it on release, which maintains the existing
behaviour of a VMA keeping the VFIO device open.
BAR zapping then happens via the existing vfio_pci_dma_buf_move()
path, which now needs to unmap PTEs in the DMABUF's address_space.
NOTE: As local LLM reviews did, Sashiko might falsely worry about
the DMABUF fd being obtained through /proc/pid/map_files and
remapped, but the file is an anon inode and doesn't support the
open op (-ENXIO) so AFAICT this is currently impossible.
NOTE: A side-effect of this is that is_mergeable_vma() will be
false between adjacent mappings of VFIO BARs; merging would require
rebuilding a larger DMABUF representing the union, not just
plugging the VMAs together.
The DMABUF backing the BAR is an implicit/internal export, even
when the CONFIG_VFIO_PCI_DMABUF feature is not included because
CONFIG_PCI_P2PDMA is not available. Without P2PDMA, it's
acceptable for the DMABUF to not have a P2PDMA provider. In this
configuration, the VFIO DMABUF .attach prohibits any import, which
avoids getting as far as a WARN in dma_buf_map_attachment(), which
would fail without P2PDMA anyway.
vfio/pci: Clean up BAR zap and revocation
In general (see NOTE!) the vfio_pci_zap_bars() is now obsolete,
since it unmaps PTEs in the VFIO device address_space which is now
unused. This consolidates all calls (e.g. around reset) with the
neighbouring vfio_pci_dma_buf_move()s into new functions, to
revoke/unrevoke (making the steps clearer).
NOTE: Because drivers can use their own vm_ops and override .mmap,
the core must conservatively assume an overridden .mmap might still
add PTEs to the VFIO device address_space and therefore still does
the zap. A new flag, zap_bars_on_revoke, enables the zap when
.mmap is overridden. A driver that does not need the zap can clear
this to opt-out, e.g. if the driver calls down to the common mmap
(and so uses DMABUFs). hisi-acc-vfio-pci does just this, and thus
sets the opt-out flag.
vfio/pci: Support mmap() of a VFIO DMABUF
Adds mmap() for a DMABUF fd exported from vfio-pci.
It was a goal to keep the VFIO device fd lifetime behaviour
unchanged with respect to the DMABUFs. An application can close
all device fds, and this will revoke/clean up all DMABUFs; then, no
mappings or other access can be performed. When enabling mmap() of
the DMABUFs, this means access through the VMA is also revoked.
This complicates the fault handler because whilst the DMABUF
exists, it has no guarantee that the corresponding VFIO device is
still alive. Adds synchronisation ensuring the vdev is available
before the locks in vdev are touched; this holds the device
registration so that even if the buffer has been cleaned up, vdev
hasn't been freed and so the locks can be safely taken.
vfio/pci: Permanently revoke a DMABUF on request
This is mostly a rename of `revoked` to an enum, `status`, and
adding a third state for a buffer: usable, revoked temporary,
revoked permanent. A new VFIO feature is added,
VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE, which takes a DMABUF (exported
from the same device) and permanently revokes it. Thus a userspace
driver can guarantee any downstream consumers of a shared fd are
prevented from accessing a BAR range, and that range can be reused.
NOTE: This might block userspace, waiting on importers to detach.
The code doing revocation in vfio_pci_dma_buf_move() is moved, to a
common function for use by ..._move() and this new feature.
Testing
=======
(The [RFC ONLY] userspace test program, which drives a QEMU
bochs-display function, can be found in the GitHub branch below. It
at least illustrates how the export, map, revoke, and close semantics
interoperate. WIP on a follow-up with a proper vfio-selftests style
test based on this -- this won't be part of this series.)
This code has been tested in mapping DMABUFs of single/multiple ranges
from multiple BARs, aliasing mmap()s, aliasing ranges across DMABUFs,
vm_pgoff > 0, revocation, shutdown/cleanup scenarios, and hugepage
mappings. No regressions observed on the VFIO selftests, or on our
internal vfio-pci applications. VFIO on i386 has been build-tested.
Thanks to Alex Mastro for building a (WIP) testcase for the
mmap_lock->memory_lock issue (which is now OK...).
Dear Reviewers,
===============
There was a lot of finessing v5->v6, all patches had
fixes/changes/cleanups/rewordings made, and so I have dropped previous
R-B tags.
Along the way several related issues came up that warrant more
eyes, and I'd be grateful for your input:
1. The mmap fault handler takes a bunch of locks non-interruptibly,
and potentially depends on a lot of DMABUF-related activities
completing. I had a go at converting them to
interruptible/killable forms, but pulling on the thread revealed
there seems to be a wider issue if move/revoke doesn't complete in
a timely fashion (due to buggy importers). Where I got to was that
just updating the fault handler won't fix the user experience of an
unkillable task, and move()/revocation will need thought too. I
don't intend to fix this here but wanted to start discussion so we
can address it in a follow up. There's now a dependency between
mmap_lock in the fault handler and the DMABUF resv (which might
take a while to resolve).
2. vfio_basic_config_write() has an error path if
vfio_default_config_write() fails that releases memory_lock but
doesn't un-revoke BARs in the case of PCI_COMMAND.MSE being
cleared. When can the write fail, in practice, perhaps surprise
removal?
The effect on this series would be: a write of MSE=0 revokes BARs,
vconfig[PCI_COMMAND]'s MSE becomes 0, but if the physical write
fails then the physical MSE remains 1 and BAR VMAs stay revoked.
This seemed a mess; fixing isn't as simple as un-revoking on the
error path since vfio_default_config_write() has already trampled
vconfig so that'd need unwinding. It felt like a catastrophic
scenario where BARs staying revoked isn't a bad outcome, but want
to hear your experience of the likelihood of this issue.
3. The exchange of a VFIO fd mmap()'s vma->vm_file with an implicitly
created DMABUF's file has implications on LSM. For example, an
mmap will be checked against the policy for a VFIO fd, but a
subsequent mprotect() relates to the DMABUF file's policy (which is
anon/unique to the mapping). This is pretty confusing.
4. If VFIO fd is opened O_RDONLY, it currently can't be mmap()ed
(because PROT_WRITE is rejected in do_mmap(), and PROT_READ alone
drops the VM_SHARED so VFIO's mmap rejects it). But it seems we
can export a DMABUF from it, and then pass the resulting fd around
for P2P writes.
I don't know if this is intentional, relied on, or a known
limitation so haven't included a change here, but:
a) We could reject export unless the device fd's f_mode has
O_RDWR, to reflect the abilities of P2P
b) Or, instead of just failing if !O_RDWR, we limit the
get_dma_buf.open_flags to the VFIO device fd's f_mode, such as:
VFIO device fd perms: Export flags: Result:
O_RDWR O_RDWR, O_RDONLY OK
O_RDONLY O_RDONLY OK
O_RDONLY O_RDWR -EPERM
O_WRONLY * -EPERM
* O_WRONLY -EPERM
(Skipping WRONLY because a PROT_WRITE-only mmap() won't work,
though it probably should be included for P2P.)
If we can do at least (a) that seems good, because the knock-on
effect in this series is that we can export a DMABUF RW from an
O_RDONLY device fd and then succeed to mmap() the DMABUF with RW.
(That said, even if one has an O_RDONLY device fd, the device state
can still be changed/reset. But it feels cleaner to least prevent
export for a O_RDONLY device fd, and match the device fd mmap()
behaviour.)
Maybe (b) is step 2: Allowing finer-grained RD/WR could be useful
if there's a future goal to tie DMABUF permissions to, say, iommufd
IOMMU_READ/IOMMU_WRITE permissions. I don't htink this is in place
today, apologies if I'm missed something.
END
===
This is based on v7.2.
These commits are on GitHub for easier browsing, along with
"[RFC ONLY] selftests: vfio: Add standalone vfio_dmabuf_mmap_test":
https://github.com/metamev/linux/compare/v7.2...dev/mev/vfio-dmabuf-mmap-v6
Thanks for reading [this astonishingly-long cover letter],
Matt
[1] https://lore.kernel.org/kvm/9a615f22-c0d4-46ae-9654-db11e94e5fec@ozlabs.org/
================================================================================
Changelog:
v6:
- Dropped patches:
PCI/P2PDMA: Split pool-related cleanup out of pci_p2pdma_release()
PCI/P2PDMA: Add CONFIG_PCI_P2PDMA_CORE
- New patch, "vfio/pci: Remove DMABUF export dependency on vdev->memory_lock":
memory_lock was overloaded to protect the vdev->dmabufs list, the
entries' revoked status, etc. Towards the goal of not needing
memory_lock in export (so not creating a mmap_lock -> memory_lock
dependency due to exporting from mmap()), add a new
vdev->dmabuf_lock which protects the list and (writing) the
revocation status. (See details in cover letter.)
- New patch, "vfio/pci: Un-revoke DMABUFs in LOW_POWER_ENTRY_WITH_WAKEUP resume":
Bugfix, as the resume path didn't un-revoke. Previously affected
"only" P2P, but since DMABUFs-everywhere it would affect BAR mmap()s
too.
- New patch, "dma-buf: Export dma_buf_set_name()":
Change in dma-buf.c's dma_buf_set_name() to take a kernel-allocated
string. The original IOCTL-related path wraps this to strdup the
user string first.
- "vfio/pci: Add a helper to look up PFNs for DMABUFs":
Commit message reworded explaining vma_pgoff_adjust, return -ERANGE
for (non-transient) pagesize failure instead of -EAGAIN. Add
dmabuf_lock annotation and assert that the DMABUF isn't revoked
(else new -ENODEV error), using a new vdev parameter (caller is
expected to safely extract it from priv).
Significant bugfix in the case of a VMA offset exceeding the 1TB
stride between VFIO_PCI_OFFSET_MASK-sized regions, which would have
silently wrapped due to being masked to 1TB. This is fixed by
including the VFIO region index in the high bits of vma_pgoff_adjust
for the traditional mmap() path, so the subtraction from
vma->vm_pgoff in ...find_pfn() removes the region index without
masking. For the DMABUF mmap() path, vma_pgoff_adjust = 0 and
(thanks to no masking) large offsets can be used.
- "vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA":
Now permits a NULL p2pdma provider to be installed in the
vfio_pci_dma_buf. The regular BAR mmap() path doesn't need a
provider for CPU access; the DMABUF backing the VMA exists primarily
for the CPU, and removing this check decouples vfio-pci from
depending on CONFIG_PCI_P2PDMA (_without_ going through the previous
hoops of splitting up P2PDMA into a _CORE subset to provide the
provider).
Only DMABUF import/map needs a provider, which is available when
P2PDMA is available; the previous series' first two patches (to
always provide a provider) are now unnecessary and have been
dropped.
Squashed "vfio/pci: Provide a user-facing name for BAR mappings"
into this patch, because it's now just very small due to the new
dma_buf_set_name(). Comment that, because of careful construction
of the name length, the setting of the name can't fail; but in case
assumptions change, a check/kfree prevents it leaking. (A new error
path doing full unwind of the DMABUF creation is overkill.)
- "vfio/pci: Convert BAR mmap() to use a DMABUF":
Instead of the previous dropped P2PDMA split commits, allows
pcim_p2pdma_provider() to return NULL if !CONFIG_VFIO_PCI_DMABUF
(meaning !CONFIG_PCI_P2PDMA). The provider isn't used unless the
DMABUF is imported; prevent import to make this fail early (rather
than at map), by creating an always-fail vfio_pci_dma_buf_attach()
stub. (Somewhat belt and braces, but makes clear you need P2PDMA to
import a BAR VMA's DMABUF!)
- "vfio/pci: Clean up BAR zap and revocation":
The zap_bars_on_revoke flag is moved back to the bitfield, and is
set from vfio_pci_core_init_dev() (instead of registration time).
Reworked the hisi driver change to update this earlier.
- "vfio/pci: Support mmap() of a VFIO DMABUF":
Previous versions introduced a build issue (which was fixed by the
next patch) due to a READ_ONCE of the priv->revoked bitfield member;
removed. It is added by the next patch in the series
(s/revoked/status/).
Added an explicit rejection for mmap of a DMABUF with
vma_pgoff_adjust > 0. Such buffers can only be created implicitly
by the VFIO BAR mmap path, and an fd can't currently be recreated
from a VMA (e.g. fished out of /proc/pid/map_files) because the anon
inode ops don't support open. But check added if a future mechanism
arises and for clarity. Rejecting an mmap of a DMABUF fd having
vma_pgoff_adjust > 0 means vfio_pci_dma_buf_find_pfn() doesn't need
to cope with the case of vma->vm_pgoff < priv->vma_pgoff_adjust and
underflow.
- "vfio/pci: Permanently revoke a DMABUF on request":
Clarified docs for returned error values; re-added
READ_ONCE(priv->status) for unlocked reads in both
vfio_pci_dma_buf_attach() and vfio_pci_dma_buf_mmap().
v5: https://lore.kernel.org/all/20260715174737.15287-1-matt@ozlabs.org/
v4: https://lore.kernel.org/all/20260701171245.90111-1-matt@ozlabs.org/
v3: https://lore.kernel.org/all/20260610154327.37758-1-matt@ozlabs.org/
v2: https://lore.kernel.org/all/20260527102319.100128-1-mattev@meta.com/
v1: https://lore.kernel.org/kvm/20260416131815.2729131-1-mattev@meta.com/
RFCv2: https://lore.kernel.org/kvm/20260312184613.3710705-1-mattev@meta.com/
RFCv1: https://lore.kernel.org/all/20260226202211.929005-1-mattev@meta.com/
Tech topic: https://lore.kernel.org/linux-iommu/20250918214425.2677057-1-amastro@fb.com/
Matt Evans (9):
vfio/pci: Remove DMABUF export dependency on vdev->memory_lock
vfio/pci: Un-revoke DMABUFs in LOW_POWER_ENTRY_WITH_WAKEUP resume
dma-buf: Export dma_buf_set_name()
vfio/pci: Add a helper to look up PFNs for DMABUFs
vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA
vfio/pci: Convert BAR mmap() to use a DMABUF
vfio/pci: Clean up BAR zap and revocation
vfio/pci: Support mmap() of a VFIO DMABUF
vfio/pci: Permanently revoke a DMABUF on request
drivers/dma-buf/dma-buf.c | 58 +-
drivers/vfio/pci/Kconfig | 4 +-
drivers/vfio/pci/Makefile | 3 +-
.../vfio/pci/hisilicon/hisi_acc_vfio_pci.c | 14 +-
drivers/vfio/pci/vfio_pci_config.c | 30 +-
drivers/vfio/pci/vfio_pci_core.c | 222 +++++--
drivers/vfio/pci/vfio_pci_dmabuf.c | 607 +++++++++++++++---
drivers/vfio/pci/vfio_pci_priv.h | 54 +-
include/linux/dma-buf.h | 2 +
include/linux/vfio_pci_core.h | 3 +
include/uapi/linux/vfio.h | 24 +
11 files changed, 851 insertions(+), 170 deletions(-)
--
2.50.1 (Apple Git-155)
PCI P2PDMA applies Request and Completion Redirect throughout both paths.
This misclassifies asymmetric and nested switches, and reports one answer
for every kind of TLP.
Three ACS controls act on TLP attributes the client chooses rather than
on the topology: Translation Blocking and Direct Translated P2P act on
a Request's Address Type, and Completion Redirect skips Completions carrying
Relaxed Ordering.
Evaluate each direction at the path divergence, decide every class from the
one walk, expose the provider to dma-buf importers, and let mlx5 ask rather
than assume.
Signed-off-by: Leon Romanovsky <leonro(a)nvidia.com>
---
Changes in v5:
- Rebase on the posted fixes.
- Dropped tags from changed patches.
- Remove Egress Control Vector interpretation and coverage.
- Keep enabled Egress Control conservative as a Request redirect.
- Use pci_dbg()/dev_dbg() for diagnostics and drop the "debug" prefix.
- Removed code comments from "Document the pdev->p2pdma lifetime and RCU
rules" patch and reduced description to actual lifetime explanation.
- Added note that Linux assumes that TLPs are in strict-ordering and
untranslated.
- Added code to calculate p2p paths per-TLP type.
- Converted mlx5 to use that new proposed API.
- Link to https://patch.msgid.link/20260821-fix-p2p-acs-v4-0-v4-0-94426b96de73@nvidia…
Changes in v4:
- Reject ACS Violations and unreadable routing state instead of treating
them as host-bridge redirects
- Added Tested-by tags from Tushar Dave
- Added support to asymmetric ACS routing
- Limited redirect checks to the two ports at the path divergence
- Added standalone ACS routing diagnostics for hardware retesting
- Dropped " PCI: Account for Direct Translated P2P in ACS isolation checks" patch
- Link to v3: https://patch.msgid.link/20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com
Changes in v3:
- Fixed pci_p2pdma_add_resource() error unwinding
- Made pdev->p2pdma teardown wait unconditionally for RCU readers
- Restricted pci_p2pmem_find_many() to pool-backed providers
- Documented the pdev->p2pdma lifetime and RCU rules
- Fixed calc_map_type_and_dist() handling of the verbose argument
- Required the ACS port and target to share a bus before indexing the
Egress Control Vector
- Gave pci_acs_enabled() and pci_acs_path_enabled() a scope, so the ACS
Direct Translated P2P rule no longer stops pci_enable_pasid() from
enabling PASID
- Dropped "Report ACS ports when the paths share no upstream bridge":
the mapping type cannot change without a shared upstream bridge, so
the pci=disable_acs_redir= hint was not actionable there and the ACS
walk only cost config space reads
- Folded the Request Redirect rule into pci_acs_rr_ineffective(), so
pci_acs_flags_enabled() and the Intel SPT PCH quirk share one copy
- Renamed pci_acs_egress_ctrl_set() to pci_acs_egress_ctrl_is_set(), it
reads the bit rather than setting it
- Reworded the blocked-path warning: ACS may also leave the direct route
indeterminate rather than blocked
- Added KUnit coverage for the shared-bus guard, a device with no ACS
capability and an unreadable ACS Control register
- Added the missing Fixes: tags, a second one on the
pci_p2pdma_add_resource() unwinding fix (the dangling devres action
dates to f58ef9d1d135) and one on the Egress Control isolation change
- Link to v2: https://patch.msgid.link/20260806-fix-p2p-acs-v2-0-0cec14812965@nvidia.com
Changes in v2:
- Added Logan's ROB tags
- Added commas in Documentation patch
- Link to v1: https://patch.msgid.link/20260802-fix-p2p-acs-v1-0-a7c5eb64fff6@nvidia.com
---
Leon Romanovsky (18):
PCI/P2PDMA: Document pdev->p2pdma lifetime rules
PCI/P2PDMA: Document the TLP attribute assumptions
PCI/P2PDMA: Derive routing from directional ACS controls
PCI: Reject unreadable ACS controls in isolation checks
PCI/P2PDMA: Evaluate ACS controls at the path divergence
PCI/P2PDMA: Document directional ACS routing
PCI/P2PDMA: Collect the path's ACS controls before deciding
PCI/P2PDMA: Answer routing per TLP class
PCI/P2PDMA: Route Relaxed Ordering Completions directly
PCI/P2PDMA: Reject Translated Requests blocked by Translation Blocking
PCI/P2PDMA: Route Translated Requests under Direct Translated P2P
PCI/P2PDMA: Log detailed ACS routing diagnostics
PCI/P2PDMA: Add KUnit tests for the ACS routing decisions
PCI/P2PDMA: Test the ACS P2P routing walk
PCI: Add KUnit coverage for ACS isolation checks
PCI/P2PDMA: Document TLP-class routing
dma-buf: Let importers ask how peer-to-peer traffic is routed
RDMA/mlx5: Ask P2PDMA whether ATS takes a direct peer-to-peer route
Documentation/admin-guide/kernel-parameters.txt | 9 +-
Documentation/driver-api/pci/p2pdma.rst | 69 +++
drivers/dma-buf/dma-buf-mapping.c | 41 +-
drivers/dma-buf/dma-buf.c | 1 +
drivers/infiniband/core/uverbs.h | 1 -
drivers/infiniband/core/uverbs_std_types_dmabuf.c | 7 +-
drivers/infiniband/hw/mlx5/mlx5_ib.h | 36 +-
drivers/infiniband/hw/mlx5/mr.c | 40 ++
drivers/pci/Kconfig | 15 +
drivers/pci/Makefile | 1 +
drivers/pci/p2pdma.c | 637 +++++++++++++++++++---
drivers/pci/pci.c | 7 +-
drivers/pci/pci.h | 26 +
drivers/pci/pci_acs_test.c | 609 +++++++++++++++++++++
drivers/pci/quirks.c | 6 +-
drivers/vfio/pci/vfio_pci_dmabuf.c | 8 +-
include/linux/dma-buf-mapping.h | 4 +-
include/linux/dma-buf.h | 5 +
include/linux/pci-p2pdma.h | 57 +-
19 files changed, 1440 insertions(+), 139 deletions(-)
---
base-commit: 08dbfad3f5040f5bdb6c529da20d6d4e81fefd72
change-id: 20260821-fix-p2p-acs-v4-0-e72455e3a261
prerequisite-message-id: <20260830-batch-p2p-fixes-v1-0-5044e8dfbe2e(a)nvidia.com>
prerequisite-patch-id: 6b25c7fcf164cdfc14e9fac5b908d97fcf6509d7
prerequisite-patch-id: 0d083c281001365aae4b35544cf28891a6ab9a96
prerequisite-patch-id: bfd9dabf271f3cc9a3a61f46387d20c20311363d
prerequisite-patch-id: fad0275efc722830fc591509506c0a5e4f581073
prerequisite-patch-id: 0c83bee688fec1f6d1564654df7c630fa6a4a978
Best regards,
--
Leon Romanovsky <leonro(a)nvidia.com>
On 9/7/26 04:00, Karl Mehltretter wrote:
> Commit 646013f513f3 ("dma-buf: enable DMABUF_DEBUG by default on DEBUG
> kernels") changed the default of DMABUF_DEBUG to "y if DEBUG", but no
> Kconfig symbol DEBUG exists, so the option has had no default since.
>
> Use DEBUG_KERNEL, the Kconfig symbol for a debug kernel.
>
> Fixes: 646013f513f3 ("dma-buf: enable DMABUF_DEBUG by default on DEBUG kernels")
> Assisted-by: LLM
> Signed-off-by: Karl Mehltretter <kmehltretter(a)gmail.com>
Reviewed and pushed to drm-misc-fixes.
Thanks,
Christian.
> ---
> This enables the dma-buf debug checks on every CONFIG_DEBUG_KERNEL
> configuration, which is what 646013f513f3 set out to do. If that is too
> broad today, the alternative is to restore the previous
> "default y if DMA_API_DEBUG".
>
> drivers/dma-buf/Kconfig | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/dma-buf/Kconfig b/drivers/dma-buf/Kconfig
> index 7efc0f0d0712..e4f078a326a4 100644
> --- a/drivers/dma-buf/Kconfig
> +++ b/drivers/dma-buf/Kconfig
> @@ -43,7 +43,7 @@ config UDMABUF
> config DMABUF_DEBUG
> bool "DMA-BUF debug checks"
> depends on DMA_SHARED_BUFFER
> - default y if DEBUG
> + default y if DEBUG_KERNEL
> help
> This option enables additional checks for DMA-BUF importers and
> exporters. Specifically it validates that importers do not peek at the
>
> base-commit: 986c24e0fe44f844b44d365b71ce831947f50298
> --
> 2.53.0
>
On Wed, Sep 09, 2026 at 02:48:17PM +0300, Dmitry Baryshkov wrote:
> On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote:
> > On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
> > > On 20-08-2026 20:17, Rob Clark wrote:
> > > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk(a)kernel.org> wrote:
> > > >>
> > > >> On 19/08/2026 17:48, Rob Clark wrote:
> > > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk(a)kernel.org> wrote:
> > > >>>> The rule of usptream development is that we do not accept duplicated
> > > >>>> code, just because a vendor wants to write something new. This is
> > > >>>> basically the concept applied all over the drivers tree, where we pushed
> > > >>>> back against all sorts of duplications all over the vendors.
> > > >>>>
> > > >>>> What I miss in this thread is why would there be any exception here. We
> > > >>>> do not grant exceptions from standard practices on "I want" reasons.
> > > >>>
> > > >>> I agree that we should not have duplicated drivers just for vendor
> > > >>> lolz. But when it comes to adopting common frameworks and integrating
> > > >>> better into the ecosystem, this doesn't seem like something we should
> > > >>> actively discourage. I don't think this is a case of vendor lolz, but
> > > >>
> > > >> No one discourages it. Following standard Linux kernel practices and
> > > >> requirements is not discouraging, do not twist the narrative here.
> > > >> Again, it is standard upstream review telling that we do not duplicate
> > > >> drivers. Ever, unless there is serious exception needed.
> > > >
> > > > I wasn't trying to twist the narrative, just trying to come up with a
> > > > path forward that isn't "no" or "improve existing driver", since
> > > > neither of those gets us towards a future using common frameworks.
> > > >
> > > >> I asked why there should be an exception granted? Is the reason for
> > > >> exception following:
> > > >> "We want to adopt common framework"
> > > >> ?
> > > >
> > > > Possibly? But I don't think we want two drivers to be any sort of
> > > > long term solution. (Ie. as long as venus/iris have co-exist.)
> > > >
> > > >>
> > > >>> rather reacting to drm/accel emerging as the standard framework for
> > > >>> this sort of driver.
> > > >>>
> > > >>> So how do we get from here to there?
> > > >>
> > > >> What is wrong with my proposal?
> > > >
> > > > Maybe I missed something, my understanding was your proposal was
> > > > "Grow/replace/improve existing driver instead of coming with a
> > > > duplicate".. grow or improve doesn't move us toward common
> > > > frameworks. Maybe "replace" is a valid option. If there is something
> > > > I missed, then I apologize.
> > > >
> > > > Options I can think of are:
> > > >
> > > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
> > > > driver
> > > > 2. Backwards compat chardev registered by new driver, providing existing
> > > > UABI. I'm not 100% sure about the feasibility/drawbacks of this..
> > > > AFAIU the fastrpc folks where planning a backwards compat layer in
> > > > userspace, so maybe it is possible.
> > > > 3. exception?
> > > >
> > > > I'd like to know what the feasibility of #2 is, since at a high level
> > > > that sounds like the best option. Possibly limit exposure of legacy
> > > > UABI to existing hw so we don't get into a place of needing to extend
> > > > the legacy UABI for new hw?
> > > >
> > > > But #1 sounds like a non-controversial place to start regardless.
> > > > Possibly with #2 coming as followup and necessary step before eventual
> > > > migration to new driver for existing hw?
> > > >
> > > > Even if we start with #2, how do we handle first-merge-window
> > > > bugs/regressions without reverting addition of new driver and removal
> > > > of old? It seems like we'd need a window of a couple release cycles
> > > > where both drivers exist?
> > > >
> > > > Maybe others have other/better options in mind?
> > > To all, I'm seeking on the approach I should follow to go ahead here. I
> > > can work on implementing #1(as per Rob's list) with hw specific
> > > compatible for v4 if it's acceptable.
> > >
> >
> > I don't see any reason for you to define a "hw specific compatible",
> > because as you have shown in this series (and as Rob point out), there's
> > no difference in the "hardware".
> >
> > The only reason for your "hw specific compatible" is to make a software
> > selection in Linux - and that's not what DeviceTree is for.
>
> That's not exactly true. There are protocol differences. For example,
> polling mode is supported only since a certain timeline in the history.
> Likewise other interface features are not supported on all the FastRPC
> devices. For the polling mode support we were already beaten by the lack
> of SoC-specific compats.
>
I can see the benefit of capturing some of the generational features in
a compatible, like the changes related to address width. But for pure
software features that doesn't have an actual bearing in the hardware,
I'd prefer if we relied on dynamic discovery.
But none of that applies to the question of "can I use compatible to
select if we should use the new or old Linux driver".
> > As such, I don't see that you have a DeviceTree problem at all, because
> > this is a Linux-internal problem.
> >
> > > #2(compat driver) is something that we are still exploring as we
> > > couldn't find any standard way to achieve it. We might start a separate
> > > discussion for that once we have few possible designs with us.
> > >
> >
> > This is the actual problem!
> >
> > We have existing user space that depends on the ioctl interface exposed
> > by the current misc driver. You must not break these.
>
> This is clear.
>
> > Hardware cutoff is not a viable solution, because that's just a
> > declaration that we'll let the old platforms rotten - or alternatively
> > you commit to maintain two drivers to the very same feature and quality
> > level.
> >
> > So the only reasonable solution is #2; from there it's a valid question
> > if you reach that point my stepwise migrating the current misc driver
> > that solution, or if you present a new driver with the fully backwards
> > compatible interface, alongside the new ABI.
>
> I think the general direction was #3 (or #2.1): implement a shim layer
> on top of the QDA driver as a separate module. Put all the historical
> over-complicated solutions into that shim module and let it die at some
> point. Current fastrpc driver lets userspace specify buffers in several
> different ways, forcing the kernel driver to perform a lot of work
> with buffer addresses. I don't think that this legacy code should be a
> part of the QDA. It can go to qda-fastrpc-shim, keeping it out of the
> kernel memory if it's not necessary.
>
> > But this does bring to a question which the cover letter should explain
> > - but doesn't: what problem does this patch series actually solve?
>
> I agree that it should be a part of the cover letter.
>
> As a person who triggered this work, I can propose my reasons:
>
> - Current driver has over-complicated memory manager (both on the
> userspace and on the kernel side).
The userspace library is spaghetti, but when I wrote my own it turned
out quite succinct. The ioctl structs certainly could have been cleaner,
and documented, but it seems to me that a fair amount of the complexity
comes from the different use cases - such as SMMU vs XPU, secure and
unsecure buffers, remnants of now unsupported options.
I'm presuming that QDA will need to adopt most of these, and that QDA
support will be bolted into the spaghetti library.
> Correspondng uAPI is not really
> suitable for virtualization. Using GEM simplifies both the kernel
> driver and uAPI. Also using handle-offset-length to specify the
> buffers makes it easy to support virtual QDA devices.
I'm looking forward to learn more about this!
>
> - Current driver predates the accel subsystem. Using common subsystem
> simplifies reviews. The QDA driver has gotten several comments about
> the usage of the DMA-BUFs. It'not unlikely that the same issues are
> present in the current FastRPC driver, just being unnoticed.
Yeah, this is unfortunate. It would certainly be nice to have a
documented and clean ABI.
>
> - The ideas present in the current FastRPC driver also predate the
> current design practices. The uAPI was created in the ad-hoc way, just
> following the momentary needs. Driver code also shows the result of
> that, having enough of the spaghetti code.
>
Yeah, again, this isn't desirable.
> Given all of that, yes, it is possible to provide an evolution of the
> FastRPC driver into the accel+shim, improve the code quality meanwhile
> and end up with the good enough split. However I think that the path
> taken would be longer and the net result might be worse.
>
> With all of that in mind, my suggestion is to continue working on the
> QDA driver, get the core of it to integrate nicely with the accel and
> DMA-BUF usage requirements and implement the (minimal?) fastrpc shim on
> top of it.
>
If you believe this is the way to reach the proper design, then I won't
object. My requirement is that my userspace continues to work when my
distro suddenly switches FASTRPC=n/QDA=m.
I have no problems with dropping the fastrpc driver once the QDA is
drop-in-compatible. I'm also open to marking the compat layer deprecated
and eventually drop it once we're certain that users have moved to a
userspace that used the accel interface.
But we can't merge the QDA driver as long as we believe that
implementing a compat layer will be hard/impossible.
Regards,
Bjorn
>
> --
> With best wishes
> Dmitry
On 9/10/26 05:57, Li Wang wrote:
...
>> AMD came up with something similar, but all those approaches are so fundamentally broken that we didn't even considered upstreaming it.
>>> And you are using this as a "bypass" for the normal accel subsystem,
>>> shouldn't this be part of that subsystem instead of a custom user/kernel
>>> api like you are creating here?
>>
>> As far as I know there is a patch set under review and even already partially merged which enables exactly that functionality as general feature for DMA-buf which is vendor independent and should at least in theory work with all drivers.
>>
>> I'm really surprised that somebody is still working on the vendor specific stuff.
> As you pointed out, every vendor has been inventing their own way and interfaces to support GDS,
> introducing custom kernel modules and proprietary UAPI interfaces, with varying performance that
> leaves developers heavily frustrated. Apologies for not making this clear enough in our commit
> messages, which understandably caused some confusion. We merely borrowed the name "GDS" to describe
> the functional purpose of fgds.
>
> In fact, we believe fgds offers four key advantages:
> (1) GPU platform independence;
> (2) POSIX/io_uring interface compatibility;
> (3) Higher performance than GDS;
> (4) Minimal kernel footprint and UAPI footprint
>
> Regarding (1), (2), and (3), please allow me to briefly explain the design mechanism of fgds:
> fgds turns a GPU memory buffer into a POSIX/io_uring-compatible user-space virtual address via
> three main steps:
>
> Step 1: Utilizing ZONE_DEVICE support, we remap the GPU memory exposed via PCIe BAR into struct pages
> using devm_memremap_pages();
>
> Step 2: Utilizing dma-buf support, the GPU memory buffer is exported as a dma-buf file descriptor (fd).
> Using this fd as a bridge, we look up the corresponding DMA addresses for the GPU memory buffer inside
> the kernel;
>
> Step 3: Through mmap, we insert the struct pages corresponding to the GPU memory buffer into the userspace
> VMA, mapping their physical/DMA addresses directly. The virtual address returned by mmap can then be directly
> passed into standard POSIX or io_uring interfaces.
Well long story short what you do here is completely broken.
Approaches like those have been suggested before and we added both documentation as well as code to prevent such hacks from working.
Please see Pavel Begunkov patch set on the LKML which adds DMA-buf support to io_uring for how to do it correctly. Just google for "Add dmabuf read/write via io_uring".
Regards,
Christian.
On Fri, Sep 04, 2026 at 12:44:55PM +0200, Thierry Reding wrote:
> From: Thierry Reding <treding(a)nvidia.com>
>
> Drivers that use this may want to be built as a module, so export them.
That's one of the worst commit log ever. No, we don't just export
core symbols dealing with the kernel direct map because
"Drivers that use this may want to be built as a module".
For one exporting this at all needs a very good justification and
not just hand waiving. But more importantly if we can't avoid
exporting it, it needs to be exported at the tightest sensible
scope. E.g. for a given module if it is so special, or a namespace
if it's not that special. But in doubt we should have a proper
core abstraction instead of opening up direct map manipulation to
random modules.
On Wed, Sep 09, 2026 at 06:42:03PM +0800, Li Wang wrote:
> Hi Greg,
> Thanks for the review!
>
> >
> > That's not really needed in a changelog text, it could be in the 0/X
> > patch :)
> >
> Sorry for the clutter. I will move most of them into the 0/X patch in v2.
>
> > Anyway, you didn't cc: the io_uring list, why?
> >
> `scripts/get_maintainer.pl` didn't output the io_uring mailing list, likely
> because this patch doesn't directly touch the io_uring codebase itself. It only
> enables remapping GPU memory buffers to CPU virtual addresses, which can then be
> consumed via standard io_uring APIs.
>
> I've added io-uring(a)vger.kernel.org to CC for this reply and will keep it in v2.
Great, as you are using that as the api, there might be some parts that
will need to be reviewed by them.
> > Nor why "fgds" is the name, that's going to be hard to remember, does it
> > stand for something?
> >
> "FGDS" stands for Fast GPUDirect Storage. GPUDirect Storage (GDS) is NVIDIA's
> technology enabling direct I/O between GPU memory and files on NVMe,
> widely used in LLM workloads to bypass CPU overhead.
That's nvidia's specific solution, but this works on other devices,
right? Or just for that one platform?
And you are using this as a "bypass" for the normal accel subsystem,
shouldn't this be part of that subsystem instead of a custom user/kernel
api like you are creating here?
thanks,
greg k-h
On Tue, Sep 08, 2026 at 09:15:45PM +0800, Li Wang wrote:
> From: Mengmeng Zhao <zhaomengmeng(a)kylinos.cn>
>
> Inspired by the paper published in SC'25 [1], we implemented a character
> device named fgds that provides two ioctl interfaces:
> `REG_BUFFER/UNREG_BUFFER`. It enables applications to perform direct I/O
> between GPU memory and NVMe via POSIX and io_uring APIs. This is
> particularly useful for LLM workloads, such as model loading, KV cache
> offloading, and checkpointing. The fgds device corresponds one-to-one with
> the PCIe GPU on the machine. The usage is straightforward: an application
> simply opens the corresponding fgds device, calls ioctl on the returned fd
> with REG_BUFFER, taking the target GPU memory buffer address (represented
> as a dma-buf fd), and the buffer length as inputs, and then invokes mmap on
> the fgds device fd, using the return value of ioctl as the input. The mmap
> call returns a CPU virtual address (call it cpu_vaddr). Afterward,
> cpu_vaddr can be passed directly to pread/pwrite, or
> io_uring_prep_read/io_uring_prep_write to perform direct I/O between files
> on NVMe and GPU memory. A minimal working example can be found in [2].
> The underlying mechanism is that, with the support of fgds device,
> cpu_vaddr is made to point directly to the GPU memory buffer corresponding
> to the dma-buf fd. This solution is loosely coupled with the GPU vendor's
> driver; the GPU vendor only needs to support exporting the allocated GPU
> memory buffer through the standard Linux kernel dma-buf framework, which
> the vast majority of mainstream GPUs already support. This allows both
> applications and the fgds device to work seamlessly with GPUs from
> different vendors without any modifications. Furthermore, applications no
> longer need to call vendor-specific proprietary APIs (such as NVIDIA's
> cuFile API) or install vendor-specific kernel modules (such as NVIDIA's
> nvidia-fs.ko) for different GPU vendors. We have tested fgds on GPU cards
> from NVIDIA, AMD, and several other vendors, and it works well.
>
> Besides the benefits in ease of use and compatibility, another key
> advantage of this solution is higher performance. [2] presents the
> performance comparison results between fgds and NVIDIA GDS. Because fgds
> eliminates the overhead of phony buffers incurred by NVIDIA GDS, it
> achieves significantly higher performance. For example, for reads, fgds
> outperforms GDS by 11% to 109%; for writes, fgds outperforms GDS by 10%
> to 71%.
>
> To further accelerate the read and write operations of large files or
> massive data volumes—which are very common in LLM scenarios—we have
> implemented library functions `fgds_read` and `fgds_write`. Under the hood,
> these interfaces split large data into chunks and submit them
> asynchronously and in parallel via io_uring, thereby further boosting I/O
> performance, with read performance improved by up to 115% and write
> performance by up to 40%. In addition, we also provide the `fgds_register`
> library interface to encapsulate the `open`, `ioctl' and `mmap` operations.
> Readers who are interested can refer to [2].
>
> In addition, we have added the LMCache backend, enabling vLLM to offload KV
> cache via LMCache using fgds, which accelerates inference performance. We
> also added PyTorch APIs, compatible with the PyTorch GDS API, to improve
> the performance of reading and writing checkpoints during LLM training.
>
> We look forward to community feedback and are fully committed to iterating
> on this series to work towards upstreaming.
That's not really needed in a changelog text, it could be in the 0/X
patch :)
Anyway, you didn't cc: the io_uring list, why?
Also, as a first cut, please see the sashiko comments on this patch:
https://sashiko.dev/#/patchset/20260908131545.105987-1-liwang@kylinos.cn
>
> [1] https://dl.acm.org/doi/10.1145/3712285.3759862
> [2] https://github.com/Storage-and-OS-for-AI/fgds
>
> Signed-off-by: Mengmeng Zhao <zhaomengmeng(a)kylinos.cn>
> Signed-off-by: Li Wang <liwang(a)kylinos.cn>
> ---
> drivers/misc/Kconfig | 9 +
> drivers/misc/Makefile | 1 +
> drivers/misc/fgds.c | 989 ++++++++++++++++++++++++++++++++++++++
> include/uapi/linux/fgds.h | 54 +++
> 4 files changed, 1053 insertions(+)
> create mode 100644 drivers/misc/fgds.c
> create mode 100644 include/uapi/linux/fgds.h
>
> diff --git a/drivers/misc/Kconfig b/drivers/misc/Kconfig
> index 7364931dad3a..2f3a5a5fd0bf 100644
> --- a/drivers/misc/Kconfig
> +++ b/drivers/misc/Kconfig
> @@ -568,6 +568,15 @@ config MCHP_LAN966X_PCI
> - lan966x-miim (MDIO_MSCC_MIIM)
> - lan966x-switch (LAN966X_SWITCH)
>
> +config FGDS
> + tristate "GPU-NVMe direct I/O control driver"
> + depends on PCI && DMA_SHARED_BUFFER && ZONE_DEVICE
> + help
> + Say Y here if you want to support GPU-NVME direct I/O
> + via POSIX/io_uring interfaces.
> +
> + If unsure, say N.
Module name is not listed here.
Nor why "fgds" is the name, that's going to be hard to remember, does it
stand for something?
> +
> source "drivers/misc/c2port/Kconfig"
> source "drivers/misc/eeprom/Kconfig"
> source "drivers/misc/cb710/Kconfig"
> diff --git a/drivers/misc/Makefile b/drivers/misc/Makefile
> index e8d8d5d88c0d..04985abe1678 100644
> --- a/drivers/misc/Makefile
> +++ b/drivers/misc/Makefile
> @@ -71,3 +71,4 @@ obj-y += keba/
> obj-y += amd-sbi/
> obj-$(CONFIG_MISC_RP1) += rp1/
> obj-$(CONFIG_INTEL_SSEI) += issei/
> +obj-$(CONFIG_FGDS) += fgds.o
> diff --git a/drivers/misc/fgds.c b/drivers/misc/fgds.c
> new file mode 100644
> index 000000000000..3aa4945f701b
> --- /dev/null
> +++ b/drivers/misc/fgds.c
> @@ -0,0 +1,989 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * Fast GPU Direct Storage via dma-buf.
> + *
> + * Copyright (C) 2026 KylinSoft. Co., Ltd. All rights reserved.
> + *
> + * Maps GPU memory into user space to enable direct NVME-to-GPU DMA
> + * pread/pwrite syscalls. BAR pages are remapped into ZONE_DEVICE via
> + * devm_memremap_pages() and populated using dma-buf backing pages.
> + */
> +#define pr_fmt(fmt) "fgds: " fmt
You are a driver, always use dev_*() print functions, not pr_()
functions, as you will loose the device information. For example:
> +/*
> + * BAR-based mapping requires device physical addresses. When using
> + * IOMMU, DMA addresses are IOVAs, which cannot be mapped directly.
> + */
> +static int fgds_check_gpu_iommu(struct pci_dev *pdev)
> +{
> + struct iommu_domain *domain;
> +
> + domain = iommu_get_domain_for_dev(&pdev->dev);
> + if (domain && domain->type != IOMMU_DOMAIN_IDENTITY) {
> + pr_warn("%s: reject attaching a translating IOMMU domain (requires iommu=pt or off\n",
> + dev_name(&pdev->dev));
Should be dev_warn(), right?
But what can userspace do with that warning, did something just break?
> + pr_info("loaded successfully: %u GPU(s) active\n", fgds_dev_count);
When drivers work, they are quiet, please remove this, and the other
pr_info() lines, as they seem to be left over from your debugging.
thanks,
greg k-h
On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
> On 20-08-2026 20:17, Rob Clark wrote:
> > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk(a)kernel.org> wrote:
> >>
> >> On 19/08/2026 17:48, Rob Clark wrote:
> >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk(a)kernel.org> wrote:
> >>>> The rule of usptream development is that we do not accept duplicated
> >>>> code, just because a vendor wants to write something new. This is
> >>>> basically the concept applied all over the drivers tree, where we pushed
> >>>> back against all sorts of duplications all over the vendors.
> >>>>
> >>>> What I miss in this thread is why would there be any exception here. We
> >>>> do not grant exceptions from standard practices on "I want" reasons.
> >>>
> >>> I agree that we should not have duplicated drivers just for vendor
> >>> lolz. But when it comes to adopting common frameworks and integrating
> >>> better into the ecosystem, this doesn't seem like something we should
> >>> actively discourage. I don't think this is a case of vendor lolz, but
> >>
> >> No one discourages it. Following standard Linux kernel practices and
> >> requirements is not discouraging, do not twist the narrative here.
> >> Again, it is standard upstream review telling that we do not duplicate
> >> drivers. Ever, unless there is serious exception needed.
> >
> > I wasn't trying to twist the narrative, just trying to come up with a
> > path forward that isn't "no" or "improve existing driver", since
> > neither of those gets us towards a future using common frameworks.
> >
> >> I asked why there should be an exception granted? Is the reason for
> >> exception following:
> >> "We want to adopt common framework"
> >> ?
> >
> > Possibly? But I don't think we want two drivers to be any sort of
> > long term solution. (Ie. as long as venus/iris have co-exist.)
> >
> >>
> >>> rather reacting to drm/accel emerging as the standard framework for
> >>> this sort of driver.
> >>>
> >>> So how do we get from here to there?
> >>
> >> What is wrong with my proposal?
> >
> > Maybe I missed something, my understanding was your proposal was
> > "Grow/replace/improve existing driver instead of coming with a
> > duplicate".. grow or improve doesn't move us toward common
> > frameworks. Maybe "replace" is a valid option. If there is something
> > I missed, then I apologize.
> >
> > Options I can think of are:
> >
> > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
> > driver
> > 2. Backwards compat chardev registered by new driver, providing existing
> > UABI. I'm not 100% sure about the feasibility/drawbacks of this..
> > AFAIU the fastrpc folks where planning a backwards compat layer in
> > userspace, so maybe it is possible.
> > 3. exception?
> >
> > I'd like to know what the feasibility of #2 is, since at a high level
> > that sounds like the best option. Possibly limit exposure of legacy
> > UABI to existing hw so we don't get into a place of needing to extend
> > the legacy UABI for new hw?
> >
> > But #1 sounds like a non-controversial place to start regardless.
> > Possibly with #2 coming as followup and necessary step before eventual
> > migration to new driver for existing hw?
> >
> > Even if we start with #2, how do we handle first-merge-window
> > bugs/regressions without reverting addition of new driver and removal
> > of old? It seems like we'd need a window of a couple release cycles
> > where both drivers exist?
> >
> > Maybe others have other/better options in mind?
> To all, I'm seeking on the approach I should follow to go ahead here. I
> can work on implementing #1(as per Rob's list) with hw specific
> compatible for v4 if it's acceptable.
>
I don't see any reason for you to define a "hw specific compatible",
because as you have shown in this series (and as Rob point out), there's
no difference in the "hardware".
The only reason for your "hw specific compatible" is to make a software
selection in Linux - and that's not what DeviceTree is for.
As such, I don't see that you have a DeviceTree problem at all, because
this is a Linux-internal problem.
> #2(compat driver) is something that we are still exploring as we
> couldn't find any standard way to achieve it. We might start a separate
> discussion for that once we have few possible designs with us.
>
This is the actual problem!
We have existing user space that depends on the ioctl interface exposed
by the current misc driver. You must not break these.
Hardware cutoff is not a viable solution, because that's just a
declaration that we'll let the old platforms rotten - or alternatively
you commit to maintain two drivers to the very same feature and quality
level.
So the only reasonable solution is #2; from there it's a valid question
if you reach that point my stepwise migrating the current misc driver
that solution, or if you present a new driver with the fully backwards
compatible interface, alongside the new ABI.
But this does bring to a question which the cover letter should explain
- but doesn't: what problem does this patch series actually solve?
Regards,
Bjorn
> Happy to take any other suggestion also.>
> > BR,
> > -R
>
>