On Wed, Sep 30, 2026 at 01:48:49PM +0100, Pavel Begunkov wrote:
On 9/30/26 11:08, Pavel Begunkov wrote:
On 9/30/26 04:00, Matthew Brost wrote:
On Mon, Sep 21, 2026 at 02:38:45PM +0100, Pavel Begunkov wrote:
...>> +struct dma_buf_io_map *dma_buf_io_create_map(struct dma_buf_io_ctx *ctx)
+{ + struct dma_buf *dmabuf = ctx->dmabuf; + struct dma_buf_io_map *map; + long ret;
+ guard(mutex)(&ctx->map_create_mutex);
+ scoped_guard(mutex, &ctx->map_mutex) { + if (ctx->maps_killed) + return ERR_PTR(-ENOENT); + /* recheck under the lock in case it has already been re-created */ + map = __dma_buf_io_get_map(ctx); + if (map) + return map; + }
+ dma_buf_io_wait_active_maps(ctx);
+ ret = dma_resv_lock_interruptible(dmabuf->resv, NULL); + if (ret) + return ERR_PTR(ret);
+ ret = dma_resv_wait_timeout(dmabuf->resv, DMA_RESV_USAGE_KERNEL, + true, MAX_SCHEDULE_TIMEOUT); + if (ret <= 0) { + if (!ret) + ret = -EAGAIN; + dma_resv_unlock(dmabuf->resv); + return ERR_PTR(ret); + }
+ map = ctx->dev_ops->map(ctx);
I'm playing around this code now.
I think you need the dma_resv_wait_timeout after the 'map'?
If a device doesn't support p2p ->map() will typically trigger an async migrate to system memory and data will be moving but the map is valid - Xe 100% does this, I checked AMDGPU and fairly confident it has the same async behavior.
There is a wait right before because I read somewhere in dma-buf comments that I need to do that, sounds a bit odd if I need to wait on fences before and after. I can add it, just curious how come that other dma_buf_map_attachment() callers don't need to do that. Or maybe they wait somewhere else?
In GPU drivers, when we receive a foreign object and map it in Xe or AMDGPU, this logic sits deep in the stack. However, the top-level call is typically ttm_bo_validate(), which in both drivers eventually resolves to the TTM ->move() vfunc. That path calls dma_buf_map_attachment(), which in turn invokes the dma-buf's ->map() callback.
That ->map() callback can trigger a asynchronous move on a different device, where the top-level call is again ttm_bo_validate(). We then end up back in the TTM ->move() vfunc, where kernel fences are installed. I realize that's a lot of layers, but the call chain can end up looking like this.
Now we have a mapping and an object with kernel fences attached. Before the object can be used by either the exec IOCTL (batch buffer submission), the VM bind IOCTL (mapping into the GPU address space), or a CPU page fault (not relevant for dma-bufs since they cannot be CPU-mapped, but a useful example of a normal BO move), we wait on those kernel fences. In the case of the exec IOCTL or VM bind IOCTL, the kernel fences are added as dependencies to a drm_sched job, delaying its execution until the fences signal. That is where the wait occurs. In the case of a CPU page fault, the kernel waits directly on the fences before installing the CPU page mappings.
So the TL;DR is if you want to immediately hand back a valid mapping, the kernel fences must be waited on before returning after ->map() call.
Matt
I can't find it, so maybe it was the comment below and I mixed sth up back then. I'll move it after ->map().
- Note that for non-dynamic exporters the driver must guarantee that
- that the memory is available for use and cleared of any old data by
- the time this function returns. Drivers which pipeline their buffer
- moves internally must wait for all moves and clears to complete.
- Dynamic exporters do not need to follow this rule: For non-dynamic
- importers the buffer is already pinned through @pin, which has the
- same requirements. Dynamic importers otoh are required to obey the
- dma_resv fences.
-- Pavel Begunkov
linaro-mm-sig@lists.linaro.org