February 2024 - Linux-stable-mirror

[tip: irq/urgent] PCI/MSI: Prevent MSI hardware interrupt number truncation

by tip-bot2 for Vidya Sagar

The following commit has been merged into the irq/urgent branch of tip: Commit-ID: db744ddd59be798c2627efbfc71f707f5a935a40 Gitweb: https://git.kernel.org/tip/db744ddd59be798c2627efbfc71f707f5a935a40 Author: Vidya Sagar <vidyas(a)nvidia.com> AuthorDate: Mon, 15 Jan 2024 19:26:49 +05:30 Committer: Thomas Gleixner <tglx(a)linutronix.de> CommitterDate: Mon, 19 Feb 2024 16:11:01 +01:00 PCI/MSI: Prevent MSI hardware interrupt number truncation While calculating the hardware interrupt number for a MSI interrupt, the higher bits (i.e. from bit-5 onwards a.k.a domain_nr >= 32) of the PCI domain number gets truncated because of the shifted value casting to return type of pci_domain_nr() which is 'int'. This for example is resulting in same hardware interrupt number for devices 0019:00:00.0 and 0039:00:00.0. To address this cast the PCI domain number to 'irq_hw_number_t' before left shifting it to calculate the hardware interrupt number. Please note that this fixes the issue only on 64-bit systems and doesn't change the behavior for 32-bit systems i.e. the 32-bit systems continue to have the issue. Since the issue surfaces only if there are too many PCIe controllers in the system which usually is the case in modern server systems and they don't tend to run 32-bit kernels. Fixes: 3878eaefb89a ("PCI/MSI: Enhance core to support hierarchy irqdomain") Signed-off-by: Vidya Sagar <vidyas(a)nvidia.com> Signed-off-by: Thomas Gleixner <tglx(a)linutronix.de> Tested-by: Shanker Donthineni <sdonthineni(a)nvidia.com> Cc: stable(a)vger.kernel.org Link: https://lore.kernel.org/r/20240115135649.708536-1-vidyas@nvidia.com --- drivers/pci/msi/irqdomain.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/pci/msi/irqdomain.c b/drivers/pci/msi/irqdomain.c index c8be056..cfd84a8 100644 --- a/drivers/pci/msi/irqdomain.c +++ b/drivers/pci/msi/irqdomain.c @@ -61,7 +61,7 @@ static irq_hw_number_t pci_msi_domain_calc_hwirq(struct msi_desc *desc) return (irq_hw_number_t)desc->msi_index | pci_dev_id(dev) << 11 | - (pci_domain_nr(dev->bus) & 0xFFFFFFFF) << 27; + ((irq_hw_number_t)(pci_domain_nr(dev->bus) & 0xFFFFFFFF)) << 27; } static void pci_msi_domain_set_desc(msi_alloc_info_t *arg,

1 year, 9 months

1
0
0 0

[PATCH] iio: imu: inv_mpu6050: fix frequency setting when chip is off

by inv.git-commit＠tdk.com

From: Jean-Baptiste Maneyrol <jean-baptiste.maneyrol(a)tdk.com> Track correctly FIFO state and apply ODR change before starting the chip. Without the fix, you cannot change ODR more than 1 time when data buffering is off. Fixes: 111e1abd0045 ("iio: imu: inv_mpu6050: use the common inv_sensors timestamp module") Signed-off-by: Jean-Baptiste Maneyrol <jean-baptiste.maneyrol(a)tdk.com> --- drivers/iio/imu/inv_mpu6050/inv_mpu_trigger.c | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/drivers/iio/imu/inv_mpu6050/inv_mpu_trigger.c b/drivers/iio/imu/inv_mpu6050/inv_mpu_trigger.c index 676704f9151f..e6e6e94452a3 100644 --- a/drivers/iio/imu/inv_mpu6050/inv_mpu_trigger.c +++ b/drivers/iio/imu/inv_mpu6050/inv_mpu_trigger.c @@ -111,6 +111,7 @@ int inv_mpu6050_prepare_fifo(struct inv_mpu6050_state *st, bool enable) if (enable) { /* reset timestamping */ inv_sensors_timestamp_reset(&st->timestamp); + inv_sensors_timestamp_apply_odr(&st->timestamp, 0, 0, 0); /* reset FIFO */ d = st->chip_config.user_ctrl | INV_MPU6050_BIT_FIFO_RST; ret = regmap_write(st->map, st->reg->user_ctrl, d); @@ -184,6 +185,10 @@ static int inv_mpu6050_set_enable(struct iio_dev *indio_dev, bool enable) if (result) goto error_power_off; } else { + st->chip_config.gyro_fifo_enable = 0; + st->chip_config.accl_fifo_enable = 0; + st->chip_config.temp_fifo_enable = 0; + st->chip_config.magn_fifo_enable = 0; result = inv_mpu6050_prepare_fifo(st, false); if (result) goto error_power_off; -- 2.34.1

1 year, 9 months

2
1
0 0

[PATCH] iio: imu: inv_mpu6050: fix FIFO parsing when empty

by inv.git-commit＠tdk.com

From: Jean-Baptiste Maneyrol <jean-baptiste.maneyrol(a)tdk.com> Now that we are reading the full FIFO in the interrupt handler, it is possible to have an emply FIFO since we are still receiving 1 interrupt per data. Handle correctly this case instead of having an error causing a reset of the FIFO. Fixes: 0829edc43e0a ("iio: imu: inv_mpu6050: read the full fifo when processing data") Signed-off-by: Jean-Baptiste Maneyrol <jean-baptiste.maneyrol(a)tdk.com> --- drivers/iio/imu/inv_mpu6050/inv_mpu_ring.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/drivers/iio/imu/inv_mpu6050/inv_mpu_ring.c b/drivers/iio/imu/inv_mpu6050/inv_mpu_ring.c index 66d4ba088e70..d4f9b5d8d28d 100644 --- a/drivers/iio/imu/inv_mpu6050/inv_mpu_ring.c +++ b/drivers/iio/imu/inv_mpu6050/inv_mpu_ring.c @@ -109,6 +109,8 @@ irqreturn_t inv_mpu6050_read_fifo(int irq, void *p) /* compute and process only all complete datum */ nb = fifo_count / bytes_per_datum; fifo_count = nb * bytes_per_datum; + if (nb == 0) + goto end_session; /* Each FIFO data contains all sensors, so same number for FIFO and sensor data */ fifo_period = NSEC_PER_SEC / INV_MPU6050_DIVIDER_TO_FIFO_RATE(st->chip_config.divider); inv_sensors_timestamp_interrupt(&st->timestamp, fifo_period, nb, nb, pf->timestamp); -- 2.34.1

1 year, 9 months

1
0
0 0

[PATCH v2 1/3] drm/buddy: fix range bias

by Matthew Auld

There is a corner case here where start/end is after/before the block range we are currently checking. If so we need to be sure that splitting the block will eventually give use the block size we need. To do that we should adjust the block range to account for the start/end, and only continue with the split if the size/alignment will fit the requested size. Not doing so can result in leaving split blocks unmerged when it eventually fails. Fixes: afea229fe102 ("drm: improve drm_buddy_alloc function") Signed-off-by: Matthew Auld <matthew.auld(a)intel.com> Cc: Arunpravin Paneer Selvam <Arunpravin.PaneerSelvam(a)amd.com> Cc: Christian König <christian.koenig(a)amd.com> Cc: <stable(a)vger.kernel.org> # v5.18+ Reviewed-by: Arunpravin Paneer Selvam <Arunpravin.PaneerSelvam(a)amd.com> --- drivers/gpu/drm/drm_buddy.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/drivers/gpu/drm/drm_buddy.c b/drivers/gpu/drm/drm_buddy.c index c4222b886db7..f3a6ac908f81 100644 --- a/drivers/gpu/drm/drm_buddy.c +++ b/drivers/gpu/drm/drm_buddy.c @@ -332,6 +332,7 @@ alloc_range_bias(struct drm_buddy *mm, u64 start, u64 end, unsigned int order) { + u64 req_size = mm->chunk_size << order; struct drm_buddy_block *block; struct drm_buddy_block *buddy; LIST_HEAD(dfs); @@ -367,6 +368,15 @@ alloc_range_bias(struct drm_buddy *mm, if (drm_buddy_block_is_allocated(block)) continue; + if (block_start < start || block_end > end) { + u64 adjusted_start = max(block_start, start); + u64 adjusted_end = min(block_end, end); + + if (round_down(adjusted_end + 1, req_size) <= + round_up(adjusted_start, req_size)) + continue; + } + if (contains(start, end, block_start, block_end) && order == drm_buddy_block_order(block)) { /* -- 2.43.0

1 year, 9 months

1
0
0 0

[PATCH v3] PCI: Increase maximum PCIe physical function number to 7 for non-ARI devices

by Bean Huo

From: Bean Huo <beanhuo(a)micron.com> As per PCIe r6.2, sec 6.13 titled "Alternative Routing-ID Interpretation (ARI)", up to 8 [fn # 0..7] Physical Functions(PFs) are allowed in an non-ARI capable device. Previously, our implementation erroneously limited the maximum number of PFs to 7 for endpoints without ARI support. This patch corrects the maximum PF count to adhere to the PCIe specification by allowing up to 8 PFs on non-ARI capable devices. This change ensures better compliance with the standard and improves compatibility with devices relying on this specification. Fixes: c3df83e01a96 ("PCI: Clean up pci_scan_slot()") Cc: stable(a)vger.kernel.org Signed-off-by: Bean Huo <beanhuo(a)micron.com> --- Changelog: v2--v3: 1. Update commit messag v1--v2: 1. Add Fixes tag 2. Modify commit message --- drivers/pci/probe.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c index ed6b7f48736a..8c3d0f63bc13 100644 --- a/drivers/pci/probe.c +++ b/drivers/pci/probe.c @@ -2630,7 +2630,8 @@ static int next_fn(struct pci_bus *bus, struct pci_dev *dev, int fn) if (pci_ari_enabled(bus)) return next_ari_fn(bus, dev, fn); - if (fn >= 7) + /* If EP does not support ARI, the maximum number of functions should be 7 */ + if (fn > 7) return -ENODEV; /* only multifunction devices may have more functions */ if (dev && !dev->multifunction) -- 2.34.1

1 year, 9 months

2
2
0 0

[PATCH][5.10, 5.15, 6.1][0/1] hrtimer: Ignore slack time for RT tasks

by Felix Moessbauer

This suggests a fix from 6.3 for stable that fixes a nasty bug in the timing behavior of periodic RT tasks w.r.t timerslack_ns. While the documentation clearly states that the slack time is ignored for RT tasks, this is not the case for the hrtimer code. This patch fixes the issue and applies to all stable kernels. Best regards, Felix Moessbauer Siemens AG Davidlohr Bueso (1): hrtimer: Ignore slack time for RT tasks in schedule_hrtimeout_range() kernel/time/hrtimer.c | 14 +++++++++++--- 1 file changed, 11 insertions(+), 3 deletions(-) -- 2.39.2

1 year, 9 months

2
2
0 0

Re: [REGRESSION] Acp5x probing regression introduced between kernel 6.7.2 -> 6.7.4

by Takashi Iwai

On Mon, 12 Feb 2024 10:13:00 +0100, Takashi Iwai wrote: > > On Sun, 11 Feb 2024 18:19:25 +0100, > Linux regression tracking (Thorsten Leemhuis) wrote: > > > > [CCing a few people] > > > > On 11.02.24 15:34, Ted Chang wrote: > > > > > > I noticed 6.7.4 has introduced a regression for the steam deck. The LCD > > > steam deck can no longer probe the acp5x audio chipset anymore. This > > > regression does not affect the 6.8.x series. I did not test kernel > > > 6.7.3 because Opensuse tumbleweed skipped the update on my machine. > > > > Thx for your report. FWIW, problems like this can be caused by all > > sorts of changes, but obviously those in the area of audio support > > are most likely to cause this. There are just a few in the > > v6.7.2..v6.7.4 range[1]. Among them a commit that is related to > > acp5x, that's why I CCed its author as well (Venkata Prasad Potturu). > > > > Maybe one of the new recipients will have an idea. If not, you most > > likely will have to bisect this and check if mainline is affected > > as well.[2] > > > > Ciao, Thorsten > > > > [1] > > $ git log --oneline v6.7.2..v6.7.4 sound/ > > f3570675bf09af ASoC: codecs: wsa883x: fix PA volume control > > 2f8e9b77ca2fea ASoC: codecs: lpass-wsa-macro: fix compander volume hack > > 5b465d6384e4eb ASoC: codecs: wcd938x: fix headphones volume controls > > 1673211a38012e ASoC: qcom: sc8280xp: limit speaker volumes > > 242b5bffa23a9c ASoC: codecs: rtq9128: Fix TDM enable and DAI format control flow > > 2c272ff9859601 ASoC: codecs: rtq9128: Fix PM_RUNTIME usage > > 4a28302b2c681e ALSA: hda/conexant: Fix headset auto detect fail in cx8070 and SN6140 > > e37a96941fdd53 ALSA: hda: intel-dspcfg: add filters for ARL-S and ARL > > ffa3eea886c6fe ALSA: hda: Intel: add HDA_ARL PCI ID support > > 4b6986b170f2f2 ASoC: amd: Add new dmi entries for acp5x platform > > This one is the only relevant change, I suppose. > The machine matches with 'Valve Jupiter'. > > Interestingly, the system seems working with 6.8-rc3, so some piece > might be missing. Or simply reverting this patch should fix. In the bugzilla entry, the reporter confirmed that the revert of the commit 4b6986b170f2f2 fixed the problem. #regzbot introduced: 4b6986b170f2f2 Takashi > > > Takashi > > > e38ad4ace20b4d ALSA: hda: Refer to correct stream index at loops > > a434c75e0671f9 soundwire: fix initializing sysfs for same devices on different buses > > > > [2] I'm working on a guide that describes what's needed: > > https://www.leemhuis.info/files/misc/How%20to%20bisect%20a%20Linux%20kernel… > > > > > Steps to reproduce the problem > > > 1. Obtain a steam deck > > > 2. Install kernel 6.7.4 > > > 3. Boot the device and you will see dummy output in gnome shell > > > > > > Observed kernel logs. > > > > > > [ 8.755614] cs35l41 spi-VLV1776:00: supply VA not found, using dummy regulator > > > [ 8.760506] cs35l41 spi-VLV1776:00: supply VP not found, using dummy regulator > > > [ 8.777148] cs35l41 spi-VLV1776:00: Cirrus Logic CS35L41 (35a40), Revision: B2 > > > [ 8.777471] cs35l41 spi-VLV1776:01: supply VA not found, using dummy regulator > > > [ 8.777532] cs35l41 spi-VLV1776:01: supply VP not found, using dummy regulator > > > [ 8.777709] cs35l41 spi-VLV1776:01: Reset line busy, assuming shared reset > > > [ 8.788465] cs35l41 spi-VLV1776:01: Cirrus Logic CS35L41 (35a40), Revision: B2 > > > [ 8.877280] snd_hda_intel 0000:04:00.1: enabling device (0000 -> 0002) > > > [ 8.877595] snd_hda_intel 0000:04:00.1: Handle vga_switcheroo audio client > > > [ 8.889913] snd_acp_pci 0000:04:00.5: enabling device (0000 -> 0002) > > > [ 8.890063] snd_acp_pci 0000:04:00.5: Unsupported device revision:0x50 > > > [ 8.890129] snd_acp_pci: probe of 0000:04:00.5 failed with error -22 > > > [ 8.906136] snd_hda_intel 0000:04:00.1: bound 0000:04:00.0 (ops amdgpu_dm_audio_component_bind_ops [amdgpu] > > > > > > > > > No kernel module in use shown. > > > > > > 04:00.5 Multimedia controller [0480]: Advanced Micro Devices, Inc. [AMD] > > > ACP/ACP3X/ACP6x Audio Coprocessor [1022:15e2] (rev 50) > > > Subsystem: Valve Software Device [1e44:1776] > > > Flags: fast devsel, IRQ 70, IOMMU group 4 > > > Memory at 80380000 (32-bit, non-prefetchable) [size=256K] > > > Capabilities: <access denied> > > > Kernel modules: snd_pci_acp3x, snd_rn_pci_acp3x, snd_pci_acp5x, > > > snd_pci_acp6x, snd_acp_pci, snd_rpl_pci_acp6x, snd_pci_ps, > > > snd_sof_amd_renoir, snd_sof_amd_rembrandt, snd_sof_amd_vangogh, > > > snd_sof_amd_acp63 > > > > > > > > > Information for package kernel-default: > > > --------------------------------------- > > > Repository : openSUSE-Tumbleweed-Oss > > > Name : kernel-default > > > Version : 6.7.4-1.1 > > > Arch : x86_64 > > > Vendor : openSUSE > > > Installed Size : 240.3 MiB > > > Installed : Yes > > > Status : up-to-date > > > Source package : kernel-default-6.7.4-1.1.nosrc > > > Upstream URL : https://www.kernel.org/ <https://www.kernel.org/> > > > Summary : The Standard Kernel > > > Description : > > > The standard kernel for both uniprocessor and multiprocessor systems. > > > > > > > > > Source Timestamp: 2024-02-06 05:32:37 +0000 > > > GIT Revision: 01735a3e65287585dd830a6a3d33d909a4f9ae7f > > > GIT Branch: stable > > > > > > Handle 0x0000, DMI type 0, 26 bytes > > > BIOS Information > > > Vendor: Valve > > > Version: F7A0120 > > > Release Date: 12/01/2023 > > > Address: 0xE0000 > > > Runtime Size: 128 kB > > > BIOS Revision: 1.20 > > > Firmware Revision: 1.16 > > > > > > #regzbot introduced: v6.7.2..v6.7.4 > > >

1 year, 9 months

3
4
0 0

[PATCH v2] PCI: Increase maximum PCIe physical function number to 7 for non-ARI devices

by Bean Huo

From: Bean Huo <beanhuo(a)micron.com> The PCIe specification allows up to 8 Physical Functions (PFs) per endpoint when ARI (Alternative Routing-ID Interpretation) is not supported. Previously, our implementation erroneously limited the maximum number of PFs to 7 for endpoints without ARI support. This patch corrects the maximum PF count to adhere to the PCIe specification by allowing up to 8 PFs on non-ARI endpoints. This change ensures better compliance with the standard and improves compatibility with devices relying on this specification. The necessity for this adjustment was verified by a thorough review of the "Alternative Routing-ID Interpretation (ARI)" section in the PCIe 3.0 Spec, which first introduced ARI. Fixes: c3df83e01a96 ("PCI: Clean up pci_scan_slot()") Cc: stable(a)vger.kernel.org Signed-off-by: Bean Huo <beanhuo(a)micron.com> --- Changelog: v1--v2: 1. Add Fixes tag 2. Modify commit message --- drivers/pci/probe.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c index ed6b7f48736a..8c3d0f63bc13 100644 --- a/drivers/pci/probe.c +++ b/drivers/pci/probe.c @@ -2630,7 +2630,8 @@ static int next_fn(struct pci_bus *bus, struct pci_dev *dev, int fn) if (pci_ari_enabled(bus)) return next_ari_fn(bus, dev, fn); - if (fn >= 7) + /* If EP does not support ARI, the maximum number of functions should be 7 */ + if (fn > 7) return -ENODEV; /* only multifunction devices may have more functions */ if (dev && !dev->multifunction) -- 2.34.1

1 year, 9 months

3
3
0 0

Re: [PATCH v2 1/2] dt-bindings: phy: mediatek: tphy: add a property for force-mode switch

by Macpaul Lin

On 12/22/23 01:15, Vinod Koul wrote: > > > External email : Please do not click links or open attachments until you > have verified the sender or the content. > > On Mon, 11 Dec 2023 10:56:23 +0800, Chunfeng Yun wrote: >> Due to some old SoCs with shared t-phy between usb3 and pcie only support >> force-mode switch, and shared and non-shared t-phy may exist at the same >> time on a SoC, can't use compatible to distinguish between shared and >> non-shared t-phy, add a property to supported it. >> Currently, only support switch from default pcie mode to usb3 mode. >> But now prefer to use "mediatek,syscon-type" on new SoC as far as possible. >> >> [...] > > Applied, thanks! > > [1/2] dt-bindings: phy: mediatek: tphy: add a property for force-mode switch > commit: cc230a4cd8e91f64c90b5494dfd76848197418ed > [2/2] phy: mediatek: tphy: add support force phy mode switch > commit: 9b27303003f5af0d378f29ccccea57c7d65cc642 > > Best regards, > -- > ~Vinod > > Is it possible to cherry-pick these 2 patches to stable branches? These 2 patches help fix USB port 1 (xhci1) for board mt8395-genio-1200-evb. The following branch has been tested. - linux-6.7.y (6.7.5): apply test, build pass, function tested OK (with corresponded dtb change). - linux-6.6.y (6.6.17): apply test, build pass. - linux-6.1.y (6.1.78): apply test, build pass. Thanks. Macpaul Lin

1 year, 9 months

1
0
0 0

[PATCH v3] mm/swap: fix race when skipping swapcache

by Kairui Song

From: Kairui Song <kasong(a)tencent.com> When skipping swapcache for SWP_SYNCHRONOUS_IO, if two or more threads swapin the same entry at the same time, they get different pages (A, B). Before one thread (T0) finishes the swapin and installs page (A) to the PTE, another thread (T1) could finish swapin of page (B), swap_free the entry, then swap out the possibly modified page reusing the same entry. It breaks the pte_same check in (T0) because PTE value is unchanged, causing ABA problem. Thread (T0) will install a stalled page (A) into the PTE and cause data corruption. One possible callstack is like this: CPU0 CPU1 ---- ---- do_swap_page() do_swap_page() with same entry <direct swapin path> <direct swapin path> <alloc page A> <alloc page B> swap_read_folio() <- read to page A swap_read_folio() <- read to page B <slow on later locks or interrupt> <finished swapin first> ... set_pte_at() swap_free() <- entry is free <write to page B, now page A stalled> <swap out page B to same swap entry> pte_same() <- Check pass, PTE seems unchanged, but page A is stalled! swap_free() <- page B content lost! set_pte_at() <- staled page A installed! And besides, for ZRAM, swap_free() allows the swap device to discard the entry content, so even if page (B) is not modified, if swap_read_folio() on CPU0 happens later than swap_free() on CPU1, it may also cause data loss. To fix this, reuse swapcache_prepare which will pin the swap entry using the cache flag, and allow only one thread to pin it. Release the pin after PT unlocked. Racers will simply wait since it's a rare and very short event. A schedule() call is added to avoid wasting too much CPU or adding too much noise to perf statistics Other methods like increasing the swap count don't seem to be a good idea after some tests, that will cause racers to fall back to use the swap cache again. Parallel swapin using different methods leads to a much more complex scenario. Reproducer: This race issue can be triggered easily using a well constructed reproducer and patched brd (with a delay in read path) [1]: With latest 6.8 mainline, race caused data loss can be observed easily: $ gcc -g -lpthread test-thread-swap-race.c && ./a.out Polulating 32MB of memory region... Keep swapping out... Starting round 0... Spawning 65536 workers... 32746 workers spawned, wait for done... Round 0: Error on 0x5aa00, expected 32746, got 32743, 3 data loss! Round 0: Error on 0x395200, expected 32746, got 32743, 3 data loss! Round 0: Error on 0x3fd000, expected 32746, got 32737, 9 data loss! Round 0 Failed, 15 data loss! This reproducer spawns multiple threads sharing the same memory region using a small swap device. Every two threads updates mapped pages one by one in opposite direction trying to create a race, with one dedicated thread keep swapping out the data out using madvise. The reproducer created a reproduce rate of about once every 5 minutes, so the race should be totally possible in production. After this patch, I ran the reproducer for over a few hundred rounds and no data loss observed. Performance overhead is minimal, microbenchmark swapin 10G from 32G zram: Before: 10934698 us After: 11157121 us Non-direct: 13155355 us (Dropping SWP_SYNCHRONOUS_IO flag) Fixes: 0bcac06f27d7 ("mm, swap: skip swapcache for swapin of synchronous device") Link: https://github.com/ryncsn/emm-test-project/tree/master/swap-stress-race [1] Reported-by: "Huang, Ying" <ying.huang(a)intel.com> Closes: https://lore.kernel.org/lkml/87bk92gqpx.fsf_-_@yhuang6-desk2.ccr.corp.intel… Signed-off-by: Kairui Song <kasong(a)tencent.com> Cc: stable(a)vger.kernel.org --- Update from V2: - Add a schedule() if raced to prevent repeated page faults wasting CPU and add noise to perf statistics. - Use a bool to state the special case instead of reusing existing variables fixing error handling [Minchan Kim]. V2: https://lore.kernel.org/all/20240206182559.32264-1-ryncsn@gmail.com/ Update from V1: - Add some words on ZRAM case, it will discard swap content on swap_free so the race window is a bit different but cure is the same. [Barry Song] - Update comments make it cleaner [Huang, Ying] - Add a function place holder to fix CONFIG_SWAP=n built [SeongJae Park] - Update the commit message and summary, refer to SWP_SYNCHRONOUS_IO instead of "direct swapin path" [Yu Zhao] - Update commit message. - Collect Review and Acks. V1: https://lore.kernel.org/all/20240205110959.4021-1-ryncsn@gmail.com/ include/linux/swap.h | 5 +++++ mm/memory.c | 20 ++++++++++++++++++++ mm/swap.h | 5 +++++ mm/swapfile.c | 13 +++++++++++++ 4 files changed, 43 insertions(+) diff --git a/include/linux/swap.h b/include/linux/swap.h index 4db00ddad261..8d28f6091a32 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -549,6 +549,11 @@ static inline int swap_duplicate(swp_entry_t swp) return 0; } +static inline int swapcache_prepare(swp_entry_t swp) +{ + return 0; +} + static inline void swap_free(swp_entry_t swp) { } diff --git a/mm/memory.c b/mm/memory.c index 7e1f4849463a..7059230d0a54 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -3799,6 +3799,7 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) struct page *page; struct swap_info_struct *si = NULL; rmap_t rmap_flags = RMAP_NONE; + bool need_clear_cache = false; bool exclusive = false; swp_entry_t entry; pte_t pte; @@ -3867,6 +3868,20 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) if (!folio) { if (data_race(si->flags & SWP_SYNCHRONOUS_IO) && __swap_count(entry) == 1) { + /* + * Prevent parallel swapin from proceeding with + * the cache flag. Otherwise, another thread may + * finish swapin first, free the entry, and swapout + * reusing the same entry. It's undetectable as + * pte_same() returns true due to entry reuse. + */ + if (swapcache_prepare(entry)) { + /* Relax a bit to prevent rapid repeated page faults */ + schedule(); + goto out; + } + need_clear_cache = true; + /* skip swapcache */ folio = vma_alloc_folio(GFP_HIGHUSER_MOVABLE, 0, vma, vmf->address, false); @@ -4117,6 +4132,9 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) if (vmf->pte) pte_unmap_unlock(vmf->pte, vmf->ptl); out: + /* Clear the swap cache pin for direct swapin after PTL unlock */ + if (need_clear_cache) + swapcache_clear(si, entry); if (si) put_swap_device(si); return ret; @@ -4131,6 +4149,8 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) folio_unlock(swapcache); folio_put(swapcache); } + if (need_clear_cache) + swapcache_clear(si, entry); if (si) put_swap_device(si); return ret; diff --git a/mm/swap.h b/mm/swap.h index 758c46ca671e..fc2f6ade7f80 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -41,6 +41,7 @@ void __delete_from_swap_cache(struct folio *folio, void delete_from_swap_cache(struct folio *folio); void clear_shadow_from_swap_cache(int type, unsigned long begin, unsigned long end); +void swapcache_clear(struct swap_info_struct *si, swp_entry_t entry); struct folio *swap_cache_get_folio(swp_entry_t entry, struct vm_area_struct *vma, unsigned long addr); struct folio *filemap_get_incore_folio(struct address_space *mapping, @@ -97,6 +98,10 @@ static inline int swap_writepage(struct page *p, struct writeback_control *wbc) return 0; } +static inline void swapcache_clear(struct swap_info_struct *si, swp_entry_t entry) +{ +} + static inline struct folio *swap_cache_get_folio(swp_entry_t entry, struct vm_area_struct *vma, unsigned long addr) { diff --git a/mm/swapfile.c b/mm/swapfile.c index 556ff7347d5f..746aa9da5302 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -3365,6 +3365,19 @@ int swapcache_prepare(swp_entry_t entry) return __swap_duplicate(entry, SWAP_HAS_CACHE); } +void swapcache_clear(struct swap_info_struct *si, swp_entry_t entry) +{ + struct swap_cluster_info *ci; + unsigned long offset = swp_offset(entry); + unsigned char usage; + + ci = lock_cluster_or_swap_info(si, offset); + usage = __swap_entry_free_locked(si, offset, SWAP_HAS_CACHE); + unlock_cluster_or_swap_info(si, ci); + if (!usage) + free_swap_slot(entry); +} + struct swap_info_struct *swp_swap_info(swp_entry_t entry) { return swap_type_to_swap_info(swp_type(entry)); -- 2.43.0

1 year, 9 months

5
19
0 0

2025

2024

2023

2022

2021

2020

2019

2018

2017

Linux-stable-mirror February 2024