Commit graph

978,566 commits

Author SHA1 Message Date
Minchan Kim
bb4e220b57
Revert "mm: protect VMA modifications using VMA sequence count"
This reverts commit 06219e6c87.

Bug: 128240262
Change-Id: If31a4c81badd891e6ca5740dbd022b5edbe47254
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:09 +00:00
Minchan Kim
99a68bfb93
Revert "mm: protect mremap() against SPF hanlder"
This reverts commit 728a43dfff.

Bug: 128240262
Change-Id: Ida9fb7e41a7905755470e20f5c72867bb3dad03f
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:08 +00:00
Minchan Kim
a43fdd54d4
Revert "mm: protect SPF handler against anon_vma changes"
This reverts commit ff6cfddf7a.

Bug: 128240262
Change-Id: I317785d46d335aab974e194e5ea0d744513baecd
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:08 +00:00
Minchan Kim
6d707879c3
Revert "mm: cache some VMA fields in the vm_fault structure"
This reverts commit 5d1dccddd7.

Bug: 128240262
Change-Id: Iea60ab5360ae83ca0efcbb47a101b3870691a93f
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:07 +00:00
Minchan Kim
c1826f5b61
Revert "mm/migrate: Pass vm_fault pointer to migrate_misplaced_page()"
This reverts commit b7efa36bf5.

Bug: 128240262
Change-Id: I21a63bb3e600beb422a8fc5240a4c3bed611fadc
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:07 +00:00
Minchan Kim
9422831755
Revert "mm: introduce __lru_cache_add_active_or_unevictable"
This reverts commit 2536f11b7d.

Bug: 128240262
Change-Id: I54aa09970c5967cfa7a93deca52f28773e27c041
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:06 +00:00
Minchan Kim
af3517b11d
Revert "mm: introduce __vm_normal_page()"
This reverts commit ec5147dea1.

Bug: 128240262
Change-Id: Iadc1c38e82431020b79194aff4bb25e7d9d66125
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:06 +00:00
Minchan Kim
7ec1f37926
Revert "mm: introduce __page_add_new_anon_rmap()"
This reverts commit d56126780c.

Bug: 128240262
Change-Id: Ia5a417e52de006fba4f8b1b51d9ae4db36fd9035
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:06 +00:00
Minchan Kim
fa0483bf62
Revert "mm: protect mm_rb tree with a rwlock"
This reverts commit 9db216b5c9.

Bug: 128240262
Change-Id: I210789054d394b929de6d9444040f1d2aac14917
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:05 +00:00
Minchan Kim
c0e5f947b4
Revert "mm: provide speculative fault infrastructure"
This reverts commit 6e9deb2ea7.

Bug: 128240262
Change-Id: I1060d8e43cebca1826a489fa3271279501c8647c
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:05 +00:00
Minchan Kim
fb8981f881
Revert "mm: adding speculative page fault failure trace events"
This reverts commit 52e2e1f805.

Bug: 128240262
Change-Id: I96fe4f7e02f5d6716e20b1f8bc03b686dd189f09
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:05 +00:00
Minchan Kim
a45b2e061e
Revert "mm: speculative page fault handler return VMA"
This reverts commit f556cd74a7.

Bug: 128240262
Change-Id: Icc3edff62de261aee10658b1567b85ff7b5d58dd
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:04 +00:00
Minchan Kim
f2f6a4239e
Revert "mm: add speculative page fault vmstats"
This reverts commit 2842d9234e.

Bug: 128240262
Change-Id: I3ba5c0ab739af15061dd376fcb15366e9644ed4f
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:04 +00:00
Martin Liu
402afac2ab
Revert "arm64/mm: define ARCH_SUPPORTS_SPECULATIVE_PAGE_FAULT"
This reverts commit bcca0756b3.

Reason for revert: remove SPF non upstream code
Bug: 140544941
Test: boot
Change-Id: Id97756b85be0a1690e000fd24f125f21915d20da
Signed-off-by: Martin Liu <liumartin@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:03 +00:00
Minchan Kim
561478b685
Revert "arm64/mm: add speculative page fault"
This reverts commit ad3023acf9.

Bug: 128240262
Change-Id: I80f12ab9b25478a13b69d9c8fe7b46b3f65b1197
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:03 +00:00
Martin Liu
37d40ca049
Revert "mm: protect against PTE changes done by dup_mmap()"
This reverts commit a9e3a1ab5d.

Reason for revert: remove SPF non upstream code
Bug: 140544941
Test: boot
Change-Id: I912a8891ac6cf3e72c7b7aa27df2922554b31491
Signed-off-by: Martin Liu <liumartin@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:03 +00:00
Minchan Kim
fb1f97dab8
Revert "mm: don't do swap readahead during speculative page fault"
This reverts commit c41742ec7c.

Bug: 128240262
Change-Id: Ide6a3bb060fd1792c1ac5880cfc2298931ad07ed
Signed-off-by: Minchan Kim <minchan@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:02 +00:00
Alexander Winkowski
6c8feaacfc
Revert "mm: Fix sleeping while atomic during speculative page fault"
This reverts commit 8fbac439e0.

Change-Id: Ifa8e76de6c80a7453969df8bd431b94f9e66e547
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:02 +00:00
Martin Liu
a01be3aa02
Revert "mm: allow vmas with vm_ops to be speculatively handled"
This reverts commit b37bae60c5.

Reason for revert: remove SPF non upstream code
Bug: 140544941
Test: boot
Change-Id: I466435dabfed767085934109a43bf7ca3da855a3
Signed-off-by: Martin Liu <liumartin@google.com>
[dereference23: Forward port to msm-5.4]
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:01 +00:00
Alexander Winkowski
a93728d35e
Revert "mm: remove the speculative page fault traces"
This reverts commit f96d3d9c9e.

Change-Id: Ib7723f11c64b05a989113078b2f3c90b00d65946
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:01 +00:00
Alexander Winkowski
d4119e8381
Revert "mm: sync rss in speculative page fault path"
This reverts commit 6011ee7ef2.

Change-Id: I9192d3ad2a32e2f284248d6ccdc936eed48a7a61
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:01 +00:00
Alexander Winkowski
8b720d1a14
Revert "mm: skip speculative path for non-anonymous COW faults"
This reverts commit d95ca82a5d.

Change-Id: I65cd50b4aab3578fdb2ae1b67d25a84775b19c4e
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:00 +00:00
Alexander Winkowski
ec928f71b9
Revert "mm: fix non-anon COW fault"
This reverts commit 1b7ca44c71.

Change-Id: I3feb16538efb3b7260cfac8404f9ad68b150a64c
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:44:00 +00:00
Alexander Winkowski
5f80d03c8c
Revert "ANDROID: mm: use raw seqcount variants in vm_write_*"
This reverts commit f2c1ef71ed.

Change-Id: I5dc242f3d5aa05eb4630426d7444c36412ed6191
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:59 +00:00
Alexander Winkowski
d74f249592
Revert "ANDROID: mm: Fix page table lookup in speculative fault path"
This reverts commit 1b3c72b43c.

Change-Id: I03356916c3b27db106e75ed07b3d53c2479926e7
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:59 +00:00
Alexander Winkowski
f3fcee6d90
Revert "ANDROID: mm: skip pte_alloc during speculative page fault"
This reverts commit b12d6d9506.

Change-Id: I0891749da9dc05c7f9a3ad909a9d054e74a45e2b
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:59 +00:00
Alexander Winkowski
a1d1eb8bd3
Revert "ANDROID: mm: prevent speculative page fault handling for in do_swap_page()"
This reverts commit cb68c255f8.

Change-Id: I3deaf837842423b91b1cca32ec974d9612c882b7
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:58 +00:00
Alexander Winkowski
9c47c0190b
Revert "ANDROID: mm: prevent reads of unstable pmd during speculation"
This reverts commit f87e6b8d45.

Change-Id: I3a896b78e77f9db5a52bcf261a35c3ec4afefb9c
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:58 +00:00
Alexander Winkowski
9ef3adc15b
Revert "BACKPORT: FROMLIST: mm: implement speculative handling in filemap_fault()"
This reverts commit 5b5bd362f1.

Change-Id: I3197cb64766cd5e102794d24c946bb88a8e73652
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:57 +00:00
Alexander Winkowski
ee757276ff
Revert "ANDROID: mm/khugepaged: add missing vm_write_{begin|end}"
This reverts commit 365a5b7af5.

Change-Id: Ifb555936b0887bab47590ccd77bdd435da10c179
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:57 +00:00
Alexander Winkowski
924ed5256e
Revert "ANDROID: mm: remove sequence counting when mmap_lock is not exclusively owned"
This reverts commit ad939deb18.

Change-Id: I210d9aa077b3b68916166d2b5265ef0600c5e3bb
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:57 +00:00
Alexander Winkowski
1185e5aca1
Revert "ANDROID: mm: assert that mmap_lock is taken exclusively in vm_write_begin"
This reverts commit 51cfccaecd.

Change-Id: Iacc9012268a5e7aedfe38afd8766e84a042a30b5
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:56 +00:00
Alexander Winkowski
d45fba04cc
Revert "ANDROID: disable page table moves when speculative page faults are enabled"
This reverts commit 78035f7a50.

Change-Id: I458224eceaa9c415120f91f64ec09a49ac6895a2
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:56 +00:00
Alexander Winkowski
12326201b5
Revert "ANDROID: mm: fix invalid backport in speculative page fault path"
This reverts commit 3aa1fadec5.

Change-Id: I5dccfc1ccc41650c9d2e5f944d92c7773093524f
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:56 +00:00
Alexander Winkowski
fa80be5b4e
Revert "ANDROID: Re-enable fast mremap and fix UAF with SPF"
This reverts commit ea5f9d7e7e.

Change-Id: Ib977f22950887f417660b60738f26289d9422c39
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:55 +00:00
Alexander Winkowski
2673354abe
Revert "ANDROID: mm/filemap: Fix missing put_page() for speculative page fault"
This reverts commit 290d702383.

Change-Id: I6d083e1e4a70a352cdf8162e72c1f3bfb1cc0b64
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:55 +00:00
Alexander Winkowski
1828476938
Revert "BACKPORT: FROMLIST: mm: protect free_pgtables with mmap_lock write lock in exit_mmap"
This reverts commit bb3bc96f35.

Change-Id: Id31ad7085d20706bbc6dc12a05b7c1e31b3edc2a
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2025-05-24 05:43:37 +00:00
Michael Bestas
4a527b3703 power: supply: qti_battery_charger: Fix charging_enabled node disabled state
The previous logic was flawed, since it was limiting charge to 1A
when charging control was disabled. Default to thermal limit to mimic
what restrict_chg node does.

Change-Id: I18fb4f18ade276b561171f3217ddafa0e48a4555
2025-05-18 09:04:39 +00:00
Lokesh Gidra
801da2ad39 ANDROID: Fix compilation error with huge_pmd_share()
There was an asterisk missing for one of the function parameters in the
upstreamed patch.

Fixes: e8ba376301a36 ("BACKPORT: FROMGIT: hugetlb: pass vma into
huge_pte_alloc() and huge_pmd_share()")

Signed-off-by: Lokesh Gidra <lokeshgidra@google.com>
Bug: 160737021
Bug: 169683130
Change-Id: I110563bc38e60a829fe7808f69dc0aa0f203a50e
2025-05-18 08:08:42 +00:00
Peter Xu
822149ba00 BACKPORT: mm/gup: Remove enfornced COW mechanism
With the more strict (but greatly simplified) page reuse logic in
do_wp_page(), we can safely go back to the world where cow is not
enforced with writes.

This essentially reverts commit 17839856fd58 ("gup: document and work
around 'COW can break either way' issue").  There are some context
differences due to some changes later on around it:

  2170ecfa7688 ("drm/i915: convert get_user_pages() --> pin_user_pages()", 2020-06-03)
  376a34efa4ee ("mm/gup: refactor and de-duplicate gup_fast() code", 2020-06-03)

Some lines moved back and forth with those, but this revert patch should
have striped out and covered all the enforced cow bits anyways.

Suggested-by: Linus Torvalds <torvalds@linux-foundation.org>
Signed-off-by: Peter Xu <peterx@redhat.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>

(cherry picked from commit a308c71bf1e6e19cc2e4ced31853ee0fc7cb439a)

Bug: 173684178
Signed-off-by: Suren Baghdasaryan <surenb@google.com>
Change-Id: Ib2308e4ff3f2e9283e5c93b096047cf276edf27f
2025-05-18 08:08:42 +00:00
Peter Xu
3d6559b609 UPSTREAM: mm/userfaultfd: selftests: fix memory corruption with thp enabled
In RHEL's gating selftests we've encountered memory corruption in the
uffd event test even with upstream kernel:

        # ./userfaultfd anon 128 4
        nr_pages: 32768, nr_pages_per_cpu: 32768
        bounces: 3, mode: rnd racing read, userfaults: 6240 missing (6240) 14729 wp (14729)
        bounces: 2, mode: racing read, userfaults: 1444 missing (1444) 28877 wp (28877)
        bounces: 1, mode: rnd read, userfaults: 6055 missing (6055) 14699 wp (14699)
        bounces: 0, mode: read, userfaults: 82 missing (82) 25196 wp (25196)
        testing uffd-wp with pagemap (pgsize=4096): done
        testing uffd-wp with pagemap (pgsize=2097152): done
        testing events (fork, remap, remove): ERROR: nr 32427 memory corruption 0 1 (errno=0, line=963)
        ERROR: faulting process failed (errno=0, line=1117)

It can be easily reproduced when global thp enabled, which is the
default for RHEL.

It's also known as a side effect of commit 0db282ba2c12 ("selftest: use
mmap instead of posix_memalign to allocate memory", 2021-07-23), which
is imho right itself on using mmap() to make sure the addresses will be
untagged even on arm.

The problem is, for each test we allocate buffers using two
allocate_area() calls.  We assumed these two buffers won't affect each
other, however they could, because mmap() could have found that the two
buffers are near each other and having the same VMA flags, so they got
merged into one VMA.

It won't be a big problem if thp is not enabled, but when thp is
agressively enabled it means when initializing the src buffer it could
accidentally setup part of the dest buffer too when there's a shared THP
that overlaps the two regions.  Then some of the dest buffer won't be
able to be trapped by userfaultfd missing mode, then it'll cause memory
corruption as described.

To fix it, do release_pages() after initializing the src buffer.

Since the previous two release_pages() calls are after
uffd_test_ctx_clear() which will unmap all the buffers anyway (which is
stronger than release pages; as unmap() also tear town pgtables), drop
them as they shouldn't really be anything useful.

We can mark the Fixes tag upon 0db282ba2c12 as it's reported to only
happen there, however the real "Fixes" IMHO should be 8ba6e8640844, as
before that commit we'll always do explicit release_pages() before
registration of uffd, and 8ba6e8640844 changed that logic by adding
extra unmap/map and we didn't release the pages at the right place.
Meanwhile I don't have a solid glue anyway on whether posix_memalign()
could always avoid triggering this bug, hence it's safer to attach this
fix to commit 8ba6e8640844.

Link: https://lkml.kernel.org/r/20210923232512.210092-1-peterx@redhat.com
Fixes: 8ba6e8640844 ("userfaultfd/selftests: reinitialize test context in each test")
Bugzilla: https://bugzilla.redhat.com/show_bug.cgi?id=1994931
Signed-off-by: Peter Xu <peterx@redhat.com>
Reported-by: Li Wang <liwan@redhat.com>
Tested-by: Li Wang <liwang@redhat.com>
Reviewed-by: Axel Rasmussen <axelrasmussen@google.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Nadav Amit <nadav.amit@gmail.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit 8913970c19915bbe773d97d42989cd85b7fdc098)
Signed-off-by: Lee Jones <lee.jones@linaro.org>
Change-Id: I1378a52661d17c5cf6e8a0e84e3216556160c1b8
2025-05-18 08:08:42 +00:00
Peter Xu
7174d6f00c UPSTREAM: mm/shmem: use page_mapping() to detect page cache for uffd continue
mfill_atomic_install_pte() checks page->mapping to detect whether one page
is used in the page cache.  However as pointed out by Matthew, the page
can logically be a tail page rather than always the head in the case of
uffd minor mode with UFFDIO_CONTINUE.  It means we could wrongly install
one pte with shmem thp tail page assuming it's an anonymous page.

It's not that clear even for anonymous page, since normally anonymous
pages also have page->mapping being setup with the anon vma.  It's safe
here only because the only such caller to mfill_atomic_install_pte() is
always passing in a newly allocated page (mcopy_atomic_pte()), whose
page->mapping is not yet setup.  However that's not extremely obvious
either.

For either of above, use page_mapping() instead.

Bug: 254441685
Link: https://lkml.kernel.org/r/Y2K+y7wnhC4vbnP2@x1n
Fixes: 153132571f02 ("userfaultfd/shmem: support UFFDIO_CONTINUE for shmem")
Signed-off-by: Peter Xu <peterx@redhat.com>
Reported-by: Matthew Wilcox <willy@infradead.org>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Hugh Dickins <hughd@google.com>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
(cherry picked from commit 93b0d9178743a68723babe8448981f658aebc58e)
Signed-off-by: Lee Jones <joneslee@google.com>
Change-Id: I03246130310cc7f3486843ed945ef92cab966cdc
2025-05-18 08:08:42 +00:00
Nadav Amit
4a02a5091b UPSTREAM: mm/userfaultfd: fix memory corruption due to writeprotect
Userfaultfd self-test fails occasionally, indicating a memory corruption.

Analyzing this problem indicates that there is a real bug since mmap_lock
is only taken for read in mwriteprotect_range() and defers flushes, and
since there is insufficient consideration of concurrent deferred TLB
flushes in wp_page_copy().  Although the PTE is flushed from the TLBs in
wp_page_copy(), this flush takes place after the copy has already been
performed, and therefore changes of the page are possible between the time
of the copy and the time in which the PTE is flushed.

To make matters worse, memory-unprotection using userfaultfd also poses a
problem.  Although memory unprotection is logically a promotion of PTE
permissions, and therefore should not require a TLB flush, the current
userrfaultfd code might actually cause a demotion of the architectural PTE
permission: when userfaultfd_writeprotect() unprotects memory region, it
unintentionally *clears* the RW-bit if it was already set.  Note that this
unprotecting a PTE that is not write-protected is a valid use-case: the
userfaultfd monitor might ask to unprotect a region that holds both
write-protected and write-unprotected PTEs.

The scenario that happens in selftests/vm/userfaultfd is as follows:

cpu0				cpu1			cpu2
----				----			----
							[ Writable PTE
							  cached in TLB ]
userfaultfd_writeprotect()
[ write-*unprotect* ]
mwriteprotect_range()
mmap_read_lock()
change_protection()

change_protection_range()
...
change_pte_range()
[ *clear* “write”-bit ]
[ defer TLB flushes ]
				[ page-fault ]
				...
				wp_page_copy()
				 cow_user_page()
				  [ copy page ]
							[ write to old
							  page ]
				...
				 set_pte_at_notify()

A similar scenario can happen:

cpu0		cpu1		cpu2		cpu3
----		----		----		----
						[ Writable PTE
				  		  cached in TLB ]
userfaultfd_writeprotect()
[ write-protect ]
[ deferred TLB flush ]
		userfaultfd_writeprotect()
		[ write-unprotect ]
		[ deferred TLB flush]
				[ page-fault ]
				wp_page_copy()
				 cow_user_page()
				 [ copy page ]
				 ...		[ write to page ]
				set_pte_at_notify()

This race exists since commit 292924b26024 ("userfaultfd: wp: apply
_PAGE_UFFD_WP bit").  Yet, as Yu Zhao pointed, these races became apparent
since commit 09854ba94c6a ("mm: do_wp_page() simplification") which made
wp_page_copy() more likely to take place, specifically if page_count(page)
> 1.

To resolve the aforementioned races, check whether there are pending
flushes on uffd-write-protected VMAs, and if there are, perform a flush
before doing the COW.

Further optimizations will follow to avoid during uffd-write-unprotect
unnecassary PTE write-protection and TLB flushes.

Bug: 254441685
Link: https://lkml.kernel.org/r/20210304095423.3825684-1-namit@vmware.com
Fixes: 09854ba94c6a ("mm: do_wp_page() simplification")
Signed-off-by: Nadav Amit <namit@vmware.com>
Suggested-by: Yu Zhao <yuzhao@google.com>
Reviewed-by: Peter Xu <peterx@redhat.com>
Tested-by: Peter Xu <peterx@redhat.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Andy Lutomirski <luto@kernel.org>
Cc: Pavel Emelyanov <xemul@openvz.org>
Cc: Mike Kravetz <mike.kravetz@oracle.com>
Cc: Mike Rapoport <rppt@linux.vnet.ibm.com>
Cc: Minchan Kim <minchan@kernel.org>
Cc: Will Deacon <will@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: <stable@vger.kernel.org>	[5.9+]
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit 6ce64428d62026a10cb5d80138ff2f90cc21d367)
Signed-off-by: Lee Jones <joneslee@google.com>
Change-Id: Ie35334aa739acfade88b74c9e5dde5c8a387925d
2025-05-18 08:08:42 +00:00
Baolin Wang
b863f7b8bb UPSTREAM: mm: hugetlb: add missing cache flushing in hugetlb_unshare_all_pmds()
Missed calling flush_cache_range() before removing the sharing PMD
entrires, otherwise data consistence issue may be occurred on some
architectures whose caches are strict and require a virtual>physical
translation to exist for a virtual address.  Thus add it.

Now no architectures enabling PMD sharing will be affected, since they do
not have a VIVT cache.  That means this issue can not be happened in
practice so far.

Bug: 254441685
Link: https://lkml.kernel.org/r/47441086affcabb6ecbe403173e9283b0d904b38.1650956489.git.baolin.wang@linux.alibaba.com
Link: https://lkml.kernel.org/r/419b0e777c9e6d1454dcd906e0f5b752a736d335.1650781755.git.baolin.wang@linux.alibaba.com
Fixes: 6dfeaff93be1 ("hugetlb/userfaultfd: unshare all pmds for hugetlbfs when register wp")
Signed-off-by: Baolin Wang <baolin.wang@linux.alibaba.com>
Reviewed-by: Muchun Song <songmuchun@bytedance.com>
Reviewed-by: Peter Xu <peterx@redhat.com>
Cc: Mike Kravetz <mike.kravetz@oracle.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
(cherry picked from commit 9c8bbfaca1bce84664403fd7dddbef6b3ff0a05a)
Signed-off-by: Lee Jones <joneslee@google.com>
Change-Id: Ib7d886d4f8bc18087771b9999bb7d9941a879581
2025-05-18 08:08:42 +00:00
Peter Xu
5fe446d24a BACKPORT: userfaultfd: wp: declare _UFFDIO_WRITEPROTECT conditionally
Only declare _UFFDIO_WRITEPROTECT if the user specified
UFFDIO_REGISTER_MODE_WP and if all the checks passed.  Then when the user
registers regions with shmem/hugetlbfs we won't expose the new ioctl to
them.  Even with complete anonymous memory range, we'll only expose the
new WP ioctl bit if the register mode has MODE_WP.

Signed-off-by: Peter Xu <peterx@redhat.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reviewed-by: Mike Rapoport <rppt@linux.vnet.ibm.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Bobby Powers <bobbypowers@gmail.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: David Hildenbrand <david@redhat.com>
Cc: Denis Plotnikov <dplotnikov@virtuozzo.com>
Cc: "Dr . David Alan Gilbert" <dgilbert@redhat.com>
Cc: Hugh Dickins <hughd@google.com>
Cc: Jerome Glisse <jglisse@redhat.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: "Kirill A . Shutemov" <kirill@shutemov.name>
Cc: Martin Cracauer <cracauer@cons.org>
Cc: Marty McFadden <mcfadden8@llnl.gov>
Cc: Maya Gokhale <gokhale2@llnl.gov>
Cc: Mel Gorman <mgorman@suse.de>
Cc: Mike Kravetz <mike.kravetz@oracle.com>
Cc: Pavel Emelyanov <xemul@openvz.org>
Cc: Rik van Riel <riel@redhat.com>
Cc: Shaohua Li <shli@fb.com>
Link: http://lkml.kernel.org/r/20200220163112.11409-18-peterx@redhat.com
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit 14819305e09fe4fda546f0dfa12134c8e5366616)

[ Kalesh Singh - resolve conflicts in fs/userfaultfd.c ]

Signed-off-by: Kalesh Singh <kaleshsingh@google.com>
Reported-by: kernel test robot <lkp@intel.com> [1]
[1] https://lore.kernel.org/r/202201170247.Cir3moOM-lkp@intel.com/

Bug: 160737021
Bug: 169683130
Change-Id: I4f205e642f5f0e5824a43303aab30626cce3ddcb
2025-05-18 08:08:42 +00:00
Michal Hocko
f222b766fa BACKPORT: mm, mempolicy: fix up gup usage in lookup_node
ba841078cd05 ("mm/mempolicy: Allow lookup_node() to handle fatal signal")
has added a special casing for 0 return value because that was a possible
gup return value when interrupted by fatal signal.  This has been fixed by
ae46d2aa6a7f ("mm/gup: Let __get_user_pages_locked() return -EINTR for
fatal signal") in the mean time so ba841078cd05 can be reverted.

This patch however doesn't go all the way to revert it because the check
for 0 is wrong and confusing here.  Firstly it is inherently unsafe to
access the page when get_user_pages_locked returns 0 (aka no page
returned).

Fortunatelly this will not happen because get_user_pages_locked will not
return 0 when nr_pages > 0 unless FOLL_NOWAIT is specified which is not
the case here.  Document this potential error code in gup code while we
are at it.

Signed-off-by: Michal Hocko <mhocko@suse.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Peter Xu <peterx@redhat.com>
Link: http://lkml.kernel.org/r/20200421071026.18394-1-mhocko@kernel.org
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit 2d3a36a47964371101d9a71691c18d59ee611e87)

[Kalesh Singh: Resolve conflict in mm/gup.c]
Bug: 176847924
Signed-off-by: Kalesh Singh <kaleshsingh@google.com>
Change-Id: Idac45e534bba1524a696993447220d12332ced05
2025-05-18 08:08:42 +00:00
Peter Xu
1034902c6e UPSTREAM: mm/mempolicy: Allow lookup_node() to handle fatal signal
lookup_node() uses gup to pin the page and get node information.  It
checks against ret>=0 assuming the page will be filled in.  However it's
also possible that gup will return zero, for example, when the thread is
quickly killed with a fatal signal.  Teach lookup_node() to gracefully
return an error -EFAULT if it happens.

Meanwhile, initialize "page" to NULL to avoid potential risk of
exploiting the pointer.

Fixes: 4426e945df58 ("mm/gup: allow VM_FAULT_RETRY for multiple times")
Reported-by: syzbot+693dc11fcb53120b5559@syzkaller.appspotmail.com
Signed-off-by: Peter Xu <peterx@redhat.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit ba841078cd0557b43b59c63f5c048b12168f0db2)

Bug: 176847924
Signed-off-by: Kalesh Singh <kaleshsingh@google.com>
Change-Id: I1cfd121c3b603db000a0bebe252c9dec6377f0b0
2025-05-18 08:08:42 +00:00
Hillf Danton
8bd06762ff UPSTREAM: mm/gup: Let __get_user_pages_locked() return -EINTR for fatal signal
__get_user_pages_locked() will return 0 instead of -EINTR after commit
4426e945df588 ("mm/gup: allow VM_FAULT_RETRY for multiple times") which
added extra code to allow gup detect fatal signal faster.

Restore the original -EINTR behavior.

Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Peter Zijlstra <peterz@infradead.org>
Fixes: 4426e945df58 ("mm/gup: allow VM_FAULT_RETRY for multiple times")
Reported-by: syzbot+3be1a33f04dc782e9fd5@syzkaller.appspotmail.com
Signed-off-by: Hillf Danton <hdanton@sina.com>
Acked-by: Michal Hocko <mhocko@suse.com>
Signed-off-by: Peter Xu <peterx@redhat.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit ae46d2aa6a7fbe8ca0946f24b061b6ccdc6c3f25)

Bug: 176847924
Signed-off-by: Kalesh Singh <kaleshsingh@google.com>
Change-Id: Ic812596b077e7b5fbfde90f2253241afa1fa42cf
2025-05-18 08:08:42 +00:00
Peter Xu
a72e5a6a8f UPSTREAM: mm/gup: fix fixup_user_fault() on multiple retries
This part was overlooked when reworking the gup code on multiple
retries.

When we get the 2nd+ retry, we'll be with TRIED flag set.  Current code
will bail out on the 2nd retry because the !TRIED check will fail so the
retry logic will be skipped.  What's worse is that, it will also return
zero which errornously hints the caller that the page is faulted in
while it's not.

The !TRIED flag check seems to not be needed even before the mutliple
retries change because if we get a VM_FAULT_RETRY, it must be the 1st
retry, and we should not have TRIED set for that.

Fix it by removing the !TRIED check, at the meantime check against fatal
signals properly before the page fault so we can still properly respond
to the user killing the process during retries.

Fixes: 4426e945df58 ("mm/gup: allow VM_FAULT_RETRY for multiple times")
Reported-by: Randy Dunlap <rdunlap@infradead.org>
Signed-off-by: Peter Xu <peterx@redhat.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Alex Williamson <alex.williamson@redhat.com>
Cc: Brian Geffon <bgeffon@google.com>
Link: http://lkml.kernel.org/r/20200502003523.8204-1-peterx@redhat.com
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit 475f4dfc021c5fde69f3b7d3287bde0a50477b05)

Bug: 176847924
Signed-off-by: Kalesh Singh <kaleshsingh@google.com>
Change-Id: I51dbc07fc845b346691f9bcf4846a1d9355cbe71
2025-05-18 08:08:42 +00:00
Peter Xu
79349abdcc UPSTREAM: mm/gup: Mark lock taken only after a successful retake
It's definitely incorrect to mark the lock as taken even if
down_read_killable() failed.

This wass overlooked when we switched from down_read() to
down_read_killable() because down_read() won't fail while
down_read_killable() could.

Fixes: 71335f37c5e8 ("mm/gup: allow to react to fatal signals")
Reported-by: syzbot+a8c70b7f3579fc0587dc@syzkaller.appspotmail.com
Signed-off-by: Peter Xu <peterx@redhat.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
(cherry picked from commit c7b6a566b98524baea6a244186e665d22b633545)

Bug: 176847924
Signed-off-by: Kalesh Singh <kaleshsingh@google.com>
Change-Id: Id0e1fe147e87ff6bba02d5e590f6346089c5e650
2025-05-18 08:08:42 +00:00