Commit graph

982,091 commits

Author SHA1 Message Date
William Bowling
16656dc06d
net: skbuff: preserve shared-frag marker during coalescing
skb_try_coalesce() can attach paged frags from @from to @to.  If @from
has SKBFL_SHARED_FRAG set, the resulting @to skb can contain the same
externally-owned or page-cache-backed frags, but the shared-frag marker
is currently lost.

That breaks the invariant relied on by later in-place writers.  In
particular, ESP input checks skb_has_shared_frag() before deciding
whether an uncloned nonlinear skb can skip skb_cow_data().  If TCP
receive coalescing has moved shared frags into an unmarked skb, ESP can
see skb_has_shared_frag() as false and decrypt in place over page-cache
backed frags.

Propagate SKBFL_SHARED_FRAG when skb_try_coalesce() transfers paged
frags.  The tailroom copy path does not need the marker because it copies
bytes into @to's linear data rather than transferring frag descriptors.

Fixes: cef401de7b ("net: fix possible wrong checksum generation")
Fixes: f4c50a4034e6 ("xfrm: esp: avoid in-place decrypt on shared skb frags")
Signed-off-by: William Bowling <vakzz@zellic.io>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Tested-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Link: https://patch.msgid.link/20260513041635.1289541-1-vakzz@zellic.io
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
(cherry picked from commit f84eca5817390257cef78013d0112481c503b4a3)

Orabug: 39368828

Clean cherry-pick build needs a build fixup, on uek6/u3, shared-frag
state is carried in tx_flags(SKBTX_SHARED_FRAG), so propagate it when
frag ownership moves, and it doesn't have skb_shared_info.flags
SKBFL_SHARED_FRAG so accordingly adapted the patch. This is due to
missing 06b4feb37e64 (net: group skb_shinfo zerocopy related bits
together.) commit in UEK6U3.

Change-Id: I71564bc3a42d7523011705022a49393b63070fd5
Signed-off-by: Harshit Mogalapalli <harshit.m.mogalapalli@oracle.com>
Reviewed-by: Joseph Salisbury <joseph.salisbury@oracle.com>
Signed-off-by: Alok Tiwari <alok.a.tiwari@oracle.com>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-07-29 03:30:59 +06:00
Davidlohr Bueso
2ec0afec5c
BACKPORT: locking/rtmutex: Skip remove_waiter() when waiter is not enqueued
syzbot triggered the following splat in remove_waiter() via
FUTEX_CMP_REQUEUE_PI:

  KASAN: null-ptr-deref in range [0x0000000000000a88-0x0000000000000a8f]
   class_raw_spinlock_constructor
   remove_waiter+0x159/0x1200 kernel/locking/rtmutex.c:1561
   rt_mutex_start_proxy_lock+0x103/0x120
   futex_requeue+0x10e4/0x20d0
   __x64_sys_futex+0x34f/0x4d0

task_blocks_on_rt_mutex() does not arm the waiter upon deadlock detection,
leaving waiter->task nil, where 3bfdc63936dd ("rtmutex: Use waiter::task instead
of current in remove_waiter()") made this fatal.

Furthermore, rt_mutex_start_proxy_lock() should not be calling into remove_waiter()
upon a successfully grabbing the rtmutex. 1a1fb985f2 ("futex: Handle early deadlock
return correctly"), moved the remove_waiter() out of __rt_mutex_start_proxy_lock()
(where 'ret' was only ever 0 or < 0) into the wrapper. Tighten this check to
account for try_to_take_rt_mutex().

Fixes: 3bfdc63936dd ("rtmutex: Use waiter::task instead of current in remove_waiter()")
Change-Id: Ib19a2f223ba67d20496c0ae95096747f5350c520
Reported-by: syzbot+78147abe6c524f183ee9@syzkaller.appspotmail.com
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: stable@vger.kernel.org
Closes: https://lore.kernel.org/all/69f114ac.050a0220.ac8b.0003.GAE@google.com/
Link: https://patch.msgid.link/20260507112913.1019537-1-dave@stgolabs.net
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-07-29 03:30:48 +06:00
Keenan Dong
6efbfed641
BACKPORT: rtmutex: Use waiter::task instead of current in remove_waiter()
remove_waiter() is used by the slowlock paths, but it is also used for
proxy-lock rollback in rt_mutex_start_proxy_lock() when invoked from
futex_requeue().

In the latter case waiter::task is not current, but remove_waiter()
operates on current for the dequeue operation. That results in several
problems:

  1) the rbtree dequeue happens without waiter::task::pi_lock being held

  2) the waiter task's pi_blocked_on state is not cleared, which leaves a
     dangling pointer primed for UAF around.

  3) rt_mutex_adjust_prio_chain() operates on the wrong top priority waiter
     task

Use waiter::task instead of current in all related operations in
remove_waiter() to cure those problems.

[ tglx: Fixup rt_mutex_adjust_prio_chain(), add a comment and amend the
  	changelog ]

Fixes: 8161239a8b ("rtmutex: Simplify PI algorithm and make highest prio task get lock")
Change-Id: I3ff8da2830773e04f55828b11c9d461ab6ee57c5
Reported-by: Yuan Tan <yuantan098@gmail.com>
Reported-by: Yifan Wu <yifanwucs@gmail.com>
Reported-by: Juefei Pu <tomapufckgml@gmail.com>
Reported-by: Xin Liu <bird@lzu.edu.cn>
Signed-off-by: Keenan Dong <keenanat2000@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: stable@vger.kernel.org
[Tashar02: Open-code scoped_guard() on msm-5.4]
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-07-29 03:30:25 +06:00
Tashfin Shakeer Rhythm
69c7394469
tcp: fix an incorrect __user annotation on tcp_use_userconfig_sysctl_handler
No user pointers for sysctls anymore.

Fixes: c3f877ec98a12 ("BACKPORT: sysctl: pass kernel pointers to ->proc_handler")
Change-Id: Ia9c27a7308aad6fb9b6476f1000d2588d0367014
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-07-29 03:30:09 +06:00
Tashfin Shakeer Rhythm
a1535b0667
tcp: fix an incorrect __user annotation on tcp_proc_delayed_ack_control
No user pointers for sysctls anymore.

Fixes: c3f877ec98a12 ("BACKPORT: sysctl: pass kernel pointers to ->proc_handler")
Change-Id: Ie7d47a3bb5ed4238cea83a493937f38a3fee2571
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-07-29 03:29:57 +06:00
Adrian Reber
ad10207c9d UPSTREAM: tty: allow TIOCSLCKTRMIOS with CAP_CHECKPOINT_RESTORE
[ Upstream commit e0f25b8992345aa5f113da2815f5add98738c611 ]

The capability CAP_CHECKPOINT_RESTORE was introduced to allow non-root
users to checkpoint and restore processes as non-root with CRIU.

This change extends CAP_CHECKPOINT_RESTORE to enable the CRIU option
'--shell-job' as non-root. CRIU's man-page describes the '--shell-job'
option like this:

  Allow one to dump shell jobs. This implies the restored task will
  inherit session and process group ID from the criu itself. This option
  also allows to migrate a single external tty connection, to migrate
  applications like top.

TIOCSLCKTRMIOS can only be done if the process has CAP_SYS_ADMIN and
this change extends it to CAP_SYS_ADMIN or CAP_CHECKPOINT_RESTORE.

With this change it is possible to checkpoint and restore processes
which have a tty connection as non-root if CAP_CHECKPOINT_RESTORE is
set.

Acked-by: Christian Brauner <brauner@kernel.org>
Change-Id: I079e010f7ab7c1dc2e829bfa4023405efd5aada8
Signed-off-by: Adrian Reber <areber@redhat.com>
Acked-by: Andrei Vagin <avagin@gmail.com>
Link: https://lore.kernel.org/r/20231208143656.1019-1-areber@redhat.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:39 -04:00
Adrian Reber
e5abb5d796 UPSTREAM: selftests: add clone3() CAP_CHECKPOINT_RESTORE test
This adds a test that changes its UID, uses capabilities to
get CAP_CHECKPOINT_RESTORE and uses clone3() with set_tid to
create a process with a given PID as non-root.

Change-Id: Iac62e21213ad1bd04f5a2762255fc924c3830ea6
Signed-off-by: Adrian Reber <areber@redhat.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-8-areber@redhat.com
[christian.brauner@ubuntu.com: use TH_LOG() instead of ksft_print_msg()]
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:39 -04:00
Nicolas Viennot
6b4aab010e UPSTREAM: prctl: exe link permission error changed from -EINVAL to -EPERM
This brings consistency with the rest of the prctl() syscall where
-EPERM is returned when failing a capability check.

Change-Id: I1b27a485d4ca9dd99376c6628572509835aabaf3
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Signed-off-by: Adrian Reber <areber@redhat.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-7-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:39 -04:00
Nicolas Viennot
c90ed82a1d UPSTREAM: prctl: Allow local CAP_CHECKPOINT_RESTORE to change /proc/self/exe
Originally, only a local CAP_SYS_ADMIN could change the exe link,
making it difficult for doing checkpoint/restore without CAP_SYS_ADMIN.
This commit adds CAP_CHECKPOINT_RESTORE in addition to CAP_SYS_ADMIN
for permitting changing the exe link.

The following describes the history of the /proc/self/exe permission
checks as it may be difficult to understand what decisions lead to this
point.

* [1] May 2012: This commit introduces the ability of changing
  /proc/self/exe if the user is CAP_SYS_RESOURCE capable.
  In the related discussion [2], no clear thread model is presented for
  what could happen if the /proc/self/exe changes multiple times, or why
  would the admin be at the mercy of userspace.

* [3] Oct 2014: This commit introduces a new API to change
  /proc/self/exe. The permission no longer checks for CAP_SYS_RESOURCE,
  but instead checks if the current user is root (uid=0) in its local
  namespace. In the related discussion [4] it is said that "Controlling
  exe_fd without privileges may turn out to be dangerous. At least
  things like tomoyo examine it for making policy decisions (see
  tomoyo_manager())."

* [5] Dec 2016: This commit removes the restriction to change
  /proc/self/exe at most once. The related discussion [6] informs that
  the audit subsystem relies on the exe symlink, presumably
  audit_log_d_path_exe() in kernel/audit.c.

* [7] May 2017: This commit changed the check from uid==0 to local
  CAP_SYS_ADMIN. No discussion.

* [8] July 2020: A PoC to spoof any program's /proc/self/exe via ptrace
  is demonstrated

Overall, the concrete points that were made to retain capability checks
around changing the exe symlink is that tomoyo_manager() and
audit_log_d_path_exe() uses the exe_file path.

Christian Brauner said that relying on /proc/<pid>/exe being immutable (or
guarded by caps) in a sake of security is a bit misleading. It can only
be used as a hint without any guarantees of what code is being executed
once execve() returns to userspace. Christian suggested that in the
future, we could call audit_log() or similar to inform the admin of all
exe link changes, instead of attempting to provide security guarantees
via permission checks. However, this proposed change requires the
understanding of the security implications in the tomoyo/audit subsystems.

[1] b32dfe3771 ("c/r: prctl: add ability to set new mm_struct::exe_file")
[2] https://lore.kernel.org/patchwork/patch/292515/
[3] f606b77f1a ("prctl: PR_SET_MM -- introduce PR_SET_MM_MAP operation")
[4] https://lore.kernel.org/patchwork/patch/479359/
[5] 3fb4afd9a5 ("prctl: remove one-shot limitation for changing exe link")
[6] https://lore.kernel.org/patchwork/patch/697304/
[7] 4d28df6152 ("prctl: Allow local CAP_SYS_ADMIN changing exe_file")
[8] https://github.com/nviennot/run_as_exe

Change-Id: I44b78ea8000443bc862178aae3f847e86785a1a1
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Signed-off-by: Adrian Reber <areber@redhat.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-6-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:39 -04:00
Adrian Reber
ef13cc913c UPSTREAM: proc: allow access in init userns for map_files with CAP_CHECKPOINT_RESTORE
Opening files in /proc/pid/map_files when the current user is
CAP_CHECKPOINT_RESTORE capable in the root namespace is useful for
checkpointing and restoring to recover files that are unreachable via
the file system such as deleted files, or memfd files.

Change-Id: I6ec3846f81a5208b66ef41e00229b7535e1e6668
Signed-off-by: Adrian Reber <areber@redhat.com>
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Reviewed-by: Cyrill Gorcunov <gorcunov@gmail.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-5-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
63b4832326 UPSTREAM: pid_namespace: use checkpoint_restore_ns_capable() for ns_last_pid
Use the newly introduced capability CAP_CHECKPOINT_RESTORE to allow
writing to ns_last_pid.

Change-Id: I95727c385f136818553260db0e9edfa654f7907e
Signed-off-by: Adrian Reber <areber@redhat.com>
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-4-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
5c348f563d UPSTREAM: pid: use checkpoint_restore_ns_capable() for set_tid
Use the newly introduced capability CAP_CHECKPOINT_RESTORE to allow
using clone3() with set_tid set.

Change-Id: Ifc86c455e500e59a62f9960ba094c4f9973891ff
Signed-off-by: Adrian Reber <areber@redhat.com>
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-3-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
ef8a66d0ee UPSTREAM: capabilities: Introduce CAP_CHECKPOINT_RESTORE
This patch introduces CAP_CHECKPOINT_RESTORE, a new capability facilitating
checkpoint/restore for non-root users.

Over the last years, The CRIU (Checkpoint/Restore In Userspace) team has
been asked numerous times if it is possible to checkpoint/restore a
process as non-root. The answer usually was: 'almost'.

The main blocker to restore a process as non-root was to control the PID
of the restored process. This feature available via the clone3 system
call, or via /proc/sys/kernel/ns_last_pid is unfortunately guarded by
CAP_SYS_ADMIN.

In the past two years, requests for non-root checkpoint/restore have
increased due to the following use cases:
* Checkpoint/Restore in an HPC environment in combination with a
  resource manager distributing jobs where users are always running as
  non-root. There is a desire to provide a way to checkpoint and
  restore long running jobs.
* Container migration as non-root
* We have been in contact with JVM developers who are integrating
  CRIU into a Java VM to decrease the startup time. These
  checkpoint/restore applications are not meant to be running with
  CAP_SYS_ADMIN.

We have seen the following workarounds:
* Use a setuid wrapper around CRIU:
  See https://github.com/FredHutch/slurm-examples/blob/master/checkpointer/lib/checkpointer/checkpointer-suid.c
* Use a setuid helper that writes to ns_last_pid.
  Unfortunately, this helper delegation technique is impossible to use
  with clone3, and is thus prone to races.
  See https://github.com/twosigma/set_ns_last_pid
* Cycle through PIDs with fork() until the desired PID is reached:
  This has been demonstrated to work with cycling rates of 100,000 PIDs/s
  See https://github.com/twosigma/set_ns_last_pid
* Patch out the CAP_SYS_ADMIN check from the kernel
* Run the desired application in a new user and PID namespace to provide
  a local CAP_SYS_ADMIN for controlling PIDs. This technique has limited
  use in typical container environments (e.g., Kubernetes) as /proc is
  typically protected with read-only layers (e.g., /proc/sys) for
  hardening purposes. Read-only layers prevent additional /proc mounts
  (due to proc's SB_I_USERNS_VISIBLE property), making the use of new
  PID namespaces limited as certain applications need access to /proc
  matching their PID namespace.

The introduced capability allows to:
* Control PIDs when the current user is CAP_CHECKPOINT_RESTORE capable
  for the corresponding PID namespace via ns_last_pid/clone3.
* Open files in /proc/pid/map_files when the current user is
  CAP_CHECKPOINT_RESTORE capable in the root namespace, useful for
  recovering files that are unreachable via the file system such as
  deleted files, or memfd files.

See corresponding selftest for an example with clone3().

Change-Id: Ia9ae3fc9a698e37774be945bac5a7d2a0d640cbc
Signed-off-by: Adrian Reber <areber@redhat.com>
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-2-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
2c395af267 UPSTREAM: selftests: add tests for clone3() with *set_tid
This tests clone3() with *set_tid to see if all desired PIDs are working
as expected. The tests are trying multiple invalid input parameters as
well as creating processes while specifying a certain PID in multiple
PID namespaces at the same time.

Additionally this moves common clone3() test code into clone3_selftests.h.

Change-Id: Ic30e083c2b8c3e5c1fe9adff28387e84431f4c71
Signed-off-by: Adrian Reber <areber@redhat.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20191115123621.142252-2-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
62ec29d7ad UPSTREAM: fork: extend clone3() to support setting a PID
The main motivation to add set_tid to clone3() is CRIU.

To restore a process with the same PID/TID CRIU currently uses
/proc/sys/kernel/ns_last_pid. It writes the desired (PID - 1) to
ns_last_pid and then (quickly) does a clone(). This works most of the
time, but it is racy. It is also slow as it requires multiple syscalls.

Extending clone3() to support *set_tid makes it possible restore a
process using CRIU without accessing /proc/sys/kernel/ns_last_pid and
race free (as long as the desired PID/TID is available).

This clone3() extension places the same restrictions (CAP_SYS_ADMIN)
on clone3() with *set_tid as they are currently in place for ns_last_pid.

The original version of this change was using a single value for
set_tid. At the 2019 LPC, after presenting set_tid, it was, however,
decided to change set_tid to an array to enable setting the PID of a
process in multiple PID namespaces at the same time. If a process is
created in a PID namespace it is possible to influence the PID inside
and outside of the PID namespace. Details also in the corresponding
selftest.

To create a process with the following PIDs:

      PID NS level         Requested PID
        0 (host)              31496
        1                        42
        2                         1

For that example the two newly introduced parameters to struct
clone_args (set_tid and set_tid_size) would need to be:

  set_tid[0] = 1;
  set_tid[1] = 42;
  set_tid[2] = 31496;
  set_tid_size = 3;

If only the PIDs of the two innermost nested PID namespaces should be
defined it would look like this:

  set_tid[0] = 1;
  set_tid[1] = 42;
  set_tid_size = 2;

The PID of the newly created process would then be the next available
free PID in the PID namespace level 0 (host) and 42 in the PID namespace
at level 1 and the PID of the process in the innermost PID namespace
would be 1.

The set_tid array is used to specify the PID of a process starting
from the innermost nested PID namespaces up to set_tid_size PID namespaces.

set_tid_size cannot be larger then the current PID namespace level.

Change-Id: I21a8a13c5a794b2a788f3e647fd639b5ae313ebf
Signed-off-by: Adrian Reber <areber@redhat.com>
Reviewed-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: Dmitry Safonov <0x7f454c46@gmail.com>
Acked-by: Andrei Vagin <avagin@gmail.com>
Link: https://lore.kernel.org/r/20191115123621.142252-1-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
0092575594 UPSTREAM: selftests: add tests for clone3()
This adds tests for clone3() with different values and sizes
of struct clone_args.

This selftest was initially part of of the clone3() with PID selftest.
After that patch was almost merged Eugene sent out a couple of patches
to fix problems with these test.

This commit now only contains the clone3() selftest after the LPC
decision to rework clone3() with PID to allow setting the PID in
multiple PID namespaces including all of Eugene's patches.

Change-Id: I32fecd920f7a820c0ca7869a70a8c5b34af583dd
Signed-off-by: Eugene Syromiatnikov <esyr@redhat.com>
Signed-off-by: Adrian Reber <areber@redhat.com>
Reviewed-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20191112095851.811884-1-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Christian Brauner
399202290e UPSTREAM: tests: test CLONE_CLEAR_SIGHAND
Test that CLONE_CLEAR_SIGHAND resets signal handlers to SIG_DFL for the
child process and that CLONE_CLEAR_SIGHAND and CLONE_SIGHAND are
mutually exclusive.

Cc: Florian Weimer <fweimer@redhat.com>
Cc: libc-alpha@sourceware.org
Cc: linux-api@vger.kernel.org
Change-Id: I07ebb8cb2e741cf1a6a43ff4dbe6ee489636f506
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20191014104538.3096-2-christian.brauner@ubuntu.com
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Christian Brauner
67da5e102b UPSTREAM: clone3: add CLONE_CLEAR_SIGHAND
Reset all signal handlers of the child not set to SIG_IGN to SIG_DFL.
Mutually exclusive with CLONE_SIGHAND to not disturb other thread's
signal handler.

In the spirit of closer cooperation between glibc developers and kernel
developers (cf. [2]) this patchset came out of a discussion on the glibc
mailing list for improving posix_spawn() (cf. [1], [3], [4]). Kernel
support for this feature has been explicitly requested by glibc and I
see no reason not to help them with this.

The child helper process on Linux posix_spawn must ensure that no signal
handlers are enabled, so the signal disposition must be either SIG_DFL
or SIG_IGN. However, it requires a sigprocmask to obtain the current
signal mask and at least _NSIG sigaction calls to reset the signal
handlers for each posix_spawn call or complex state tracking that might
lead to data corruption in glibc. Adding this flags lets glibc avoid
these problems.

[1]: https://www.sourceware.org/ml/libc-alpha/2019-10/msg00149.html
[3]: https://www.sourceware.org/ml/libc-alpha/2019-10/msg00158.html
[4]: https://www.sourceware.org/ml/libc-alpha/2019-10/msg00160.html
[2]: https://lwn.net/Articles/799331/
     '[...] by asking for better cooperation with the C-library projects
     in general. They should be copied on patches containing ABI
     changes, for example. I noted that there are often times where
     C-library developers wish the kernel community had done things
     differently; how could those be avoided in the future? Members of
     the audience suggested that more glibc developers should perhaps
     join the linux-api list. The other suggestion was to "copy Florian
     on everything".'
Cc: Florian Weimer <fweimer@redhat.com>
Cc: libc-alpha@sourceware.org
Cc: linux-api@vger.kernel.org
Change-Id: Iced02ef93ed4dcda5ca5d3bd475cbfcae928fb19
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/20191014104538.3096-1-christian.brauner@ubuntu.com
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Nolen Johnson
8d80a9da0f drivers: staging: qca-wifi-host-cmn: Fix implicit-enum-enum-cast errors
Change-Id: I2e72f2fe56aed7c42db7b5ad52937080acb20343
2026-06-24 21:58:50 -04:00
Nolen Johnson
33304e5fe0 techpack: display: msm: sde: Fix uninitialized errors
Change-Id: I481190f2a03bc49976241e4c0b60304ca634f89a
2026-06-24 21:58:49 -04:00
Nolen Johnson
ce01923678 drivers: hwtracing: coresight: Fix sometimes-uninitialized errors
Change-Id: Ia0a6242bec8ceaa27ef623a668e032a08e5fab0d
2026-06-24 21:27:17 -04:00
Nathan Chancellor
ae1cc25c5a
kbuild: Remove support for Clang's ThinLTO caching
There is an issue in clang's ThinLTO caching (enabled for the kernel via
'--thinlto-cache-dir') with .incbin, which the kernel occasionally uses
to include data within the kernel, such as the .config file for
/proc/config.gz. For example, when changing the .config and rebuilding
vmlinux, the copy of .config in vmlinux does not match the copy of
.config in the build folder:

  $ echo 'CONFIG_LTO_NONE=n
  CONFIG_LTO_CLANG_THIN=y
  CONFIG_IKCONFIG=y
  CONFIG_HEADERS_INSTALL=y' >kernel/configs/repro.config

  $ make -skj"$(nproc)" ARCH=x86_64 LLVM=1 clean defconfig repro.config vmlinux
  ...

  $ grep CONFIG_HEADERS_INSTALL .config
  CONFIG_HEADERS_INSTALL=y

  $ scripts/extract-ikconfig vmlinux | grep CONFIG_HEADERS_INSTALL
  CONFIG_HEADERS_INSTALL=y

  $ scripts/config -d HEADERS_INSTALL

  $ make -kj"$(nproc)" ARCH=x86_64 LLVM=1 vmlinux
  ...
    UPD     kernel/config_data
    GZIP    kernel/config_data.gz
    CC      kernel/configs.o
  ...
    LD      vmlinux
  ...

  $ grep CONFIG_HEADERS_INSTALL .config
  # CONFIG_HEADERS_INSTALL is not set

  $ scripts/extract-ikconfig vmlinux | grep CONFIG_HEADERS_INSTALL
  CONFIG_HEADERS_INSTALL=y

Without '--thinlto-cache-dir' or when using full LTO, this issue does
not occur.

Benchmarking incremental builds on a few different machines with and
without the cache shows a 20% increase in incremental build time without
the cache when measured by touching init/main.c and running 'make all'.

ARCH=arm64 defconfig + CONFIG_LTO_CLANG_THIN=y on an arm64 host:

  Benchmark 1: With ThinLTO cache
    Time (mean ± σ):     56.347 s ±  0.163 s    [User: 83.768 s, System: 24.661 s]
    Range (min … max):   56.109 s … 56.594 s    10 runs

  Benchmark 2: Without ThinLTO cache
    Time (mean ± σ):     67.740 s ±  0.479 s    [User: 718.458 s, System: 31.797 s]
    Range (min … max):   67.059 s … 68.556 s    10 runs

  Summary
    With ThinLTO cache ran
      1.20 ± 0.01 times faster than Without ThinLTO cache

ARCH=x86_64 defconfig + CONFIG_LTO_CLANG_THIN=y on an x86_64 host:

  Benchmark 1: With ThinLTO cache
    Time (mean ± σ):     85.772 s ±  0.252 s    [User: 91.505 s, System: 8.408 s]
    Range (min … max):   85.447 s … 86.244 s    10 runs

  Benchmark 2: Without ThinLTO cache
    Time (mean ± σ):     103.833 s ±  0.288 s    [User: 232.058 s, System: 8.569 s]
    Range (min … max):   103.286 s … 104.124 s    10 runs

  Summary
    With ThinLTO cache ran
      1.21 ± 0.00 times faster than Without ThinLTO cache

While it is unfortunate to take this performance improvement off the
table, correctness is more important. If/when this is fixed in LLVM, it
can potentially be brought back in a conditional manner. Alternatively,
a developer can just disable LTO if doing incremental compiles quickly
is important, as a full compile cycle can still take over a minute even
with the cache and it is unlikely that LTO will result in functional
differences for a kernel change.

Cc: stable@vger.kernel.org
Fixes: dc5723b02e52 ("kbuild: add support for Clang LTO")
Reported-by: Yifan Hong <elsk@google.com>
Closes: https://github.com/ClangBuiltLinux/linux/issues/2021
Reported-by: Masami Hiramatsu <mhiramat@kernel.org>
Closes: https://lore.kernel.org/r/20220327115526.cc4b0ff55fc53c97683c3e4d@kernel.org/
Signed-off-by: Nathan Chancellor <nathan@kernel.org>
Signed-off-by: Masahiro Yamada <masahiroy@kernel.org>
Change-Id: I5e416fd740e48f2e3920d9d6917a600f179e4885
2026-05-29 22:07:02 +00:00
Sami Tolvanen
99876597c2
kbuild: lto: force rebuilds when switching CONFIG_LTO
When doing non-clean builds and switching between CONFIG_LTO=n and
CONFIG_LTO=y, the build system (correctly) didn't notice that assembly
and LTO-excluded C object files were rewritten in place by objtool (to
add the .orc_unwind* sections), since their build command lines were the
same between CONFIG_LTO=y and CONFIG_LTO=n. The objtool step would fail:

vmlinux.o: warning: objtool: file already has .orc_unwind section, skipping
make: *** [Makefile:1194: vmlinux] Error 255

Avoid this by making sure the build will see a difference between an LTO
and non-LTO build (by including "-fno-lto" in KBUILD_*FLAGS). This will
get ignored when CC_FLAGS_LTO is present, and will not be included at
all when CONFIG_LTO=n.

Change-Id: I7b6a3221d0972ad13db0d138c66f3789eec60062
Signed-off-by: Sami Tolvanen <samitolvanen@google.com>
Signed-off-by: Kees Cook <keescook@chromium.org>
2026-05-29 22:07:01 +00:00
Fangrui Song
823ac02460
Makefile: use -z pack-relative-relocs
Commit 27f2a4db76e8 ("Makefile: fix GDB warning with CONFIG_RELR")
added --use-android-relr-tags to fix a GDB warning

BFD: /android0/linux-next/vmlinux: unknown type [0x13] section `.relr.dyn'

The GDB warning has been fixed in version 11.2.

The DT_ANDROID_RELR tag was deprecated since DT_RELR was standardized.
Thus, --use-android-relr-tags should be removed. While making the
change, try -z pack-relative-relocs, which is supported since LLD 15.
Keep supporting --pack-dyn-relocs=relr as well for older LLD versions.
There is no indication of obsolescence for --pack-dyn-relocs=relr.

As of today, GNU ld supports the latter option for x86 and powerpc64
ports and has no intention to support --pack-dyn-relocs=relr. In the
absence of the glibc symbol version GLIBC_ABI_DT_RELR,
--pack-dyn-relocs=relr and -z pack-relative-relocs are identical in
ld.lld.

GNU ld and newer versions of LLD report warnings (instead of errors) for
unknown -z options. Only errors lead to non-zero exit codes. Therefore,
we should test --pack-dyn-relocs=relr before testing
-z pack-relative-relocs.

Link: https://github.com/ClangBuiltLinux/linux/issues/1057
Link: https://sourceware.org/git/?p=binutils-gdb.git;a=commit;h=a619b58721f0a03fd91c27670d3e4c2fb0d88f1e
Signed-off-by: Fangrui Song <maskray@google.com>
Reviewed-by: Nick Desaulniers <ndesaulniers@google.com>
Acked-by: Will Deacon <will@kernel.org>
Signed-off-by: Masahiro Yamada <masahiroy@kernel.org>
Change-Id: Ib4ca450ed863e16651a8725c721e7e0389c992ae
2026-05-29 22:07:01 +00:00
Nick Desaulniers
312d011421
Revert "kbuild: disable clang's default use of -fmerge-all-constants"
This reverts commit 87e0d4f0f3.

-fno-merge-all-constants has been the default since clang-6; the minimum
supported version of clang in the kernel is clang-10 (10.0.1).

Suggested-by: Nathan Chancellor <natechancellor@gmail.com>
Signed-off-by: Nick Desaulniers <ndesaulniers@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Tested-by: Sedat Dilek <sedat.dilek@gmail.com>
Reviewed-by: Fangrui Song <maskray@google.com>
Reviewed-by: Nathan Chancellor <natechancellor@gmail.com>
Reviewed-by: Sedat Dilek <sedat.dilek@gmail.com>
Reviewed-by: Kees Cook <keescook@chromium.org>
Cc: Andrey Konovalov <andreyknvl@google.com>
Cc: Marco Elver <elver@google.com>
Cc: Miguel Ojeda <miguel.ojeda.sandonis@gmail.com>
Cc: Alexei Starovoitov <ast@kernel.org>
Cc: Daniel Borkmann <daniel@iogearbox.net>
Cc: Masahiro Yamada <masahiroy@kernel.org>
Cc: Vincenzo Frascino <vincenzo.frascino@arm.com>
Cc: Will Deacon <will@kernel.org>
Link: https://lkml.kernel.org/r/20200902225911.209899-3-ndesaulniers@google.com
Link: https://reviews.llvm.org/rL329300.
Link: https://github.com/ClangBuiltLinux/linux/issues/9
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Change-Id: Idf61fb41074a182131efff0acf5f7bd958a0a875
2026-05-29 22:06:46 +00:00
Srinivasarao Pathipati
55e0fe8b63
BACKPORT: dma-buf: return success for cmo of dummy clients
kgsl-3d0 is dummy device which don't need cmo
return success along with suppressing warnings for dummy clients.

Change-Id: I4ddcee898931cf017e21c8ecbfec863b4566d962
Signed-off-by: Srinivasarao Pathipati <quic_c_spathi@quicinc.com>
2026-05-29 17:59:14 +00:00
Pankaj Gupta
7393ba39db
Revert "msm: kgsl: Call dma_buf_unmap_attachment() early"
This reverts msm-5.10 commit b84bd97e37f8ac3a0f3194ebcbb1fa962e42cd73.

Warnings for direct dma clients during cache operations are now being
handled by dma-buf driver. Instead of doing dma_buf_unmap_attachment
early if we do it in destroy path, it helps in saving cycles and
improving performance during app launch.

Change-Id: Ic66dd3b66136318abf59685f95fee17890377fd4
Signed-off-by: Pankaj Gupta <quic_gpankaj@quicinc.com>
Signed-off-by: Archana Sriram <quic_c_apsrir@quicinc.com>
2026-05-29 17:59:03 +00:00
Alexander Winkowski
c404d8ddf7
disp: msm: Avoid UB in VBIF register shift calculation
Left-shifting a 32-bit integer by 32 bits or more results in UB. The
common values for vbif_xin_id[] are {10, 11} which means reg_shift
becomes 40/44. For IDs >= 8, the shift is meant to be relative to the
second 32-bit register, not to the first. Fix this issue by masking the
ID so that it will properly describe the intended shift.

Change-Id: Icd530a3bb7cfd9d087e2aecc763d24f775e20466
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2026-05-26 12:07:55 +00:00
Alexander Winkowski
3a757c1357
disp: msm: Fix division by zero during ESD recovery
During ESD recovery, num_mixers becomes 0 for a few moments.

[   34.758513] ================================================================================
[   34.758519] UBSAN: division-overflow in ../techpack/display/msm/sde/sde_crtc.h:522:32
[   34.758521] division by zero
[   34.758524] CPU: 6 PID: 1039 Comm: SDM_EventThread Tainted: G S                5.4.302~positron/4565e293 #17
[   34.758525] Hardware name: Qualcomm Technologies, Inc. Blair QRD NOPMI (DT)
[   34.758526] Call trace:
[   34.758531] dump_backtrace+0x0/0x298
[   34.758532] __dump_stack+0x20/0x28
[   34.758533] dump_stack+0x74/0x9c
[   34.758535] ubsan_epilogue+0xc/0x48
[   34.758538] __ubsan_handle_divrem_overflow+0x184/0x198
[   34.758540] _sde_crtc_setup_lm_bounds+0x300/0x318
[   34.758541] sde_crtc_atomic_check+0x260/0x253c
[   34.758543] drm_atomic_helper_check_planes+0x1b0/0x224
[   34.758544] sde_kms_atomic_check+0xb4/0x2f4
[   34.758545] drm_atomic_check_only+0x410/0x778
[   34.758546] drm_mode_atomic_ioctl+0xb24/0x1020
[   34.758548] drm_ioctl_kernel+0x21c/0x2c0
[   34.758549] drm_ioctl+0x2d0/0x434
[   34.758551] do_vfs_ioctl+0x9b4/0xf9c
[   34.758552] __arm64_sys_ioctl+0x128/0x144
[   34.758553] el0_svc_common+0xdc/0x158
[   34.758554] el0_svc+0x8/0x700
[   34.758555] ================================================================================

Fixes: 4fef803aff ("disp: msm: sde: increase max number of mixers to 4")
Change-Id: Idb60e329ffb47d842bb0b4a16b3d086cc4ec16bc
Signed-off-by: Alexander Winkowski <dereference23@outlook.com>
2026-05-26 12:07:55 +00:00
Vedant Yevale
8c29510d44
disp: msm: dsi: Nullify display modes after kfree
The 'display->modes' pointer must be set to NULL immediately
after calling kfree on it to prevent potential double-free
vulnerabilities or use-after-free issues.

This change ensures robust memory management by clearing the
pointer to previously freed memory, aligning with best practices
for kernel memory deallocation.

Change-Id: I4fd939bc6ee5f3d9c506f0c311a9400f366e7fd2
Signed-off-by: Vedant Yevale <vyevale@qti.qualcomm.com>
(cherry picked from commit 1d3c5a477ea8f6aca1158209e1b87209c21c61f5)
2026-05-26 12:07:54 +00:00
Shreyas K K
0192aaa6cd
devfreq: Fix suspend callback for non-zero min_freq
With commit commit 921884ca1498 ("devfreq: Fix PM
callbacks to support zero frequency"), suspend_freq
of zero is considered valid.

However, the devfreq_set_target is called with zero frequency.
To avoid this, do not call devfreq_set_target if the
suspend_freq is less than the minimun allowed frequency.

Change-Id: I506165d28395eb67138ae2dac53e842357809bbb
Signed-off-by: Shreyas K K <shrekk@codeaurora.org>
2026-05-25 19:48:35 +00:00
Shreyas K K
6024e3974b
devfreq: Fix PM callbacks to support zero frequency
It is perfectly valid for a device to have a zero frequency, usually
it denotes a power collapse. To support this, fix the PM suspend
callback to allow the suspend_freq to be zero.

Change-Id: Id10f95b59b219e82e0695f80361aa6b35cdc5137
Signed-off-by: Shreyas K K <shrekk@codeaurora.org>
2026-05-25 19:48:35 +00:00
Abhijeet Dharmapurikar
01674892a6
devfreq: Allow zero values in opp table
commit ab8f58ad72 ("PM / devfreq: Set
min/max_freq when adding the devfreq device") introduces a change
where an error in finding the ceil or floor is returned as a
zero frequency. This return of zero is treated as error condition
from the callsites.

It is perfectly valid for a device to have a zero frequency, usually
it denotes a power collapse. This valid return of zero frequency ends
up being treated as an error condition.

Fix it.

Change-Id: I5eb6fa622d6fe09207746b96dce05d0b58b233dc
Signed-off-by: Abhijeet Dharmapurikar <adharmap@codeaurora.org>
Signed-off-by: Shreyas K K <shrekk@codeaurora.org>
2026-05-25 19:48:34 +00:00
Abhishek Shah
d8ea382c8a
devfreq: govener_memlat: fix cpu_hotplug_lock recursive lock warning
Possible unsafe locking scenario:

      CPU0
      ----
 lock(cpu_hotplug_lock.rw_sem);
 lock(cpu_hotplug_lock.rw_sem);

*** DEADLOCK ***

Below code under start_hwmon may cause this recursive locking.

	get_online_cpus();
	for_each_cpu(cpu, cpu_possible_mask) {
		if (!cpumask_test_cpu(cpu, cpu_online_mask))
			per_cpu(cpu_is_hp, cpu) = true;
	}
	ret = memlat_event_cpu_hp_init();
	put_online_cpus();

get_online_cpus() acquires cpu_hotplug_lock.rw_sem lock.
Then memlat_event_cpu_hp_init() -> __cpuhp_setup_state() tries
to acquire the same lock again.
Use cpuslocked version of __cpuhp_setup_state() to avoid this warning.

Change-Id: Ied9fe53d02c74816f38c1efe954cea91f9831cc7
Signed-off-by: Abhishek Shah <abhshah@codeaurora.org>
2026-05-25 19:48:34 +00:00
Abhishek Shah
63898d5378
devfreq: governor_memlat: avoid deadlock due to cpu_grp->mons_lock usage
lockdep is detecting possible circular locking dependency
due to cpu_grp->mons_lock mutex as shown below:

Chain exists of:
  &cpu_grp->mons_lock --> state_lock#3 --> devfreq_list_lock

Possible unsafe locking scenario:

      CPU0                    CPU1
      ----                    ----
 lock(devfreq_list_lock);
                              lock(state_lock#3);
                              lock(devfreq_list_lock);
 lock(&cpu_grp->mons_lock);

*** DEADLOCK ***

Below is partial call stacks (in reverse order) showing relevant
locking paths:

Call stack for CPU0:
start_hwmon+0x6c/0x5d0			[may acquire &cpu_grp->mons_lock]
devfreq_memlat_ev_handler+0x2f4/0x3f8
devfreq_add_device+0x418/0x538		[may acquire devfreq_list_lock]
devfreq_add_icc+0x41c/0x528
devfreq_icc_probe+0x20/0x30

Call stack for CPU1:
devfreq_add_governor+0x3c/0x260		[may acquire devfreq_list_lock]
register_memlat+0x84/0x100
memlat_mon_probe+0x414/0x550		[may acquire &cpu_grp->mons_lock]
arm_memlat_mon_driver_probe+0x120/0x3b8

Practically, this race is not possible, since devfreq_add_device first
tries to find the governor, and only if it is succeeds,
devfreq_memlat_ev_handler is called. And the governor would be found
only if devfreq_add_governor has added it priorly.
We have mechanism(initcall_level) in place to make sure that
governor is added before devfreq_add_device happens.

In attempt to quiet the lockdep warning, below fix is adopted:
cpu_grp gets allocated by memlat_cpu_grp_probe, but it is used by its
multiple memlat_mon children's memlat_mon_probe. cpu_grp->mons_lock is
used to prevent race between them for access to cpu_grp and members.
Use a new lock - cpu_grp->init_mons_lock - for the probe routine,
and continue using cpu_grp->mons_lock in other routines.

Change-Id: Ie59c28de0a40914fb120a305067ef0fe504a4455
Signed-off-by: Abhishek Shah <abhshah@codeaurora.org>
2026-05-25 19:48:34 +00:00
Abhishek Shah
ae1efd457b
devfreq: governor_bw_hwmon: fix deadlock warning due to state_lock usage
lockdep is detecting possible circular locking dependency
due to state_lock mutex as shown below:

Possible unsafe locking scenario:

      CPU0                    CPU1
      ----                    ----
 lock(devfreq_list_lock);
                              lock(state_lock#2);
                              lock(devfreq_list_lock);
 lock(state_lock#2);

*** DEADLOCK ***

Below is partial call stacks (in reverse order) showing relevant
locking paths:

Call stack for CPU0:
devfreq_bw_hwmon_ev_handler+0x4c/0x5e0	[may acquire &state_lock]
devfreq_add_device+0x418/0x538		[may acquire &devfreq_list_lock]
devfreq_add_icc+0x41c/0x528
devfreq_icc_probe+0x20/0x30

Call stack for CPU1:
devfreq_add_governor+0x3c/0x260 	[may acquire &devfreq_list_lock]
register_bw_hwmon+0x1e8/0x248		[may acquire &state_lock]
bimc_bwmon_driver_probe+0x310/0x408

Practically, this race is not possible, since devfreq_add_device first
tries to find the governor, and only if it is succeeds,
devfreq_bw_hwmon_ev_handler is called. And the governor would be found
only if devfreq_add_governor has added it priorly.
We have mechanism(initcall_level) in place to make sure that
governor is added before devfreq_add_device happens.

In attempt to quiet the lockdep warning, below fix is adopted:
Since state_lock mutex has different purpose, introduce a new
event_handle_lock mutex for devfreq_bw_hwmon_ev_handler
to avoid this warning.

Change-Id: I53c4514aa2357bf39e9405683f5e73ada0160ba5
Signed-off-by: Abhishek Shah <abhshah@codeaurora.org>
2026-05-25 19:48:33 +00:00
Jeyaprabu J
726d78ef8a
msm: kgsl: Fix UBSAN warnings
Fix possible division by zero error.

Change-Id: I64eb6a5ee1247dafea6701b8244eefab1c40eea6
Signed-off-by: Jeyaprabu J <quic_jeyaprab@quicinc.com>
2026-05-25 19:48:33 +00:00
Pankaj Gupta
e97ce5ceb5
msm: kgsl: Use kthread instead of workqueue for event work
Currently a workqueue is being used to process the event work. In
certain scenarios like when most of CPU cores are busy, there can be a
significant delay between the actual timestamp retire event and when the
work is processed by the events workqueue as workqueues cannot have RT
priority. Hence use kthread instead of workqueue for event work.

Change-Id: Ib1ec7fa1ec3a133d03104c9a029dcc4c06180609
Signed-off-by: Puranam V G Tejaswi <quic_pvgtejas@quicinc.com>
Signed-off-by: Pankaj Gupta <quic_gpankaj@quicinc.com>
2026-05-18 19:21:32 +03:00
Akhil P Oommen
bf39effde5
BACKPORT: msm: kgsl: Avoid unmap after kgsl system suspend
Smmu driver cannot safely handle unmap request after smmu device is
system suspended. So ensure kgsl doesn't initiate an unmap request
after kgsl system suspend as smmu device suspend happens after gpu's.

Kgsl does unmap from the following paths:
  1. Userspace unmap calls
  2. Mementry workqueue
  3. During reclaim

(1) is not a concern as userspace will be collapse before driver
suspend. So we need to ensure that mementry/events workqueues are
flushed and we don't participate in reclaim to take care off
(2) & (3) before kgsl system suspend completes.

[mkbestas]: Ignore kgsl_reclaim parts that don't exist in 5.4
Change-Id: Ibe2c8f5a90fd4d8d4cf212f17c04c52873738e36
Signed-off-by: Akhil P Oommen <quic_akhilpo@quicinc.com>
2026-05-18 19:20:01 +03:00
fadlyas07
6754b27a37
fixup! BACKPORT: kbuild: check the minimum assembler version in Kconfig
Since minimum assembler version checking is not required in 5.10, it should not be required in 4.19 either.

Change-Id: I333454a3dffb4798295adf9b3038ebc9acda948e
2026-05-18 02:59:56 +03:00
duckyduckg
16c281182b msm: vidc/cvp: fix callback type for msm_vidc/cvp_callback
Change-Id: Id31de9aa86d78a046480fdbdbefba5e7b83b2a56
Signed-off-by: duckyduckg <duckyduckg65@gmail.com>
2026-05-17 00:17:24 +05:00
Flicker372
605ff4b4b8 msm: vidc/cvp: fix function type for hfi_cmd_response_callback
Change-Id: I5ae2a546fa08343a7e16dc6aff8183c08a5f3ea2
Signed-off-by: Flicker372 flicker372@outlook.com
2026-05-16 23:45:18 +05:00
Miguel de Dios
4545cdce64 msm: vidc/cvp: Fix handle_cmd_response parameter type
Bug: 117299373
Change-Id: Iffee5fd84fb1340bed7585d3dfb438e7955f61c5
Signed-off-by: Miguel de Dios <migueldedios@google.com>
Signed-off-by: Flicker372 <flicker372@outlook.com>
2026-05-16 23:42:13 +05:00
Daniel Rosenberg
8e41e2a724 ANDROID: Add FUSE_BPF to gki_defconfig
Bug: 202785178
Test: test_fuse passes on linux, feature works on cuttlefish
Signed-off-by: Paul Lawrence <paullawrence@google.com>
Signed-off-by: Daniel Rosenberg <drosen@google.com>
Change-Id: If5d56aa5ff5bf5ee660073ac8f1ef1573a74cfd1
2026-05-13 04:24:06 +00:00
Alexander Martinz
63de6916f0
Reapply "UPSTREAM: seccomp: Remove bogus __user annotations"
This reverts commit cb8dc8a108.

As pointed out on gerrit[1]:
> This revert is wrong. `android12-5.4` doesn't have
> "sysctl: pass kernel pointers to ->proc_handler" but this kernel does.

[1] - https://review.lineageos.org/c/LineageOS/android_kernel_qcom_sm8350/+/482457

Change-Id: I72cd996f6739a67cf78141843e2d933f19e2e92b
Signed-off-by: Alexander Martinz <amartinz@shiftphones.com>
2026-05-12 12:10:45 +02:00
Sandeep Dhavale
4884b8ae4c
ANDROID: fuse: Open-code vma_set_file() logic in fuse_backing_mmap()
Manual manipulation of vma->vm_file and file reference counts in
fuse_backing_mmap() is error-prone.

Open-code the logic of vma_set_file() using swap() to safely
transfer the file reference and ensure proper reference counting
during VMA setup, as vma_set_file() is not available in this
kernel version.

Bug: 498749316
Signed-off-by: Sandeep Dhavale <dhavale@google.com>
Cherrypick-From: https://android-review.googlesource.com/q/commit:5c3d7b8761264e2a2deb121020bb933273b240ba
Merged-In: I22ba0618f9fde6b698aeeeff2005830c4392969c
Change-Id: I22ba0618f9fde6b698aeeeff2005830c4392969c
[dhavale: as mentioned in updated commit message, this patch is adjusted
to account for lack of vma_set_file() helper ]
2026-05-08 21:46:20 +02:00
Darrick J. Wong
e3b7b58b0a
BACKPORT: fuse: fix livelock in synchronous file put from fuseblk workers
[ Upstream commit 26e5c67deb2e1f42a951f022fdf5b9f7eb747b01 ]

I observed a hang when running generic/323 against a fuseblk server.
This test opens a file, initiates a lot of AIO writes to that file
descriptor, and closes the file descriptor before the writes complete.
Unsurprisingly, the AIO exerciser threads are mostly stuck waiting for
responses from the fuseblk server:

[<0>] request_wait_answer+0x1fe/0x2a0 [fuse]
[<0>] __fuse_simple_request+0xd3/0x2b0 [fuse]
[<0>] fuse_do_getattr+0xfc/0x1f0 [fuse]
[<0>] fuse_file_read_iter+0xbe/0x1c0 [fuse]
[<0>] aio_read+0x130/0x1e0
[<0>] io_submit_one+0x542/0x860
[<0>] __x64_sys_io_submit+0x98/0x1a0
[<0>] do_syscall_64+0x37/0xf0
[<0>] entry_SYSCALL_64_after_hwframe+0x4b/0x53

But the /weird/ part is that the fuseblk server threads are waiting for
responses from itself:

[<0>] request_wait_answer+0x1fe/0x2a0 [fuse]
[<0>] __fuse_simple_request+0xd3/0x2b0 [fuse]
[<0>] fuse_file_put+0x9a/0xd0 [fuse]
[<0>] fuse_release+0x36/0x50 [fuse]
[<0>] __fput+0xec/0x2b0
[<0>] task_work_run+0x55/0x90
[<0>] syscall_exit_to_user_mode+0xe9/0x100
[<0>] do_syscall_64+0x43/0xf0
[<0>] entry_SYSCALL_64_after_hwframe+0x4b/0x53

The fuseblk server is fuse2fs so there's nothing all that exciting in
the server itself.  So why is the fuse server calling fuse_file_put?
The commit message for the fstest sheds some light on that:

"By closing the file descriptor before calling io_destroy, you pretty
much guarantee that the last put on the ioctx will be done in interrupt
context (during I/O completion).

Aha.  AIO fgets a new struct file from the fd when it queues the ioctx.
The completion of the FUSE_WRITE command from userspace causes the fuse
server to call the AIO completion function.  The completion puts the
struct file, queuing a delayed fput to the fuse server task.  When the
fuse server task returns to userspace, it has to run the delayed fput,
which in the case of a fuseblk server, it does synchronously.

Sending the FUSE_RELEASE command sychronously from fuse server threads
is a bad idea because a client program can initiate enough simultaneous
AIOs such that all the fuse server threads end up in delayed_fput, and
now there aren't any threads left to handle the queued fuse commands.

Fix this by only using asynchronous fputs when closing files, and leave
a comment explaining why.

Change-Id: I78be67f615f1318447e2fdad9ebfb34e60857b95
Cc: stable@vger.kernel.org # v2.6.38
Fixes: 5a18ec176c ("fuse: fix hang of single threaded fuseblk filesystem")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
[ added isdir parameter to fuse_file_put() call ]
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2026-05-08 21:46:20 +02:00
Arnd Bergmann
1ec19d7588
BACKPORT: compat_ioctl: move more drivers to compat_ptr_ioctl
The .ioctl and .compat_ioctl file operations have the same prototype so
they can both point to the same function, which works great almost all
the time when all the commands are compatible.

One exception is the s390 architecture, where a compat pointer is only
31 bit wide, and converting it into a 64-bit pointer requires calling
compat_ptr(). Most drivers here will never run in s390, but since we now
have a generic helper for it, it's easy enough to use it consistently.

I double-checked all these drivers to ensure that all ioctl arguments
are used as pointers or are ignored, but are not interpreted as integer
values.

Acked-by: Jason Gunthorpe <jgg@mellanox.com>
Acked-by: Daniel Vetter <daniel.vetter@ffwll.ch>
Acked-by: Mauro Carvalho Chehab <mchehab+samsung@kernel.org>
Acked-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Acked-by: David Sterba <dsterba@suse.com>
Acked-by: Darren Hart (VMware) <dvhart@infradead.org>
Acked-by: Jonathan Cameron <Jonathan.Cameron@huawei.com>
Acked-by: Bjorn Andersson <bjorn.andersson@linaro.org>
Acked-by: Dan Williams <dan.j.williams@intel.com>
Change-Id: I7c19d15ea2d2e48a30e726ca0a5b8ee915e6bbe7
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
2026-05-08 21:46:20 +02:00
Miklos Szeredi
389625e5d7
UPSTREAM: fuse: make sure reclaim doesn't write the inode
In writeback cache mode mtime/ctime updates are cached, and flushed to the
server using the ->write_inode() callback.

Closing the file will result in a dirty inode being immediately written,
but in other cases the inode can remain dirty after all references are
dropped.  This result in the inode being written back from reclaim, which
can deadlock on a regular allocation while the request is being served.

The usual mechanisms (GFP_NOFS/PF_MEMALLOC*) don't work for FUSE, because
serving a request involves unrelated userspace process(es).

Instead do the same as for dirty pages: make sure the inode is written
before the last reference is gone.

 - fallocate(2)/copy_file_range(2): these call file_update_time() or
   file_modified(), so flush the inode before returning from the call

 - unlink(2), link(2) and rename(2): these call fuse_update_ctime(), so
   flush the ctime directly from this helper

fuse_flush_time_update(inode) was skipped to call in __fuse_copy_file_range()
because of huge dependent changes.

Change-Id: I102dab1992c9ed2b5e89606265b3d3aa9c1cdb8a
Reported-by: chenguanyou <chenguanyou@xiaomi.com>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
Git-commit: 5c791fe1e2a4f401f819065ea4fc0450849f1818
Git-repo: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
Signed-off-by: Pradeep P V K <quic_pragalla@quicinc.com>
2026-05-08 21:46:20 +02:00
Daniel Rosenberg
1c9bc1da8e
UPSTREAM: ANDROID: fuse-bpf: Correct fuse bpf feature flag
The feature flag should only advertise fuse-bpf if fuse-bpf is a
supported feature

Bug: 372951405
Test: Compile with CONFIG_FUSE_BPF unset
Change-Id: I0049a3075f78576499168b8ebb6e833ccd18db0f
Signed-off-by: Daniel Rosenberg <drosen@google.com>
2026-05-08 21:46:20 +02:00