Commit graph android_kernel_motorola_sm6375/kernel
Author SHA1 Message Date
Michael Bestas
d9d80047e5
Merge remote-tracking branch 'sm8350/lineage-20' into lineage-23.2
* sm8350/lineage-20:
  drivers: staging: fw-api: Correct mis-merge
  techpack: datarmnet-ext: shs: fix CFI failure on module init with newer Clang versions
  ipv6: icmp: clear skb2->cb[] in ip6_err_gen_icmpv6_unreach()
  net: skbuff: fix missing zerocopy reference in pskb_carve helpers
  net: fix fanout UAF in packet_release() via NETDEV_UP race
  ip6_tunnel: clear skb2->cb[] in ip4ip6_err()
  net: skbuff: propagate shared-frag marker through frag-transfer helpers
  net: skbuff: preserve shared-frag marker during coalescing
  BACKPORT: locking/rtmutex: Skip remove_waiter() when waiter is not enqueued
  BACKPORT: rtmutex: Use waiter::task instead of current in remove_waiter()
  tcp: fix an incorrect __user annotation on tcp_use_userconfig_sysctl_handler
  tcp: fix an incorrect __user annotation on tcp_proc_delayed_ack_control
  msm: camera: common: Fix OOB access in bandwidth path handling
  fw-api: CL 32870364 - update fw common interface files
  fw-api: CL 32860533 - update fw common interface files
  fw-api: CL 32860532 - update fw common interface files
  fw-api: CL 32835676 - update fw common interface files
  qcacld-3.0: Avoid phymode and puncture mismatch
  qcacld-3.0: Update phymode along with ch_width update to FW
  fw-api: CL 32817922 - update fw common interface files
  fw-api: CL 32798720 - update fw common interface files
  fw-api: CL 32796759 - update fw common interface files
  fw-api: CL 32796472 - update fw common interface files
  msm:adsprpc: Fix UAF of ctx->perf in async invoke perf counter
  fw-api: CL 32762973 - update fw common interface files
  fw-api: CL 32742658 - update fw common interface files
  fw-api: CL 32731178 - update fw common interface files
  fw-api: CL 32697676 - update fw common interface files
  fw-api: CL 32697669 - update fw common interface files
  fw-api: CL 32692322 - update fw common interface files
  fw-api: CL 32688516 - update fw common interface files
  fw-api: CL 32668717 - update fw common interface files
  fw-api: CL 32640276 - update fw common interface files
  fw-api: CL 32636152 - update fw common interface files
  fw-api: CL 32626000 - update fw common interface files
  fw-api: CL 32623461 - update fw common interface files
  fw-api: CL 32623458 - update fw common interface files
  fw-api: CL 32578922 - update fw common interface files
  qcacld-3.0: Add print log for new roam scan type
  fw-api: CL 32560454 - update fw common interface files
  fw-api: CL 32559793 - update fw common interface files
  fw-api: CL 32559532 - update fw common interface files
  fw-api: CL 32554999 - update fw common interface files
  fw-api: CL 32544807 - update fw common interface files
  fw-api: CL 32516836 - update fw common interface files
  qcacld-3.0:  Disable SAP bringup in DFS when DFSmastercap is disabled
  qcacld-3.0: Set “WIPHY_FLAG_DFS_OFFLOAD” always in wiphy flag
  fw-api: CL 32489576 - update fw common interface files
  fw-api: CL 32479166 - update fw common interface files
  fw-api: CL 32477405 - update fw common interface files
  qcacld-3.0: Add A_INT64 typedef
  fw-api: CL 32465518 - update fw common interface files
  fw-api: CL 32464183 - update fw common interface files
  fw-api: CL 32437017 - update fw common interface files
  fw-api: CL 32415720 - update fw common interface files
  fw-api: CL 32412612 - update fw common interface files
  fw-api: CL 32402817 - update fw common interface files
  fw-api: CL 32402810 - update fw common interface files
  fw-api: CL 32397731 - update fw common interface files
  fw-api: CL 32397726 - update fw common interface files
  tty: msm: Update timer handling for interrupt-safe context
  fw-api: CL 32387418 - update fw common interface files
  fw-api: CL 32362944 - update fw common interface files
  fw-api: CL 32356331 - update fw common interface files
  fw-api: CL 32335202 - update fw common interface files
  fw-api: CL 32323534 - update fw common interface files
  fw-api: CL 32322205 - update fw common interface files
  fw-api: CL 32320410 - update fw common interface files
  disp: msm: dsi: Nullify display modes after kfree
  fw-api: CL 32306378 - update fw common interface files
  fw-api: CL 32296279 - update fw common interface files
  fw-api: CL 32294648 - update fw common interface files
  fw-api: CL 32293073 - update fw common interface files
  fw-api: CL 32292001 - update fw common interface files
  fw-api: CL 32290054 - update fw common interface files
  fw-api: CL 32280322 - update fw common interface files
  fw-api: CL 32280311 - update fw common interface files
  fw-api: CL 32267210 - update fw common interface files
  fw-api: CL 32239632 - update fw common interface files
  fw-api: CL 32229597 - update fw common interface files
  fw-api: CL 32227711 - update fw common interface files
  fw-api: CL 32227291 - update fw common interface files
  fw-api: CL 32227070 - update fw common interface files
  fw-api: CL 32226903 - update fw common interface files
  fw-api: CL 32168619 - update fw common interface files
  qcacld-3.0: Use orig_key_mgmt vdev param for AKM negotiation
  fw-api: CL 32115955 - update fw common interface files
  fw-api: CL 32091641 - update fw common interface files
  fw-api: CL 32088692 - update fw common interface files
  fw-api: CL 32083726 - update fw common interface files
  fw-api: CL 32059381 - update fw common interface files
  fw-api: CL 32054464 - update fw common interface files
  fw-api: CL 32045293 - update fw common interface files
  fw-api: hw_headers: Add v1 e3r50 hw headers for wcn8750
  fw-api: CL 32033412 - update fw common interface files
  fw-api: CL 32023116 - update fw common interface files
  fw-api: CL 32000241 - update fw common interface files
  fw-api: CL 31990638 - update fw common interface files
  qcacld-3.0: Reset SPMK cache before every connection
  fw-api: CL 31978468 - update fw common interface files
  fw-api: CL 31959202 - update fw common interface files
  fw-api: CL 31945200 - update fw common interface files
  fw-api: CL 31944819 - update fw common interface files
  fw-api: CL 31933928 - update fw common interface files
  fw-api: CL 31932286 - update fw common interface files
  fw-api: CL 31924680 - update fw common interface files
  fw-api: CL 31897484 - update fw common interface files
  fw-api: CL 31882288 - update fw common interface files
  fw-api: CL 31863503 - update fw common interface files
  fw-api: CL 31845524 - update fw common interface files
  fw-api: CL 31837430 - update fw common interface files
  fw-api: CL 31823693 - update fw common interface files
  fw-api: CL 31823686 - update fw common interface files
  fw-api: CL 31740103 - update fw common interface files
  fw-api: CL 31732805 - update fw common interface files
  fw-api: CL 31732804 - update fw common interface files
  disp: msm: dp: handle panel_init for eDP case
  fw-api: CL 31721572 - update fw common interface files
  fw-api: CL 31721565 - update fw common interface files
  fw-api: CL 31710587 - update fw common interface files
  qcacld-3.0: Self rsn cap intersect with AP rsn cap
  fw-api: CL 31670251 - update fw common interface files
  fw-api: CL 31670248 - update fw common interface files
  fw-api: CL 31660856 - update fw common interface files
  fw-api: CL 31636084 - update fw common interface files
  fw-api: CL 31634761 - update fw common interface files
  fw-api: CL 31633543 - update fw common interface files
  fw-api: CL 31632679 - update fw common interface files
  fw-api: CL 31620095 - update fw common interface files
  msm: camera: isp: Fix potential illegal access in Acquire HW. Fix for above
  msm: camera: isp: Fix potential illegal access in Acquire HW
  msm: camera: ois: Copy packet header in kernel
  msm: camera: common: Add missing put_cpu_buf calls
  msm: camera: sensor: handling condition for random read
  msm: camera: icp: io buf config num validation
  msm: camera: ope: check cpu buffer offset and cmd buf idx
  fw-api: CL 31606165 - update fw common interface files
  fw-api: CL 31594124 - update fw common interface files
  fw-api: CL 31593473 - update fw common interface files
  fw-api: CL 31593469 - update fw common interface files
  fw-api: CL 31580442 - update fw common interface files
  fw-api: CL 31580439 - update fw common interface files
  fw-api: CL 31579265 - update fw common interface files
  fw-api: Add PMM spare registers address related header files
  fw-api: CL 31562446 - update fw common interface files
  fw-api: CL 31562430 - update fw common interface files
  fw-api: hw_headers: Add v2 3181 hw headers for fig
  fw-api: CL 31548376 - update fw common interface files
  fw-api: CL 31548353 - update fw common interface files
  fw-api: CL 31532680 - update fw common interface files
  fw-api: CL 31531034 - update fw common interface files
  fw-api: CL 31530723 - update fw common interface files
  fw-api: CL 31528807 - update fw common interface files
  fw-api: CL 31516598 - update fw common interface files
  fw-api: CL 31516597 - update fw common interface files
  fw-api: CL 31516593 - update fw common interface files
  qcacld-3.0: Prevent ping loss due to runtime suspend
  qcacld-3.0: Add runtime pm lock during roaming
  fw-api: CL 31493387 - update fw common interface files
  fw-api: CL 31480148 - update fw common interface files
  fw-api: CL 31472949 - update fw common interface files
  fw-api: CL 31472689 - update fw common interface files
  fw-api: CL 31460336 - update fw common interface files
  fw-api: CL 31459767 - update fw common interface files
  fw-api: CL 31440102 - update fw common interface files
  fw-api: CL 31436976 - update fw common interface files
  fw-api: CL 31424696 - update fw common interface files
  fw-api: CL 31413468 - update fw common interface files
  fw-api: CL 31412194 - update fw common interface files
  fw-api: CL 31404904 - update fw common interface files
  fw-api: CL 31399985 - update fw common interface files
  fw-api: CL 31397472 - update fw common interface files
  fw-api: CL 31390056 - update fw common interface files
  fw-api: CL 31386508 - update fw common interface files
  fw-api: CL 31376775 - update fw common interface files
  qcacld-3.0: Add buffer overflow check in wma_fill_rx_stats function
  fw-api: CL 31362795 - update fw common interface files
  fw-api: CL 31360546 - update fw common interface files
  fw-api: CL 31351405 - update fw common interface files
  fw-api: CL 31348242 - update fw common interface files
  fw-api: CL 31341426 - update fw common interface files
  fw-api: CL 31338231 - update fw common interface files
  fw-api: CL 31328053 - update fw common interface files
  fw-api: CL 31316506 - update fw common interface files
  fw-api: CL 31305348 - update fw common interface files
  fw-api: CL 31295115 - update fw common interface files
  fw-api: CL 31282590 - update fw common interface files
  fw-api: CL 31281768 - update fw common interface files
  fw-api: CL 31274787 - update fw common interface files
  fw-api: CL 31274786 - update fw common interface files
  fw-api: CL 31250105 - update fw common interface files
  fw-api: CL 31244431 - update fw common interface files
  fw-api: CL 31233937 - update fw common interface files
  fw-api: CL 31233932 - update fw common interface files
  fw-api: CL 31221789 - update fw common interface files
  qcacmn: Remove 6G domain check on Maschan list
  qcacmn: Replace ap reg rules with client reg rules
  qcacmn: Add reo_mismatch stats for FISA path
  fw-api: CL 31209257 - update fw common interface files
  fw-api: CL 31209256 - update fw common interface files
  fw-api: CL 31184750 - update fw common interface files
  fw-api: CL 31184747 - update fw common interface files
  qcacld-3.0: Validate fse metadata before aggregation of FISA flow
  fw-api: CL 31181068 - update fw common interface files
  fw-api: hw_headers: Add v2 3175 hw headers for fig
  fw-api: CL 31152020 - update fw common interface files
  fw-api: CL 31139090 - update fw common interface files
  fw-api: CL 31139082 - update fw common interface files
  fw-api: CL 31129784 - update fw common interface files
  fw-api: CL 31127648 - update fw common interface files
  fw-api: CL 31108122 - update fw common interface files
  fw-api: CL 31100347 - update fw common interface files
  fw-api: CL 31060090 - update fw common interface files
  fw-api: CL 31059684 - update fw common interface files
  fw-api: CL 31023126 - update fw common interface files
  fw-api: CL 31021605 - update fw common interface files
  fw-api: CL 31012511 - update fw common interface files
  fw-api: CL 30998406 - update fw common interface files
  qcacld-3.0: Consider intersected AKM for association
  fw-api: fig: v1 remove unused reo hw data structure files
  fw-api: CL 30960339 - update fw common interface files
  fw-api: CL 30949206 - update fw common interface files
  fw-api: CL 30939043 - update fw common interface files
  fw-api: CL 30933603 - update fw common interface files
  fw-api: CL 30924456 - update fw common interface files
  fw-api: CL 30906519 - update fw common interface files
  fw-api: CL 30905405 - update fw common interface files
  fw-api: CL 30894242 - update fw common interface files
  fw-api: CL 30855107 - update fw common interface files
  fw-api: CL 30855102 - update fw common interface files
  fw-api: CL 30839654 - update fw common interface files
  fw-api: CL 30839650 - update fw common interface files
  fw-api: CL 30825844 - update fw common interface files
  fw-api: CL 30811728 - update fw common interface files
  fw-api: CL 30797342 - update fw common interface files
  fw-api: fig: Update config macros for REO_R0_RBM_DESTINATION_RING_CTRL
  fw-api: CL 30786115 - update fw common interface files
  fw-api: CL 30768408 - update fw common interface files
  fw-api: CL 30768407 - update fw common interface files
  fw-api: CL 30757804 - update fw common interface files
  fw-api: E3.0 WCSS_VERSION 3150 fig headers
  fw-api: CL 30735748 - update fw common interface files
  fw-api: CL 30720606 - update fw common interface files
  fw-api: CL 30714212 - update fw common interface files
  fw-api: CL 30697325 - update fw common interface files
  fw-api: CL 30678803 - update fw common interface files
  fw-api: CL 30669089 - update fw common interface files
  fw-api: CL 30669084 - update fw common interface files
  fw-api: CL 30584319 - update fw common interface files
  fw-api: CL 30583992 - update fw common interface files
  fw-api: CL 30580039 - update fw common interface files
  fw-api: CL 30544126 - update fw common interface files
  fw-api: CL 30541340 - update fw common interface files
  fw-api: CL 30538540 - update fw common interface files
  fw-api: CL 30538539 - update fw common interface files
  fw-api: CL 30538536 - update fw common interface files
  fw-api: CL 30519258 - update fw common interface files
  fw-api: CL 30519235 - update fw common interface files
  fw-api: CL 30501923 - update fw common interface files
  fw-api: CL 30443155 - update fw common interface files
  fw-api: CL 30391907 - update fw common interface files
  fw-api: CL 30387120 - update fw common interface files
  fw-api: CL 30358561 - update fw common interface files
  fw-api: CL 30297480 - update fw common interface files
  fw-api: CL 30281577 - update fw common interface files
  fw-api: CL 30273614 - update fw common interface files
  fw-api: CL 30273613 - update fw common interface files
  fw-api: CL 30269010 - update fw common interface files
  fw-api: CL 30243121 - update fw common interface files
  fw-api: CL 30231044 - update fw common interface files
  fw-api: CL 30227097 - update fw common interface files
  fw-api: CL 30201837 - update fw common interface files
  fw-api: CL 30201834 - update fw common interface files
  fw-api: CL 30198227 - update fw common interface files
  fw-api: CL 30152986 - update fw common interface files
  fw-api: CL 30143205 - update fw common interface files
  fw-api: CL 30056543 - update fw common interface files
  fw-api: CL 30026411 - update fw common interface files
  fw-api: CL 30026410 - update fw common interface files
  fw-api: CL 29948685 - update fw common interface files
  fw-api: CL 29930068 - update fw common interface files
  fw-api: CL 29930067 - update fw common interface files
  fw-api: CL 29873461 - update fw common interface files
  fw-api: CL 29858563 - update fw common interface files
  fw-api: CL 29858560 - update fw common interface files
  fw-api: CL 29845167 - update fw common interface files
  fw-api: CL 29832217 - update fw common interface files
  fw-api: CL 29804280 - update fw common interface files
  fw-api: CL 29789406 - update fw common interface files
  msm: camera: sensor: TOCTOU error handling

Change-Id: I93a709ae7fd15531830ebd257a9aaa98cac1dde8
2026-08-11 03:25:50 +03:00
Davidlohr Bueso
2ec0afec5c
BACKPORT: locking/rtmutex: Skip remove_waiter() when waiter is not enqueued
syzbot triggered the following splat in remove_waiter() via
FUTEX_CMP_REQUEUE_PI:

  KASAN: null-ptr-deref in range [0x0000000000000a88-0x0000000000000a8f]
   class_raw_spinlock_constructor
   remove_waiter+0x159/0x1200 kernel/locking/rtmutex.c:1561
   rt_mutex_start_proxy_lock+0x103/0x120
   futex_requeue+0x10e4/0x20d0
   __x64_sys_futex+0x34f/0x4d0

task_blocks_on_rt_mutex() does not arm the waiter upon deadlock detection,
leaving waiter->task nil, where 3bfdc63936dd ("rtmutex: Use waiter::task instead
of current in remove_waiter()") made this fatal.

Furthermore, rt_mutex_start_proxy_lock() should not be calling into remove_waiter()
upon a successfully grabbing the rtmutex. 1a1fb985f2 ("futex: Handle early deadlock
return correctly"), moved the remove_waiter() out of __rt_mutex_start_proxy_lock()
(where 'ret' was only ever 0 or < 0) into the wrapper. Tighten this check to
account for try_to_take_rt_mutex().

Fixes: 3bfdc63936dd ("rtmutex: Use waiter::task instead of current in remove_waiter()")
Change-Id: Ib19a2f223ba67d20496c0ae95096747f5350c520
Reported-by: syzbot+78147abe6c524f183ee9@syzkaller.appspotmail.com
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: stable@vger.kernel.org
Closes: https://lore.kernel.org/all/69f114ac.050a0220.ac8b.0003.GAE@google.com/
Link: https://patch.msgid.link/20260507112913.1019537-1-dave@stgolabs.net
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-07-29 03:30:48 +06:00
Keenan Dong
6efbfed641
BACKPORT: rtmutex: Use waiter::task instead of current in remove_waiter()
remove_waiter() is used by the slowlock paths, but it is also used for
proxy-lock rollback in rt_mutex_start_proxy_lock() when invoked from
futex_requeue().

In the latter case waiter::task is not current, but remove_waiter()
operates on current for the dequeue operation. That results in several
problems:

  1) the rbtree dequeue happens without waiter::task::pi_lock being held

  2) the waiter task's pi_blocked_on state is not cleared, which leaves a
     dangling pointer primed for UAF around.

  3) rt_mutex_adjust_prio_chain() operates on the wrong top priority waiter
     task

Use waiter::task instead of current in all related operations in
remove_waiter() to cure those problems.

[ tglx: Fixup rt_mutex_adjust_prio_chain(), add a comment and amend the
  	changelog ]

Fixes: 8161239a8b ("rtmutex: Simplify PI algorithm and make highest prio task get lock")
Change-Id: I3ff8da2830773e04f55828b11c9d461ab6ee57c5
Reported-by: Yuan Tan <yuantan098@gmail.com>
Reported-by: Yifan Wu <yifanwucs@gmail.com>
Reported-by: Juefei Pu <tomapufckgml@gmail.com>
Reported-by: Xin Liu <bird@lzu.edu.cn>
Signed-off-by: Keenan Dong <keenanat2000@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: stable@vger.kernel.org
[Tashar02: Open-code scoped_guard() on msm-5.4]
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-07-29 03:30:25 +06:00
Michael Bestas
5048089773
Merge remote-tracking branch 'sm8350/lineage-20' into lineage-23.2
* sm8350/lineage-20:
  UPSTREAM: tty: allow TIOCSLCKTRMIOS with CAP_CHECKPOINT_RESTORE
  UPSTREAM: selftests: add clone3() CAP_CHECKPOINT_RESTORE test
  UPSTREAM: prctl: exe link permission error changed from -EINVAL to -EPERM
  UPSTREAM: prctl: Allow local CAP_CHECKPOINT_RESTORE to change /proc/self/exe
  UPSTREAM: proc: allow access in init userns for map_files with CAP_CHECKPOINT_RESTORE
  UPSTREAM: pid_namespace: use checkpoint_restore_ns_capable() for ns_last_pid
  UPSTREAM: pid: use checkpoint_restore_ns_capable() for set_tid
  UPSTREAM: capabilities: Introduce CAP_CHECKPOINT_RESTORE
  UPSTREAM: selftests: add tests for clone3() with *set_tid
  UPSTREAM: fork: extend clone3() to support setting a PID
  UPSTREAM: selftests: add tests for clone3()
  UPSTREAM: tests: test CLONE_CLEAR_SIGHAND
  UPSTREAM: clone3: add CLONE_CLEAR_SIGHAND

Change-Id: I7890fb550e85b9fd4c0c3a27144d5bfd8963d496
2026-07-25 17:05:51 +03:00
Nicolas Viennot
6b4aab010e UPSTREAM: prctl: exe link permission error changed from -EINVAL to -EPERM
This brings consistency with the rest of the prctl() syscall where
-EPERM is returned when failing a capability check.

Change-Id: I1b27a485d4ca9dd99376c6628572509835aabaf3
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Signed-off-by: Adrian Reber <areber@redhat.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-7-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:39 -04:00
Nicolas Viennot
c90ed82a1d UPSTREAM: prctl: Allow local CAP_CHECKPOINT_RESTORE to change /proc/self/exe
Originally, only a local CAP_SYS_ADMIN could change the exe link,
making it difficult for doing checkpoint/restore without CAP_SYS_ADMIN.
This commit adds CAP_CHECKPOINT_RESTORE in addition to CAP_SYS_ADMIN
for permitting changing the exe link.

The following describes the history of the /proc/self/exe permission
checks as it may be difficult to understand what decisions lead to this
point.

* [1] May 2012: This commit introduces the ability of changing
  /proc/self/exe if the user is CAP_SYS_RESOURCE capable.
  In the related discussion [2], no clear thread model is presented for
  what could happen if the /proc/self/exe changes multiple times, or why
  would the admin be at the mercy of userspace.

* [3] Oct 2014: This commit introduces a new API to change
  /proc/self/exe. The permission no longer checks for CAP_SYS_RESOURCE,
  but instead checks if the current user is root (uid=0) in its local
  namespace. In the related discussion [4] it is said that "Controlling
  exe_fd without privileges may turn out to be dangerous. At least
  things like tomoyo examine it for making policy decisions (see
  tomoyo_manager())."

* [5] Dec 2016: This commit removes the restriction to change
  /proc/self/exe at most once. The related discussion [6] informs that
  the audit subsystem relies on the exe symlink, presumably
  audit_log_d_path_exe() in kernel/audit.c.

* [7] May 2017: This commit changed the check from uid==0 to local
  CAP_SYS_ADMIN. No discussion.

* [8] July 2020: A PoC to spoof any program's /proc/self/exe via ptrace
  is demonstrated

Overall, the concrete points that were made to retain capability checks
around changing the exe symlink is that tomoyo_manager() and
audit_log_d_path_exe() uses the exe_file path.

Christian Brauner said that relying on /proc/<pid>/exe being immutable (or
guarded by caps) in a sake of security is a bit misleading. It can only
be used as a hint without any guarantees of what code is being executed
once execve() returns to userspace. Christian suggested that in the
future, we could call audit_log() or similar to inform the admin of all
exe link changes, instead of attempting to provide security guarantees
via permission checks. However, this proposed change requires the
understanding of the security implications in the tomoyo/audit subsystems.

[1] b32dfe3771 ("c/r: prctl: add ability to set new mm_struct::exe_file")
[2] https://lore.kernel.org/patchwork/patch/292515/
[3] f606b77f1a ("prctl: PR_SET_MM -- introduce PR_SET_MM_MAP operation")
[4] https://lore.kernel.org/patchwork/patch/479359/
[5] 3fb4afd9a5 ("prctl: remove one-shot limitation for changing exe link")
[6] https://lore.kernel.org/patchwork/patch/697304/
[7] 4d28df6152 ("prctl: Allow local CAP_SYS_ADMIN changing exe_file")
[8] https://github.com/nviennot/run_as_exe

Change-Id: I44b78ea8000443bc862178aae3f847e86785a1a1
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Signed-off-by: Adrian Reber <areber@redhat.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-6-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:39 -04:00
Adrian Reber
63b4832326 UPSTREAM: pid_namespace: use checkpoint_restore_ns_capable() for ns_last_pid
Use the newly introduced capability CAP_CHECKPOINT_RESTORE to allow
writing to ns_last_pid.

Change-Id: I95727c385f136818553260db0e9edfa654f7907e
Signed-off-by: Adrian Reber <areber@redhat.com>
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-4-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
5c348f563d UPSTREAM: pid: use checkpoint_restore_ns_capable() for set_tid
Use the newly introduced capability CAP_CHECKPOINT_RESTORE to allow
using clone3() with set_tid set.

Change-Id: Ifc86c455e500e59a62f9960ba094c4f9973891ff
Signed-off-by: Adrian Reber <areber@redhat.com>
Signed-off-by: Nicolas Viennot <Nicolas.Viennot@twosigma.com>
Reviewed-by: Serge Hallyn <serge@hallyn.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-3-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Adrian Reber
62ec29d7ad UPSTREAM: fork: extend clone3() to support setting a PID
The main motivation to add set_tid to clone3() is CRIU.

To restore a process with the same PID/TID CRIU currently uses
/proc/sys/kernel/ns_last_pid. It writes the desired (PID - 1) to
ns_last_pid and then (quickly) does a clone(). This works most of the
time, but it is racy. It is also slow as it requires multiple syscalls.

Extending clone3() to support *set_tid makes it possible restore a
process using CRIU without accessing /proc/sys/kernel/ns_last_pid and
race free (as long as the desired PID/TID is available).

This clone3() extension places the same restrictions (CAP_SYS_ADMIN)
on clone3() with *set_tid as they are currently in place for ns_last_pid.

The original version of this change was using a single value for
set_tid. At the 2019 LPC, after presenting set_tid, it was, however,
decided to change set_tid to an array to enable setting the PID of a
process in multiple PID namespaces at the same time. If a process is
created in a PID namespace it is possible to influence the PID inside
and outside of the PID namespace. Details also in the corresponding
selftest.

To create a process with the following PIDs:

      PID NS level         Requested PID
        0 (host)              31496
        1                        42
        2                         1

For that example the two newly introduced parameters to struct
clone_args (set_tid and set_tid_size) would need to be:

  set_tid[0] = 1;
  set_tid[1] = 42;
  set_tid[2] = 31496;
  set_tid_size = 3;

If only the PIDs of the two innermost nested PID namespaces should be
defined it would look like this:

  set_tid[0] = 1;
  set_tid[1] = 42;
  set_tid_size = 2;

The PID of the newly created process would then be the next available
free PID in the PID namespace level 0 (host) and 42 in the PID namespace
at level 1 and the PID of the process in the innermost PID namespace
would be 1.

The set_tid array is used to specify the PID of a process starting
from the innermost nested PID namespaces up to set_tid_size PID namespaces.

set_tid_size cannot be larger then the current PID namespace level.

Change-Id: I21a8a13c5a794b2a788f3e647fd639b5ae313ebf
Signed-off-by: Adrian Reber <areber@redhat.com>
Reviewed-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: Dmitry Safonov <0x7f454c46@gmail.com>
Acked-by: Andrei Vagin <avagin@gmail.com>
Link: https://lore.kernel.org/r/20191115123621.142252-1-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Christian Brauner
67da5e102b UPSTREAM: clone3: add CLONE_CLEAR_SIGHAND
Reset all signal handlers of the child not set to SIG_IGN to SIG_DFL.
Mutually exclusive with CLONE_SIGHAND to not disturb other thread's
signal handler.

In the spirit of closer cooperation between glibc developers and kernel
developers (cf. [2]) this patchset came out of a discussion on the glibc
mailing list for improving posix_spawn() (cf. [1], [3], [4]). Kernel
support for this feature has been explicitly requested by glibc and I
see no reason not to help them with this.

The child helper process on Linux posix_spawn must ensure that no signal
handlers are enabled, so the signal disposition must be either SIG_DFL
or SIG_IGN. However, it requires a sigprocmask to obtain the current
signal mask and at least _NSIG sigaction calls to reset the signal
handlers for each posix_spawn call or complex state tracking that might
lead to data corruption in glibc. Adding this flags lets glibc avoid
these problems.

[1]: https://www.sourceware.org/ml/libc-alpha/2019-10/msg00149.html
[3]: https://www.sourceware.org/ml/libc-alpha/2019-10/msg00158.html
[4]: https://www.sourceware.org/ml/libc-alpha/2019-10/msg00160.html
[2]: https://lwn.net/Articles/799331/
     '[...] by asking for better cooperation with the C-library projects
     in general. They should be copied on patches containing ABI
     changes, for example. I noted that there are often times where
     C-library developers wish the kernel community had done things
     differently; how could those be avoided in the future? Members of
     the audience suggested that more glibc developers should perhaps
     join the linux-api list. The other suggestion was to "copy Florian
     on everything".'
Cc: Florian Weimer <fweimer@redhat.com>
Cc: libc-alpha@sourceware.org
Cc: linux-api@vger.kernel.org
Change-Id: Iced02ef93ed4dcda5ca5d3bd475cbfcae928fb19
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lore.kernel.org/r/20191014104538.3096-1-christian.brauner@ubuntu.com
Signed-off-by: ralph950412 <ralph950412@gmail.com>
2026-07-09 09:06:38 -04:00
Michael Bestas
e14e0696e2
Merge remote-tracking branch 'sm8350/lineage-20' into lineage-23.2
* sm8350/lineage-20: (148 commits)
  ANDROID: Add FUSE_BPF to gki_defconfig
  Reapply "UPSTREAM: seccomp: Remove bogus __user annotations"
  ANDROID: fuse: Open-code vma_set_file() logic in fuse_backing_mmap()
  BACKPORT: fuse: fix livelock in synchronous file put from fuseblk workers
  BACKPORT: compat_ioctl: move more drivers to compat_ptr_ioctl
  UPSTREAM: fuse: make sure reclaim doesn't write the inode
  UPSTREAM: ANDROID: fuse-bpf: Correct fuse bpf feature flag
  UPSTREAM: fuse: verify {g,u}id mount options correctly
  fixup! BACKPORT: fuse: name fs_context consistently
  BACKPORT: fuse: name fs_context consistently
  UPSTREAM: fuse: fix root lookup with nonzero generation
  UPSTREAM: ANDROID: fuse-bpf: Fix the issue of abnormal lseek system calls
  UPSTREAM: ANDROID: fuse-bpf: Follow mounts in lookups
  UPSTREAM: fuse: dax: set fc->dax to NULL in fuse_dax_conn_free()
  BACKPORT: ANDROID: fuse-bpf: Ignore readaheads unless they go to the daemon
  UPSTREAM: ANDROID: fuse-bpf: Add NULL pointer check in fuse_release_in
  UPSTREAM: ANDROID: fs/passthrough: Fix compatibility with R/O file system
  UPSTREAM: ANDROID: fuse-bpf: Add NULL pointer check in fuse_entry_revalidate
  UPSTREAM: ANDROID: fuse-bpf: Get correct inode in mkdir
  UPSTREAM: ANDROID: fuse-bpf: Use stored bpf for create_open
  ...

Change-Id: Ice6b8c09d19db311c827d0a9f48d46795be998ee
2026-05-14 13:35:30 +02:00
Alexander Martinz
63de6916f0
Reapply "UPSTREAM: seccomp: Remove bogus __user annotations"
This reverts commit cb8dc8a108.

As pointed out on gerrit[1]:
> This revert is wrong. `android12-5.4` doesn't have
> "sysctl: pass kernel pointers to ->proc_handler" but this kernel does.

[1] - https://review.lineageos.org/c/LineageOS/android_kernel_qcom_sm8350/+/482457

Change-Id: I72cd996f6739a67cf78141843e2d933f19e2e92b
Signed-off-by: Alexander Martinz <amartinz@shiftphones.com>
2026-05-12 12:10:45 +02:00
Daniel Rosenberg
29efcdd0db
BACKPORT: ANDROID: fuse-bpf: Use fuse_bpf_args in uapi
fuse_args is not suitable for use in the uapi - it is not stable, and
contains internal pointers. Replace with stable equivalent.

The end_offset values are currently unused and unset, but will be used
in a follow up patch by the verifier.

Test: fuse_test, atest ScopedStorageDeviceTest pass
Bug: 202785178
Signed-off-by: Daniel Rosenberg <drosen@google.com>
Change-Id: Ic1c12f9706aeae233cc30a0d68ed2533030e485b
2026-05-08 21:46:17 +02:00
Daniel Rosenberg
b3fbcd3ca5
BACKPORT: ANDROID: fuse-bpf v1
Bug: 202785178
Test: test_fuse passes on linux, feature works on cuttlefish
Signed-off-by: Paul Lawrence <paullawrence@google.com>
Signed-off-by: Daniel Rosenberg <drosen@google.com>
Change-Id: I987684b799b07391ccde350e98fde7976f5601aa
2026-05-08 21:46:14 +02:00
Michael Bestas
376925d57f
Merge remote-tracking branch 'sm8350/lineage-20' into lineage-23.2
* sm8350/lineage-20: (306 commits)
  ANDROID: gki_defconfig: reduce KFENCE pool size
  ANDROID: GKI: enable KFENCE by setting the sample interval to 500ms
  ANDROID: GKI: Enable KFENCE
  UPSTREAM: random: split initialization into early step and later step
  UPSTREAM: kfence: avoid passing -g for test
  UPSTREAM: net: add and use skb_unclone_keeptruesize() helper
  UPSTREAM: kfence: fix memory leak when cat kfence objects
  UPSTREAM: kfence: unconditionally use unbound work queue
  UPSTREAM: kfence: use TASK_IDLE when awaiting allocation
  UPSTREAM: kfence: fix is_kfence_address() for addresses below KFENCE_POOL_SIZE
  FROMLIST: kfence: skip all GFP_ZONEMASK allocations
  FROMLIST: kfence: move the size check to the beginning of __kfence_alloc()
  ANDROID: kasan: fix interoperability with KFENCE
  UPSTREAM: arm64: mm: don't use CON and BLK mapping if KFENCE is enabled
  ANDROID: kfence: clean up unused variables
  FROMGIT: kfence: use power-efficient work queue to run delayed work
  FROMGIT: kfence: maximize allocation wait timeout duration
  FROMGIT: kfence: await for allocation using wait_event
  FROMGIT: kfence: zero guard page after out-of-bounds access
  UPSTREAM: kfence: make compatible with kmemleak
  ...

Change-Id: Id1436b77692839dc2394f500bba42487fac1b2b4
2026-05-07 18:19:43 +03:00
Alexander Potapenko
741d9ecacc
FROMGIT: tracing: add error_report_end trace point
Patch series "Add error_report_end tracepoint to KFENCE and KASAN", v3.

This patchset adds a tracepoint, error_repor_end, that is to be used by
KFENCE, KASAN, and potentially other bug detection tools, when they print
an error report.  One of the possible use cases is userspace collection of
kernel error reports: interested parties can subscribe to the tracing
event via tracefs, and get notified when an error report occurs.

This patch (of 3):

Introduce error_report_end tracepoint.  It can be used in debugging tools
like KASAN, KFENCE, etc.  to provide extensions to the error reporting
mechanisms (e.g.  allow tests hook into error reporting, ease error report
collection from production kernels).  Another benefit would be making use
of ftrace for debugging or benchmarking the tools themselves.

Should we need it, the tracepoint name leaves us with the possibility to
introduce a complementary error_report_start tracepoint in the future.

Link: https://lkml.kernel.org/r/20210121131915.1331302-1-glider@google.com
Link: https://lkml.kernel.org/r/20210121131915.1331302-2-glider@google.com
Signed-off-by: Alexander Potapenko <glider@google.com>
Suggested-by: Marco Elver <elver@google.com>
Cc: Andrey Konovalov <andreyknvl@google.com>
Cc: Dmitry Vyukov <dvyukov@google.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Petr Mladek <pmladek@suse.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Sergey Senozhatsky <sergey.senozhatsky@gmail.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Vlastimil Babka <vbabka@suse.cz>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>

Bug: 177201466
(cherry picked from commit ba7612c00686f204f7bca4ceb7394a9e705e84bd
    https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10)
Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish
Signed-off-by: Alexander Potapenko <glider@google.com>
Change-Id: Ic86e29982c04dad4b3b7889a424f37b22cc5f22b
2026-05-07 17:24:05 +03:00
Jeff Vander Stoep
cb8dc8a108 Revert "UPSTREAM: seccomp: Remove bogus __user annotations"
This reverts commit 5444477e8a4d31f6e6ff720c2d018d06e405bcc1.
Bug: 176068146
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Change-Id: Ic35b23f2f3ad99093b7df5e82633bba90acbe82a
2026-05-07 10:17:18 -04:00
Hsuan-Chi Kuo
c83d6f169d UPSTREAM: seccomp: Fix setting loaded filter count during TSYNC
The desired behavior is to set the caller's filter count to thread's.
This value is reported via /proc, so this fixes the inaccurate count
exposed to userspace; it is not used for reference counting, etc.

Signed-off-by: Hsuan-Chi Kuo <hsuanchikuo@gmail.com>
Link: https://lore.kernel.org/r/20210304233708.420597-1-hsuanchikuo@gmail.com
Co-developed-by: Wiktor Garbacz <wiktorg@google.com>
Signed-off-by: Wiktor Garbacz <wiktorg@google.com>
Link: https://lore.kernel.org/lkml/20210810125158.329849-1-wiktorg@google.com
Signed-off-by: Kees Cook <keescook@chromium.org>
Cc: stable@vger.kernel.org
Fixes: c818c03b661c ("seccomp: Report number of loaded filters in /proc/$pid/status")
(cherry picked from commit b4d8a58f8dcfcc890f296696cadb76e77be44b5f)
Bug: 187129171
Signed-off-by: Connor O'Brien <connoro@google.com>
Change-Id: Ia3ee7ec71e9fdbb8d958f9b42b1c6e02c761503f
2026-05-07 10:17:18 -04:00
YiFei Zhu
c639e51eef UPSTREAM: seccomp/cache: Add "emulator" to check if filter is constant allow
SECCOMP_CACHE will only operate on syscalls that do not access
any syscall arguments or instruction pointer. To facilitate
this we need a static analyser to know whether a filter will
return allow regardless of syscall arguments for a given
architecture number / syscall number pair. This is implemented
here with a pseudo-emulator, and stored in a per-filter bitmap.

In order to build this bitmap at filter attach time, each filter is
emulated for every syscall (under each possible architecture), and
checked for any accesses of struct seccomp_data that are not the "arch"
nor "nr" (syscall) members. If only "arch" and "nr" are examined, and
the program returns allow, then we can be sure that the filter must
return allow independent from syscall arguments.

Nearly all seccomp filters are built from these cBPF instructions:

BPF_LD  | BPF_W    | BPF_ABS
BPF_JMP | BPF_JEQ  | BPF_K
BPF_JMP | BPF_JGE  | BPF_K
BPF_JMP | BPF_JGT  | BPF_K
BPF_JMP | BPF_JSET | BPF_K
BPF_JMP | BPF_JA
BPF_RET | BPF_K
BPF_ALU | BPF_AND  | BPF_K

Each of these instructions are emulated. Any weirdness or loading
from a syscall argument will cause the emulator to bail.

The emulation is also halted if it reaches a return. In that case,
if it returns an SECCOMP_RET_ALLOW, the syscall is marked as good.

Emulator structure and comments are from Kees [1] and Jann [2].

Emulation is done at attach time. If a filter depends on more
filters, and if the dependee does not guarantee to allow the
syscall, then we skip the emulation of this syscall.

[1] https://lore.kernel.org/lkml/20200923232923.3142503-5-keescook@chromium.org/
[2] https://lore.kernel.org/lkml/CAG48ez1p=dR_2ikKq=xVxkoGg0fYpTBpkhJSv1w-6BG=76PAvw@mail.gmail.com/

Suggested-by: Jann Horn <jannh@google.com>
Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu>
Reviewed-by: Jann Horn <jannh@google.com>
Co-developed-by: Kees Cook <keescook@chromium.org>
Signed-off-by: Kees Cook <keescook@chromium.org>
Link: https://lore.kernel.org/r/71c7be2db5ee08905f41c3be5c1ad6e2601ce88f.1602431034.git.yifeifz2@illinois.edu
(cherry picked from commit 8e01b51a31a1e08e2c3e8fcc0ef6790441be2f61)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Change-Id: I5047f7f0d6502e5de6c047743f1053fda3025a6e
Bug: 176068146
2026-05-07 10:17:17 -04:00
YiFei Zhu
993d91ceea UPSTREAM: seccomp/cache: Lookup syscall allowlist bitmap for fast path
The overhead of running Seccomp filters has been part of some past
discussions [1][2][3]. Oftentimes, the filters have a large number
of instructions that check syscall numbers one by one and jump based
on that. Some users chain BPF filters which further enlarge the
overhead. A recent work [6] comprehensively measures the Seccomp
overhead and shows that the overhead is non-negligible and has a
non-trivial impact on application performance.

We observed some common filters, such as docker's [4] or
systemd's [5], will make most decisions based only on the syscall
numbers, and as past discussions considered, a bitmap where each bit
represents a syscall makes most sense for these filters.

The fast (common) path for seccomp should be that the filter permits
the syscall to pass through, and failing seccomp is expected to be
an exceptional case; it is not expected for userspace to call a
denylisted syscall over and over.

When it can be concluded that an allow must occur for the given
architecture and syscall pair (this determination is introduced in
the next commit), seccomp will immediately allow the syscall,
bypassing further BPF execution.

Each architecture number has its own bitmap. The architecture
number in seccomp_data is checked against the defined architecture
number constant before proceeding to test the bit against the
bitmap with the syscall number as the index of the bit in the
bitmap, and if the bit is set, seccomp returns allow. The bitmaps
are all clear in this patch and will be initialized in the next
commit.

When only one architecture exists, the check against architecture
number is skipped, suggested by Kees Cook [7].

[1] https://lore.kernel.org/linux-security-module/c22a6c3cefc2412cad00ae14c1371711@huawei.com/T/
[2] https://lore.kernel.org/lkml/202005181120.971232B7B@keescook/T/
[3] https://github.com/seccomp/libseccomp/issues/116
[4] ae0ef82b90/profiles/seccomp/default.json
[5] 6743a1caf4/src/shared/seccomp-util.c (L270)
[6] Draco: Architectural and Operating System Support for System Call Security
    https://tianyin.github.io/pub/draco.pdf, MICRO-53, Oct. 2020
[7] https://lore.kernel.org/bpf/202010091614.8BB0EB64@keescook/

Co-developed-by: Dimitrios Skarlatos <dskarlat@cs.cmu.edu>
Signed-off-by: Dimitrios Skarlatos <dskarlat@cs.cmu.edu>
Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu>
Reviewed-by: Jann Horn <jannh@google.com>
Signed-off-by: Kees Cook <keescook@chromium.org>
Link: https://lore.kernel.org/r/10f91a367ec4fcdea7fc3f086de3f5f13a4a7436.1602431034.git.yifeifz2@illinois.edu
(cherry picked from commit f9d480b6ffbeb336bf7f6ce44825c00f61b3abae)A
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Change-Id: I50b6682e17dc6e91b5e92017361200d722282825
Bug: 176068146
2026-05-07 10:17:17 -04:00
Denis Efremov
e24e8cf802 UPSTREAM: seccomp: Use current_pt_regs() instead of task_pt_regs(current)
As described in commit a3460a5974 ("new helper: current_pt_regs()"):
- arch versions are "optimized versions".
- some architectures have task_pt_regs() working only for traced tasks
  blocked on signal delivery. current_pt_regs() needs to work for *all*
  processes.

In preparation for adding a coccinelle rule for using current_*(), instead
of raw accesses to current members, modify seccomp_do_user_notification(),
__seccomp_filter(), __secure_computing() to use current_pt_regs().

Signed-off-by: Denis Efremov <efremov@linux.com>
Link: https://lore.kernel.org/r/20200824125921.488311-1-efremov@linux.com
[kees: Reworded commit log, add comment to populate_seccomp_data()]
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit 2d9ca267a944c2b6ed5b4d750b1cf0407b6262b4)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: Ib66bcc8cfa077c7a493fb9d14501279f32f48550
2026-05-07 10:17:17 -04:00
Rich Felker
c6279c1cb7 UPSTREAM: seccomp: kill process instead of thread for unknown actions
Asynchronous termination of a thread outside of the userspace thread
library's knowledge is an unsafe operation that leaves the process in
an inconsistent, corrupt, and possibly unrecoverable state. In order
to make new actions that may be added in the future safe on kernels
not aware of them, change the default action from
SECCOMP_RET_KILL_THREAD to SECCOMP_RET_KILL_PROCESS.

Signed-off-by: Rich Felker <dalias@libc.org>
Link: https://lore.kernel.org/r/20200829015609.GA32566@brightrain.aerifal.cx
[kees: Fixed up coredump selection logic to match]
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit 4d671d922d51907bc41f1f7f2dc737c928ae78fd)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I23140e1efbb4346de8566421c35d2810de26b209
2026-05-07 10:17:17 -04:00
Tycho Andersen
43c11dbfa8 UPSTREAM: seccomp: don't leave dangling ->notif if file allocation fails
Christian and Kees both pointed out that this is a bit sloppy to open-code
both places, and Christian points out that we leave a dangling pointer to
->notif if file allocation fails. Since we check ->notif for null in order
to determine if it's ok to install a filter, this means people won't be
able to install a filter if the file allocation fails for some reason, even
if they subsequently should be able to.

To fix this, let's hoist this free+null into its own little helper and use
it.

Reported-by: Kees Cook <keescook@chromium.org>
Reported-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: Tycho Andersen <tycho@tycho.pizza>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200902140953.1201956-1-tycho@tycho.pizza
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit e839317900e9f13c83d8711d684de88c625b307a)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: Ie20e79fa15b6891895b7ada1f6ef7d08fdf81e01
2026-05-07 10:17:17 -04:00
Tycho Andersen
0506d711a3 UPSTREAM: seccomp: don't leak memory when filter install races
In seccomp_set_mode_filter() with TSYNC | NEW_LISTENER, we first initialize
the listener fd, then check to see if we can actually use it later in
seccomp_may_assign_mode(), which can fail if anyone else in our thread
group has installed a filter and caused some divergence. If we can't, we
partially clean up the newly allocated file: we put the fd, put the file,
but don't actually clean up the *memory* that was allocated at
filter->notif. Let's clean that up too.

To accomplish this, let's hoist the actual "detach a notifier from a
filter" code to its own helper out of seccomp_notify_release(), so that in
case anyone adds stuff to init_listener(), they only have to add the
cleanup code in one spot. This does a bit of extra locking and such on the
failure path when the filter is not attached, but it's a slow failure path
anyway.

Fixes: 51891498f2da ("seccomp: allow TSYNC and USER_NOTIF together")
Reported-by: syzbot+3ad9614a12f80994c32e@syzkaller.appspotmail.com
Signed-off-by: Tycho Andersen <tycho@tycho.pizza>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200902014017.934315-1-tycho@tycho.pizza
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit a566a9012acd7c9a4be7e30dc7acb7a811ec2260)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I8e6ce0f1646ff997623458f25b733ba79c0a47a2
2026-05-07 10:17:17 -04:00
Christian Brauner
bcce8defc3 UPSTREAM: seccomp: add SECCOMP_USER_NOTIF_FLAG_CONTINUE
This allows the seccomp notifier to continue a syscall. A positive
discussion about this feature was triggered by a post to the
ksummit-discuss mailing list (cf. [3]) and took place during KSummit
(cf. [1]) and again at the containers/checkpoint-restore
micro-conference at Linux Plumbers.

Recently we landed seccomp support for SECCOMP_RET_USER_NOTIF (cf. [4])
which enables a process (watchee) to retrieve an fd for its seccomp
filter. This fd can then be handed to another (usually more privileged)
process (watcher). The watcher will then be able to receive seccomp
messages about the syscalls having been performed by the watchee.

This feature is heavily used in some userspace workloads. For example,
it is currently used to intercept mknod() syscalls in user namespaces
aka in containers.
The mknod() syscall can be easily filtered based on dev_t. This allows
us to only intercept a very specific subset of mknod() syscalls.
Furthermore, mknod() is not possible in user namespaces toto coelo and
so intercepting and denying syscalls that are not in the whitelist on
accident is not a big deal. The watchee won't notice a difference.

In contrast to mknod(), a lot of other syscall we intercept (e.g.
setxattr()) cannot be easily filtered like mknod() because they have
pointer arguments. Additionally, some of them might actually succeed in
user namespaces (e.g. setxattr() for all "user.*" xattrs). Since we
currently cannot tell seccomp to continue from a user notifier we are
stuck with performing all of the syscalls in lieu of the container. This
is a huge security liability since it is extremely difficult to
correctly assume all of the necessary privileges of the calling task
such that the syscall can be successfully emulated without escaping
other additional security restrictions (think missing CAP_MKNOD for
mknod(), or MS_NODEV on a filesystem etc.). This can be solved by
telling seccomp to resume the syscall.

One thing that came up in the discussion was the problem that another
thread could change the memory after userspace has decided to let the
syscall continue which is a well known TOCTOU with seccomp which is
present in other ways already.
The discussion showed that this feature is already very useful for any
syscall without pointer arguments. For any accidentally intercepted
non-pointer syscall it is safe to continue.
For syscalls with pointer arguments there is a race but for any cautious
userspace and the main usec cases the race doesn't matter. The notifier
is intended to be used in a scenario where a more privileged watcher
supervises the syscalls of lesser privileged watchee to allow it to get
around kernel-enforced limitations by performing the syscall for it
whenever deemed save by the watcher. Hence, if a user tricks the watcher
into allowing a syscall they will either get a deny based on
kernel-enforced restrictions later or they will have changed the
arguments in such a way that they manage to perform a syscall with
arguments that they would've been allowed to do anyway.
In general, it is good to point out again, that the notifier fd was not
intended to allow userspace to implement a security policy but rather to
work around kernel security mechanisms in cases where the watcher knows
that a given action is safe to perform.

/* References */
[1]: https://linuxplumbersconf.org/event/4/contributions/560
[2]: https://linuxplumbersconf.org/event/4/contributions/477
[3]: https://lore.kernel.org/r/20190719093538.dhyopljyr5ns33qx@brauner.io
[4]: commit 6a21cc50f0 ("seccomp: add a return code to trap to userspace")

Co-developed-by: Kees Cook <keescook@chromium.org>
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Tycho Andersen <tycho@tycho.ws>
Cc: Andy Lutomirski <luto@amacapital.net>
Cc: Will Drewry <wad@chromium.org>
CC: Tyler Hicks <tyhicks@canonical.com>
Link: https://lore.kernel.org/r/20190920083007.11475-2-christian.brauner@ubuntu.com
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit fb3c5386b382d4097476ce9647260fc89b34afdb)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: Ifd5de971a0da6a507cb8ca1178381ca715693e07
2026-05-07 10:17:17 -04:00
Kees Cook
2e686a7319 UPSTREAM: seccomp: Use -1 marker for end of mode 1 syscall list
The terminator for the mode 1 syscalls list was a 0, but that could be
a valid syscall number (e.g. x86_64 __NR_read). By luck, __NR_read was
listed first and the loop construct would not test it, so there was no
bug. However, this is fragile. Replace the terminator with -1 instead,
and make the variable name for mode 1 syscall lists more descriptive.

Cc: Andy Lutomirski <luto@amacapital.net>
Cc: Will Drewry <wad@chromium.org>
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit fe4bfff86ec54773df3db79e8112e3b0f820c799)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I3d91bf57236f2c20d71f22fa9c0ef7b0b8869bcf
2026-05-07 10:17:17 -04:00
Kees Cook
d3720689a9 UPSTREAM: seccomp: Use pr_fmt
Avoid open-coding "seccomp: " prefixes for pr_*() calls.

Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit e68f9d49dda1744d548426c8b4335a8d693a36d0)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I2fa08d91c11b59dbb962cc767aa4cd25a42f761d
2026-05-07 10:17:17 -04:00
Christian Brauner
246a019c84 UPSTREAM: seccomp: notify about unused filter
We've been making heavy use of the seccomp notifier to intercept and
handle certain syscalls for containers. This patch allows a syscall
supervisor listening on a given notifier to be notified when a seccomp
filter has become unused.

A container is often managed by a singleton supervisor process the
so-called "monitor". This monitor process has an event loop which has
various event handlers registered. If the user specified a seccomp
profile that included a notifier for various syscalls then we also
register a seccomp notify even handler. For any container using a
separate pid namespace the lifecycle of the seccomp notifier is bound to
the init process of the pid namespace, i.e. when the init process exits
the filter must be unused.

If a new process attaches to a container we force it to assume a seccomp
profile. This can either be the same seccomp profile as the container
was started with or a modified one. If the attaching process makes use
of the seccomp notifier we will register a new seccomp notifier handler
in the monitor's event loop. However, when the attaching process exits
we can't simply delete the handler since other child processes could've
been created (daemons spawned etc.) that have inherited the seccomp
filter and so we need to keep the seccomp notifier fd alive in the event
loop. But this is problematic since we don't get a notification when the
seccomp filter has become unused and so we currently never remove the
seccomp notifier fd from the event loop and just keep accumulating fds
in the event loop. We've had this issue for a while but it has recently
become more pressing as more and larger users make use of this.

To fix this, we introduce a new "users" reference counter that tracks any
tasks and dependent filters making use of a filter. When a notifier is
registered waiting tasks will be notified that the filter is now empty
by receiving a (E)POLLHUP event.

The concept in this patch introduces is the same as for signal_struct,
i.e. reference counting for life-cycle management is decoupled from
reference counting taks using the object. There's probably some trickery
possible but the second counter is just the correct way of doing this
IMHO and has precedence.

Cc: Tycho Andersen <tycho@tycho.ws>
Cc: Kees Cook <keescook@chromium.org>
Cc: Matt Denton <mpdenton@google.com>
Cc: Sargun Dhillon <sargun@sargun.me>
Cc: Jann Horn <jannh@google.com>
Cc: Chris Palmer <palmer@google.com>
Cc: Aleksa Sarai <cyphar@cyphar.com>
Cc: Robert Sesek <rsesek@google.com>
Cc: Jeffrey Vander Stoep <jeffv@google.com>
Cc: Linux Containers <containers@lists.linux-foundation.org>
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200531115031.391515-3-christian.brauner@ubuntu.com
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit 99cdb8b9a57393b5978e7a6310a2cba511dd179b)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I1501f5134bf738fa641c8b35ed65a4bed01e6fef
2026-05-07 10:17:17 -04:00
Christian Brauner
902f009956 UPSTREAM: seccomp: Lift wait_queue into struct seccomp_filter
Lift the wait_queue from struct notification into struct seccomp_filter.
This is cleaner overall and lets us avoid having to take the notifier
mutex in the future for EPOLLHUP notifications since we need to neither
read nor modify the notifier specific aspects of the seccomp filter. In
the exit path I'd very much like to avoid having to take the notifier mutex
for each filter in the task's filter hierarchy.

Cc: Tycho Andersen <tycho@tycho.ws>
Cc: Kees Cook <keescook@chromium.org>
Cc: Matt Denton <mpdenton@google.com>
Cc: Sargun Dhillon <sargun@sargun.me>
Cc: Jann Horn <jannh@google.com>
Cc: Chris Palmer <palmer@google.com>
Cc: Aleksa Sarai <cyphar@cyphar.com>
Cc: Robert Sesek <rsesek@google.com>
Cc: Jeffrey Vander Stoep <jeffv@google.com>
Cc: Linux Containers <containers@lists.linux-foundation.org>
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit 76194c4e830d570d9e369d637bb907591d2b3111)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: Id87b359f1155d26068cd62d5295dc39849d46495
2026-05-07 10:17:17 -04:00
Christian Brauner
31e7b9c9d8 BACKPORT: seccomp: release filter after task is fully dead
The seccomp filter used to be released in free_task() which is called
asynchronously via call_rcu() and assorted mechanisms. Since we need
to inform tasks waiting on the seccomp notifier when a filter goes empty
we will notify them as soon as a task has been marked fully dead in
release_task(). To not split seccomp cleanup into two parts, move
filter release out of free_task() and into release_task() after we've
unhashed struct task from struct pid, exited signals, and unlinked it
from the threadgroups' thread list. We'll put the empty filter
notification infrastructure into it in a follow up patch.

This also renames put_seccomp_filter() to seccomp_filter_release() which
is a more descriptive name of what we're doing here especially once
we've added the empty filter notification mechanism in there.

We're also NULL-ing the task's filter tree entrypoint which seems
cleaner than leaving a dangling pointer in there. Note that this shouldn't
need any memory barriers since we're calling this when the task is in
release_task() which means it's EXIT_DEAD. So it can't modify its seccomp
filters anymore. You can also see this from the point where we're calling
seccomp_filter_release(). It's after __exit_signal() and at this point,
tsk->sighand will already have been NULLed which is required for
thread-sync and filter installation alike.

Cc: Tycho Andersen <tycho@tycho.ws>
Cc: Kees Cook <keescook@chromium.org>
Cc: Matt Denton <mpdenton@google.com>
Cc: Sargun Dhillon <sargun@sargun.me>
Cc: Jann Horn <jannh@google.com>
Cc: Chris Palmer <palmer@google.com>
Cc: Aleksa Sarai <cyphar@cyphar.com>
Cc: Robert Sesek <rsesek@google.com>
Cc: Jeffrey Vander Stoep <jeffv@google.com>
Cc: Linux Containers <containers@lists.linux-foundation.org>
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200531115031.391515-2-christian.brauner@ubuntu.com
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit 3a15fb6ed92cb32b0a83f406aa4a96f28c9adbc3)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Resolved minor merge conflict in kernel/exit.c where 5.4 does not have
commits 7bc3e6e55acf0 and 6ade99ec6175a.
Change-Id: I4a0113f3f64a86937ba5c9ac6e2537926e2827be
2026-05-07 10:17:17 -04:00
Christian Brauner
07bd1e5332 UPSTREAM: seccomp: rename "usage" to "refs" and document
Naming the lifetime counter of a seccomp filter "usage" suggests a
little too strongly that its about tasks that are using this filter
while it also tracks other references such as the user notifier or
ptrace. This also updates the documentation to note this fact.

We'll be introducing an actual usage counter in a follow-up patch.

Cc: Tycho Andersen <tycho@tycho.ws>
Cc: Kees Cook <keescook@chromium.org>
Cc: Matt Denton <mpdenton@google.com>
Cc: Sargun Dhillon <sargun@sargun.me>
Cc: Jann Horn <jannh@google.com>
Cc: Chris Palmer <palmer@google.com>
Cc: Aleksa Sarai <cyphar@cyphar.com>
Cc: Robert Sesek <rsesek@google.com>
Cc: Jeffrey Vander Stoep <jeffv@google.com>
Cc: Linux Containers <containers@lists.linux-foundation.org>
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200531115031.391515-1-christian.brauner@ubuntu.com
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit b707ddee11d1dc4518ab7f1aa5e7af9ceaa23317)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: If4c78f85b9873aea2dffa1c8afb250f534d0fedc
2026-05-07 10:17:17 -04:00
Sargun Dhillon
3c6d93041e UPSTREAM: seccomp: Add find_notification helper
This adds a helper which can iterate through a seccomp_filter to
find a notification matching an ID. It removes several replicated
chunks of code.

Signed-off-by: Sargun Dhillon <sargun@sargun.me>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Tycho Andersen <tycho@tycho.ws>
Cc: Matt Denton <mpdenton@google.com>
Cc: Kees Cook <keescook@google.com>,
Cc: Jann Horn <jannh@google.com>,
Cc: Robert Sesek <rsesek@google.com>,
Cc: Chris Palmer <palmer@google.com>
Cc: Christian Brauner <christian.brauner@ubuntu.com>
Cc: Tycho Andersen <tycho@tycho.ws>
Link: https://lore.kernel.org/r/20200601112532.150158-1-sargun@sargun.me
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit 9f87dcf14b82b05ff0e26970439b372ae135de0c)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I328edd8717fc06a415f7d4cf4fe561d86b2d4682
2026-05-07 10:17:17 -04:00
Kees Cook
9db11ce119 UPSTREAM: seccomp: Report number of loaded filters in /proc/$pid/status
A common question asked when debugging seccomp filters is "how many
filters are attached to your process?" Provide a way to easily answer
this question through /proc/$pid/status with a "Seccomp_filters" line.

Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit c818c03b661cd769e035e41673d5543ba2ebda64)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: Ib189ea77511a6088d5d328706c77c88709d8d63e
2026-05-07 10:17:07 -04:00
Michael Bestas
ee679763c2
Merge remote-tracking branch 'sm8350/lineage-20' into lineage-23.2
* sm8350/lineage-20: (60 commits)
  dsp-kernel: Avoid overflow in ALIGN macro usage
  fixup! BACKPORT: treewide: Use fallthrough pseudo-keyword
  Revert "net, sctp, filter: remap copy_from_user failure error"
  UPSTREAM: bpf: Fix L4 csum update on IPv6 in CHECKSUM_COMPLETE
  UPSTREAM: net: Fix checksum update for ILA adj-transport
  UPSTREAM: lwt_bpf: Replace preempt_disable() with migrate_disable()
  ANDROID: add kabi padding for bpf_verifier structure
  UPSTREAM: net/bpfilter: Initialize pos in __bpfilter_process_sockopt
  UPSTREAM: bpfilter: switch bpfilter_ip_set_sockopt to sockptr_t
  UPSTREAM: bpfilter: reject kernel addresses
  UPSTREAM: net/bpfilter: split __bpfilter_process_sockopt
  UPSTREAM: bpfilter: fix up a sparse annotation
  UPSTREAM: bpfilter: Allow to build bpfilter_umh as a module without static library
  UPSTREAM: bpfilter: Initialize pos variable
  UPSTREAM: bpfilter: switch to kernel_write
  UPSTREAM: bpfilter: Take advantage of the facilities of struct pid
  UPSTREAM: bpfilter: Move bpfilter_umh back into init data
  UPSTREAM: bpfilter: document build requirements for bpfilter_umh
  UPSTREAM: bpfilter: use 'userprogs' syntax to build bpfilter_umh
  UPSTREAM: exit: Factor thread_group_exited out of pidfd_poll
  ...

Change-Id: Ia34a9cf685cbded88f9b0463ca77d3cca36a2053
2026-04-18 13:39:10 +03:00
Eric W. Biederman
6782cb9b3f
UPSTREAM: exit: Factor thread_group_exited out of pidfd_poll
Create an independent helper thread_group_exited which returns true
when all threads have passed exit_notify in do_exit.  AKA all of the
threads are at least zombies and might be dead or completely gone.

Create this helper by taking the logic out of pidfd_poll where it is
already tested, and adding a READ_ONCE on the read of
task->exit_state.

I will be changing the user mode driver code to use this same logic
to know when a user mode driver needs to be restarted.

Place the new helper thread_group_exited in kernel/exit.c and
EXPORT it so it can be used by modules.

Link: https://lkml.kernel.org/r/20200702164140.4468-13-ebiederm@xmission.com
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Acked-by: Alexei Starovoitov <ast@kernel.org>
Tested-by: Alexei Starovoitov <ast@kernel.org>
Change-Id: I26813e9dad6a0bb78528cf1cff3fbc6c2369231f
Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:48 -08:00
Wang Qing
81d35fca80
UPSTREAM: bpf: Fix passing zero to PTR_ERR() in bpf_btf_printf_prepare
There is a bug when passing zero to PTR_ERR() and return.
Fix the smatch error.

Fixes: c4d0bfb45068 ("bpf: Add bpf_snprintf_btf helper")
Change-Id: Ibddaa08258ecc2f9887f84445abac7db55a90537
Signed-off-by: Wang Qing <wangqing@vivo.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Yonghong Song <yhs@fb.com>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Link: https://lore.kernel.org/bpf/1604735144-686-1-git-send-email-wangqing@vivo.com
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:47 -08:00
Yonghong Song
d34a994203
UPSTREAM: bpf: Fix a possible task gone issue with bpf_send_signal[_thread]() helpers
[ Upstream commit bdb7fdb0aca8b96cef9995d3a57e251c2289322f ]

In current bpf_send_signal() and bpf_send_signal_thread() helper
implementation, irq_work is used to handle nmi context. Hao Sun
reported in [1] that the current task at the entry of the helper
might be gone during irq_work callback processing. To fix the issue,
a reference is acquired for the current task before enqueuing into
the irq_work so that the queued task is still available during
irq_work callback processing.

  [1] https://lore.kernel.org/bpf/20230109074425.12556-1-sunhao.th@gmail.com/

Fixes: 8b401f9ed2 ("bpf: implement bpf_send_signal() helper")
Tested-by: Hao Sun <sunhao.th@gmail.com>
Reported-by: Hao Sun <sunhao.th@gmail.com>
Change-Id: I9caf795beeffc0041ced83cf2d7c960ce2e3a206
Signed-off-by: Yonghong Song <yhs@fb.com>
Link: https://lore.kernel.org/r/20230118204815.3331855-1-yhs@fb.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:46 -08:00
Jiri Olsa
75696eea7d
UPSTREAM: bpf: Disable preemption in bpf_event_output
commit d62cc390c2e99ae267ffe4b8d7e2e08b6c758c32 upstream.

We received report [1] of kernel crash, which is caused by
using nesting protection without disabled preemption.

The bpf_event_output can be called by programs executed by
bpf_prog_run_array_cg function that disabled migration but
keeps preemption enabled.

This can cause task to be preempted by another one inside the
nesting protection and lead eventually to two tasks using same
perf_sample_data buffer and cause crashes like:

  BUG: kernel NULL pointer dereference, address: 0000000000000001
  #PF: supervisor instruction fetch in kernel mode
  #PF: error_code(0x0010) - not-present page
  ...
  ? perf_output_sample+0x12a/0x9a0
  ? finish_task_switch.isra.0+0x81/0x280
  ? perf_event_output+0x66/0xa0
  ? bpf_event_output+0x13a/0x190
  ? bpf_event_output_data+0x22/0x40
  ? bpf_prog_dfc84bbde731b257_cil_sock4_connect+0x40a/0xacb
  ? xa_load+0x87/0xe0
  ? __cgroup_bpf_run_filter_sock_addr+0xc1/0x1a0
  ? release_sock+0x3e/0x90
  ? sk_setsockopt+0x1a1/0x12f0
  ? udp_pre_connect+0x36/0x50
  ? inet_dgram_connect+0x93/0xa0
  ? __sys_connect+0xb4/0xe0
  ? udp_setsockopt+0x27/0x40
  ? __pfx_udp_push_pending_frames+0x10/0x10
  ? __sys_setsockopt+0xdf/0x1a0
  ? __x64_sys_connect+0xf/0x20
  ? do_syscall_64+0x3a/0x90
  ? entry_SYSCALL_64_after_hwframe+0x72/0xdc

Fixing this by disabling preemption in bpf_event_output.

[1] https://github.com/cilium/cilium/issues/26756
Cc: stable@vger.kernel.org
Reported-by: Oleg "livelace" Popov <o.popov@livelace.ru>
Closes: https://github.com/cilium/cilium/issues/26756
Fixes: 2a916f2f546c ("bpf: Use migrate_disable/enable in array macros and cgroup/lirc code.")
Acked-by: Hou Tao <houtao1@huawei.com>
Change-Id: Icaf207af7dfb071f0adb10fe9bb7f367e25c1860
Signed-off-by: Jiri Olsa <jolsa@kernel.org>
Link: https://lore.kernel.org/r/20230725084206.580930-3-jolsa@kernel.org
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:46 -08:00
Yonghong Song
3150a87b86
UPSTREAM: bpf: Support 'X' in bpf_seq_printf() helper
'X' tells kernel to print hex with upper case letters.
/proc/net/tcp{4,6} seq_file show() used this, and
supports it in bpf_seq_printf() helper too.

Change-Id: I7c88f9f2c9088a6acb3b2085988de50535b970e6
Signed-off-by: Yonghong Song <yhs@fb.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20200623230807.3988014-1-yhs@fb.com
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:45 -08:00
Song Liu
b23f48db3c
UPSTREAM: bpf: Allow %pB in bpf_seq_printf() and bpf_trace_printk()
This makes it easy to dump stack trace in text.

Change-Id: Ia29115428a1dadf823881bfbbc793628b2fc9715
Signed-off-by: Song Liu <songliubraving@fb.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Yonghong Song <yhs@fb.com>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Link: https://lore.kernel.org/bpf/20200630062846.664389-4-songliubraving@fb.com
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:45 -08:00
Stanislav Fomichev
f2a2b97107
UPSTREAM: bpf: Remove inline from bpf_do_trace_printk
I get the following error during compilation on my side:
kernel/trace/bpf_trace.c: In function 'bpf_do_trace_printk':
kernel/trace/bpf_trace.c:386:34: error: function 'bpf_do_trace_printk' can never be inlined because it uses variable argument lists
 static inline __printf(1, 0) int bpf_do_trace_printk(const char *fmt, ...)
                                  ^

Fixes: ac5a72ea5c89 ("bpf: Use dedicated bpf_trace_printk event instead of trace_printk()")
Change-Id: I105192533535b3f3cbc78c7a54f1fbbc2eb6450b
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200806182612.1390883-1-sdf@google.com
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:44 -08:00
Alan Maguire
b338e77ab0
UPSTREAM: bpf: Use dedicated bpf_trace_printk event instead of trace_printk()
The bpf helper bpf_trace_printk() uses trace_printk() under the hood.
This leads to an alarming warning message originating from trace
buffer allocation which occurs the first time a program using
bpf_trace_printk() is loaded.

We can instead create a trace event for bpf_trace_printk() and enable
it in-kernel when/if we encounter a program using the
bpf_trace_printk() helper.  With this approach, trace_printk()
is not used directly and no warning message appears.

This work was started by Steven (see Link) and finished by Alan; added
Steven's Signed-off-by with his permission.

Change-Id: I3644e434a2f1e99d1fa523d247ac0b19911d041b
Signed-off-by: Steven Rostedt (VMware) <rostedt@goodmis.org>
Signed-off-by: Alan Maguire <alan.maguire@oracle.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Link: https://lore.kernel.org/r/20200628194334.6238b933@oasis.local.home
Link: https://lore.kernel.org/bpf/1594641154-18897-2-git-send-email-alan.maguire@oracle.com
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:43 -08:00
Daniel Xu
5e5bb29653
UPSTREAM: lib/strncpy_from_user.c: Mask out bytes after NUL terminator.
do_strncpy_from_user() may copy some extra bytes after the NUL
terminator into the destination buffer. This usually does not matter for
normal string operations. However, when BPF programs key BPF maps with
strings, this matters a lot.

A BPF program may read strings from user memory by calling the
bpf_probe_read_user_str() helper which eventually calls
do_strncpy_from_user(). The program can then key a map with the
destination buffer. BPF map keys are fixed-width and string-agnostic,
meaning that map keys are treated as a set of bytes.

The issue is when do_strncpy_from_user() overcopies bytes after the NUL
terminator, it can result in seemingly identical strings occupying
multiple slots in a BPF map. This behavior is subtle and totally
unexpected by the user.

This commit masks out the bytes following the NUL while preserving
long-sized stride in the fast path.

Fixes: 6ae08ae3dea2 ("bpf: Add probe_read_{user, kernel} and probe_read_{user, kernel}_str helpers")
Change-Id: I10d009c0e2db1ca242fbc0076fa2933f17c4a9a9
Signed-off-by: Daniel Xu <dxu@dxuuu.xyz>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/21efc982b3e9f2f7b0379eed642294caaa0c27a7.1605642949.git.dxu@dxuuu.xyz
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:43 -08:00
Daniel Borkmann
be5cfab9c4
BACKPORT: bpf: Add lockdown check for probe_write_user helper
commit 51e1bb9eeaf7868db56e58f47848e364ab4c4129 upstream.

Back then, commit 96ae522795 ("bpf: Add bpf_probe_write_user BPF helper
to be called in tracers") added the bpf_probe_write_user() helper in order
to allow to override user space memory. Its original goal was to have a
facility to "debug, divert, and manipulate execution of semi-cooperative
processes" under CAP_SYS_ADMIN. Write to kernel was explicitly disallowed
since it would otherwise tamper with its integrity.

One use case was shown in cf9b1199de ("samples/bpf: Add test/example of
using bpf_probe_write_user bpf helper") where the program DNATs traffic
at the time of connect(2) syscall, meaning, it rewrites the arguments to
a syscall while they're still in userspace, and before the syscall has a
chance to copy the argument into kernel space. These days we have better
mechanisms in BPF for achieving the same (e.g. for load-balancers), but
without having to write to userspace memory.

Of course the bpf_probe_write_user() helper can also be used to abuse
many other things for both good or bad purpose. Outside of BPF, there is
a similar mechanism for ptrace(2) such as PTRACE_PEEK{TEXT,DATA} and
PTRACE_POKE{TEXT,DATA}, but would likely require some more effort.
Commit 96ae522795 explicitly dedicated the helper for experimentation
purpose only. Thus, move the helper's availability behind a newly added
LOCKDOWN_BPF_WRITE_USER lockdown knob so that the helper is disabled under
the "integrity" mode. More fine-grained control can be implemented also
from LSM side with this change.

Fixes: 96ae522795 ("bpf: Add bpf_probe_write_user BPF helper to be called in tracers")
Change-Id: I824213c9e94b9c0c3ce0fe6188ed82ddebab018c
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andrii@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:42 -08:00
Alexei Starovoitov
c285b52590
UPSTREAM: bpf: Unbreak BPF_PROG_TYPE_KPROBE when kprobe is called via do_int3
[ Upstream commit 548f1191d86ccb9bde2a5305988877b7584c01eb ]

The commit 0d00449c7a28 ("x86: Replace ist_enter() with nmi_enter()")
converted do_int3 handler to be "NMI-like".
That made old if (in_nmi()) check abort execution of bpf programs
attached to kprobe when kprobe is firing via int3
(For example when kprobe is placed in the middle of the function).
Remove the check to restore user visible behavior.

Fixes: 0d00449c7a28 ("x86: Replace ist_enter() with nmi_enter()")
Reported-by: Nikolay Borisov <nborisov@suse.com>
Change-Id: I29ee95686ea9a51b2b59f8b06462510253036360
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Tested-by: Nikolay Borisov <nborisov@suse.com>
Reviewed-by: Masami Hiramatsu <mhiramat@kernel.org>
Link: https://lore.kernel.org/bpf/20210203070636.70926-1-alexei.starovoitov@gmail.com
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:42 -08:00
Thomas Gleixner
7ad9996931
UPSTREAM: bpf/trace: Remove redundant preempt_disable from trace_call_bpf()
Similar to __bpf_trace_run this is redundant because __bpf_trace_run() is
invoked from a trace point via __DO_TRACE() which already disables
preemption _before_ invoking any of the functions which are attached to a
trace point.

Remove it and add a cant_sleep() check.

Change-Id: I6106e5ed79cf84ae2b6e6032c7b081209a68da6c
Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200224145643.059995527@linutronix.de
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:41 -08:00
Jiri Olsa
19b8833055
UPSTREAM: bpf: Add extra path pointer check to d_path helper
[ Upstream commit f46fab0e36e611a2389d3843f34658c849b6bd60 ]

Anastasios reported crash on stable 5.15 kernel with following
BPF attached to lsm hook:

  SEC("lsm.s/bprm_creds_for_exec")
  int BPF_PROG(bprm_creds_for_exec, struct linux_binprm *bprm)
  {
          struct path *path = &bprm->executable->f_path;
          char p[128] = { 0 };

          bpf_d_path(path, p, 128);
          return 0;
  }

But bprm->executable can be NULL, so bpf_d_path call will crash:

  BUG: kernel NULL pointer dereference, address: 0000000000000018
  #PF: supervisor read access in kernel mode
  #PF: error_code(0x0000) - not-present page
  PGD 0 P4D 0
  Oops: 0000 [#1] PREEMPT SMP DEBUG_PAGEALLOC NOPTI
  ...
  RIP: 0010:d_path+0x22/0x280
  ...
  Call Trace:
   <TASK>
   bpf_d_path+0x21/0x60
   bpf_prog_db9cf176e84498d9_bprm_creds_for_exec+0x94/0x99
   bpf_trampoline_6442506293_0+0x55/0x1000
   bpf_lsm_bprm_creds_for_exec+0x5/0x10
   security_bprm_creds_for_exec+0x29/0x40
   bprm_execve+0x1c1/0x900
   do_execveat_common.isra.0+0x1af/0x260
   __x64_sys_execve+0x32/0x40

It's problem for all stable trees with bpf_d_path helper, which was
added in 5.9.

This issue is fixed in current bpf code, where we identify and mark
trusted pointers, so the above code would fail even to load.

For the sake of the stable trees and to workaround potentially broken
verifier in the future, adding the code that reads the path object from
the passed pointer and verifies it's valid in kernel space.

Fixes: 6e22ab9da793 ("bpf: Add d_path helper")
Reported-by: Anastasios Papagiannis <tasos.papagiannnis@gmail.com>
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Change-Id: I865e5c3a9c4d0268e9cbb35a74a162898c1c0a7e
Signed-off-by: Jiri Olsa <jolsa@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Stanislav Fomichev <sdf@google.com>
Acked-by: Yonghong Song <yhs@fb.com>
Link: https://lore.kernel.org/bpf/20230606181714.532998-1-jolsa@kernel.org
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:41 -08:00
Kajol Jain
506d2a3bbf
UPSTREAM: bpf: Remove config check to enable bpf support for branch records
[ Upstream commit db52f57211b4e45f0ebb274e2c877b211dc18591 ]

Branch data available to BPF programs can be very useful to get stack traces
out of userspace application.

Commit fff7b64355ea ("bpf: Add bpf_read_branch_records() helper") added BPF
support to capture branch records in x86. Enable this feature also for other
architectures as well by removing checks specific to x86.

If an architecture doesn't support branch records, bpf_read_branch_records()
still has appropriate checks and it will return an -EINVAL in that scenario.
Based on UAPI helper doc in include/uapi/linux/bpf.h, unsupported architectures
should return -ENOENT in such case. Hence, update the appropriate check to
return -ENOENT instead.

Selftest 'perf_branches' result on power9 machine which has the branch stacks
support:

 - Before this patch:

  [command]# ./test_progs -t perf_branches
   #88/1 perf_branches/perf_branches_hw:FAIL
   #88/2 perf_branches/perf_branches_no_hw:OK
   #88 perf_branches:FAIL
  Summary: 0/1 PASSED, 0 SKIPPED, 1 FAILED

 - After this patch:

  [command]# ./test_progs -t perf_branches
   #88/1 perf_branches/perf_branches_hw:OK
   #88/2 perf_branches/perf_branches_no_hw:OK
   #88 perf_branches:OK
  Summary: 1/2 PASSED, 0 SKIPPED, 0 FAILED

Selftest 'perf_branches' result on power9 machine which doesn't have branch
stack report:

 - After this patch:

  [command]# ./test_progs -t perf_branches
   #88/1 perf_branches/perf_branches_hw:SKIP
   #88/2 perf_branches/perf_branches_no_hw:OK
   #88 perf_branches:OK
  Summary: 1/1 PASSED, 1 SKIPPED, 0 FAILED

Fixes: fff7b64355eac ("bpf: Add bpf_read_branch_records() helper")
Suggested-by: Peter Zijlstra <peterz@infradead.org>
Change-Id: Ide9c32fdd0c825e46d7a9c6872d78122e4bb9b79
Signed-off-by: Kajol Jain <kjain@linux.ibm.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20211206073315.77432-1-kjain@linux.ibm.com
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:40 -08:00
KP Singh
2e3ca6ceb2
UPSTREAM: bpf: Fix bpf_prog_test_run_tracing for !CONFIG_NET
test_run.o is not built when CONFIG_NET is not set and
bpf_prog_test_run_tracing being referenced in bpf_trace.o causes the
linker error:

ld: kernel/trace/bpf_trace.o:(.rodata+0x38): undefined reference to
 `bpf_prog_test_run_tracing'

Add a __weak function in bpf_trace.c to handle this.

Fixes: da00d2f117a0 ("bpf: Add test ops for BPF_PROG_TYPE_TRACING")
Change-Id: Icd0d76d484c459c235066dc28046dd6997b3d434
Signed-off-by: KP Singh <kpsingh@google.com>
Reported-by: Randy Dunlap <rdunlap@infradead.org>
Acked-by: Randy Dunlap <rdunlap@infradead.org>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200305220127.29109-1-kpsingh@chromium.org
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:40 -08:00
Yonghong Song
cbecc486db
UPSTREAM: bpf: Fix build failure for kernel/trace/bpf_trace.c with CONFIG_NET=n
When CONFIG_NET is not defined, I hit the following build error:
    kernel/trace/bpf_trace.o:(.rodata+0x110): undefined reference to `bpf_prog_test_run_raw_tp'

Commit 1b4d60ec162f ("bpf: Enable BPF_PROG_TEST_RUN for raw_tracepoint")
added test_run support for raw_tracepoint in /kernel/trace/bpf_trace.c.
But the test_run function bpf_prog_test_run_raw_tp is defined in
net/bpf/test_run.c, only available with CONFIG_NET=y.

Adding a CONFIG_NET guard for
    .test_run = bpf_prog_test_run_raw_tp;
fixed the above build issue.

Fixes: 1b4d60ec162f ("bpf: Enable BPF_PROG_TEST_RUN for raw_tracepoint")
Change-Id: Ifdaa9a4b0e6501fd432bccbd1184cd0993cb3c02
Signed-off-by: Yonghong Song <yhs@fb.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20201007062933.3425899-1-yhs@fb.com
Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com>
2026-02-01 20:44:40 -08:00