Commit graph

980,574 commits

Author SHA1 Message Date
Daniel Borkmann
3944aaeb2c
UPSTREAM: uaccess: Add strict non-pagefault kernel-space read function
Add two new probe_kernel_read_strict() and strncpy_from_unsafe_strict()
helpers which by default alias to the __probe_kernel_read() and the
__strncpy_from_unsafe(), respectively, but can be overridden by archs
which have non-overlapping address ranges for kernel space and user
space in order to bail out with -EFAULT when attempting to probe user
memory including non-canonical user access addresses [0]:

  4-level page tables:
    user-space mem: 0x0000000000000000 - 0x00007fffffffffff
    non-canonical:  0x0000800000000000 - 0xffff7fffffffffff

  5-level page tables:
    user-space mem: 0x0000000000000000 - 0x00ffffffffffffff
    non-canonical:  0x0100000000000000 - 0xfeffffffffffffff

The idea is that these helpers are complementary to the probe_user_read()
and strncpy_from_unsafe_user() which probe user-only memory. Both added
helpers here do the same, but for kernel-only addresses.

Both set of helpers are going to be used for BPF tracing. They also
explicitly avoid throwing the splat for non-canonical user addresses from
00c42373d3 ("x86-64: add warning for non-canonical user access address
dereferences").

For compat, the current probe_kernel_read() and strncpy_from_unsafe() are
left as-is.

  [0] Documentation/x86/x86_64/mm.txt

Change-Id: I82fb94bc1f76ff318c4c7d71746f583bd711b830
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: x86@kernel.org
Link: https://lore.kernel.org/bpf/eefeefd769aa5a013531f491a71f0936779e916b.1572649915.git.daniel@iogearbox.net
2025-12-23 13:35:39 -08:00
Nadav Amit
f1632e4deb
UPSTREAM: mm/tlb: Provide default nmi_uaccess_okay()
x86 has an nmi_uaccess_okay(), but other architectures do not.
Arch-independent code might need to know whether access to user
addresses is ok in an NMI context or in other code whose execution
context is unknown.  Specifically, this function is needed for
bpf_probe_write_user().

Add a default implementation of nmi_uaccess_okay() for architectures
that do not have such a function.

Change-Id: Ie710d06700a2c5b392bff0dbc82798f9c0d985cc
Signed-off-by: Nadav Amit <namit@vmware.com>
Signed-off-by: Rick Edgecombe <rick.p.edgecombe@intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Cc: <akpm@linux-foundation.org>
Cc: <ard.biesheuvel@linaro.org>
Cc: <deneen.t.dock@intel.com>
Cc: <kernel-hardening@lists.openwall.com>
Cc: <kristen@linux.intel.com>
Cc: <linux_dti@icloud.com>
Cc: <will.deacon@arm.com>
Cc: Andy Lutomirski <luto@kernel.org>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: H. Peter Anvin <hpa@zytor.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Rik van Riel <riel@surriel.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Link: https://lkml.kernel.org/r/20190426001143.4983-23-namit@vmware.com
Signed-off-by: Ingo Molnar <mingo@kernel.org>
2025-12-23 13:35:38 -08:00
Björn Töpel
8035f90173
UPSTREAM: xsk: Restructure/inline XSKMAP lookup/redirect/flush
In this commit the XSKMAP entry lookup function used by the XDP
redirect code is moved from the xskmap.c file to the xdp_sock.h
header, so the lookup can be inlined from, e.g., the
bpf_xdp_redirect_map() function.

Further the __xsk_map_redirect() and __xsk_map_flush() is moved to the
xsk.c, which lets the compiler inline the xsk_rcv() and xsk_flush()
functions.

Finally, all the XDP socket functions were moved from linux/bpf.h to
net/xdp_sock.h, where most of the XDP sockets functions are anyway.

This yields a ~2% performance boost for the xdpsock "rx_drop"
scenario.

Change-Id: I045aa7c44454340e6af3cd93c45f62323f4c7406
Signed-off-by: Björn Töpel <bjorn.topel@intel.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20191101110346.15004-4-bjorn.topel@gmail.com
2025-12-23 13:35:38 -08:00
Maciej Fijalkowski
8a13992407
UPSTREAM: bpf: Implement map_gen_lookup() callback for XSKMAP
Inline the xsk_map_lookup_elem() via implementing the map_gen_lookup()
callback. This results in emitting the bpf instructions in place of
bpf_map_lookup_elem() helper call and better performance of bpf
programs.

Change-Id: Ica4d9a58854a3ac2b9da819e84eccabef7d8adeb
Signed-off-by: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Jonathan Lemon <jonathan.lemon@gmail.com>
Link: https://lore.kernel.org/bpf/20191101110346.15004-3-bjorn.topel@gmail.com
2025-12-23 13:35:38 -08:00
Björn Töpel
e898f409b1
UPSTREAM: xsk: Store struct xdp_sock as a flexible array member of the XSKMAP
Prior this commit, the array storing XDP socket instances were stored
in a separate allocated array of the XSKMAP. Now, we store the sockets
as a flexible array member in a similar fashion as the arraymap. Doing
so, we do less pointer chasing in the lookup.

Change-Id: Iae24fe35cf8f2dce2ec41ad58968a26645fbed9e
Signed-off-by: Björn Töpel <bjorn.topel@intel.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Jonathan Lemon <jonathan.lemon@gmail.com>
Link: https://lore.kernel.org/bpf/20191101110346.15004-2-bjorn.topel@gmail.com
2025-12-23 13:35:38 -08:00
Alexei Starovoitov
9df1a7412e
BACKPORT: bpf: Replace prog_raw_tp+btf_id with prog_tracing
The bpf program type raw_tp together with 'expected_attach_type'
was the most appropriate api to indicate BTF-enabled raw_tp programs.
But during development it became apparent that 'expected_attach_type'
cannot be used and new 'attach_btf_id' field had to be introduced.
Which means that the information is duplicated in two fields where
one of them is ignored.
Clean it up by introducing new program type where both
'expected_attach_type' and 'attach_btf_id' fields have
specific meaning.
In the future 'expected_attach_type' will be extended
with other attach points that have similar semantics to raw_tp.
This patch is replacing BTF-enabled BPF_PROG_TYPE_RAW_TRACEPOINT with
prog_type = BPF_RPOG_TYPE_TRACING
expected_attach_type = BPF_TRACE_RAW_TP
attach_btf_id = btf_id of raw tracepoint inside the kernel
Future patches will add
expected_attach_type = BPF_TRACE_FENTRY or BPF_TRACE_FEXIT
where programs have the same input context and the same helpers,
but different attach points.

Change-Id: If19aaa2fc33fc4923931e4fb589e9b066d9d5695
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191030223212.953010-2-ast@kernel.org
2025-12-23 13:35:38 -08:00
Alexei Starovoitov
45a15724f0
UPSTREAM: bpf: Fix bpf jit kallsym access
Jiri reported crash when JIT is on, but net.core.bpf_jit_kallsyms is off.
bpf_prog_kallsyms_find() was skipping addr->bpf_prog resolution
logic in oops and stack traces. That's incorrect.
It should only skip addr->name resolution for 'cat /proc/kallsyms'.
That's what bpf_jit_kallsyms and bpf_jit_harden protect.

Fixes: 3dec541b2e63 ("bpf: Add support for BTF pointers to x86 JIT")
Reported-by: Jiri Olsa <jolsa@redhat.com>
Change-Id: I97955f0f9aa51745493f2399c3154f8a2ade6f56
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20191030233019.1187404-1-ast@kernel.org
2025-12-23 13:35:38 -08:00
Paul E. McKenney
1c4e2fa042
UPSTREAM: bpf/cgroup: Replace rcu_swap_protected() with rcu_replace_pointer()
This commit replaces the use of rcu_swap_protected() with the more
intuitively appealing rcu_replace_pointer() as a step towards removing
rcu_swap_protected().

Link: https://lore.kernel.org/lkml/CAHk-=wiAsJLw1egFEE=Z7-GGtM6wcvtyytXZA1+BHqta4gg6Hw@mail.gmail.com/
Reported-by: Linus Torvalds <torvalds@linux-foundation.org>
[ paulmck: From rcu_replace() to rcu_replace_pointer() per Ingo Molnar. ]
Change-Id: I49e16f57052bd11695036615d36bcd5ec6f20266
Signed-off-by: Paul E. McKenney <paulmck@kernel.org>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Song Liu <songliubraving@fb.com>
Cc: Alexei Starovoitov <ast@kernel.org>
Cc: Daniel Borkmann <daniel@iogearbox.net>
Cc: Martin KaFai Lau <kafai@fb.com>
Cc: Yonghong Song <yhs@fb.com>
Cc: <netdev@vger.kernel.org>
Cc: <bpf@vger.kernel.org>
2025-12-23 13:35:37 -08:00
Alexei Starovoitov
eb5ed883a2
UPSTREAM: bpf: Enforce 'return 0' in BTF-enabled raw_tp programs
The return value of raw_tp programs is ignored by __bpf_trace_run()
that calls them. The verifier also allows any value to be returned.
For BTF-enabled raw_tp lets enforce 'return 0', so that return value
can be used for something in the future.

Change-Id: I11bae02f217565a577dccbaf5a7c5b105be83cd9
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Link: https://lore.kernel.org/bpf/20191029032426.1206762-1-ast@kernel.org
2025-12-23 13:35:37 -08:00
Martin KaFai Lau
462a97a99f
UPSTREAM: bpf: Prepare btf_ctx_access for non raw_tp use case
This patch makes a few changes to btf_ctx_access() to prepare
it for non raw_tp use case where the attach_btf_id is not
necessary a BTF_KIND_TYPEDEF.

It moves the "btf_trace_" prefix check and typedef-follow logic to a new
function "check_attach_btf_id()" which is called only once during
bpf_check().  btf_ctx_access() only operates on a BTF_KIND_FUNC_PROTO
type now. That should also be more efficient since it is done only
one instead of every-time check_ctx_access() is called.

"check_attach_btf_id()" needs to find the func_proto type from
the attach_btf_id.  It needs to store the result into the
newly added prog->aux->attach_func_proto.  func_proto
btf type has no name, so a proper name should be stored into
"attach_func_name" also.

v2:
- Move the "btf_trace_" check to an earlier verifier phase (Alexei)

Change-Id: I2716756b4f85b1e068b8f8f2245e992d7b97cc6b
Signed-off-by: Martin KaFai Lau <kafai@fb.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20191025001811.1718491-1-kafai@fb.com
2025-12-23 13:35:37 -08:00
Alexei Starovoitov
fa3ab43405
UPSTREAM: bpf: Fix bpf_attr.attach_btf_id check
Only raw_tracepoint program type can have bpf_attr.attach_btf_id >= 0.
Make sure to reject other program types that accidentally set it to non-zero.

Fixes: ccfe29eb29c2 ("bpf: Add attach_btf_id attribute to program load")
Reported-by: Andrii Nakryiko <andriin@fb.com>
Change-Id: If05e54cff182417e681379c8c6adb0b85cd8d984
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Yonghong Song <yhs@fb.com>
Link: https://lore.kernel.org/bpf/20191018060933.2950231-1-ast@kernel.org
2025-12-23 13:35:37 -08:00
Alexei Starovoitov
ce02bc7b7d
BACKPORT: bpf: Check types of arguments passed into helpers
Introduce new helper that reuses existing skb perf_event output
implementation, but can be called from raw_tracepoint programs
that receive 'struct sk_buff *' as tracepoint argument or
can walk other kernel data structures to skb pointer.

In order to do that teach verifier to resolve true C types
of bpf helpers into in-kernel BTF ids.
The type of kernel pointer passed by raw tracepoint into bpf
program will be tracked by the verifier all the way until
it's passed into helper function.
For example:
kfree_skb() kernel function calls trace_kfree_skb(skb, loc);
bpf programs receives that skb pointer and may eventually
pass it into bpf_skb_output() bpf helper which in-kernel is
implemented via bpf_skb_event_output() kernel function.
Its first argument in the kernel is 'struct sk_buff *'.
The verifier makes sure that types match all the way.

Change-Id: I2782993220c91f02cd53321bac032c8049c80273
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191016032505.2089704-11-ast@kernel.org
2025-12-23 13:35:37 -08:00
Alexei Starovoitov
988dc8126d
BACKPORT: bpf: Add support for BTF pointers to x86 JIT
Pointer to BTF object is a pointer to kernel object or NULL.
Such pointers can only be used by BPF_LDX instructions.
The verifier changed their opcode from LDX|MEM|size
to LDX|PROBE_MEM|size to make JITing easier.
The number of entries in extable is the number of BPF_LDX insns
that access kernel memory via "pointer to BTF type".
Only these load instructions can fault.
Since x86 extable is relative it has to be allocated in the same
memory region as JITed code.
Allocate it prior to last pass of JITing and let the last pass populate it.
Pointer to extable in bpf_prog_aux is necessary to make page fault
handling fast.
Page fault handling is done in two steps:
1. bpf_prog_kallsyms_find() finds BPF program that page faulted.
   It's done by walking rb tree.
2. then extable for given bpf program is binary searched.
This process is similar to how page faulting is done for kernel modules.
The exception handler skips over faulting x86 instruction and
initializes destination register with zero. This mimics exact
behavior of bpf_probe_read (when probe_kernel_read faults dest is zeroed).

JITs for other architectures can add support in similar way.
Until then they will reject unknown opcode and fallback to interpreter.

Since extable should be aligned and placed near JITed code
make bpf_jit_binary_alloc() return 4 byte aligned image offset,
so that extable aligning formula in bpf_int_jit_compile() doesn't need
to rely on internal implementation of bpf_jit_binary_alloc().
On x86 gcc defaults to 16-byte alignment for regular kernel functions
due to better performance. JITed code may be aligned to 16 in the future,
but it will use 4 in the meantime.

Change-Id: Ic2a23af52b7b88482f524c9b98a62c14ea2a52c1
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191016032505.2089704-10-ast@kernel.org
2025-12-23 13:35:36 -08:00
Alexei Starovoitov
506ea3d519
BACKPORT: bpf: Add support for BTF pointers to interpreter
Pointer to BTF object is a pointer to kernel object or NULL.
The memory access in the interpreter has to be done via probe_kernel_read
to avoid page faults.

Change-Id: Ief5b8d67dbcc988362dfb302b486e99ca6eec7b6
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191016032505.2089704-9-ast@kernel.org
2025-12-23 13:35:36 -08:00
Alexei Starovoitov
75c620fba0
UPSTREAM: bpf: Attach raw_tp program with BTF via type name
BTF type id specified at program load time has all
necessary information to attach that program to raw tracepoint.
Use kernel type name to find raw tracepoint.

Add missing CHECK_ATTR() condition.

Change-Id: I545792a45a62a9d2086e474a6857340215842407
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191016032505.2089704-8-ast@kernel.org
2025-12-23 13:35:36 -08:00
Alexei Starovoitov
00a0fd127c
BACKPORT: bpf: Implement accurate raw_tp context access via BTF
libbpf analyzes bpf C program, searches in-kernel BTF for given type name
and stores it into expected_attach_type.
The kernel verifier expects this btf_id to point to something like:
typedef void (*btf_trace_kfree_skb)(void *, struct sk_buff *skb, void *loc);
which represents signature of raw_tracepoint "kfree_skb".

Then btf_ctx_access() matches ctx+0 access in bpf program with 'skb'
and 'ctx+8' access with 'loc' arguments of "kfree_skb" tracepoint.
In first case it passes btf_id of 'struct sk_buff *' back to the verifier core
and 'void *' in second case.

Then the verifier tracks PTR_TO_BTF_ID as any other pointer type.
Like PTR_TO_SOCKET points to 'struct bpf_sock',
PTR_TO_TCP_SOCK points to 'struct bpf_tcp_sock', and so on.
PTR_TO_BTF_ID points to in-kernel structs.
If 1234 is btf_id of 'struct sk_buff' in vmlinux's BTF
then PTR_TO_BTF_ID#1234 points to one of in kernel skbs.

When PTR_TO_BTF_ID#1234 is dereferenced (like r2 = *(u64 *)r1 + 32)
the btf_struct_access() checks which field of 'struct sk_buff' is
at offset 32. Checks that size of access matches type definition
of the field and continues to track the dereferenced type.
If that field was a pointer to 'struct net_device' the r2's type
will be PTR_TO_BTF_ID#456. Where 456 is btf_id of 'struct net_device'
in vmlinux's BTF.

Such verifier analysis prevents "cheating" in BPF C program.
The program cannot cast arbitrary pointer to 'struct sk_buff *'
and access it. C compiler would allow type cast, of course,
but the verifier will notice type mismatch based on BPF assembly
and in-kernel BTF.

Change-Id: Icd510a80078dca9aceb6ee7e7db93644e08ea773
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191016032505.2089704-7-ast@kernel.org
2025-12-23 13:35:36 -08:00
Alexei Starovoitov
171e4b06de
UPSTREAM: bpf: Add attach_btf_id attribute to program load
Add attach_btf_id attribute to prog_load command.
It's similar to existing expected_attach_type attribute which is
used in several cgroup based program types.
Unfortunately expected_attach_type is ignored for
tracing programs and cannot be reused for new purpose.
Hence introduce attach_btf_id to verify bpf programs against
given in-kernel BTF type id at load time.
It is strictly checked to be valid for raw_tp programs only.
In a later patches it will become:
btf_id == 0 semantics of existing raw_tp progs.
btd_id > 0 raw_tp with BTF and additional type safety.

Change-Id: I16a13a9f199ba3923608b610a61b58a26d757034
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191016032505.2089704-5-ast@kernel.org
2025-12-23 13:35:36 -08:00
Alexei Starovoitov
d61fa57de5
UPSTREAM: bpf: Process in-kernel BTF
If in-kernel BTF exists parse it and prepare 'struct btf *btf_vmlinux'
for further use by the verifier.
In-kernel BTF is trusted just like kallsyms and other build artifacts
embedded into vmlinux.
Yet run this BTF image through BTF verifier to make sure
that it is valid and it wasn't mangled during the build.
If in-kernel BTF is incorrect it means either gcc or pahole or kernel
are buggy. In such case disallow loading BPF programs.

Change-Id: I80a64018a8ce4390c8c3d4284c4cc1fdcd0fd81f
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191016032505.2089704-4-ast@kernel.org
2025-12-23 13:35:36 -08:00
Stanislav Fomichev
ca51d66a1e
UPSTREAM: bpf: Allow __sk_buff tstamp in BPF_PROG_TEST_RUN
It's useful for implementing EDT related tests (set tstamp, run the
test, see how the tstamp is changed or observe some other parameter).

Note that bpf_ktime_get_ns() helper is using monotonic clock, so for
the BPF programs that compare tstamp against it, tstamp should be
derived from clock_gettime(CLOCK_MONOTONIC, ...).

Change-Id: I6fdeccb4baac4c4ead931dc2d3c0730163d1c470
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191015183125.124413-1-sdf@google.com
2025-12-23 13:35:35 -08:00
Eric Dumazet
96bdef8718
UPSTREAM: bpf: Align struct bpf_prog_stats
Do not risk spanning these small structures on two cache lines.

Change-Id: Ie13182cdff956b270296b5999b78c41d77133e23
Signed-off-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20191011181140.2898-1-edumazet@google.com
2025-12-23 13:35:35 -08:00
Jakub Kicinski
1cea6f1d77
BACKPORT: net: sockmap: use bitmap for copy info
Don't use bool array in struct sk_msg_sg, save 12 bytes.

Change-Id: I50359c5cd6101a1c4f2c78f7d83877fbbaa16063
Signed-off-by: Jakub Kicinski <jakub.kicinski@netronome.com>
Reviewed-by: Dirk van der Merwe <dirk.vandermerwe@netronome.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2025-12-23 13:35:35 -08:00
Andrii Nakryiko
5418db0090
UPSTREAM: uapi/bpf: fix helper docs
Various small fixes to BPF helper documentation comments, enabling
automatic header generation with a list of BPF helpers.

Change-Id: If4c1db80f20522013d5823e1c5928e4f4b8193b9
Signed-off-by: Andrii Nakryiko <andriin@fb.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
2025-12-23 13:35:35 -08:00
wuzhe
3156cf63cc ANDROID: gki_defconfig: Enable HID_BETOP_FF JOYSTICK_XPAD_FF and JOYSTICK_XPAD_LEDS
Enable several configs to support gamepad

Bug: 188463121
Change-Id: Iba12760edd79d45466fddb481877d5ccf148fb53
Signed-off-by: Zhe Wu <wuzhe@oppo.com>
2025-12-23 06:55:24 +00:00
Heiko Carstens
605a96c07d
BACKPORT: epoll: fix compat syscall wire up of epoll_pwait2
Commit b0a0c2615f6f ("epoll: wire up syscall epoll_pwait2") wired up
the 64 bit syscall instead of the compat variant in a couple of places.

Fixes: b0a0c2615f6f ("epoll: wire up syscall epoll_pwait2")
Change-Id: Ida2267e953240c036d64bc61177ee72bb20f35a3
Signed-off-by: Heiko Carstens <hca@linux.ibm.com>
Acked-by: Arnd Bergmann <arnd@arndb.de>
Cc: Willem de Bruijn <willemb@google.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Thomas Bogendoerfer <tsbogend@alpha.franken.de>
Cc: Vasily Gorbik <gor@linux.ibm.com>
Cc: Christian Borntraeger <borntraeger@de.ibm.com>
Cc: "David S. Miller" <davem@davemloft.net>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2025-12-22 07:42:51 +02:00
Willem de Bruijn
b0f56c9401
BACKPORT: epoll: wire up syscall epoll_pwait2
Split off from prev patch in the series that implements the syscall.

Link: https://lkml.kernel.org/r/20201121144401.3727659-4-willemdebruijn.kernel@gmail.com
Change-Id: I48dfae6f721b24ebc53de603e393289954a95908
Signed-off-by: Willem de Bruijn <willemb@google.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2025-12-22 07:42:50 +02:00
Willem de Bruijn
9eb868290d
UPSTREAM: epoll: add syscall epoll_pwait2
Add syscall epoll_pwait2, an epoll_wait variant with nsec resolution that
replaces int timeout with struct timespec.  It is equivalent otherwise.

    int epoll_pwait2(int fd, struct epoll_event *events,
                     int maxevents,
                     const struct timespec *timeout,
                     const sigset_t *sigset);

The underlying hrtimer is already programmed with nsec resolution.
pselect and ppoll also set nsec resolution timeout with timespec.

The sigset_t in epoll_pwait has a compat variant. epoll_pwait2 needs
the same.

For timespec, only support this new interface on 2038 aware platforms
that define __kernel_timespec_t. So no CONFIG_COMPAT_32BIT_TIME.

Link: https://lkml.kernel.org/r/20201121144401.3727659-3-willemdebruijn.kernel@gmail.com
Change-Id: I8cb4e756aacb4bee7cbe1c2cb5320a59e07626f8
Signed-off-by: Willem de Bruijn <willemb@google.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2025-12-22 04:28:19 +02:00
Willem de Bruijn
4889a679d0
UPSTREAM: epoll: convert internal api to timespec64
Patch series "add epoll_pwait2 syscall", v4.

Enable nanosecond timeouts for epoll.

Analogous to pselect and ppoll, introduce an epoll_wait syscall
variant that takes a struct timespec instead of int timeout.

This patch (of 4):

Make epoll more consistent with select/poll: pass along the timeout as
timespec64 pointer.

In anticipation of additional changes affecting all three polling
mechanisms:

- add epoll_pwait2 syscall with timespec semantics,
  and share poll_select_set_timeout implementation.
- compute slack before conversion to absolute time,
  to save one ktime_get_ts64 call.

Link: https://lkml.kernel.org/r/20201121144401.3727659-1-willemdebruijn.kernel@gmail.com
Link: https://lkml.kernel.org/r/20201121144401.3727659-2-willemdebruijn.kernel@gmail.com
Change-Id: Id5851ad620849d202d18a99351a98b8bc3b6820a
Signed-off-by: Willem de Bruijn <willemb@google.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Arnd Bergmann <arnd@arndb.de>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2025-12-22 04:28:18 +02:00
Al Viro
ec5ff5dc24
UPSTREAM: close_range(): fix the logics in descriptor table trimming
commit 678379e1d4f7443b170939525d3312cfc37bf86b upstream.

Cloning a descriptor table picks the size that would cover all currently
opened files.  That's fine for clone() and unshare(), but for close_range()
there's an additional twist - we clone before we close, and it would be
a shame to have
	close_range(3, ~0U, CLOSE_RANGE_UNSHARE)
leave us with a huge descriptor table when we are not going to keep
anything past stderr, just because some large file descriptor used to
be open before our call has taken it out.

Unfortunately, it had been dealt with in an inherently racy way -
sane_fdtable_size() gets a "don't copy anything past that" argument
(passed via unshare_fd() and dup_fd()), close_range() decides how much
should be trimmed and passes that to unshare_fd().

The problem is, a range that used to extend to the end of descriptor
table back when close_range() had looked at it might very well have stuff
grown after it by the time dup_fd() has allocated a new files_struct
and started to figure out the capacity of fdtable to be attached to that.

That leads to interesting pathological cases; at the very least it's a
QoI issue, since unshare(CLONE_FILES) is atomic in a sense that it takes
a snapshot of descriptor table one might have observed at some point.
Since CLOSE_RANGE_UNSHARE close_range() is supposed to be a combination
of unshare(CLONE_FILES) with plain close_range(), ending up with a
weird state that would never occur with unshare(2) is confusing, to put
it mildly.

It's not hard to get rid of - all it takes is passing both ends of the
range down to sane_fdtable_size().  There we are under ->files_lock,
so the race is trivially avoided.

So we do the following:
	* switch close_files() from calling unshare_fd() to calling
dup_fd().
	* undo the calling convention change done to unshare_fd() in
60997c3d45d9 "close_range: add CLOSE_RANGE_UNSHARE"
	* introduce struct fd_range, pass a pointer to that to dup_fd()
and sane_fdtable_size() instead of "trim everything past that point"
they are currently getting.  NULL means "we are not going to be punching
any holes"; NR_OPEN_MAX is gone.
	* make sane_fdtable_size() use find_last_bit() instead of
open-coding it; it's easier to follow that way.
	* while we are at it, have dup_fd() report errors by returning
ERR_PTR(), no need to use a separate int *errorp argument.

Fixes: 60997c3d45d9 "close_range: add CLOSE_RANGE_UNSHARE"
Cc: stable@vger.kernel.org
Change-Id: I6782a2edf98970b6c2d662048061e28f7e57b9c9
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-22 04:28:18 +02:00
Linus Torvalds
dc0e6903ac
UPSTREAM: fs: fix fd table size alignment properly
[ Upstream commit d888c83fcec75194a8a48ccd283953bdba7b2550 ]

Jason Donenfeld reports that my commit 1c24a186398f ("fs: fd tables have
to be multiples of BITS_PER_LONG") doesn't work, and the reason is an
embarrassing brown-paper-bag bug.

Yes, we want to align the number of fds to BITS_PER_LONG, and yes, the
reason they might not be aligned is because the incoming 'max_fd'
argument might not be aligned.

But aligining the argument - while simple - will cause a "infinitely
big" maxfd (eg NR_OPEN_MAX) to just overflow to zero.  Which most
definitely isn't what we want either.

The obvious fix was always just to do the alignment last, but I had
moved it earlier just to make the patch smaller and the code look
simpler.  Duh.  It certainly made _me_ look simple.

Fixes: 1c24a186398f ("fs: fd tables have to be multiples of BITS_PER_LONG")
Reported-and-tested-by: Jason A. Donenfeld <Jason@zx2c4.com>
Cc: Fedor Pchelkin <aissur0002@gmail.com>
Cc: Alexey Khoroshilov <khoroshilov@ispras.ru>
Cc: Christian Brauner <brauner@kernel.org>
Change-Id: I6d3a0c28896e8cf8cab30f60679aab785aeee193
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
2025-12-22 04:28:18 +02:00
Linus Torvalds
c9f8bc1026
UPSTREAM: fs: fd tables have to be multiples of BITS_PER_LONG
[ Upstream commit 1c24a186398f59c80adb9a967486b65c1423a59d ]

This has always been the rule: fdtables have several bitmaps in them,
and as a result they have to be sized properly for bitmaps.  We walk
those bitmaps in chunks of 'unsigned long' in serveral cases, but even
when we don't, we use the regular kernel bitops that are defined to work
on arrays of 'unsigned long', not on some byte array.

Now, the distinction between arrays of bytes and 'unsigned long'
normally only really ends up being noticeable on big-endian systems, but
Fedor Pchelkin and Alexey Khoroshilov reported that copy_fd_bitmaps()
could be called with an argument that wasn't even a multiple of
BITS_PER_BYTE.  And then it fails to do the proper copy even on
little-endian machines.

The bug wasn't in copy_fd_bitmap(), but in sane_fdtable_size(), which
didn't actually sanitize the fdtable size sufficiently, and never made
sure it had the proper BITS_PER_LONG alignment.

That's partly because the alignment historically came not from having to
explicitly align things, but simply from previous fdtable sizes, and
from count_open_files(), which counts the file descriptors by walking
them one 'unsigned long' word at a time and thus naturally ends up doing
sizing in the proper 'chunks of unsigned long'.

But with the introduction of close_range(), we now have an external
source of "this is how many files we want to have", and so
sane_fdtable_size() needs to do a better job.

This also adds that explicit alignment to alloc_fdtable(), although
there it is mainly just for documentation at a source code level.  The
arithmetic we do there to pick a reasonable fdtable size already aligns
the result sufficiently.

In fact,clang notices that the added ALIGN() in that function doesn't
actually do anything, and does not generate any extra code for it.

It turns out that gcc ends up confusing itself by combining a previous
constant-sized shift operation with the variable-sized shift operations
in roundup_pow_of_two().  And probably due to that doesn't notice that
the ALIGN() is a no-op.  But that's a (tiny) gcc misfeature that doesn't
matter.  Having the explicit alignment makes sense, and would actually
matter on a 128-bit architecture if we ever go there.

This also adds big comments above both functions about how fdtable sizes
have to have that BITS_PER_LONG alignment.

Fixes: 60997c3d45d9 ("close_range: add CLOSE_RANGE_UNSHARE")
Reported-by: Fedor Pchelkin <aissur0002@gmail.com>
Reported-by: Alexey Khoroshilov <khoroshilov@ispras.ru>
Link: https://lore.kernel.org/all/20220326114009.1690-1-aissur0002@gmail.com/
Tested-and-acked-by: Christian Brauner <brauner@kernel.org>
Change-Id: Ib64319f20fecec1c367bf0167e1b14b47751537a
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
2025-12-22 04:28:18 +02:00
Christian Brauner
011100cdcc
UPSTREAM: file: simplify logic in __close_range()
It never looked too pleasant and it doesn't really buy us anything
anymore now that CLOSE_RANGE_CLOEXEC exists and we need to retake the
current maximum under the lock for it anyway. This also makes the logic
easier to follow.

Cc: Christoph Hellwig <hch@lst.de>
Cc: Giuseppe Scrivano <gscrivan@redhat.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: linux-fsdevel@vger.kernel.org
Change-Id: I72be5b078d7edc061318b0eda36a351e1dc01e56
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
2025-12-22 04:28:18 +02:00
Christian Brauner
f5d2458f94
UPSTREAM: file: fix close_range() for unshare+cloexec
syzbot reported a bug when putting the last reference to a tasks file
descriptor table. Debugging this showed we didn't recalculate the
current maximum fd number for CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC
after we unshared the file descriptors table. So max_fd could exceed the
current fdtable maximum causing us to set excessive bits. As a concrete
example, let's say the user requested everything from fd 4 to ~0UL to be
closed and their current fdtable size is 256 with their highest open fd
being 4. With CLOSE_RANGE_UNSHARE the caller will end up with a new
fdtable which has room for 64 file descriptors since that is the lowest
fdtable size we accept. But now max_fd will still point to 255 and needs
to be adjusted. Fix this by retrieving the correct maximum fd value in
__range_cloexec().

Reported-by: syzbot+283ce5a46486d6acdbaf@syzkaller.appspotmail.com
Fixes: 582f1fb6b721 ("fs, close_range: add flag CLOSE_RANGE_CLOEXEC")
Fixes: fec8a6a69103 ("close_range: unshare all fds for CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC")
Cc: Christoph Hellwig <hch@lst.de>
Cc: Giuseppe Scrivano <gscrivan@redhat.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: linux-fsdevel@vger.kernel.org
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
(cherry picked from commit 9b5b872215fe6d1ca6a1ef411f130bd58e269012)
Bug: 216276716
Signed-off-by: Maciej Żenczykowski <maze@google.com>
Change-Id: Id04f75bfeb49dcff0457136f74b47a2d05bd9d51
2025-12-22 03:57:46 +02:00
Christian Brauner
b6060beee9
UPSTREAM: close_range: unshare all fds for CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC
After introducing CLOSE_RANGE_CLOEXEC syzbot reported a crash when
CLOSE_RANGE_CLOEXEC is specified in conjunction with CLOSE_RANGE_UNSHARE.
When CLOSE_RANGE_UNSHARE is specified the caller will receive a private
file descriptor table in case their file descriptor table is currently
shared.

For the case where the caller has requested all file descriptors to be
actually closed via e.g. close_range(3, ~0U, 0) the kernel knows that
the caller does not need any of the file descriptors anymore and will
optimize the close operation by only copying all files in the range from
0 to 3 and no others.

However, if the caller requested CLOSE_RANGE_CLOEXEC together with
CLOSE_RANGE_UNSHARE the caller wants to still make use of the file
descriptors so the kernel needs to copy all of them and can't optimize.

The original patch didn't account for this and thus could cause oopses
as evidenced by the syzbot report because it assumed that all fds had
been copied. Fix this by handling the CLOSE_RANGE_CLOEXEC case.

syzbot reported
==================================================================
BUG: KASAN: null-ptr-deref in instrument_atomic_read include/linux/instrumented.h:71 [inline]
BUG: KASAN: null-ptr-deref in atomic64_read include/asm-generic/atomic-instrumented.h:837 [inline]
BUG: KASAN: null-ptr-deref in atomic_long_read include/asm-generic/atomic-long.h:29 [inline]
BUG: KASAN: null-ptr-deref in filp_close+0x22/0x170 fs/open.c:1274
Read of size 8 at addr 0000000000000077 by task syz-executor511/8522

CPU: 1 PID: 8522 Comm: syz-executor511 Not tainted 5.10.0-syzkaller #0
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 01/01/2011
Call Trace:
 __dump_stack lib/dump_stack.c:79 [inline]
 dump_stack+0x107/0x163 lib/dump_stack.c:120
 __kasan_report mm/kasan/report.c:549 [inline]
 kasan_report.cold+0x5/0x37 mm/kasan/report.c:562
 check_memory_region_inline mm/kasan/generic.c:186 [inline]
 check_memory_region+0x13d/0x180 mm/kasan/generic.c:192
 instrument_atomic_read include/linux/instrumented.h:71 [inline]
 atomic64_read include/asm-generic/atomic-instrumented.h:837 [inline]
 atomic_long_read include/asm-generic/atomic-long.h:29 [inline]
 filp_close+0x22/0x170 fs/open.c:1274
 close_files fs/file.c:402 [inline]
 put_files_struct fs/file.c:417 [inline]
 put_files_struct+0x1cc/0x350 fs/file.c:414
 exit_files+0x12a/0x170 fs/file.c:435
 do_exit+0xb4f/0x2a00 kernel/exit.c:818
 do_group_exit+0x125/0x310 kernel/exit.c:920
 get_signal+0x428/0x2100 kernel/signal.c:2792
 arch_do_signal_or_restart+0x2a8/0x1eb0 arch/x86/kernel/signal.c:811
 handle_signal_work kernel/entry/common.c:147 [inline]
 exit_to_user_mode_loop kernel/entry/common.c:171 [inline]
 exit_to_user_mode_prepare+0x124/0x200 kernel/entry/common.c:201
 __syscall_exit_to_user_mode_work kernel/entry/common.c:291 [inline]
 syscall_exit_to_user_mode+0x19/0x50 kernel/entry/common.c:302
 entry_SYSCALL_64_after_hwframe+0x44/0xa9
RIP: 0033:0x447039
Code: Unable to access opcode bytes at RIP 0x44700f.
RSP: 002b:00007f1b1225cdb8 EFLAGS: 00000246 ORIG_RAX: 00000000000000ca
RAX: 0000000000000001 RBX: 00000000006dbc28 RCX: 0000000000447039
RDX: 00000000000f4240 RSI: 0000000000000081 RDI: 00000000006dbc2c
RBP: 00000000006dbc20 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 00000000006dbc2c
R13: 00007fff223b6bef R14: 00007f1b1225d9c0 R15: 00000000006dbc2c
==================================================================

syzbot has tested the proposed patch and the reproducer did not trigger any issue:

Reported-and-tested-by: syzbot+96cfd2b22b3213646a93@syzkaller.appspotmail.com

Tested on:

commit:         10f7cddd selftests/core: add regression test for CLOSE_RAN..
git tree:       git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux.git vfs
kernel config:  https://syzkaller.appspot.com/x/.config?x=5d42216b510180e3
dashboard link: https://syzkaller.appspot.com/bug?extid=96cfd2b22b3213646a93
compiler:       gcc (GCC) 10.1.0-syz 20200507

Reported-by: syzbot+96cfd2b22b3213646a93@syzkaller.appspotmail.com
Fixes: 582f1fb6b721 ("fs, close_range: add flag CLOSE_RANGE_CLOEXEC")
Cc: Giuseppe Scrivano <gscrivan@redhat.com>
Cc: linux-fsdevel@vger.kernel.org
Link: https://lore.kernel.org/r/20201217213303.722643-1-christian.brauner@ubuntu.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
(cherry picked from commit fec8a6a691033f2538cd46848f17f337f0739923)
Bug: 216276716
Signed-off-by: Maciej Żenczykowski <maze@google.com>
Change-Id: I0c6bbd8a46b93293c212fff08728514915d5c690
2025-12-22 03:57:39 +02:00
Giuseppe Scrivano
e5c9a44a83
UPSTREAM: fs, close_range: add flag CLOSE_RANGE_CLOEXEC
When the flag CLOSE_RANGE_CLOEXEC is set, close_range doesn't
immediately close the files but it sets the close-on-exec bit.

It is useful for e.g. container runtimes that usually install a
seccomp profile "as late as possible" before execv'ing the container
process itself.  The container runtime could either do:
  1                                  2
- install_seccomp_profile();       - close_range(MIN_FD, MAX_INT, 0);
- close_range(MIN_FD, MAX_INT, 0); - install_seccomp_profile();
- execve(...);                     - execve(...);

Both alternative have some disadvantages.

In the first variant the seccomp_profile cannot block the close_range
syscall, as well as opendir/read/close/... for the fallback on older
kernels.
In the second variant, close_range() can be used only on the fds
that are not going to be needed by the runtime anymore, and it must be
potentially called multiple times to account for the different ranges
that must be closed.

Using close_range(..., ..., CLOSE_RANGE_CLOEXEC) solves these issues.
The runtime is able to use the existing open fds, the seccomp profile
can block close_range() and the syscalls used for its fallback.

Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>
Link: https://lore.kernel.org/r/20201118104746.873084-2-gscrivan@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
(cherry picked from commit 582f1fb6b721facf04848d2ca57f34468da1813e)
Bug: 216276716
Signed-off-by: Maciej Żenczykowski <maze@google.com>
Change-Id: Ib2d44f9760a80e3febdb25925d17ce5ffd14910e
2025-12-22 03:57:34 +02:00
Christian Brauner
c9353ba21b
BACKPORT: arch: wire-up close_range()
This wires up the close_range() syscall into all arches at once.

Suggested-by: Arnd Bergmann <arnd@arndb.de>
Change-Id: Ib962b01f3a490c901b0e0f51e436df1a26887ca4
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Acked-by: Arnd Bergmann <arnd@arndb.de>
Acked-by: Michael Ellerman <mpe@ellerman.id.au> (powerpc)
Cc: Jann Horn <jannh@google.com>
Cc: David Howells <dhowells@redhat.com>
Cc: Dmitry V. Levin <ldv@altlinux.org>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Florian Weimer <fweimer@redhat.com>
Cc: linux-api@vger.kernel.org
Cc: linux-alpha@vger.kernel.org
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-ia64@vger.kernel.org
Cc: linux-m68k@lists.linux-m68k.org
Cc: linux-mips@vger.kernel.org
Cc: linux-parisc@vger.kernel.org
Cc: linuxppc-dev@lists.ozlabs.org
Cc: linux-s390@vger.kernel.org
Cc: linux-sh@vger.kernel.org
Cc: sparclinux@vger.kernel.org
Cc: linux-xtensa@linux-xtensa.org
Cc: linux-arch@vger.kernel.org
Cc: x86@kernel.org
2025-12-22 03:56:25 +02:00
Christian Brauner
857abe7dd6
UPSTREAM: close_range: add CLOSE_RANGE_UNSHARE
One of the use-cases of close_range() is to drop file descriptors just before
execve(). This would usually be expressed in the sequence:

unshare(CLONE_FILES);
close_range(3, ~0U);

as pointed out by Linus it might be desirable to have this be a part of
close_range() itself under a new flag CLOSE_RANGE_UNSHARE.

This expands {dup,unshare)_fd() to take a max_fds argument that indicates the
maximum number of file descriptors to copy from the old struct files. When the
user requests that all file descriptors are supposed to be closed via
close_range(min, max) then we can cap via unshare_fd(min) and hence don't need
to do any of the heavy fput() work for everything above min.

The patch makes it so that if CLOSE_RANGE_UNSHARE is requested and we do in
fact currently share our file descriptor table we create a new private copy.
We then close all fds in the requested range and finally after we're done we
install the new fd table.

Suggested-by: Linus Torvalds <torvalds@linux-foundation.org>
Change-Id: I0813045886501e40a45693ee1edad50bdf2b66e5
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
2025-12-22 03:48:31 +02:00
Christian Brauner
90fe243e9a
UPSTREAM: open: add close_range()
This adds the close_range() syscall. It allows to efficiently close a range
of file descriptors up to all file descriptors of a calling task.

I was contacted by FreeBSD as they wanted to have the same close_range()
syscall as we proposed here. We've coordinated this and in the meantime, Kyle
was fast enough to merge close_range() into FreeBSD already in April:
https://reviews.freebsd.org/D21627
https://svnweb.freebsd.org/base?view=revision&revision=359836
and the current plan is to backport close_range() to FreeBSD 12.2 (cf. [2])
once its merged in Linux too. Python is in the process of switching to
close_range() on FreeBSD and they are waiting on us to merge this to switch on
Linux as well: https://bugs.python.org/issue38061

The syscall came up in a recent discussion around the new mount API and
making new file descriptor types cloexec by default. During this
discussion, Al suggested the close_range() syscall (cf. [1]). Note, a
syscall in this manner has been requested by various people over time.

First, it helps to close all file descriptors of an exec()ing task. This
can be done safely via (quoting Al's example from [1] verbatim):

        /* that exec is sensitive */
        unshare(CLONE_FILES);
        /* we don't want anything past stderr here */
        close_range(3, ~0U);
        execve(....);

The code snippet above is one way of working around the problem that file
descriptors are not cloexec by default. This is aggravated by the fact that
we can't just switch them over without massively regressing userspace. For
a whole class of programs having an in-kernel method of closing all file
descriptors is very helpful (e.g. demons, service managers, programming
language standard libraries, container managers etc.).
(Please note, unshare(CLONE_FILES) should only be needed if the calling
task is multi-threaded and shares the file descriptor table with another
thread in which case two threads could race with one thread allocating file
descriptors and the other one closing them via close_range(). For the
general case close_range() before the execve() is sufficient.)

Second, it allows userspace to avoid implementing closing all file
descriptors by parsing through /proc/<pid>/fd/* and calling close() on each
file descriptor. From looking at various large(ish) userspace code bases
this or similar patterns are very common in:
- service managers (cf. [4])
- libcs (cf. [6])
- container runtimes (cf. [5])
- programming language runtimes/standard libraries
  - Python (cf. [2])
  - Rust (cf. [7], [8])
As Dmitry pointed out there's even a long-standing glibc bug about missing
kernel support for this task (cf. [3]).
In addition, the syscall will also work for tasks that do not have procfs
mounted and on kernels that do not have procfs support compiled in. In such
situations the only way to make sure that all file descriptors are closed
is to call close() on each file descriptor up to UINT_MAX or RLIMIT_NOFILE,
OPEN_MAX trickery (cf. comment [8] on Rust).

The performance is striking. For good measure, comparing the following
simple close_all_fds() userspace implementation that is essentially just
glibc's version in [6]:

static int close_all_fds(void)
{
        int dir_fd;
        DIR *dir;
        struct dirent *direntp;

        dir = opendir("/proc/self/fd");
        if (!dir)
                return -1;
        dir_fd = dirfd(dir);
        while ((direntp = readdir(dir))) {
                int fd;
                if (strcmp(direntp->d_name, ".") == 0)
                        continue;
                if (strcmp(direntp->d_name, "..") == 0)
                        continue;
                fd = atoi(direntp->d_name);
                if (fd == dir_fd || fd == 0 || fd == 1 || fd == 2)
                        continue;
                close(fd);
        }
        closedir(dir);
        return 0;
}

to close_range() yields:
1. closing 4 open files:
   - close_all_fds(): ~280 us
   - close_range():    ~24 us

2. closing 1000 open files:
   - close_all_fds(): ~5000 us
   - close_range():   ~800 us

close_range() is designed to allow for some flexibility. Specifically, it
does not simply always close all open file descriptors of a task. Instead,
callers can specify an upper bound.
This is e.g. useful for scenarios where specific file descriptors are
created with well-known numbers that are supposed to be excluded from
getting closed.
For extra paranoia close_range() comes with a flags argument. This can e.g.
be used to implement extension. Once can imagine userspace wanting to stop
at the first error instead of ignoring errors under certain circumstances.
There might be other valid ideas in the future. In any case, a flag
argument doesn't hurt and keeps us on the safe side.

From an implementation side this is kept rather dumb. It saw some input
from David and Jann but all nonsense is obviously my own!
- Errors to close file descriptors are currently ignored. (Could be changed
  by setting a flag in the future if needed.)
- __close_range() is a rather simplistic wrapper around __close_fd().
  My reasoning behind this is based on the nature of how __close_fd() needs
  to release an fd. But maybe I misunderstood specifics:
  We take the files_lock and rcu-dereference the fdtable of the calling
  task, we find the entry in the fdtable, get the file and need to release
  files_lock before calling filp_close().
  In the meantime the fdtable might have been altered so we can't just
  retake the spinlock and keep the old rcu-reference of the fdtable
  around. Instead we need to grab a fresh reference to the fdtable.
  If my reasoning is correct then there's really no point in fancyfying
  __close_range(): We just need to rcu-dereference the fdtable of the
  calling task once to cap the max_fd value correctly and then go on
  calling __close_fd() in a loop.

/* References */
[1]: https://lore.kernel.org/lkml/20190516165021.GD17978@ZenIV.linux.org.uk/
[2]: 9e4f2f3a6b/Modules/_posixsubprocess.c (L220)
[3]: https://sourceware.org/bugzilla/show_bug.cgi?id=10353#c7
[4]: 5238e95759/src/basic/fd-util.c (L217)
[5]: ddf4b77e11/src/lxc/start.c (L236)
[6]: https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/unix/sysv/linux/grantpt.c;h=2030e07fa6e652aac32c775b8c6e005844c3c4eb;hb=HEAD#l17
     Note that this is an internal implementation that is not exported.
     Currently, libc seems to not provide an exported version of this
     because of missing kernel support to do this.
     Note, in a recent patch series Florian made grantpt() a nop thereby
     removing the code referenced here.
[7]: https://github.com/rust-lang/rust/issues/12148
[8]: 5f47c0613e/src/libstd/sys/unix/process2.rs (L303-L308)
     Rust's solution is slightly different but is equally unperformant.
     Rust calls getdtablesize() which is a glibc library function that
     simply returns the current RLIMIT_NOFILE or OPEN_MAX values. Rust then
     goes on to call close() on each fd. That's obviously overkill for most
     tasks. Rarely, tasks - especially non-demons - hit RLIMIT_NOFILE or
     OPEN_MAX.
     Let's be nice and assume an unprivileged user with RLIMIT_NOFILE set
     to 1024. Even in this case, there's a very high chance that in the
     common case Rust is calling the close() syscall 1021 times pointlessly
     if the task just has 0, 1, and 2 open.

Suggested-by: Al Viro <viro@zeniv.linux.org.uk>
Change-Id: I2f7abcf9210a1e79855837d1b7c580cc7f7a38e2
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Kyle Evans <self@kyle-evans.net>
Cc: Jann Horn <jannh@google.com>
Cc: David Howells <dhowells@redhat.com>
Cc: Dmitry V. Levin <ldv@altlinux.org>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Florian Weimer <fweimer@redhat.com>
Cc: linux-api@vger.kernel.org
2025-12-22 03:47:50 +02:00
basamaryan
612434f5cc
tcp: Fix compiler error due to type confusion
tcp_tso_should_defer() used the min() macro on incompatible types (u64 vs long)

  error: comparison of distinct pointer types
  ('typeof (srtt_in_ns >> 1) *' (aka 'unsigned long long *')
  and 'typeof (1000000L) *' (aka 'long *'))

Fix this up by using min_t to ensure the comparison is performed as
unsigned 64-bit integers

Fixes: 8b7ac7af3a
Change-Id: Iccd7ae5bdcc3fd77fea126389761cc30c5558315
2025-12-11 02:37:07 -08:00
Thomas Turner
09ce0d8408
fixup! net: usb: rtl8150: Fix frame padding
Fixes commit d4000407ca9b94779b86567198f74a3b34d3c2fc upstream.

The fix was based on commit 2f5281ba2a
upstream.

/home/thomas/a21s/kernel/qcom/sm8250/drivers/net/usb/rtl8150.c:707:10: error: comparison of distinct pointer types ('typeof (skb->len) *' (aka 'unsigned int *') and 'typeof (60) *' (aka 'int *')) [-Werror,-Wcompare-distinct-pointer-types]
  707 |         count = max(skb->len, ETH_ZLEN);
      |                 ^~~~~~~~~~~~~~~~~~~~~~~

Change-Id: If4a118c06c11ee66b80042534300158ec1f2039b
2025-12-11 02:36:44 -08:00
basamaryan
9181ada660
Merge branch 'android11-5.4-lts' of https://android.googlesource.com/kernel/common into android13-5.4-lahaina
* 'android11-5.4-lts' of https://android.googlesource.com/kernel/common:
  Linux 5.4.302
  Input: pegasus-notetaker - fix potential out-of-bounds access
  Input: remove third argument of usb_maxpacket()
  usb: deprecate the third argument of usb_maxpacket()
  ata: libata-scsi: Fix system suspend for a security locked drive
  fs/proc: fix uaf in proc_readdir_de()
  pmdomain: imx: Fix reference count leak in imx_gpc_remove
  pmdomain: arm: scmi: Fix genpd leak on provider registration failure
  net: netpoll: fix incorrect refcount handling causing incorrect cleanup
  net: qede: Initialize qede_ll_ops with designated initializer
  uio_hv_generic: Set event for all channels on the device
  net: ethernet: ti: netcp: Standardize knav_dma_open_channel to return NULL on error
  ALSA: usb-audio: fix uac2 clock source at terminal parser
  mm/page_alloc: fix hash table order logging in alloc_large_system_hash()
  kconfig/nconf: Initialize the default locale at startup
  kconfig/mconf: Initialize the default locale at startup
  vsock: Ignore signal/timeout on connect() if already established
  s390/ctcm: Fix double-kfree
  net: openvswitch: remove never-working support for setting nsh fields
  mlxsw: spectrum: Fix memory leak in mlxsw_sp_flower_stats()
  MIPS: Malta: Fix !EVA SOC-it PCI MMIO
  scsi: target: tcm_loop: Fix segfault in tcm_loop_tpg_address_show()
  scsi: sg: Do not sleep in atomic context
  Input: cros_ec_keyb - fix an invalid memory access
  be2net: pass wrb_params in case of OS2BMC
  HID: quirks: work around VID/PID conflict for 0x4c4a/0x4155
  isdn: mISDN: hfcsusb: fix memory leak in hfcsusb_probe()
  EDAC/altera: Use INTTEST register for Ethernet and USB SBE injection
  EDAC/altera: Handle OCRAM ECC enable after warm reset
  spi: Try to get ACPI GPIO IRQ earlier
  ipv4: route: Prevent rt_bind_exception() from rebinding stale fnhe
  strparser: Fix signed/unsigned mismatch bug
  gcov: add support for GCC 15
  mm/ksm: fix flag-dropping behavior in ksm_madvise
  ALSA: usb-audio: Fix NULL pointer dereference in snd_usb_mixer_controls_badd
  drm/vmwgfx: Validate command header size against SVGA_CMD_MAX_DATASIZE
  ASoC: cs4271: Fix regulator leak on probe failure
  regulator: fixed: fix GPIO descriptor leak on register failure
  regulator: fixed: use dev_err_probe for register
  Bluetooth: L2CAP: export l2cap_chan_hold for modules
  net_sched: limit try_bulk_dequeue_skb() batches
  net_sched: remove need_resched() from qdisc_run()
  net/mlx5e: Fix wraparound in rate limiting for values above 255 Gbps
  net/mlx5e: Fix maxrate wraparound in threshold between units
  net: sched: act_ife: initialize struct tc_ife to fix KMSAN kernel-infoleak
  wifi: mac80211: skip rate verification for not captured PSDUs
  net: mdio: fix resource leak in mdiobus_register_device()
  tipc: Fix use-after-free in tipc_mon_reinit_self().
  tipc: simplify the finalize work queue
  sctp: prevent possible shift-out-of-bounds in sctp_transport_update_rto
  sctp: get netns from asoc and ep base
  Bluetooth: 6lowpan: Don't hold spin lock over sleeping functions
  Bluetooth: 6lowpan: fix BDADDR_LE vs ADDR_LE_DEV address type confusion
  Bluetooth: 6lowpan: reset link-local header on ipv6 recv path
  Bluetooth: btusb: reorder cleanup in btusb_disconnect to avoid UAF
  net: fec: correct rx_bytes statistic for the case SHIFT16 is set
  ASoC: max98090/91: fixed max98091 ALSA widget powering up/down
  HID: quirks: avoid Cooler Master MM712 dongle wakeup bug
  NFS4: Fix state renewals missing after boot
  compiler_types: Move unused static inline functions warning to W=2
  extcon: adc-jack: Cleanup wakeup source only if it was enabled
  tracing: Fix memory leaks in create_field_var()
  net: usb: qmi_wwan: initialize MAC header offset in qmimux_rx_fixup
  sctp: Prevent TOCTOU out-of-bounds write
  sctp: Hold RCU read lock while iterating over address list
  net: dsa: b53: stop reading ARL entries if search is done
  net: dsa: b53: fix enabling ip multicast
  net: dsa: b53: fix resetting speed and pause on forced link
  net: dsa: b53: prevent GMII_PORT_OVERRIDE_CTRL access on BCM5325
  net: dsa/b53: change b53_force_port_config() pause argument
  net: vlan: sync VLAN features with lower device
  ceph: add checking of wait_for_completion_killable() return value
  fbdev: Add bounds checking in bit_putcs to fix vmalloc-out-of-bounds
  ACPI: property: Return present device nodes only on fwnode interface
  9p: sysfs_init: don't hardcode error to ENOMEM
  9p: fix /sys/fs/9p/caches overwriting itself
  fs/hpfs: Fix error code for new_inode() failure in mkdir/create/mknod/symlink
  ACPICA: Update dsmethod.c to get rid of unused variable warning
  orangefs: fix xattr related buffer overflow...
  page_pool: Clamp pool size to max 16K pages
  Bluetooth: bcsp: receive data only if registered
  Bluetooth: SCO: Fix UAF on sco_conn_free
  net: macb: avoid dealing with endianness in macb_set_hwaddr()
  nfs4_setup_readdir(): insufficient locking for ->d_parent->d_inode dereferencing
  NFSv4.1: fix mount hang after CREATE_SESSION failure
  NFSv4: handle ERR_GRACE on delegation recalls
  remoteproc: qcom: q6v5: Avoid handling handover twice
  sparc/module: Add R_SPARC_UA64 relocation handling
  net: intel: fm10k: Fix parameter idx set but not used
  jfs: fix uninitialized waitqueue in transaction manager
  jfs: Verify inode mode when loading from disk
  ipv6: np->rxpmtu race annotation
  usb: xhci: plat: Facilitate using autosuspend for xhci plat devices
  usb: mon: Increase BUFF_MAX to 64 MiB to support multi-MB URBs
  allow finish_no_open(file, ERR_PTR(-E...))
  scsi: lpfc: Define size of debugfs entry for xri rebalancing
  scsi: lpfc: Check return status of lpfc_reset_flush_io_context during TGT_RESET
  selftests/Makefile: include $(INSTALL_DEP_TARGETS) in clean target to clean net/lib dependency
  net/cls_cgroup: Fix task_get_classid() during qdisc run
  selftests: Replace sleep with slowwait
  selftests: Disable dad for ipv6 in fcnal-test.sh
  media: redrat3: use int type to store negative error codes
  net: sh_eth: Disable WoL if system can not suspend
  phy: cadence: cdns-dphy: Enable lower resolutions in dphy
  usb: gadget: f_hid: Fix zero length packet transfer
  net: call cond_resched() less often in __release_sock()
  ALSA: usb-audio: apply quirk for MOONDROP Quark2
  net: nfc: nci: Increase NCI_DATA_TIMEOUT to 3000 ms
  dmaengine: dw-edma: Set status for callback_result
  dmaengine: mv_xor: match alloc_wc and free_wc
  dmaengine: sh: setup_xref error handling
  scsi: pm8001: Use int instead of u32 to store error codes
  mips: lantiq: xway: sysctrl: rename stp clock
  mips: lantiq: danube: add missing device_type in pci node
  mips: lantiq: danube: add missing properties to cpu node
  media: fix uninitialized symbol warnings
  drm/amdkfd: Tie UNMAP_LATENCY to queue_preemption
  extcon: adc-jack: Fix wakeup source leaks on device unbind
  rds: Fix endianness annotation for RDS_MPATH_HASH
  PCI/P2PDMA: Fix incorrect pointer usage in devm_kfree() call
  net: Call trace_sock_exceed_buf_limit() for memcg failure with SK_MEM_RECV.
  net: When removing nexthops, don't call synchronize_net if it is not necessary
  char: misc: Does not request module for miscdevice with dynamic minor
  usb: gadget: f_ncm: Fix MAC assignment NCM ethernet
  iio: adc: spear_adc: mask SPEAR_ADC_STATUS channel and avg sample before setting register
  media: imon: make send_packet() more robust
  net: ipv6: fix field-spanning memcpy warning in AH output
  bridge: Redirect to backup port when port is administratively down
  powerpc/eeh: Use result of error_detected() in uevent
  x86/vsyscall: Do not require X86_PF_INSTR to emulate vsyscall
  media: pci: ivtv: Don't create fake v4l2_fh
  drm/amdkfd: return -ENOTTY for unsupported IOCTLs
  selftests/net: Ensure assert() triggers in psock_tpacket.c
  selftests/net: Replace non-standard __WORDSIZE with sizeof(long) * 8
  PCI: Disable MSI on RDC PCI to PCIe bridges
  drm/nouveau: replace snprintf() with scnprintf() in nvkm_snprintbf()
  mfd: madera: Work around false-positive -Wininitialized warning
  mfd: stmpe-i2c: Add missing MODULE_LICENSE
  mfd: stmpe: Remove IRQ domain upon removal
  tools/power x86_energy_perf_policy: Prefer driver HWP limits
  tools/power x86_energy_perf_policy: Enhance HWP enable
  tools/cpupower: Fix incorrect size in cpuidle_state_disable()
  hwmon: (dell-smm) Add support for Dell OptiPlex 7040
  uprobe: Do not emulate/sstep original instruction when ip is changed
  clocksource/drivers/vf-pit: Replace raw_readl/writel to readl/writel
  video: backlight: lp855x_bl: Set correct EPROM start for LP8556
  tee: allow a driver to allocate a tee_device without a pool
  ACPICA: dispatcher: Use acpi_ds_clear_operands() in acpi_ds_call_control_method()
  mmc: sdhci-msm: Enable tuning for SDR50 mode for SD card
  irqchip/gic-v2m: Handle Multiple MSI base IRQ Alignment
  arc: Fix __fls() const-foldability via __builtin_clzl()
  cpufreq/longhaul: handle NULL policy in longhaul_exit
  selftests/bpf: Fix bpf_prog_detach2 usage in test_lirc_mode2
  ACPI: video: force native for Lenovo 82K8
  memstick: Add timeout to prevent indefinite waiting
  mmc: host: renesas_sdhi: Fix the actual clock
  bpf: Don't use %pK through printk
  spi: loopback-test: Don't use %pK through printk
  soc: qcom: smem: Fix endian-unaware access of num_entries
  usb: gadget: f_fs: Fix epfile null pointer access after ep enable.
  serial: 8250_dw: handle reset control deassert error
  serial: 8250_dw: Use devm_add_action_or_reset()
  serial: 8250_dw: Use devm_clk_get_optional() to get the input clock
  can: gs_usb: increase max interface to U8_MAX
  devcoredump: Fix circular locking dependency with devcd->mutex.
  net: ravb: Enforce descriptor type ordering
  x86/resctrl: Fix miscount of bandwidth event when reactivating previously unavailable RMID
  wifi: brcmfmac: fix crash while sending Action Frames in standalone AP Mode
  net: phy: dp83867: Disable EEE support as not implemented
  regmap: slimbus: fix bus_context pointer in regmap init calls
  drm/etnaviv: fix flush sequence logic
  usbnet: Prevents free active kevent
  wifi: ath10k: Fix memory leak on unsupported WMI command
  ASoC: qdsp6: q6asm: do not sleep while atomic
  fbdev: valkyriefb: Fix reference count leak in valkyriefb_init
  fbdev: pvr2fb: Fix leftover reference to ONCHIP_NR_DMA_CHANNELS
  fbdev: bitblit: bound-check glyph index in bit_putcs*
  ACPI: video: Fix use-after-free in acpi_video_switch_brightness()
  fbdev: atyfb: Check if pll_ops->init_pll failed
  net: usb: asix_devices: Check return value of usbnet_get_endpoints
  btrfs: use smp_mb__after_atomic() when forcing COW in create_pending_snapshot()
  x86/bugs: Fix reporting of LFENCE retpoline
  net/sched: sch_qfq: Fix null-deref in agg_dequeue

 Conflicts:
	drivers/mmc/host/sdhci-msm.c
	drivers/usb/host/xhci-plat.c

Change-Id: I30739fea8840ace3825ea8955598d9efa687bb27
2025-12-11 01:41:51 -08:00
Michael Bestas
350c4ccfe1
Merge tag 'ASB-2025-12-01_11-5.4' of https://android.googlesource.com/kernel/common into android13-5.4-lahaina
https://source.android.com/docs/security/bulletin/2025-12-01
CVE-2025-48623
CVE-2025-48624
CVE-2025-48637
CVE-2025-48638
CVE-2024-35970
CVE-2025-38236
CVE-2025-38349
CVE-2025-48610
CVE-2025-38500

* tag 'ASB-2025-12-01_11-5.4' of https://android.googlesource.com/kernel/common:
  UPSTREAM: crypto: essiv - Check ssize for decryption and in-place encryption
  ANDROID: GKI: fix up build break where timer_delete_sync() was used
  Revert "net: rtnetlink: remove redundant assignment to variable err"
  Revert "net: rtnetlink: add msg kind names"
  Revert "net: rtnetlink: add helper to extract msg type's kind"
  Revert "net: rtnetlink: use BIT for flag values"
  Revert "net: netlink: add NLM_F_BULK delete request modifier"
  Revert "net: rtnetlink: add bulk delete support flag"
  Revert "net: rtnetlink: fix module reference count leak issue in rtnetlink_rcv_msg"
  Revert "net: add ndo_fdb_del_bulk"
  Revert "net: rtnetlink: add NLM_F_BULK support to rtnl_fdb_del"
  Revert "rtnetlink: Allow deleting FDB entries in user namespace"
  Linux 5.4.301
  net: rtnetlink: fix module reference count leak issue in rtnetlink_rcv_msg
  media: s5p-mfc: remove an unused/uninitialized variable
  NFSD: Fix last write offset handling in layoutcommit
  NFSD: Minor cleanup in layoutcommit processing
  padata: Reset next CPU when reorder sequence wraps around
  KEYS: trusted_tpm1: Compare HMAC values in constant time
  NFSD: Define a proc_layoutcommit for the FlexFiles layout type
  vfs: Don't leak disconnected dentries on umount
  jbd2: ensure that all ongoing I/O complete before freeing blocks
  ext4: detect invalid INLINE_DATA + EXTENTS flag combination
  drm/amdgpu: use atomic functions with memory barriers for vm fault info
  ext4: avoid potential buffer over-read in parse_apply_sb_mount_options()
  spi: cadence-quadspi: Flush posted register writes before DAC access
  spi: cadence-quadspi: Flush posted register writes before INDAC access
  memory: samsung: exynos-srom: Fix of_iomap leak in exynos_srom_probe
  memory: samsung: exynos-srom: Correct alignment
  arm64: errata: Apply workarounds for Neoverse-V3AE
  arm64: cputype: Add Neoverse-V3AE definitions
  comedi: fix divide-by-zero in comedi_buf_munge()
  binder: remove "invalid inc weak" check
  xhci: dbc: enable back DbC in resume if it was enabled before suspend
  usb/core/quirks: Add Huawei ME906S to wakeup quirk
  USB: serial: option: add Telit FN920C04 ECM compositions
  USB: serial: option: add Quectel RG255C
  USB: serial: option: add UNISOC UIS7720
  net: ravb: Ensure memory write completes before ringing TX doorbell
  net: usb: rtl8150: Fix frame padding
  ocfs2: clear extent cache after moving/defragmenting extents
  MIPS: Malta: Fix keyboard resource preventing i8042 driver from registering
  Revert "cpuidle: menu: Avoid discarding useful information"
  net: bonding: fix possible peer notify event loss or dup issue
  sctp: avoid NULL dereference when chunk data buffer is missing
  arm64, mm: avoid always making PTE dirty in pte_mkwrite()
  net: enetc: correct the value of ENETC_RXB_TRUESIZE
  rtnetlink: Allow deleting FDB entries in user namespace
  net: rtnetlink: add NLM_F_BULK support to rtnl_fdb_del
  net: add ndo_fdb_del_bulk
  net: rtnetlink: add bulk delete support flag
  net: netlink: add NLM_F_BULK delete request modifier
  net: rtnetlink: use BIT for flag values
  net: rtnetlink: add helper to extract msg type's kind
  net: rtnetlink: add msg kind names
  net: rtnetlink: remove redundant assignment to variable err
  m68k: bitops: Fix find_*_bit() signatures
  hfsplus: return EIO when type of hidden directory mismatch in hfsplus_fill_super()
  hfs: fix KMSAN uninit-value issue in hfs_find_set_zero_bits()
  dlm: check for defined force value in dlm_lockspace_release
  hfsplus: fix KMSAN uninit-value issue in hfsplus_delete_cat()
  hfs: validate record offset in hfsplus_bmap_alloc
  hfsplus: fix KMSAN uninit-value issue in __hfsplus_ext_cache_extent()
  hfs: make proper initalization of struct hfs_find_data
  hfs: clear offset and space out of valid records in b-tree node
  exec: Fix incorrect type for ret
  hfsplus: fix slab-out-of-bounds read in hfsplus_strcasecmp()
  ALSA: firewire: amdtp-stream: fix enum kernel-doc warnings
  sched/fair: Fix pelt lost idle time detection
  sched/balancing: Rename newidle_balance() => sched_balance_newidle()
  sched/fair: Trivial correction of the newidle_balance() comment
  sched: Make newidle_balance() static again
  tls: don't rely on tx_work during send()
  tls: always set record_type in tls_process_cmsg
  tg3: prevent use of uninitialized remote_adv and local_adv variables
  tcp: fix tcp_tso_should_defer() vs large RTT
  amd-xgbe: Avoid spurious link down messages during interface toggle
  net/ip6_tunnel: Prevent perpetual tunnel growth
  net: dlink: handle dma_map_single() failure properly
  net: dl2k: switch from 'pci_' to 'dma_' API
  media: pci: ivtv: Add missing check after DMA map
  media: pci/ivtv: switch from 'pci_' to 'dma_' API
  xen/events: Update virq_to_irq on migration
  media: lirc: Fix error handling in lirc_register()
  media: rc: Directly use ida_free()
  drm/exynos: exynos7_drm_decon: remove ctx->suspended
  btrfs: avoid potential out-of-bounds in btrfs_encode_fh()
  pwm: berlin: Fix wrong register in suspend/resume
  media: cx18: Add missing check after DMA map
  xen/events: Cleanup find_virq() return codes
  cramfs: Verify inode mode when loading from disk
  fs: Add 'initramfs_options' to set initramfs mount options
  pid: Add a judgment for ns null in pid_nr_ns
  minixfs: Verify inode mode when loading from disk
  tracing: Fix race condition in kprobe initialization causing NULL pointer dereference
  dm: fix NULL pointer dereference in __dm_suspend()
  mfd: intel_soc_pmic_chtdc_ti: Set use_single_read regmap_config flag
  mfd: intel_soc_pmic_chtdc_ti: Drop unneeded assignment for cache_type
  mfd: intel_soc_pmic_chtdc_ti: Fix invalid regmap-config max_register value
  Squashfs: reject negative file sizes in squashfs_read_inode()
  Squashfs: add additional inode sanity checking
  media: mc: Clear minor number before put device
  mfd: vexpress-sysreg: Check the return value of devm_gpiochip_add_data()
  fs: udf: fix OOB read in lengthAllocDescs handling
  KVM: x86: Don't (re)check L1 intercepts when completing userspace I/O
  net/9p: fix double req put in p9_fd_cancelled
  ext4: guard against EA inode refcount underflow in xattr update
  ext4: correctly handle queries for metadata mappings
  ext4: increase i_disksize to offset + len in ext4_update_disksize_before_punch()
  nfsd: nfserr_jukebox in nlm_fopen should lead to a retry
  x86/umip: Fix decoding of register forms of 0F 01 (SGDT and SIDT aliases)
  x86/umip: Check that the instruction opcode is at least two bytes
  PCI: keystone: Use devm_request_irq() to free "ks-pcie-error-irq" on exit
  PCI/AER: Fix missing uevent on recovery when a reset is requested
  PCI/IOV: Add PCI rescan-remove locking when enabling/disabling SR-IOV
  rseq/selftests: Use weak symbol reference, not definition, to link with glibc
  rtc: interface: Fix long-standing race when setting alarm
  rtc: interface: Ensure alarm irq is enabled when UIE is enabled
  mmc: core: SPI mode remove cmd7
  mtd: rawnand: fsmc: Default to autodetect buswidth
  sparc: fix error handling in scan_one_device()
  sparc64: fix hugetlb for sun4u
  sctp: Fix MAC comparison to be constant-time
  scsi: hpsa: Fix potential memory leak in hpsa_big_passthru_ioctl()
  parisc: don't reference obsolete termio struct for TC* constants
  lib/genalloc: fix device leak in of_gen_pool_get()
  iio: frequency: adf4350: Fix prescaler usage.
  iio: dac: ad5421: use int type to store negative error codes
  iio: dac: ad5360: use int type to store negative error codes
  crypto: atmel - Fix dma_unmap_sg() direction
  cpufreq: intel_pstate: Fix object lifecycle issue in update_qos_request()
  drm/nouveau: fix bad ret code in nouveau_bo_move_prep
  media: i2c: mt9v111: fix incorrect type for ret
  firmware: meson_sm: fix device leak at probe
  xen/manage: Fix suspend error path
  arm64: dts: qcom: msm8916: Add missing MDSS reset
  ACPI: debug: fix signedness issues in read/write helpers
  ACPI: TAD: Add missing sysfs_remove_group() for ACPI_TAD_RT
  tpm_tis: Fix incorrect arguments in tpm_tis_probe_irq_single
  tpm, tpm_tis: Claim locality before writing interrupt registers
  crypto: essiv - Check ssize for decryption and in-place encryption
  mailbox: zynqmp-ipi: Remove dev.parent check in zynqmp_ipi_free_mboxes
  mailbox: zynqmp-ipi: Remove redundant mbox_controller_unregister() call
  tools build: Align warning options with perf
  net: fsl_pq_mdio: Fix device node reference leak in fsl_pq_mdio_probe
  tcp: Don't call reqsk_fastopen_remove() in tcp_conn_request().
  net/sctp: fix a null dereference in sctp_disposition sctp_sf_do_5_1D_ce()
  drm/vmwgfx: Fix Use-after-free in validation
  net/mlx4: prevent potential use after free in mlx4_en_do_uc_filter()
  scsi: mvsas: Fix use-after-free bugs in mvs_work_queue
  scsi: mvsas: Use sas_task_find_rq() for tagging
  scsi: mvsas: Delete mvs_tag_init()
  scsi: libsas: Add sas_task_find_rq()
  clk: nxp: Fix pll0 rate check condition in LPC18xx CGU driver
  clk: nxp: lpc18xx-cgu: convert from round_rate() to determine_rate()
  perf session: Fix handling when buffer exceeds 2 GiB
  rtc: x1205: Fix Xicor X1205 vendor prefix
  perf util: Fix compression checks returning -1 as bool
  iio: frequency: adf4350: Fix ADF4350_REG3_12BIT_CLKDIV_MODE
  clocksource/drivers/clps711x: Fix resource leaks in error paths
  pinctrl: check the return value of pinmux_ops::get_function_name()
  Input: uinput - zero-initialize uinput_ff_upload_compat to avoid info leak
  mm: hugetlb: avoid soft lockup when mprotect to large memory area
  uio_hv_generic: Let userspace take care of interrupt mask
  Squashfs: fix uninit-value in squashfs_get_parent
  Revert "net/mlx5e: Update and set Xon/Xoff upon MTU set"
  net: ena: return 0 in ena_get_rxfh_key_size() when RSS hash key is not configurable
  nfp: fix RSS hash key size when RSS is not supported
  drivers/base/node: fix double free in register_one_node()
  ocfs2: fix double free in user_cluster_connect()
  net: usb: Remove disruptive netif_wake_queue in rtl8150_set_multicast
  RDMA/siw: Always report immediate post SQ errors
  usb: vhci-hcd: Prevent suspending virtually attached devices
  scsi: mpt3sas: Fix crash in transport port remove by using ioc_info()
  ipvs: Defer ip_vs_ftp unregister during netns cleanup
  NFSv4.1: fix backchannel max_resp_sz verification check
  remoteproc: qcom: q6v5: Avoid disabling handover IRQ twice
  sparc: fix accurate exception reporting in copy_{from,to}_user for M7
  sparc: fix accurate exception reporting in copy_to_user for Niagara 4
  sparc: fix accurate exception reporting in copy_{from_to}_user for Niagara
  sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC III
  sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC
  IB/sa: Fix sa_local_svc_timeout_ms read race
  RDMA/core: Resolve MAC of next-hop device without ARP support
  wifi: mt76: fix potential memory leak in mt76_wmac_probe()
  drivers/base/node: handle error properly in register_one_node()
  watchdog: mpc8xxx_wdt: Reload the watchdog timer when enabling the watchdog
  netfilter: ipset: Remove unused htable_bits in macro ahash_region
  iio: consumers: Fix offset handling in iio_convert_raw_to_processed()
  ASoC: Intel: bytcr_rt5651: Fix invalid quirk input mapping
  ASoC: Intel: bytcr_rt5640: Fix invalid quirk input mapping
  ASoC: Intel: bytcht_es8316: Fix invalid quirk input mapping
  pps: fix warning in pps_register_cdev when register device fail
  misc: genwqe: Fix incorrect cmd field being reported in error
  usb: gadget: configfs: Correctly set use_os_string at bind
  usb: phy: twl6030: Fix incorrect type for ret
  tcp: fix __tcp_close() to only send RST when required
  PCI: tegra: Fix devm_kcalloc() argument order for port->phys allocation
  wifi: mwifiex: send world regulatory domain to driver
  ALSA: lx_core: use int type to store negative error codes
  media: rj54n1cb0c: Fix memleak in rj54n1_probe()
  scsi: myrs: Fix dma_alloc_coherent() error check
  scsi: pm80xx: Fix array-index-out-of-of-bounds on rmmod
  serial: max310x: Add error checking in probe()
  usb: host: max3421-hcd: Fix error pointer dereference in probe cleanup
  drm/radeon/r600_cs: clean up of dead code in r600_cs
  i2c: designware: Add disabling clocks when probe fails
  i2c: mediatek: fix potential incorrect use of I2C_MASTER_WRRD
  bpf: Explicitly check accesses to bpf_sock_addr
  selftests: watchdog: skip ping loop if WDIOF_KEEPALIVEPING not supported
  pwm: tiehrpwm: Fix corner case in clock divisor calculation
  block: use int to store blk_stack_limits() return value
  blk-mq: check kobject state_in_sysfs before deleting in blk_mq_unregister_hctx
  pinctrl: meson-gxl: add missing i2c_d pinmux
  soc: qcom: rpmh-rsc: Unconditionally clear _TRIGGER bit for TCS
  ACPI: processor: idle: Fix memory leak when register cpuidle device failed
  regmap: Remove superfluous check for !config in __regmap_init()
  x86/vdso: Fix output operand size of RDPID
  perf: arm_spe: Prevent overflow in PERF_IDX2OFF()
  driver core/PM: Set power.no_callbacks along with power.no_pm
  staging: axis-fifo: flush RX FIFO on read errors
  staging: axis-fifo: fix maximum TX packet length check
  perf subcmd: avoid crash in exclude_cmds when excludes is empty
  dm-integrity: limit MAX_TAG_SIZE to 255
  wifi: rtlwifi: rtl8192cu: Don't claim USB ID 07b8:8188
  USB: serial: option: add SIMCom 8230C compositions
  media: rc: fix races with imon_disconnect()
  media: imon: grab lock earlier in imon_ir_change_protocol()
  media: imon: reorganize serialization
  media: rc: Add support for another iMON 0xffdc device
  media: i2c: tc358743: Fix use-after-free bugs caused by orphan timer in probe
  media: tuner: xc5000: Fix use-after-free in xc5000_release
  media: tunner: xc5000: Refactor firmware load
  udp: Fix memory accounting leak.
  media: b2c2: Fix use-after-free causing by irq_check_work in flexcop_pci_remove
  scsi: target: target_core_configfs: Add length check to avoid buffer overflow

 Conflicts:
	drivers/soc/qcom/rpmh-rsc.c
	kernel/sched/fair.c

Change-Id: I58ab24a3db8be4c698c41fd47daeb1f1fb7884ee
2025-12-04 19:21:35 +02:00
Greg Kroah-Hartman
ca00e0f525 This is the 5.4.302 stable release
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEZH8oZUiU471FcZm+ONu9yGCSaT4FAmkwI4sACgkQONu9yGCS
 aT46Mg//f9a0IiDkO2ybqt7JAStVCkQ5MM2CgPjlgHGyns6hWyxUES5twrVrTO0v
 mdYmeXLztyFSArUHMnoWcUK1O4IVUVK32SX3eEMFy81ojX+LpYm/m5TZg3tU1rvq
 jaTE0i6ihmwG48ciB63i28TxQfhY8QuVJTEV400Ro+ILY2hs1l6c6DYf9i0S/v4g
 gKDQpzuwR0AnGNFI+D6D6D0D8jbLVcBYnAndyvDrLYTIILczf7nJ66ZePYTBvlg3
 rIiIMjG0BX0V2ctPmez3mz0BDTpnZY4pwIwIG8K/bX4UZrgOKJPSZWh1SzbASaGN
 vTJQwRs9nGHvM/kK9CUUsi1+hhaAl7UmQEiRJ06BkXlR7chFCcL/fI/VEQjf12wL
 dyw6/RR/rXPxLbLxCYk+9C9ANTE/ByirLKSeJkNT/yoeiCSpaFuKxtoYkNslBK/B
 i0/Qweez0IzJuVUtUp/2/PrkNKyCaEbbqhDerbnLhU/0lcnT0fkFlFoLLLNzxNdt
 hj6RRJrcpoQB+Qrkf0+5Sxx6W1feP5clZJ2nRSL+rtUNEzCWk7HVh4GprCyvDn6i
 xjINN9Yh40suJkuKaE0+IrPF1tcL3/128OaT3e0VlL57wK5M4YmJTwrb7g4mlR2/
 31Z4bL1IFYju9WYPW2fWHrdrZUH4h22vLcpW6Q/VFdG0wTC6wVc=
 =Wt2i
 -----END PGP SIGNATURE-----

Merge 5.4.302 into android11-5.4-lts

Changes in 5.4.302
	net/sched: sch_qfq: Fix null-deref in agg_dequeue
	x86/bugs: Fix reporting of LFENCE retpoline
	btrfs: use smp_mb__after_atomic() when forcing COW in create_pending_snapshot()
	net: usb: asix_devices: Check return value of usbnet_get_endpoints
	fbdev: atyfb: Check if pll_ops->init_pll failed
	ACPI: video: Fix use-after-free in acpi_video_switch_brightness()
	fbdev: bitblit: bound-check glyph index in bit_putcs*
	fbdev: pvr2fb: Fix leftover reference to ONCHIP_NR_DMA_CHANNELS
	fbdev: valkyriefb: Fix reference count leak in valkyriefb_init
	ASoC: qdsp6: q6asm: do not sleep while atomic
	wifi: ath10k: Fix memory leak on unsupported WMI command
	usbnet: Prevents free active kevent
	drm/etnaviv: fix flush sequence logic
	regmap: slimbus: fix bus_context pointer in regmap init calls
	net: phy: dp83867: Disable EEE support as not implemented
	wifi: brcmfmac: fix crash while sending Action Frames in standalone AP Mode
	x86/resctrl: Fix miscount of bandwidth event when reactivating previously unavailable RMID
	net: ravb: Enforce descriptor type ordering
	devcoredump: Fix circular locking dependency with devcd->mutex.
	can: gs_usb: increase max interface to U8_MAX
	serial: 8250_dw: Use devm_clk_get_optional() to get the input clock
	serial: 8250_dw: Use devm_add_action_or_reset()
	serial: 8250_dw: handle reset control deassert error
	usb: gadget: f_fs: Fix epfile null pointer access after ep enable.
	soc: qcom: smem: Fix endian-unaware access of num_entries
	spi: loopback-test: Don't use %pK through printk
	bpf: Don't use %pK through printk
	mmc: host: renesas_sdhi: Fix the actual clock
	memstick: Add timeout to prevent indefinite waiting
	ACPI: video: force native for Lenovo 82K8
	selftests/bpf: Fix bpf_prog_detach2 usage in test_lirc_mode2
	cpufreq/longhaul: handle NULL policy in longhaul_exit
	arc: Fix __fls() const-foldability via __builtin_clzl()
	irqchip/gic-v2m: Handle Multiple MSI base IRQ Alignment
	mmc: sdhci-msm: Enable tuning for SDR50 mode for SD card
	ACPICA: dispatcher: Use acpi_ds_clear_operands() in acpi_ds_call_control_method()
	tee: allow a driver to allocate a tee_device without a pool
	video: backlight: lp855x_bl: Set correct EPROM start for LP8556
	clocksource/drivers/vf-pit: Replace raw_readl/writel to readl/writel
	uprobe: Do not emulate/sstep original instruction when ip is changed
	hwmon: (dell-smm) Add support for Dell OptiPlex 7040
	tools/cpupower: Fix incorrect size in cpuidle_state_disable()
	tools/power x86_energy_perf_policy: Enhance HWP enable
	tools/power x86_energy_perf_policy: Prefer driver HWP limits
	mfd: stmpe: Remove IRQ domain upon removal
	mfd: stmpe-i2c: Add missing MODULE_LICENSE
	mfd: madera: Work around false-positive -Wininitialized warning
	drm/nouveau: replace snprintf() with scnprintf() in nvkm_snprintbf()
	PCI: Disable MSI on RDC PCI to PCIe bridges
	selftests/net: Replace non-standard __WORDSIZE with sizeof(long) * 8
	selftests/net: Ensure assert() triggers in psock_tpacket.c
	drm/amdkfd: return -ENOTTY for unsupported IOCTLs
	media: pci: ivtv: Don't create fake v4l2_fh
	x86/vsyscall: Do not require X86_PF_INSTR to emulate vsyscall
	powerpc/eeh: Use result of error_detected() in uevent
	bridge: Redirect to backup port when port is administratively down
	net: ipv6: fix field-spanning memcpy warning in AH output
	media: imon: make send_packet() more robust
	iio: adc: spear_adc: mask SPEAR_ADC_STATUS channel and avg sample before setting register
	usb: gadget: f_ncm: Fix MAC assignment NCM ethernet
	char: misc: Does not request module for miscdevice with dynamic minor
	net: When removing nexthops, don't call synchronize_net if it is not necessary
	net: Call trace_sock_exceed_buf_limit() for memcg failure with SK_MEM_RECV.
	PCI/P2PDMA: Fix incorrect pointer usage in devm_kfree() call
	rds: Fix endianness annotation for RDS_MPATH_HASH
	extcon: adc-jack: Fix wakeup source leaks on device unbind
	drm/amdkfd: Tie UNMAP_LATENCY to queue_preemption
	media: fix uninitialized symbol warnings
	mips: lantiq: danube: add missing properties to cpu node
	mips: lantiq: danube: add missing device_type in pci node
	mips: lantiq: xway: sysctrl: rename stp clock
	scsi: pm8001: Use int instead of u32 to store error codes
	dmaengine: sh: setup_xref error handling
	dmaengine: mv_xor: match alloc_wc and free_wc
	dmaengine: dw-edma: Set status for callback_result
	net: nfc: nci: Increase NCI_DATA_TIMEOUT to 3000 ms
	ALSA: usb-audio: apply quirk for MOONDROP Quark2
	net: call cond_resched() less often in __release_sock()
	usb: gadget: f_hid: Fix zero length packet transfer
	phy: cadence: cdns-dphy: Enable lower resolutions in dphy
	net: sh_eth: Disable WoL if system can not suspend
	media: redrat3: use int type to store negative error codes
	selftests: Disable dad for ipv6 in fcnal-test.sh
	selftests: Replace sleep with slowwait
	net/cls_cgroup: Fix task_get_classid() during qdisc run
	selftests/Makefile: include $(INSTALL_DEP_TARGETS) in clean target to clean net/lib dependency
	scsi: lpfc: Check return status of lpfc_reset_flush_io_context during TGT_RESET
	scsi: lpfc: Define size of debugfs entry for xri rebalancing
	allow finish_no_open(file, ERR_PTR(-E...))
	usb: mon: Increase BUFF_MAX to 64 MiB to support multi-MB URBs
	usb: xhci: plat: Facilitate using autosuspend for xhci plat devices
	ipv6: np->rxpmtu race annotation
	jfs: Verify inode mode when loading from disk
	jfs: fix uninitialized waitqueue in transaction manager
	net: intel: fm10k: Fix parameter idx set but not used
	sparc/module: Add R_SPARC_UA64 relocation handling
	remoteproc: qcom: q6v5: Avoid handling handover twice
	NFSv4: handle ERR_GRACE on delegation recalls
	NFSv4.1: fix mount hang after CREATE_SESSION failure
	nfs4_setup_readdir(): insufficient locking for ->d_parent->d_inode dereferencing
	net: macb: avoid dealing with endianness in macb_set_hwaddr()
	Bluetooth: SCO: Fix UAF on sco_conn_free
	Bluetooth: bcsp: receive data only if registered
	page_pool: Clamp pool size to max 16K pages
	orangefs: fix xattr related buffer overflow...
	ACPICA: Update dsmethod.c to get rid of unused variable warning
	fs/hpfs: Fix error code for new_inode() failure in mkdir/create/mknod/symlink
	9p: fix /sys/fs/9p/caches overwriting itself
	9p: sysfs_init: don't hardcode error to ENOMEM
	ACPI: property: Return present device nodes only on fwnode interface
	fbdev: Add bounds checking in bit_putcs to fix vmalloc-out-of-bounds
	ceph: add checking of wait_for_completion_killable() return value
	net: vlan: sync VLAN features with lower device
	net: dsa/b53: change b53_force_port_config() pause argument
	net: dsa: b53: prevent GMII_PORT_OVERRIDE_CTRL access on BCM5325
	net: dsa: b53: fix resetting speed and pause on forced link
	net: dsa: b53: fix enabling ip multicast
	net: dsa: b53: stop reading ARL entries if search is done
	sctp: Hold RCU read lock while iterating over address list
	sctp: Prevent TOCTOU out-of-bounds write
	net: usb: qmi_wwan: initialize MAC header offset in qmimux_rx_fixup
	tracing: Fix memory leaks in create_field_var()
	extcon: adc-jack: Cleanup wakeup source only if it was enabled
	compiler_types: Move unused static inline functions warning to W=2
	NFS4: Fix state renewals missing after boot
	HID: quirks: avoid Cooler Master MM712 dongle wakeup bug
	ASoC: max98090/91: fixed max98091 ALSA widget powering up/down
	net: fec: correct rx_bytes statistic for the case SHIFT16 is set
	Bluetooth: btusb: reorder cleanup in btusb_disconnect to avoid UAF
	Bluetooth: 6lowpan: reset link-local header on ipv6 recv path
	Bluetooth: 6lowpan: fix BDADDR_LE vs ADDR_LE_DEV address type confusion
	Bluetooth: 6lowpan: Don't hold spin lock over sleeping functions
	sctp: get netns from asoc and ep base
	sctp: prevent possible shift-out-of-bounds in sctp_transport_update_rto
	tipc: simplify the finalize work queue
	tipc: Fix use-after-free in tipc_mon_reinit_self().
	net: mdio: fix resource leak in mdiobus_register_device()
	wifi: mac80211: skip rate verification for not captured PSDUs
	net: sched: act_ife: initialize struct tc_ife to fix KMSAN kernel-infoleak
	net/mlx5e: Fix maxrate wraparound in threshold between units
	net/mlx5e: Fix wraparound in rate limiting for values above 255 Gbps
	net_sched: remove need_resched() from qdisc_run()
	net_sched: limit try_bulk_dequeue_skb() batches
	Bluetooth: L2CAP: export l2cap_chan_hold for modules
	regulator: fixed: use dev_err_probe for register
	regulator: fixed: fix GPIO descriptor leak on register failure
	ASoC: cs4271: Fix regulator leak on probe failure
	drm/vmwgfx: Validate command header size against SVGA_CMD_MAX_DATASIZE
	ALSA: usb-audio: Fix NULL pointer dereference in snd_usb_mixer_controls_badd
	mm/ksm: fix flag-dropping behavior in ksm_madvise
	gcov: add support for GCC 15
	strparser: Fix signed/unsigned mismatch bug
	ipv4: route: Prevent rt_bind_exception() from rebinding stale fnhe
	spi: Try to get ACPI GPIO IRQ earlier
	EDAC/altera: Handle OCRAM ECC enable after warm reset
	EDAC/altera: Use INTTEST register for Ethernet and USB SBE injection
	isdn: mISDN: hfcsusb: fix memory leak in hfcsusb_probe()
	HID: quirks: work around VID/PID conflict for 0x4c4a/0x4155
	be2net: pass wrb_params in case of OS2BMC
	Input: cros_ec_keyb - fix an invalid memory access
	scsi: sg: Do not sleep in atomic context
	scsi: target: tcm_loop: Fix segfault in tcm_loop_tpg_address_show()
	MIPS: Malta: Fix !EVA SOC-it PCI MMIO
	mlxsw: spectrum: Fix memory leak in mlxsw_sp_flower_stats()
	net: openvswitch: remove never-working support for setting nsh fields
	s390/ctcm: Fix double-kfree
	vsock: Ignore signal/timeout on connect() if already established
	kconfig/mconf: Initialize the default locale at startup
	kconfig/nconf: Initialize the default locale at startup
	mm/page_alloc: fix hash table order logging in alloc_large_system_hash()
	ALSA: usb-audio: fix uac2 clock source at terminal parser
	net: ethernet: ti: netcp: Standardize knav_dma_open_channel to return NULL on error
	uio_hv_generic: Set event for all channels on the device
	net: qede: Initialize qede_ll_ops with designated initializer
	net: netpoll: fix incorrect refcount handling causing incorrect cleanup
	pmdomain: arm: scmi: Fix genpd leak on provider registration failure
	pmdomain: imx: Fix reference count leak in imx_gpc_remove
	fs/proc: fix uaf in proc_readdir_de()
	ata: libata-scsi: Fix system suspend for a security locked drive
	usb: deprecate the third argument of usb_maxpacket()
	Input: remove third argument of usb_maxpacket()
	Input: pegasus-notetaker - fix potential out-of-bounds access
	Linux 5.4.302

Change-Id: I7291d845c3cfde8a154957356156fadcc4b96b80
Signed-off-by: Greg Kroah-Hartman <gregkh@google.com>
2025-12-03 14:36:44 +00:00
Greg Kroah-Hartman
9e3157c56e Linux 5.4.302
Last release of the 5.4.y branch.

This branch is now end-of-life, do not use anymore, please move to a
newer kernel release.  As of this point in time, there are over 1500
known unfixed CVEs for this branch, and that number will only increase
over time.

Link: https://lore.kernel.org/r/20251201112241.242614045@linuxfoundation.org
Tested-by: Brett A C Sheffield <bacs@librecast.net>
Tested-by: Florian Fainelli <florian.fainelli@broadcom.com>
Tested-by: Slade Watkins <sr@sladewatkins.com>
Tested-by: Shuah Khan <skhan@linuxfoundation.org>
Link: https://lore.kernel.org/r/20251202095448.089783651@linuxfoundation.org
Tested-by: Brett A C Sheffield <bacs@librecast.net>
Tested-by: Jon Hunter <jonathanh@nvidia.com>
Tested-by: Alok Tiwari <alok.a.tiwari@oracle.com>
Link: https://lore.kernel.org/r/20251202152903.637577865@linuxfoundation.org
Tested-by: Brett A C Sheffield <bacs@librecast.net>
Tested-by: Jon Hunter <jonathanh@nvidia.com>
Tested-by: Florian Fainelli <florian.fainelli@broadcom.com>
Tested-by: Linux Kernel Functional Testing <lkft@linaro.org>
Tested-by: Pavel Machek (CIP) <pavel@denx.de>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:22 +01:00
Seungjin Bae
c4e746651b Input: pegasus-notetaker - fix potential out-of-bounds access
[ Upstream commit 69aeb507312306f73495598a055293fa749d454e ]

In the pegasus_notetaker driver, the pegasus_probe() function allocates
the URB transfer buffer using the wMaxPacketSize value from
the endpoint descriptor. An attacker can use a malicious USB descriptor
to force the allocation of a very small buffer.

Subsequently, if the device sends an interrupt packet with a specific
pattern (e.g., where the first byte is 0x80 or 0x42),
the pegasus_parse_packet() function parses the packet without checking
the allocated buffer size. This leads to an out-of-bounds memory access.

Fixes: 1afca2b66a ("Input: add Pegasus Notetaker tablet driver")
Signed-off-by: Seungjin Bae <eeodqql09@gmail.com>
Link: https://lore.kernel.org/r/20251007214131.3737115-2-eeodqql09@gmail.com
Cc: stable@vger.kernel.org
Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:22 +01:00
Vincent Mailhol
a643fecbca Input: remove third argument of usb_maxpacket()
[ Upstream commit 948bf187694fc1f4c20cf972fa18b1a6fb3d7603 ]

The third argument of usb_maxpacket(): in_out has been deprecated
because it could be derived from the second argument (e.g. using
usb_pipeout(pipe)).

N.B. function usb_maxpacket() was made variadic to accommodate the
transition from the old prototype with three arguments to the new one
with only two arguments (so that no renaming is needed). The variadic
argument is to be removed once all users of usb_maxpacket() get
migrated.

CC: Ville Syrjala <syrjala@sci.fi>
CC: Dmitry Torokhov <dmitry.torokhov@gmail.com>
CC: Henk Vergonet <Henk.Vergonet@gmail.com>
Signed-off-by: Vincent Mailhol <mailhol.vincent@wanadoo.fr>
Link: https://lore.kernel.org/r/20220317035514.6378-4-mailhol.vincent@wanadoo.fr
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Stable-dep-of: 69aeb5073123 ("Input: pegasus-notetaker - fix potential out-of-bounds access")
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:22 +01:00
Vincent Mailhol
2c6d503287 usb: deprecate the third argument of usb_maxpacket()
[ Upstream commit 0f08c2e7458e25c967d844170f8ad1aac3b57a02 ]

This is a transitional patch with the ultimate goal of changing the
prototype of usb_maxpacket() from:
| static inline __u16
| usb_maxpacket(struct usb_device *udev, int pipe, int is_out)

into:
| static inline u16 usb_maxpacket(struct usb_device *udev, int pipe)

The third argument of usb_maxpacket(): is_out gets removed because it
can be derived from its second argument: pipe using
usb_pipeout(pipe). Furthermore, in the current version,
ubs_pipeout(pipe) is called regardless in order to sanitize the is_out
parameter.

In order to make a smooth change, we first deprecate the is_out
parameter by simply ignoring it (using a variadic function) and will
remove it later, once all the callers get updated.

The body of the function is reworked accordingly and is_out is
replaced by usb_pipeout(pipe). The WARN_ON() calls become unnecessary
and get removed.

Finally, the return type is changed from __u16 to u16 because this is
not a UAPI function.

Signed-off-by: Vincent Mailhol <mailhol.vincent@wanadoo.fr>
Link: https://lore.kernel.org/r/20220317035514.6378-2-mailhol.vincent@wanadoo.fr
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Stable-dep-of: 69aeb5073123 ("Input: pegasus-notetaker - fix potential out-of-bounds access")
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:21 +01:00
Niklas Cassel
ce9e6c5f2a ata: libata-scsi: Fix system suspend for a security locked drive
[ Upstream commit b11890683380a36b8488229f818d5e76e8204587 ]

Commit cf3fc037623c ("ata: libata-scsi: Fix ata_to_sense_error() status
handling") fixed ata_to_sense_error() to properly generate sense key
ABORTED COMMAND (without any additional sense code), instead of the
previous bogus sense key ILLEGAL REQUEST with the additional sense code
UNALIGNED WRITE COMMAND, for a failed command.

However, this broke suspend for Security locked drives (drives that have
Security enabled, and have not been Security unlocked by boot firmware).

The reason for this is that the SCSI disk driver, for the Synchronize
Cache command only, treats any sense data with sense key ILLEGAL REQUEST
as a successful command (regardless of ASC / ASCQ).

After commit cf3fc037623c ("ata: libata-scsi: Fix ata_to_sense_error()
status handling") the code that treats any sense data with sense key
ILLEGAL REQUEST as a successful command is no longer applicable, so the
command fails, which causes the system suspend to be aborted:

  sd 1:0:0:0: PM: dpm_run_callback(): scsi_bus_suspend returns -5
  sd 1:0:0:0: PM: failed to suspend async: error -5
  PM: Some devices failed to suspend, or early wake event detected

To make suspend work once again, for a Security locked device only,
return sense data LOGICAL UNIT ACCESS NOT AUTHORIZED, the actual sense
data which a real SCSI device would have returned if locked.
The SCSI disk driver treats this sense data as a successful command.

Cc: stable@vger.kernel.org
Reported-by: Ilia Baryshnikov <qwelias@gmail.com>
Closes: https://bugzilla.kernel.org/show_bug.cgi?id=220704
Fixes: cf3fc037623c ("ata: libata-scsi: Fix ata_to_sense_error() status handling")
Reviewed-by: Hannes Reinecke <hare@suse.de>
Reviewed-by: Martin K. Petersen <martin.petersen@oracle.com>
Reviewed-by: Damien Le Moal <dlemoal@kernel.org>
Signed-off-by: Niklas Cassel <cassel@kernel.org>
[ Adjust context ]
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:21 +01:00
Wei Yang
1d1596d68a fs/proc: fix uaf in proc_readdir_de()
[ Upstream commit 895b4c0c79b092d732544011c3cecaf7322c36a1 ]

Pde is erased from subdir rbtree through rb_erase(), but not set the node
to EMPTY, which may result in uaf access.  We should use RB_CLEAR_NODE()
set the erased node to EMPTY, then pde_subdir_next() will return NULL to
avoid uaf access.

We found an uaf issue while using stress-ng testing, need to run testcase
getdent and tun in the same time.  The steps of the issue is as follows:

1) use getdent to traverse dir /proc/pid/net/dev_snmp6/, and current
   pde is tun3;

2) in the [time windows] unregister netdevice tun3 and tun2, and erase
   them from rbtree.  erase tun3 first, and then erase tun2.  the
   pde(tun2) will be released to slab;

3) continue to getdent process, then pde_subdir_next() will return
   pde(tun2) which is released, it will case uaf access.

CPU 0                                      |    CPU 1
-------------------------------------------------------------------------
traverse dir /proc/pid/net/dev_snmp6/      |   unregister_netdevice(tun->dev)   //tun3 tun2
sys_getdents64()                           |
  iterate_dir()                            |
    proc_readdir()                         |
      proc_readdir_de()                    |     snmp6_unregister_dev()
        pde_get(de);                       |       proc_remove()
        read_unlock(&proc_subdir_lock);    |         remove_proc_subtree()
                                           |           write_lock(&proc_subdir_lock);
        [time window]                      |           rb_erase(&root->subdir_node, &parent->subdir);
                                           |           write_unlock(&proc_subdir_lock);
        read_lock(&proc_subdir_lock);      |
        next = pde_subdir_next(de);        |
        pde_put(de);                       |
        de = next;    //UAF                |

rbtree of dev_snmp6
                        |
                    pde(tun3)
                     /    \
                  NULL  pde(tun2)

Link: https://lkml.kernel.org/r/20251025024233.158363-1-albin_yang@163.com
Signed-off-by: Wei Yang <albinwyang@tencent.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Christian Brauner <brauner@kernel.org>
Cc: wangzijie <wangzijie1@honor.com>
Cc: Alexey Dobriyan <adobriyan@gmail.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:21 +01:00
Miaoqian Lin
e45dfc5368 pmdomain: imx: Fix reference count leak in imx_gpc_remove
[ Upstream commit bbde14682eba21d86f5f3d6fe2d371b1f97f1e61 ]

of_get_child_by_name() returns a node pointer with refcount incremented, we
should use of_node_put() on it when not needed anymore. Add the missing
of_node_put() to avoid refcount leak.

Fixes: 721cabf6c6 ("soc: imx: move PGC handling to a new GPC driver")
Cc: stable@vger.kernel.org
Signed-off-by: Miaoqian Lin <linmq006@gmail.com>
Signed-off-by: Ulf Hansson <ulf.hansson@linaro.org>
[ drivers/pmdomain/imx/gpc.c -> drivers/soc/imx/gpc.c ]
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:21 +01:00
Sudeep Holla
18249a167f pmdomain: arm: scmi: Fix genpd leak on provider registration failure
[ Upstream commit 7458f72cc28f9eb0de811effcb5376d0ec19094a ]

If of_genpd_add_provider_onecell() fails during probe, the previously
created generic power domains are not removed, leading to a memory leak
and potential kernel crash later in genpd_debug_add().

Add proper error handling to unwind the initialized domains before
returning from probe to ensure all resources are correctly released on
failure.

Example crash trace observed without this fix:

  | Unable to handle kernel paging request at virtual address fffffffffffffc70
  | CPU: 1 UID: 0 PID: 1 Comm: swapper/0 Not tainted 6.18.0-rc1 #405 PREEMPT
  | Hardware name: ARM LTD ARM Juno Development Platform/ARM Juno Development Platform
  | pstate: 00000005 (nzcv daif -PAN -UAO -TCO -DIT -SSBS BTYPE=--)
  | pc : genpd_debug_add+0x2c/0x160
  | lr : genpd_debug_init+0x74/0x98
  | Call trace:
  |  genpd_debug_add+0x2c/0x160 (P)
  |  genpd_debug_init+0x74/0x98
  |  do_one_initcall+0xd0/0x2d8
  |  do_initcall_level+0xa0/0x140
  |  do_initcalls+0x60/0xa8
  |  do_basic_setup+0x28/0x40
  |  kernel_init_freeable+0xe8/0x170
  |  kernel_init+0x2c/0x140
  |  ret_from_fork+0x10/0x20

Fixes: 898216c97e ("firmware: arm_scmi: add device power domain support using genpd")
Signed-off-by: Sudeep Holla <sudeep.holla@arm.com>
Reviewed-by: Peng Fan <peng.fan@nxp.com>
Cc: stable@vger.kernel.org
Signed-off-by: Ulf Hansson <ulf.hansson@linaro.org>
[ drivers/pmdomain/arm/scmi_pm_domain.c -> drivers/firmware/arm_scmi/scmi_pm_domain.c ]
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2025-12-03 12:45:21 +01:00