This adds a test that changes its UID, uses capabilities to
get CAP_CHECKPOINT_RESTORE and uses clone3() with set_tid to
create a process with a given PID as non-root.
Change-Id: Iac62e21213ad1bd04f5a2762255fc924c3830ea6
Signed-off-by: Adrian Reber <areber@redhat.com>
Link: https://lore.kernel.org/r/20200719100418.2112740-8-areber@redhat.com
[christian.brauner@ubuntu.com: use TH_LOG() instead of ksft_print_msg()]
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
This tests clone3() with *set_tid to see if all desired PIDs are working
as expected. The tests are trying multiple invalid input parameters as
well as creating processes while specifying a certain PID in multiple
PID namespaces at the same time.
Additionally this moves common clone3() test code into clone3_selftests.h.
Change-Id: Ic30e083c2b8c3e5c1fe9adff28387e84431f4c71
Signed-off-by: Adrian Reber <areber@redhat.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20191115123621.142252-2-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
This adds tests for clone3() with different values and sizes
of struct clone_args.
This selftest was initially part of of the clone3() with PID selftest.
After that patch was almost merged Eugene sent out a couple of patches
to fix problems with these test.
This commit now only contains the clone3() selftest after the LPC
decision to rework clone3() with PID to allow setting the PID in
multiple PID namespaces including all of Eugene's patches.
Change-Id: I32fecd920f7a820c0ca7869a70a8c5b34af583dd
Signed-off-by: Eugene Syromiatnikov <esyr@redhat.com>
Signed-off-by: Adrian Reber <areber@redhat.com>
Reviewed-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20191112095851.811884-1-areber@redhat.com
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Signed-off-by: ralph950412 <ralph950412@gmail.com>
Test that CLONE_CLEAR_SIGHAND resets signal handlers to SIG_DFL for the
child process and that CLONE_CLEAR_SIGHAND and CLONE_SIGHAND are
mutually exclusive.
Cc: Florian Weimer <fweimer@redhat.com>
Cc: libc-alpha@sourceware.org
Cc: linux-api@vger.kernel.org
Change-Id: I07ebb8cb2e741cf1a6a43ff4dbe6ee489636f506
Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20191014104538.3096-2-christian.brauner@ubuntu.com
Signed-off-by: ralph950412 <ralph950412@gmail.com>
Update the kselftest framework to allow client drivers to
specify that some tests were skipped.
Signed-off-by: Timur Tabi <timur@kernel.org>
Reviewed-by: Petr Mladek <pmladek@suse.com>
Tested-by: Petr Mladek <pmladek@suse.com>
Acked-by: Marco Elver <elver@google.com>
Signed-off-by: Petr Mladek <pmladek@suse.com>
Link: https://lore.kernel.org/r/20210214161348.369023-3-timur@kernel.org
Bug: 177201466
Bug: 180086542
(cherry picked from commit d9d4de2309cd1721421c6488f1bb5744d2c83a39)
Signed-off-by: Alexander Potapenko <glider@google.com>
Change-Id: Ifbfd5c75c207d9491dd8d7d5133ea30d12a85b12
As seccomp_benchmark tries to calibrate how many samples will take more
than 5 seconds to execute, it may end up picking up a number of samples
that take 10 (but up to 12) seconds. As the calibration will take double
that time, it takes around 20 seconds. Then, it executes the whole thing
again, and then once more, with some added overhead. So, the thing might
take more than 40 seconds, which is too close to the 45s timeout.
That is very dependent on the system where it's executed, so may not be
observed always, but it has been observed on x86 VMs. Using a 90s timeout
seems safe enough.
Signed-off-by: Thadeu Lima de Souza Cascardo <cascardo@canonical.com>
Link: https://lore.kernel.org/r/20200601123202.1183526-1-cascardo@canonical.com
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit bc32c9c86581abf7baacf71342df3b0affe367db)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I52379d62bddc9163c6cbbb934483ceabb10cc909
The seccomp benchmark calibration loop did not need to take so long.
Instead, use a simple 1 second timeout and multiply up to target. It
does not need to be accurate.
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit 81a0c8bc82be7c15dbf3e54832334552e6b76e2b)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I5583b28bb635c01413548dbe05b2bed8d50def00
It's useful to see how much (at a minimum) each filter adds to the
syscall overhead. Add additional calculations.
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit d3a37ea9f6e548388b83fe895c7a037bc2ec3f7f)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: Ibc36b007b580cc129b2880b12b151a659590abb9
commit 6715df8d5d24655b9fd368e904028112b54c7de1 upstream.
This commits updates the following functions to allow reads from
uninitialized stack locations when env->allow_uninit_stack option is
enabled:
- check_stack_read_fixed_off()
- check_stack_range_initialized(), called from:
- check_stack_read_var_off()
- check_helper_mem_access()
Such change allows to relax logic in stacksafe() to treat STACK_MISC
and STACK_INVALID in a same way and make the following stack slot
configurations equivalent:
| Cached state | Current state |
| stack slot | stack slot |
|------------------+------------------|
| STACK_INVALID or | STACK_INVALID or |
| STACK_MISC | STACK_SPILL or |
| | STACK_MISC or |
| | STACK_ZERO or |
| | STACK_DYNPTR |
This leads to significant verification speed gains (see below).
The idea was suggested by Andrii Nakryiko [1] and initial patch was
created by Alexei Starovoitov [2].
Currently the env->allow_uninit_stack is allowed for programs loaded
by users with CAP_PERFMON or CAP_SYS_ADMIN capabilities.
A number of test cases from verifier/*.c were expecting uninitialized
stack access to be an error. These test cases were updated to execute
in unprivileged mode (thus preserving the tests).
The test progs/test_global_func10.c expected "invalid indirect read
from stack" error message because of the access to uninitialized
memory region. This error is no longer possible in privileged mode.
The test is updated to provoke an error "invalid indirect access to
stack" because of access to invalid stack address (such error is not
verified by progs/test_global_func*.c series of tests).
The following tests had to be removed because these can't be made
unprivileged:
- verifier/sock.c:
- "sk_storage_get(map, skb->sk, &stack_value, 1): partially init
stack_value"
BPF_PROG_TYPE_SCHED_CLS programs are not executed in unprivileged mode.
- verifier/var_off.c:
- "indirect variable-offset stack access, max_off+size > max_initialized"
- "indirect variable-offset stack access, uninitialized"
These tests verify that access to uninitialized stack values is
detected when stack offset is not a constant. However, variable
stack access is prohibited in unprivileged mode, thus these tests
are no longer valid.
* * *
Here is veristat log comparing this patch with current master on a
set of selftest binaries listed in tools/testing/selftests/bpf/veristat.cfg
and cilium BPF binaries (see [3]):
$ ./veristat -e file,prog,states -C -f 'states_pct<-30' master.log current.log
File Program States (A) States (B) States (DIFF)
-------------------------- -------------------------- ---------- ---------- ----------------
bpf_host.o tail_handle_ipv6_from_host 349 244 -105 (-30.09%)
bpf_host.o tail_handle_nat_fwd_ipv4 1320 895 -425 (-32.20%)
bpf_lxc.o tail_handle_nat_fwd_ipv4 1320 895 -425 (-32.20%)
bpf_sock.o cil_sock4_connect 70 48 -22 (-31.43%)
bpf_sock.o cil_sock4_sendmsg 68 46 -22 (-32.35%)
bpf_xdp.o tail_handle_nat_fwd_ipv4 1554 803 -751 (-48.33%)
bpf_xdp.o tail_lb_ipv4 6457 2473 -3984 (-61.70%)
bpf_xdp.o tail_lb_ipv6 7249 3908 -3341 (-46.09%)
pyperf600_bpf_loop.bpf.o on_event 287 145 -142 (-49.48%)
strobemeta.bpf.o on_event 15915 4772 -11143 (-70.02%)
strobemeta_nounroll2.bpf.o on_event 17087 3820 -13267 (-77.64%)
xdp_synproxy_kern.bpf.o syncookie_tc 21271 6635 -14636 (-68.81%)
xdp_synproxy_kern.bpf.o syncookie_xdp 23122 6024 -17098 (-73.95%)
-------------------------- -------------------------- ---------- ---------- ----------------
Note: I limited selection by states_pct<-30%.
Inspection of differences in pyperf600_bpf_loop behavior shows that
the following patch for the test removes almost all differences:
- a/tools/testing/selftests/bpf/progs/pyperf.h
+ b/tools/testing/selftests/bpf/progs/pyperf.h
@ -266,8 +266,8 @ int __on_event(struct bpf_raw_tracepoint_args *ctx)
}
if (event->pthread_match || !pidData->use_tls) {
- void* frame_ptr;
- FrameData frame;
+ void* frame_ptr = 0;
+ FrameData frame = {};
Symbol sym = {};
int cur_cpu = bpf_get_smp_processor_id();
W/o this patch the difference comes from the following pattern
(for different variables):
static bool get_frame_data(... FrameData *frame ...)
{
...
bpf_probe_read_user(&frame->f_code, ...);
if (!frame->f_code)
return false;
...
bpf_probe_read_user(&frame->co_name, ...);
if (frame->co_name)
...;
}
int __on_event(struct bpf_raw_tracepoint_args *ctx)
{
FrameData frame;
...
get_frame_data(... &frame ...) // indirectly via a bpf_loop & callback
...
}
SEC("raw_tracepoint/kfree_skb")
int on_event(struct bpf_raw_tracepoint_args* ctx)
{
...
ret |= __on_event(ctx);
ret |= __on_event(ctx);
...
}
With regards to value `frame->co_name` the following is important:
- Because of the conditional `if (!frame->f_code)` each call to
__on_event() produces two states, one with `frame->co_name` marked
as STACK_MISC, another with it as is (and marked STACK_INVALID on a
first call).
- The call to bpf_probe_read_user() does not mark stack slots
corresponding to `&frame->co_name` as REG_LIVE_WRITTEN but it marks
these slots as BPF_MISC, this happens because of the following loop
in the check_helper_call():
for (i = 0; i < meta.access_size; i++) {
err = check_mem_access(env, insn_idx, meta.regno, i, BPF_B,
BPF_WRITE, -1, false);
if (err)
return err;
}
Note the size of the write, it is a one byte write for each byte
touched by a helper. The BPF_B write does not lead to write marks
for the target stack slot.
- Which means that w/o this patch when second __on_event() call is
verified `if (frame->co_name)` will propagate read marks first to a
stack slot with STACK_MISC marks and second to a stack slot with
STACK_INVALID marks and these states would be considered different.
[1] https://lore.kernel.org/bpf/CAEf4BzY3e+ZuC6HUa8dCiUovQRg2SzEk7M-dSkqNZyn=xEmnPA@mail.gmail.com/
[2] https://lore.kernel.org/bpf/CAADnVQKs2i1iuZ5SUGuJtxWVfGYR9kDgYKhq3rNV+kBLQCu7rA@mail.gmail.com/
[3] git@github.com:anakryiko/cilium.git
Suggested-by: Andrii Nakryiko <andrii@kernel.org>
Co-developed-by: Alexei Starovoitov <ast@kernel.org>
Change-Id: I80629cd663a28da7499da586c34adb97417233bb
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Acked-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/r/20230219200427.606541-2-eddyz87@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Maxim Mikityanskiy <maxim@isovalent.com>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
[ Upstream commit ecdf985d7615356b78241fdb159c091830ed0380 ]
For aligned stack writes using BPF_ST instruction track stored values
in a same way BPF_STX is handled, e.g. make sure that the following
commands produce similar verifier knowledge:
fp[-8] = 42; r1 = 42;
fp[-8] = r1;
This covers two cases:
- non-null values written to stack are stored as spill of fake
registers;
- null values written to stack are stored as STACK_ZERO marks.
Previously both cases above used STACK_MISC marks instead.
Some verifier test cases relied on the old logic to obtain STACK_MISC
marks for some stack values. These test cases are updated in the same
commit to avoid failures during bisect.
Change-Id: I51cb563cc8c0fbc51434abe82a79afaa822fe5b1
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Link: https://lore.kernel.org/r/20230214232030.1502829-2-eddyz87@gmail.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Stable-dep-of: 713274f1f2c8 ("bpf: Fix verifier id tracking of scalars on spill")
Signed-off-by: Sasha Levin <sashal@kernel.org>
[ Upstream commit b9979db8340154526d9ab38a1883d6f6ba9b6d47 ]
Before this fix:
166: (b5) if r2 <= 0x1 goto pc+22
from 166 to 189: R2=invP(id=1,umax_value=1,var_off=(0x0; 0xffffffff))
After this fix:
166: (b5) if r2 <= 0x1 goto pc+22
from 166 to 189: R2=invP(id=1,umax_value=1,var_off=(0x0; 0x1))
While processing BPF_JLE the reg_set_min_max() would set true_reg->umax_value = 1
and call __reg_combine_64_into_32(true_reg).
Without the fix it would not pass the condition:
if (__reg64_bound_u32(reg->umin_value) && __reg64_bound_u32(reg->umax_value))
since umin_value == 0 at this point.
Before commit 10bf4e83167c the umin was incorrectly ingored.
The commit 10bf4e83167c fixed the correctness issue, but pessimized
propagation of 64-bit min max into 32-bit min max and corresponding var_off.
Fixes: 10bf4e83167c ("bpf: Fix propagation of 32 bit unsigned bounds from 64 bit bounds")
Change-Id: I3687ccc00a02b1177414ae042636695817a94037
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Acked-by: Yonghong Song <yhs@fb.com>
Link: https://lore.kernel.org/bpf/20211101222153.78759-1-alexei.starovoitov@gmail.com
Signed-off-by: Sasha Levin <sashal@kernel.org>
Some kernels builds might inline vfs_getattr call within fstat
syscall code path, so fentry/vfs_getattr trampoline is not called.
Add security_inode_getattr to allowlist and switch the d_path test stat
trampoline to security_inode_getattr.
Keeping dentry_open and filp_close, because they are in their own
files, so unlikely to be inlined, but in case they are, adding
security_file_open.
Adding flags that indicate trampolines were called and failing
the test if any of them got missed, so it's easier to identify
the issue next time.
Fixes: e4d1af4b16f8 ("selftests/bpf: Add test for d_path helper")
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Change-Id: Ifa431c01d768bb64a3cd99742ddc22313bf209a0
Signed-off-by: Jiri Olsa <jolsa@redhat.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200918112338.2618444-1-jolsa@kernel.org
There is a spelling mistake in a check error message. Fix it.
Change-Id: I20e23a8c33d722f75f3ef7b7ad0038372e1a8a14
Signed-off-by: Colin Ian King <colin.king@canonical.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200826085907.43095-1-colin.king@canonical.com
Alexei reported compile breakage on newer systems with
following error:
In file included from /usr/include/fcntl.h:290:0,
4814 from ./test_progs.h:29,
4815 from
.../bpf-next/tools/testing/selftests/bpf/prog_tests/d_path.c:3:
4816In function ‘open’,
4817 inlined from ‘trigger_fstat_events’ at
.../bpf-next/tools/testing/selftests/bpf/prog_tests/d_path.c:50:10,
4818 inlined from ‘test_d_path’ at
.../bpf-next/tools/testing/selftests/bpf/prog_tests/d_path.c:119:6:
4819/usr/include/x86_64-linux-gnu/bits/fcntl2.h:50:4: error: call to
‘__open_missing_mode’ declared with attribute error: open with O_CREAT
or O_TMPFILE in second argument needs 3 arguments
4820 __open_missing_mode ();
4821 ^~~~~~~~~~~~~~~~~~~~~~
We're missing permission bits as 3rd argument
for open call with O_CREAT flag specified.
Fixes: e4d1af4b16f8 ("selftests/bpf: Add test for d_path helper")
Reported-by: Alexei Starovoitov <ast@kernel.org>
Change-Id: Ia8711a4f77158b797fa34e2bd33040e10dec4314
Signed-off-by: Jiri Olsa <jolsa@kernel.org>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200826101845.747617-1-jolsa@kernel.org
Adding test for d_path helper which is pretty much
copied from Wenbo Zhang's test for bpf_get_fd_path,
which never made it in.
The test is doing fstat/close on several fd types,
and verifies we got the d_path helper working on
kernel probes for vfs_getattr/filp_close functions.
Original-patch-by: Wenbo Zhang <ethercflow@gmail.com>
Change-Id: Ifd3a25c7db0503cc316288016269c7fc2b0d0873
Signed-off-by: Jiri Olsa <jolsa@kernel.org>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Link: https://lore.kernel.org/bpf/20200825192124.710397-14-jolsa@kernel.org
[ Upstream commit 10bf4e83167cc68595b85fd73bb91e8f2c086e36 ]
Similarly as b02709587ea3 ("bpf: Fix propagation of 32-bit signed bounds
from 64-bit bounds."), we also need to fix the propagation of 32 bit
unsigned bounds from 64 bit counterparts. That is, really only set the
u32_{min,max}_value when /both/ {umin,umax}_value safely fit in 32 bit
space. For example, the register with a umin_value == 1 does /not/ imply
that u32_min_value is also equal to 1, since umax_value could be much
larger than 32 bit subregister can hold, and thus u32_min_value is in
the interval [0,1] instead.
Before fix, invalid tracking result of R2_w=inv1:
[...]
5: R0_w=inv1337 R1=ctx(id=0,off=0,imm=0) R2_w=inv(id=0) R10=fp0
5: (35) if r2 >= 0x1 goto pc+1
[...] // goto path
7: R0=inv1337 R1=ctx(id=0,off=0,imm=0) R2=inv(id=0,umin_value=1) R10=fp0
7: (b6) if w2 <= 0x1 goto pc+1
[...] // goto path
9: R0=inv1337 R1=ctx(id=0,off=0,imm=0) R2=inv(id=0,smin_value=-9223372036854775807,smax_value=9223372032559808513,umin_value=1,umax_value=18446744069414584321,var_off=(0x1; 0xffffffff00000000),s32_min_value=1,s32_max_value=1,u32_max_value=1) R10=fp0
9: (bc) w2 = w2
10: R0=inv1337 R1=ctx(id=0,off=0,imm=0) R2_w=inv1 R10=fp0
[...]
After fix, correct tracking result of R2_w=inv(id=0,umax_value=1,var_off=(0x0; 0x1)):
[...]
5: R0_w=inv1337 R1=ctx(id=0,off=0,imm=0) R2_w=inv(id=0) R10=fp0
5: (35) if r2 >= 0x1 goto pc+1
[...] // goto path
7: R0=inv1337 R1=ctx(id=0,off=0,imm=0) R2=inv(id=0,umin_value=1) R10=fp0
7: (b6) if w2 <= 0x1 goto pc+1
[...] // goto path
9: R0=inv1337 R1=ctx(id=0,off=0,imm=0) R2=inv(id=0,smax_value=9223372032559808513,umax_value=18446744069414584321,var_off=(0x0; 0xffffffff00000001),s32_min_value=0,s32_max_value=1,u32_max_value=1) R10=fp0
9: (bc) w2 = w2
10: R0=inv1337 R1=ctx(id=0,off=0,imm=0) R2_w=inv(id=0,umax_value=1,var_off=(0x0; 0x1)) R10=fp0
[...]
Thus, same issue as in b02709587ea3 holds for unsigned subregister tracking.
Also, align __reg64_bound_u32() similarly to __reg64_bound_s32() as done in
b02709587ea3 to make them uniform again.
Fixes: 3f50f132d840 ("bpf: Verifier, do explicit ALU32 bounds tracking")
Reported-by: Manfred Paul (@_manfp)
Change-Id: I2f52adb7651878cdba6bc39049d201fb7668a411
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: John Fastabend <john.fastabend@gmail.com>
Acked-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Currently verifier enforces return code checks for subprograms in the
same manner as it does for program entry points. This prevents returning
arbitrary scalar values from subprograms. Scalar type of returned values
is checked by btf_prepare_func_args() and hence it should be safe to
allow only scalars for now. Relax return code checks for subprograms and
allow any correct scalar values.
Fixes: 51c39bb1d5d10 (bpf: Introduce function-by-function verification)
Change-Id: I903a762fbc0b9d273512babb287ce29880d240f7
Signed-off-by: Dmitrii Banshchikov <me@ubique.spb.ru>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20201113171756.90594-1-me@ubique.spb.ru
The error path in libbpf.c:load_program() has calls to pr_warn()
which ends up for global_funcs tests to
test_global_funcs.c:libbpf_debug_print().
For the tests with no struct test_def::err_str initialized with a
string, it causes call of strstr() with NULL as the second argument
and it segfaults.
Fix it by calling strstr() only for non-NULL err_str.
Change-Id: I919e29818dafd5cf35c7e2a8f6216d28c80915f7
Signed-off-by: Yauheni Kaliuta <yauheni.kaliuta@redhat.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Yonghong Song <yhs@fb.com>
Link: https://lore.kernel.org/bpf/20200820115843.39454-1-yauheni.kaliuta@redhat.com
test_global_func[12] - check 512 stack limit.
test_global_func[34] - check 8 frame call chain limit.
test_global_func5 - check that non-ctx pointer cannot be passed into
a function that expects context.
test_global_func6 - check that ctx pointer is unmodified.
test_global_func7 - check that global function returns scalar.
Change-Id: Ia3963c47a6202a1440de22cd69cd17ede200e8a9
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Song Liu <songliubraving@fb.com>
Link: https://lore.kernel.org/bpf/20200110064124.1760511-7-ast@kernel.org
Zero-fill element values for all other cpus than current, just as
when not using prealloc. This is the only way the bpf program can
ensure known initial values for all cpus ('onallcpus' cannot be
set when coming from the bpf program).
The scenario is: bpf program inserts some elements in a per-cpu
map, then deletes some (or userspace does). When later adding
new elements using bpf_map_update_elem(), the bpf program can
only set the value of the new elements for the current cpu.
When prealloc is enabled, previously deleted elements are re-used.
Without the fix, values for other cpus remain whatever they were
when the re-used entry was previously freed.
A selftest is added to validate correct operation in above
scenario as well as in case of LRU per-cpu map element re-use.
Fixes: 6c90598174 ("bpf: pre-allocate hash map elements")
Change-Id: I5ff28a3ee950b7733ab5eb0a5ca1fa1bcd898228
Signed-off-by: David Verbeiren <david.verbeiren@tessares.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Matthieu Baerts <matthieu.baerts@tessares.net>
Acked-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20201104112332.15191-1-david.verbeiren@tessares.net
The 64-bit JEQ/JNE handling in reg_set_min_max() was clearing reg->id in either
true or false branch. In the case 'if (reg->id)' check was done on the other
branch the counter part register would have reg->id == 0 when called into
find_equal_scalars(). In such case the helper would incorrectly identify other
registers with id == 0 as equivalent and propagate the state incorrectly.
Fix it by preserving ID across reg_set_min_max().
In other words any kind of comparison operator on the scalar register
should preserve its ID to recognize:
r1 = r2
if (r1 == 20) {
#1 here both r1 and r2 == 20
} else if (r2 < 20) {
#2 here both r1 and r2 < 20
}
The patch is addressing #1 case. The #2 was working correctly already.
Fixes: 75748837b7e5 ("bpf: Propagate scalar ranges through register assignments.")
Change-Id: Id5737c3392daa0d695a768206b3d3e4bda187aaf
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andrii@kernel.org>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Tested-by: Yonghong Song <yhs@fb.com>
Link: https://lore.kernel.org/bpf/20201014175608.1416-1-alexei.starovoitov@gmail.com
The llvm register allocator may use two different registers representing the
same virtual register. In such case the following pattern can be observed:
1047: (bf) r9 = r6
1048: (a5) if r6 < 0x1000 goto pc+1
1050: ...
1051: (a5) if r9 < 0x2 goto pc+66
1052: ...
1053: (bf) r2 = r9 /* r2 needs to have upper and lower bounds */
This is normal behavior of greedy register allocator.
The slides 137+ explain why regalloc introduces such register copy:
http://llvm.org/devmtg/2018-04/slides/Yatsina-LLVM%20Greedy%20Register%20Allocator.pdf
There is no way to tell llvm 'not to do this'.
Hence the verifier has to recognize such patterns.
In order to track this information without backtracking allocate ID
for scalars in a similar way as it's done for find_good_pkt_pointers().
When the verifier encounters r9 = r6 assignment it will assign the same ID
to both registers. Later if either register range is narrowed via conditional
jump propagate the register state into the other register.
Clear register ID in adjust_reg_min_max_vals() for any alu instruction. The
register ID is ignored for scalars in regsafe() and doesn't affect state
pruning. mark_reg_unknown() clears the ID. It's used to process call, endian
and other instructions. Hence ID is explicitly cleared only in
adjust_reg_min_max_vals() and in 32-bit mov.
Change-Id: Iecc5a1ad365e63064d963e66fcbbf70bb38ebdd3
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrii Nakryiko <andrii@kernel.org>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Link: https://lore.kernel.org/bpf/20201009011240.48506-2-alexei.starovoitov@gmail.com
There is a much higher chance we can see the regressions if the
test is part of test_progs.
Change-Id: I376f8cf6541b11e110af291503a33652f3722c4a
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20200515194904.229296-2-sdf@google.com
Commit 294f2fc6da27 ("bpf: Verifer, adjust_scalar_min_max_vals to always
call update_reg_bounds()") changed the way verifier logs some of its state,
adjust the test_align accordingly. Where possible, I tried to not copy-paste
the entire log line and resorted to dropping the last closing brace instead.
Fixes: 294f2fc6da27 ("bpf: Verifer, adjust_scalar_min_max_vals to always call update_reg_bounds()")
Change-Id: If037a9198b047fd8d1b34f66454d2f86f1a33235
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20200515194904.229296-1-sdf@google.com
LD_[ABS|IND] instructions may return from the function early. bpf_tail_call
pseudo instruction is either fallthrough or return. Allow them in the
subprograms only when subprograms are BTF annotated and have scalar return
types. Allow ld_abs and tail_call in the main program even if it calls into
subprograms. In the past that was not ok to do for ld_abs, since it was JITed
with special exit sequence. Since bpf_gen_ld_abs() was introduced the ld_abs
looks like normal exit insn from JIT point of view, so it's safe to allow them
in the main program.
Change-Id: I46141c0f33c94a5f84430820b1ce0d977bea0428
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
A purely mechanical change to split the renaming from the actual
generalization.
Flags/consts:
SK_STORAGE_CREATE_FLAG_MASK BPF_LOCAL_STORAGE_CREATE_FLAG_MASK
BPF_SK_STORAGE_CACHE_SIZE BPF_LOCAL_STORAGE_CACHE_SIZE
MAX_VALUE_SIZE BPF_LOCAL_STORAGE_MAX_VALUE_SIZE
Structs:
bucket bpf_local_storage_map_bucket
bpf_sk_storage_map bpf_local_storage_map
bpf_sk_storage_data bpf_local_storage_data
bpf_sk_storage_elem bpf_local_storage_elem
bpf_sk_storage bpf_local_storage
The "sk" member in bpf_local_storage is also updated to "owner"
in preparation for changing the type to void * in a subsequent patch.
Functions:
selem_linked_to_sk selem_linked_to_storage
selem_alloc bpf_selem_alloc
__selem_unlink_sk bpf_selem_unlink_storage_nolock
__selem_link_sk bpf_selem_link_storage_nolock
selem_unlink_sk __bpf_selem_unlink_storage
sk_storage_update bpf_local_storage_update
__sk_storage_lookup bpf_local_storage_lookup
bpf_sk_storage_map_free bpf_local_storage_map_free
bpf_sk_storage_map_alloc bpf_local_storage_map_alloc
bpf_sk_storage_map_alloc_check bpf_local_storage_map_alloc_check
bpf_sk_storage_map_check_btf bpf_local_storage_map_check_btf
Change-Id: I282ab51390db10efbfccf54cd25896abbd6582cb
Signed-off-by: KP Singh <kpsingh@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20200825182919.1118197-2-kpsingh@chromium.org
Add selftests to test access to map pointers from bpf program for all
map types except struct_ops (that one would need additional work).
verifier test focuses mostly on scenarios that must be rejected.
prog_tests test focuses on accessing multiple fields both scalar and a
nested struct from bpf program and verifies that those fields have
expected values.
Change-Id: I4b0c119c8d686871551bbc4011e11cc68858f775
Signed-off-by: Andrey Ignatov <rdna@fb.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/139a6a17f8016491e39347849b951525335c6eb4.1592600985.git.rdna@fb.com
Now skb->dev is unconditionally set to the loopback device in current net
namespace. But if we want to test bpf program which contains code branch
based on ifindex condition (eg filters out localhost packets) it is useful
to allow specifying of ifindex from userspace. This patch adds such option
through ctx_in (__sk_buff) parameter.
Change-Id: Ia56479984ad24f3e0deb878c30dc42a6377bfb0b
Signed-off-by: Dmitry Yakunin <zeil@yandex-team.ru>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20200803090545.82046-3-zeil@yandex-team.ru
Add tests to verify ability to add an XDP program to a
entry in a DEVMAP.
Add negative tests to show DEVMAP programs can not be
attached to devices as a normal XDP program, and accesses
to egress_ifindex require BPF_XDP_DEVMAP attach type.
Change-Id: Ic91c2835212b87d397dc33161c718619f26a8ac7
Signed-off-by: David Ahern <dsahern@kernel.org>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Toke Høiland-Jørgensen <toke@redhat.com>
Link: https://lore.kernel.org/bpf/20200529220716.75383-6-dsahern@kernel.org
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
There are multiple use-cases when it's convenient to have access to bpf
map fields, both `struct bpf_map` and map type specific struct-s such as
`struct bpf_array`, `struct bpf_htab`, etc.
For example while working with sock arrays it can be necessary to
calculate the key based on map->max_entries (some_hash % max_entries).
Currently this is solved by communicating max_entries via "out-of-band"
channel, e.g. via additional map with known key to get info about target
map. That works, but is not very convenient and error-prone while
working with many maps.
In other cases necessary data is dynamic (i.e. unknown at loading time)
and it's impossible to get it at all. For example while working with a
hash table it can be convenient to know how much capacity is already
used (bpf_htab.count.counter for BPF_F_NO_PREALLOC case).
At the same time kernel knows this info and can provide it to bpf
program.
Fill this gap by adding support to access bpf map fields from bpf
program for both `struct bpf_map` and map type specific fields.
Support is implemented via btf_struct_access() so that a user can define
their own `struct bpf_map` or map type specific struct in their program
with only necessary fields and preserve_access_index attribute, cast a
map to this struct and use a field.
For example:
struct bpf_map {
__u32 max_entries;
} __attribute__((preserve_access_index));
struct bpf_array {
struct bpf_map map;
__u32 elem_size;
} __attribute__((preserve_access_index));
struct {
__uint(type, BPF_MAP_TYPE_ARRAY);
__uint(max_entries, 4);
__type(key, __u32);
__type(value, __u32);
} m_array SEC(".maps");
SEC("cgroup_skb/egress")
int cg_skb(void *ctx)
{
struct bpf_array *array = (struct bpf_array *)&m_array;
struct bpf_map *map = (struct bpf_map *)&m_array;
/* .. use map->max_entries or array->map.max_entries .. */
}
Similarly to other btf_struct_access() use-cases (e.g. struct tcp_sock
in net/ipv4/bpf_tcp_ca.c) the patch allows access to any fields of
corresponding struct. Only reading from map fields is supported.
For btf_struct_access() to work there should be a way to know btf id of
a struct that corresponds to a map type. To get btf id there should be a
way to get a stringified name of map-specific struct, such as
"bpf_array", "bpf_htab", etc for a map type. Two new fields are added to
`struct bpf_map_ops` to handle it:
* .map_btf_name keeps a btf name of a struct returned by map_alloc();
* .map_btf_id is used to cache btf id of that struct.
To make btf ids calculation cheaper they're calculated once while
preparing btf_vmlinux and cached same way as it's done for btf_id field
of `struct bpf_func_proto`
While calculating btf ids, struct names are NOT checked for collision.
Collisions will be checked as a part of the work to prepare btf ids used
in verifier in compile time that should land soon. The only known
collision for `struct bpf_htab` (kernel/bpf/hashtab.c vs
net/core/sock_map.c) was fixed earlier.
Both new fields .map_btf_name and .map_btf_id must be set for a map type
for the feature to work. If neither is set for a map type, verifier will
return ENOTSUPP on a try to access map_ptr of corresponding type. If
just one of them set, it's verifier misconfiguration.
Only `struct bpf_array` for BPF_MAP_TYPE_ARRAY and `struct bpf_htab` for
BPF_MAP_TYPE_HASH are supported by this patch. Other map types will be
supported separately.
The feature is available only for CONFIG_DEBUG_INFO_BTF=y and gated by
perfmon_capable() so that unpriv programs won't have access to bpf map
fields.
Change-Id: Ibd44c4fe4f296e0488dff536e1ec56e90d542f56
Signed-off-by: Andrey Ignatov <rdna@fb.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/6479686a0cd1e9067993df57b4c3eef0e276fec9.1592600985.git.rdna@fb.com
This commit adds a new MPSC ring buffer implementation into BPF ecosystem,
which allows multiple CPUs to submit data to a single shared ring buffer. On
the consumption side, only single consumer is assumed.
Motivation
----------
There are two distinctive motivators for this work, which are not satisfied by
existing perf buffer, which prompted creation of a new ring buffer
implementation.
- more efficient memory utilization by sharing ring buffer across CPUs;
- preserving ordering of events that happen sequentially in time, even
across multiple CPUs (e.g., fork/exec/exit events for a task).
These two problems are independent, but perf buffer fails to satisfy both.
Both are a result of a choice to have per-CPU perf ring buffer. Both can be
also solved by having an MPSC implementation of ring buffer. The ordering
problem could technically be solved for perf buffer with some in-kernel
counting, but given the first one requires an MPSC buffer, the same solution
would solve the second problem automatically.
Semantics and APIs
------------------
Single ring buffer is presented to BPF programs as an instance of BPF map of
type BPF_MAP_TYPE_RINGBUF. Two other alternatives considered, but ultimately
rejected.
One way would be to, similar to BPF_MAP_TYPE_PERF_EVENT_ARRAY, make
BPF_MAP_TYPE_RINGBUF could represent an array of ring buffers, but not enforce
"same CPU only" rule. This would be more familiar interface compatible with
existing perf buffer use in BPF, but would fail if application needed more
advanced logic to lookup ring buffer by arbitrary key. HASH_OF_MAPS addresses
this with current approach. Additionally, given the performance of BPF
ringbuf, many use cases would just opt into a simple single ring buffer shared
among all CPUs, for which current approach would be an overkill.
Another approach could introduce a new concept, alongside BPF map, to
represent generic "container" object, which doesn't necessarily have key/value
interface with lookup/update/delete operations. This approach would add a lot
of extra infrastructure that has to be built for observability and verifier
support. It would also add another concept that BPF developers would have to
familiarize themselves with, new syntax in libbpf, etc. But then would really
provide no additional benefits over the approach of using a map.
BPF_MAP_TYPE_RINGBUF doesn't support lookup/update/delete operations, but so
doesn't few other map types (e.g., queue and stack; array doesn't support
delete, etc).
The approach chosen has an advantage of re-using existing BPF map
infrastructure (introspection APIs in kernel, libbpf support, etc), being
familiar concept (no need to teach users a new type of object in BPF program),
and utilizing existing tooling (bpftool). For common scenario of using
a single ring buffer for all CPUs, it's as simple and straightforward, as
would be with a dedicated "container" object. On the other hand, by being
a map, it can be combined with ARRAY_OF_MAPS and HASH_OF_MAPS map-in-maps to
implement a wide variety of topologies, from one ring buffer for each CPU
(e.g., as a replacement for perf buffer use cases), to a complicated
application hashing/sharding of ring buffers (e.g., having a small pool of
ring buffers with hashed task's tgid being a look up key to preserve order,
but reduce contention).
Key and value sizes are enforced to be zero. max_entries is used to specify
the size of ring buffer and has to be a power of 2 value.
There are a bunch of similarities between perf buffer
(BPF_MAP_TYPE_PERF_EVENT_ARRAY) and new BPF ring buffer semantics:
- variable-length records;
- if there is no more space left in ring buffer, reservation fails, no
blocking;
- memory-mappable data area for user-space applications for ease of
consumption and high performance;
- epoll notifications for new incoming data;
- but still the ability to do busy polling for new data to achieve the
lowest latency, if necessary.
BPF ringbuf provides two sets of APIs to BPF programs:
- bpf_ringbuf_output() allows to *copy* data from one place to a ring
buffer, similarly to bpf_perf_event_output();
- bpf_ringbuf_reserve()/bpf_ringbuf_commit()/bpf_ringbuf_discard() APIs
split the whole process into two steps. First, a fixed amount of space is
reserved. If successful, a pointer to a data inside ring buffer data area
is returned, which BPF programs can use similarly to a data inside
array/hash maps. Once ready, this piece of memory is either committed or
discarded. Discard is similar to commit, but makes consumer ignore the
record.
bpf_ringbuf_output() has disadvantage of incurring extra memory copy, because
record has to be prepared in some other place first. But it allows to submit
records of the length that's not known to verifier beforehand. It also closely
matches bpf_perf_event_output(), so will simplify migration significantly.
bpf_ringbuf_reserve() avoids the extra copy of memory by providing a memory
pointer directly to ring buffer memory. In a lot of cases records are larger
than BPF stack space allows, so many programs have use extra per-CPU array as
a temporary heap for preparing sample. bpf_ringbuf_reserve() avoid this needs
completely. But in exchange, it only allows a known constant size of memory to
be reserved, such that verifier can verify that BPF program can't access
memory outside its reserved record space. bpf_ringbuf_output(), while slightly
slower due to extra memory copy, covers some use cases that are not suitable
for bpf_ringbuf_reserve().
The difference between commit and discard is very small. Discard just marks
a record as discarded, and such records are supposed to be ignored by consumer
code. Discard is useful for some advanced use-cases, such as ensuring
all-or-nothing multi-record submission, or emulating temporary malloc()/free()
within single BPF program invocation.
Each reserved record is tracked by verifier through existing
reference-tracking logic, similar to socket ref-tracking. It is thus
impossible to reserve a record, but forget to submit (or discard) it.
bpf_ringbuf_query() helper allows to query various properties of ring buffer.
Currently 4 are supported:
- BPF_RB_AVAIL_DATA returns amount of unconsumed data in ring buffer;
- BPF_RB_RING_SIZE returns the size of ring buffer;
- BPF_RB_CONS_POS/BPF_RB_PROD_POS returns current logical possition of
consumer/producer, respectively.
Returned values are momentarily snapshots of ring buffer state and could be
off by the time helper returns, so this should be used only for
debugging/reporting reasons or for implementing various heuristics, that take
into account highly-changeable nature of some of those characteristics.
One such heuristic might involve more fine-grained control over poll/epoll
notifications about new data availability in ring buffer. Together with
BPF_RB_NO_WAKEUP/BPF_RB_FORCE_WAKEUP flags for output/commit/discard helpers,
it allows BPF program a high degree of control and, e.g., more efficient
batched notifications. Default self-balancing strategy, though, should be
adequate for most applications and will work reliable and efficiently already.
Design and implementation
-------------------------
This reserve/commit schema allows a natural way for multiple producers, either
on different CPUs or even on the same CPU/in the same BPF program, to reserve
independent records and work with them without blocking other producers. This
means that if BPF program was interruped by another BPF program sharing the
same ring buffer, they will both get a record reserved (provided there is
enough space left) and can work with it and submit it independently. This
applies to NMI context as well, except that due to using a spinlock during
reservation, in NMI context, bpf_ringbuf_reserve() might fail to get a lock,
in which case reservation will fail even if ring buffer is not full.
The ring buffer itself internally is implemented as a power-of-2 sized
circular buffer, with two logical and ever-increasing counters (which might
wrap around on 32-bit architectures, that's not a problem):
- consumer counter shows up to which logical position consumer consumed the
data;
- producer counter denotes amount of data reserved by all producers.
Each time a record is reserved, producer that "owns" the record will
successfully advance producer counter. At that point, data is still not yet
ready to be consumed, though. Each record has 8 byte header, which contains
the length of reserved record, as well as two extra bits: busy bit to denote
that record is still being worked on, and discard bit, which might be set at
commit time if record is discarded. In the latter case, consumer is supposed
to skip the record and move on to the next one. Record header also encodes
record's relative offset from the beginning of ring buffer data area (in
pages). This allows bpf_ringbuf_commit()/bpf_ringbuf_discard() to accept only
the pointer to the record itself, without requiring also the pointer to ring
buffer itself. Ring buffer memory location will be restored from record
metadata header. This significantly simplifies verifier, as well as improving
API usability.
Producer counter increments are serialized under spinlock, so there is
a strict ordering between reservations. Commits, on the other hand, are
completely lockless and independent. All records become available to consumer
in the order of reservations, but only after all previous records where
already committed. It is thus possible for slow producers to temporarily hold
off submitted records, that were reserved later.
Reservation/commit/consumer protocol is verified by litmus tests in
Documentation/litmus-test/bpf-rb.
One interesting implementation bit, that significantly simplifies (and thus
speeds up as well) implementation of both producers and consumers is how data
area is mapped twice contiguously back-to-back in the virtual memory. This
allows to not take any special measures for samples that have to wrap around
at the end of the circular buffer data area, because the next page after the
last data page would be first data page again, and thus the sample will still
appear completely contiguous in virtual memory. See comment and a simple ASCII
diagram showing this visually in bpf_ringbuf_area_alloc().
Another feature that distinguishes BPF ringbuf from perf ring buffer is
a self-pacing notifications of new data being availability.
bpf_ringbuf_commit() implementation will send a notification of new record
being available after commit only if consumer has already caught up right up
to the record being committed. If not, consumer still has to catch up and thus
will see new data anyways without needing an extra poll notification.
Benchmarks (see tools/testing/selftests/bpf/benchs/bench_ringbuf.c) show that
this allows to achieve a very high throughput without having to resort to
tricks like "notify only every Nth sample", which are necessary with perf
buffer. For extreme cases, when BPF program wants more manual control of
notifications, commit/discard/output helpers accept BPF_RB_NO_WAKEUP and
BPF_RB_FORCE_WAKEUP flags, which give full control over notifications of data
availability, but require extra caution and diligence in using this API.
Comparison to alternatives
--------------------------
Before considering implementing BPF ring buffer from scratch existing
alternatives in kernel were evaluated, but didn't seem to meet the needs. They
largely fell into few categores:
- per-CPU buffers (perf, ftrace, etc), which don't satisfy two motivations
outlined above (ordering and memory consumption);
- linked list-based implementations; while some were multi-producer designs,
consuming these from user-space would be very complicated and most
probably not performant; memory-mapping contiguous piece of memory is
simpler and more performant for user-space consumers;
- io_uring is SPSC, but also requires fixed-sized elements. Naively turning
SPSC queue into MPSC w/ lock would have subpar performance compared to
locked reserve + lockless commit, as with BPF ring buffer. Fixed sized
elements would be too limiting for BPF programs, given existing BPF
programs heavily rely on variable-sized perf buffer already;
- specialized implementations (like a new printk ring buffer, [0]) with lots
of printk-specific limitations and implications, that didn't seem to fit
well for intended use with BPF programs.
[0] https://lwn.net/Articles/779550/
Change-Id: I0da8e403a2ccc0d786a38da207557060300ab1b2
Signed-off-by: Andrii Nakryiko <andriin@fb.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20200529075424.3139988-2-andriin@fb.com
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
We want to have a tighter control on what ports we bind to in
the BPF_CGROUP_INET{4,6}_CONNECT hooks even if it means
connect() becomes slightly more expensive. The expensive part
comes from the fact that we now need to call inet_csk_get_port()
that verifies that the port is not used and allocates an entry
in the hash table for it.
Since we can't rely on "snum || !bind_address_no_port" to prevent
us from calling POST_BIND hook anymore, let's add another bind flag
to indicate that the call site is BPF program.
v5:
* fix wrong AF_INET (should be AF_INET6) in the bpf program for v6
v3:
* More bpf_bind documentation refinements (Martin KaFai Lau)
* Add UDP tests as well (Martin KaFai Lau)
* Don't start the thread, just do socket+bind+listen (Martin KaFai Lau)
v2:
* Update documentation (Andrey Ignatov)
* Pass BIND_FORCE_ADDRESS_NO_PORT conditionally (Andrey Ignatov)
Change-Id: Ia0d7052ae78cb82da3a0a7eea21afef9b174d26e
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Andrey Ignatov <rdna@fb.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20200508174611.228805-5-sdf@google.com
Currently, bpf_getsockopt and bpf_setsockopt helpers operate on the
'struct bpf_sock_ops' context in BPF_PROG_TYPE_SOCK_OPS program.
Let's generalize them and make them available for 'struct bpf_sock_addr'.
That way, in the future, we can allow those helpers in more places.
As an example, let's expose those 'struct bpf_sock_addr' based helpers to
BPF_CGROUP_INET{4,6}_CONNECT hooks. That way we can override CC before the
connection is made.
v3:
* Expose custom helpers for bpf_sock_addr context instead of doing
generic bpf_sock argument (as suggested by Daniel). Even with
try_socket_lock that doesn't sleep we have a problem where context sk
is already locked and socket lock is non-nestable.
v2:
* s/BPF_PROG_TYPE_CGROUP_SOCKOPT/BPF_PROG_TYPE_SOCK_OPS/
Change-Id: Ic1ec1a886b1c822be6652075f0f85549a0594266
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Link: https://lore.kernel.org/bpf/20200430233152.199403-1-sdf@google.com
* Load/attach a BPF program that hooks to file_mprotect (int)
and bprm_committed_creds (void).
* Perform an action that triggers the hook.
* Verify if the audit event was received using the shared global
variables for the process executed.
* Verify if the mprotect returns a -EPERM.
Change-Id: I0e041365f359a29ffb2d68b178ed090af54fc6f7
Signed-off-by: KP Singh <kpsingh@google.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Reviewed-by: Brendan Jackman <jackmanb@google.com>
Reviewed-by: Florent Revest <revest@google.com>
Reviewed-by: Thomas Garnier <thgarnie@google.com>
Reviewed-by: James Morris <jamorris@linux.microsoft.com>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Link: https://lore.kernel.org/bpf/20200329004356.27286-8-kpsingh@chromium.org
To make BPF verifier verbose log more releavant and easier to use to debug
verification failures, "pop" parts of log that were successfully verified.
This has effect of leaving only verifier logs that correspond to code branches
that lead to verification failure, which in practice should result in much
shorter and more relevant verifier log dumps. This behavior is made the
default behavior and can be overriden to do exhaustive logging by specifying
BPF_LOG_LEVEL2 log level.
Using BPF_LOG_LEVEL2 to disable this behavior is not ideal, because in some
cases it's good to have BPF_LOG_LEVEL2 per-instruction register dump
verbosity, but still have only relevant verifier branches logged. But for this
patch, I didn't want to add any new flags. It might be worth-while to just
rethink how BPF verifier logging is performed and requested and streamline it
a bit. But this trimming of successfully verified branches seems to be useful
and a good default behavior.
To test this, I modified runqslower slightly to introduce read of
uninitialized stack variable. Log (**truncated in the middle** to save many
lines out of this commit message) BEFORE this change:
; int handle__sched_switch(u64 *ctx)
0: (bf) r6 = r1
; struct task_struct *prev = (struct task_struct *)ctx[1];
1: (79) r1 = *(u64 *)(r6 +8)
func 'sched_switch' arg1 has btf_id 151 type STRUCT 'task_struct'
2: (b7) r2 = 0
; struct event event = {};
3: (7b) *(u64 *)(r10 -24) = r2
last_idx 3 first_idx 0
regs=4 stack=0 before 2: (b7) r2 = 0
4: (7b) *(u64 *)(r10 -32) = r2
5: (7b) *(u64 *)(r10 -40) = r2
6: (7b) *(u64 *)(r10 -48) = r2
; if (prev->state == TASK_RUNNING)
[ ... instruction dump from insn #7 through #50 are cut out ... ]
51: (b7) r2 = 16
52: (85) call bpf_get_current_comm#16
last_idx 52 first_idx 42
regs=4 stack=0 before 51: (b7) r2 = 16
; bpf_perf_event_output(ctx, &events, BPF_F_CURRENT_CPU,
53: (bf) r1 = r6
54: (18) r2 = 0xffff8881f3868800
56: (18) r3 = 0xffffffff
58: (bf) r4 = r7
59: (b7) r5 = 32
60: (85) call bpf_perf_event_output#25
last_idx 60 first_idx 53
regs=20 stack=0 before 59: (b7) r5 = 32
61: (bf) r2 = r10
; event.pid = pid;
62: (07) r2 += -16
; bpf_map_delete_elem(&start, &pid);
63: (18) r1 = 0xffff8881f3868000
65: (85) call bpf_map_delete_elem#3
; }
66: (b7) r0 = 0
67: (95) exit
from 44 to 66: safe
from 34 to 66: safe
from 11 to 28: R1_w=inv0 R2_w=inv0 R6_w=ctx(id=0,off=0,imm=0) R10=fp0 fp-8=mmmm???? fp-24_w=00000000 fp-32_w=00000000 fp-40_w=00000000 fp-48_w=00000000
; bpf_map_update_elem(&start, &pid, &ts, 0);
28: (bf) r2 = r10
;
29: (07) r2 += -16
; tsp = bpf_map_lookup_elem(&start, &pid);
30: (18) r1 = 0xffff8881f3868000
32: (85) call bpf_map_lookup_elem#1
invalid indirect read from stack off -16+0 size 4
processed 65 insns (limit 1000000) max_states_per_insn 1 total_states 5 peak_states 5 mark_read 4
Notice how there is a successful code path from instruction 0 through 67, few
successfully verified jumps (44->66, 34->66), and only after that 11->28 jump
plus error on instruction #32.
AFTER this change (full verifier log, **no truncation**):
; int handle__sched_switch(u64 *ctx)
0: (bf) r6 = r1
; struct task_struct *prev = (struct task_struct *)ctx[1];
1: (79) r1 = *(u64 *)(r6 +8)
func 'sched_switch' arg1 has btf_id 151 type STRUCT 'task_struct'
2: (b7) r2 = 0
; struct event event = {};
3: (7b) *(u64 *)(r10 -24) = r2
last_idx 3 first_idx 0
regs=4 stack=0 before 2: (b7) r2 = 0
4: (7b) *(u64 *)(r10 -32) = r2
5: (7b) *(u64 *)(r10 -40) = r2
6: (7b) *(u64 *)(r10 -48) = r2
; if (prev->state == TASK_RUNNING)
7: (79) r2 = *(u64 *)(r1 +16)
; if (prev->state == TASK_RUNNING)
8: (55) if r2 != 0x0 goto pc+19
R1_w=ptr_task_struct(id=0,off=0,imm=0) R2_w=inv0 R6_w=ctx(id=0,off=0,imm=0) R10=fp0 fp-24_w=00000000 fp-32_w=00000000 fp-40_w=00000000 fp-48_w=00000000
; trace_enqueue(prev->tgid, prev->pid);
9: (61) r1 = *(u32 *)(r1 +1184)
10: (63) *(u32 *)(r10 -4) = r1
; if (!pid || (targ_pid && targ_pid != pid))
11: (15) if r1 == 0x0 goto pc+16
from 11 to 28: R1_w=inv0 R2_w=inv0 R6_w=ctx(id=0,off=0,imm=0) R10=fp0 fp-8=mmmm???? fp-24_w=00000000 fp-32_w=00000000 fp-40_w=00000000 fp-48_w=00000000
; bpf_map_update_elem(&start, &pid, &ts, 0);
28: (bf) r2 = r10
;
29: (07) r2 += -16
; tsp = bpf_map_lookup_elem(&start, &pid);
30: (18) r1 = 0xffff8881db3ce800
32: (85) call bpf_map_lookup_elem#1
invalid indirect read from stack off -16+0 size 4
processed 65 insns (limit 1000000) max_states_per_insn 1 total_states 5 peak_states 5 mark_read 4
Notice how in this case, there are 0-11 instructions + jump from 11 to
28 is recorded + 28-32 instructions with error on insn #32.
test_verifier test runner was updated to specify BPF_LOG_LEVEL2 for
VERBOSE_ACCEPT expected result due to potentially "incomplete" success verbose
log at BPF_LOG_LEVEL1.
On success, verbose log will only have a summary of number of processed
instructions, etc, but no branch tracing log. Having just a last succesful
branch tracing seemed weird and confusing. Having small and clean summary log
in success case seems quite logical and nice, though.
Change-Id: I3666f731c62f71ceb6baa6f758aaed3a30f8825c
Signed-off-by: Andrii Nakryiko <andriin@fb.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200423195850.1259827-1-andriin@fb.com
Currently the following prog types don't fall back to bpf_base_func_proto()
(instead they have cgroup_base_func_proto which has a limited set of
helpers from bpf_base_func_proto):
* BPF_PROG_TYPE_CGROUP_DEVICE
* BPF_PROG_TYPE_CGROUP_SYSCTL
* BPF_PROG_TYPE_CGROUP_SOCKOPT
I don't see any specific reason why we shouldn't use bpf_base_func_proto(),
every other type of program (except bpf-lirc and, understandably, tracing)
use it, so let's fall back to bpf_base_func_proto for those prog types
as well.
This basically boils down to adding access to the following helpers:
* BPF_FUNC_get_prandom_u32
* BPF_FUNC_get_smp_processor_id
* BPF_FUNC_get_numa_node_id
* BPF_FUNC_tail_call
* BPF_FUNC_ktime_get_ns
* BPF_FUNC_spin_lock (CAP_SYS_ADMIN)
* BPF_FUNC_spin_unlock (CAP_SYS_ADMIN)
* BPF_FUNC_jiffies64 (CAP_SYS_ADMIN)
I've also added bpf_perf_event_output() because it's really handy for
logging and debugging.
Change-Id: I1d1bb9f918ac89ac20a045e7428e01b676d989a5
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Link: https://lore.kernel.org/bpf/20200420174610.77494-1-sdf@google.com
Introduce new helper that reuses existing xdp perf_event output
implementation, but can be called from raw_tracepoint programs
that receive 'struct xdp_buff *' as a tracepoint argument.
Change-Id: I418e7a0142b7f7e9f44febf5b1d6bdd5519ef5ca
Signed-off-by: Eelco Chaudron <echaudro@redhat.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Acked-by: Toke Høiland-Jørgensen <toke@redhat.com>
Link: https://lore.kernel.org/bpf/158348514556.2239.11050972434793741444.stgit@xdp-tutorial
Use the new bpf_program__set_attach_target() API in the xdp_bpf2bpf
selftest so it can be referenced as an example on how to use it.
Change-Id: I21e608200f21a01a504b4b40cf24f5461161989f
Signed-off-by: Eelco Chaudron <echaudro@redhat.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Toke Høiland-Jørgensen <toke@redhat.com>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Link: https://lore.kernel.org/bpf/158220520562.127661.14289388017034825841.stgit@xdp-tutorial
Add a test that will attach a FENTRY and FEXIT program to the XDP test
program. It will also verify data from the XDP context on FENTRY and
verifies the return code on exit.
Change-Id: If0116a206c1776fef4094acbad42784b0b5a1f03
Signed-off-by: Eelco Chaudron <echaudro@redhat.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Link: https://lore.kernel.org/bpf/157909410480.47481.11202505690938004673.stgit@xdp-tutorial
In order for sockmap/sockhash types to become generic collections for
storing TCP sockets we need to loosen the checks during map update, while
tightening the checks in redirect helpers.
Currently sock{map,hash} require the TCP socket to be in established state,
which prevents inserting listening sockets.
Change the update pre-checks so the socket can also be in listening state.
Since it doesn't make sense to redirect with sock{map,hash} to listening
sockets, add appropriate socket state checks to BPF redirect helpers too.
Change-Id: I3a1bf4a5cda6f09c023d24a1c8653a4aa4e5a5c8
Signed-off-by: Jakub Sitnicki <jakub@cloudflare.com>
Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: John Fastabend <john.fastabend@gmail.com>
Link: https://lore.kernel.org/bpf/20200218171023.844439-5-jakub@cloudflare.com
Test for two scenarios:
* When the fmod_ret program returns 0, the original function should
be called along with fentry and fexit programs.
* When the fmod_ret program returns a non-zero value, the original
function should not be called, no side effect should be observed and
fentry and fexit programs should be called.
The result from the kernel function call and whether a side-effect is
observed is returned via the retval attr of the BPF_PROG_TEST_RUN (bpf)
syscall.
Change-Id: Ifd5cfbe80986c6650cf7cc95e8ef190fb01ed04a
Signed-off-by: KP Singh <kpsingh@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Andrii Nakryiko <andriin@fb.com>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Link: https://lore.kernel.org/bpf/20200304191853.1529-8-kpsingh@chromium.org
allow to pass skb's mark field into bpf_prog_test_run ctx
for BPF_PROG_TYPE_SCHED_CLS prog type. that would allow
to test bpf programs which are doing decision based on this
field
Change-Id: Id05ce69ad66d1635c4c839e28d389d7c80b38a02
Signed-off-by: Nikita V. Shirokov <tehnerd@tehnerd.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Make sure we can pass arbitrary data in wire_len/gso_segs.
Change-Id: I67098cd27a7b31a86c175ba42a79fd5e26dd5ef7
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191213223028.161282-2-sdf@google.com
Make sure BPF_PROG_TEST_RUN accepts tstamp and exports any
modifications that BPF program does.
Change-Id: Ib03df9539df637f1dc2595798e96aff87e04b695
Signed-off-by: Stanislav Fomichev <sdf@google.com>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Link: https://lore.kernel.org/bpf/20191015183125.124413-2-sdf@google.com
* 'android11-5.4-lts' of https://android.googlesource.com/kernel/common:
Linux 5.4.302
Input: pegasus-notetaker - fix potential out-of-bounds access
Input: remove third argument of usb_maxpacket()
usb: deprecate the third argument of usb_maxpacket()
ata: libata-scsi: Fix system suspend for a security locked drive
fs/proc: fix uaf in proc_readdir_de()
pmdomain: imx: Fix reference count leak in imx_gpc_remove
pmdomain: arm: scmi: Fix genpd leak on provider registration failure
net: netpoll: fix incorrect refcount handling causing incorrect cleanup
net: qede: Initialize qede_ll_ops with designated initializer
uio_hv_generic: Set event for all channels on the device
net: ethernet: ti: netcp: Standardize knav_dma_open_channel to return NULL on error
ALSA: usb-audio: fix uac2 clock source at terminal parser
mm/page_alloc: fix hash table order logging in alloc_large_system_hash()
kconfig/nconf: Initialize the default locale at startup
kconfig/mconf: Initialize the default locale at startup
vsock: Ignore signal/timeout on connect() if already established
s390/ctcm: Fix double-kfree
net: openvswitch: remove never-working support for setting nsh fields
mlxsw: spectrum: Fix memory leak in mlxsw_sp_flower_stats()
MIPS: Malta: Fix !EVA SOC-it PCI MMIO
scsi: target: tcm_loop: Fix segfault in tcm_loop_tpg_address_show()
scsi: sg: Do not sleep in atomic context
Input: cros_ec_keyb - fix an invalid memory access
be2net: pass wrb_params in case of OS2BMC
HID: quirks: work around VID/PID conflict for 0x4c4a/0x4155
isdn: mISDN: hfcsusb: fix memory leak in hfcsusb_probe()
EDAC/altera: Use INTTEST register for Ethernet and USB SBE injection
EDAC/altera: Handle OCRAM ECC enable after warm reset
spi: Try to get ACPI GPIO IRQ earlier
ipv4: route: Prevent rt_bind_exception() from rebinding stale fnhe
strparser: Fix signed/unsigned mismatch bug
gcov: add support for GCC 15
mm/ksm: fix flag-dropping behavior in ksm_madvise
ALSA: usb-audio: Fix NULL pointer dereference in snd_usb_mixer_controls_badd
drm/vmwgfx: Validate command header size against SVGA_CMD_MAX_DATASIZE
ASoC: cs4271: Fix regulator leak on probe failure
regulator: fixed: fix GPIO descriptor leak on register failure
regulator: fixed: use dev_err_probe for register
Bluetooth: L2CAP: export l2cap_chan_hold for modules
net_sched: limit try_bulk_dequeue_skb() batches
net_sched: remove need_resched() from qdisc_run()
net/mlx5e: Fix wraparound in rate limiting for values above 255 Gbps
net/mlx5e: Fix maxrate wraparound in threshold between units
net: sched: act_ife: initialize struct tc_ife to fix KMSAN kernel-infoleak
wifi: mac80211: skip rate verification for not captured PSDUs
net: mdio: fix resource leak in mdiobus_register_device()
tipc: Fix use-after-free in tipc_mon_reinit_self().
tipc: simplify the finalize work queue
sctp: prevent possible shift-out-of-bounds in sctp_transport_update_rto
sctp: get netns from asoc and ep base
Bluetooth: 6lowpan: Don't hold spin lock over sleeping functions
Bluetooth: 6lowpan: fix BDADDR_LE vs ADDR_LE_DEV address type confusion
Bluetooth: 6lowpan: reset link-local header on ipv6 recv path
Bluetooth: btusb: reorder cleanup in btusb_disconnect to avoid UAF
net: fec: correct rx_bytes statistic for the case SHIFT16 is set
ASoC: max98090/91: fixed max98091 ALSA widget powering up/down
HID: quirks: avoid Cooler Master MM712 dongle wakeup bug
NFS4: Fix state renewals missing after boot
compiler_types: Move unused static inline functions warning to W=2
extcon: adc-jack: Cleanup wakeup source only if it was enabled
tracing: Fix memory leaks in create_field_var()
net: usb: qmi_wwan: initialize MAC header offset in qmimux_rx_fixup
sctp: Prevent TOCTOU out-of-bounds write
sctp: Hold RCU read lock while iterating over address list
net: dsa: b53: stop reading ARL entries if search is done
net: dsa: b53: fix enabling ip multicast
net: dsa: b53: fix resetting speed and pause on forced link
net: dsa: b53: prevent GMII_PORT_OVERRIDE_CTRL access on BCM5325
net: dsa/b53: change b53_force_port_config() pause argument
net: vlan: sync VLAN features with lower device
ceph: add checking of wait_for_completion_killable() return value
fbdev: Add bounds checking in bit_putcs to fix vmalloc-out-of-bounds
ACPI: property: Return present device nodes only on fwnode interface
9p: sysfs_init: don't hardcode error to ENOMEM
9p: fix /sys/fs/9p/caches overwriting itself
fs/hpfs: Fix error code for new_inode() failure in mkdir/create/mknod/symlink
ACPICA: Update dsmethod.c to get rid of unused variable warning
orangefs: fix xattr related buffer overflow...
page_pool: Clamp pool size to max 16K pages
Bluetooth: bcsp: receive data only if registered
Bluetooth: SCO: Fix UAF on sco_conn_free
net: macb: avoid dealing with endianness in macb_set_hwaddr()
nfs4_setup_readdir(): insufficient locking for ->d_parent->d_inode dereferencing
NFSv4.1: fix mount hang after CREATE_SESSION failure
NFSv4: handle ERR_GRACE on delegation recalls
remoteproc: qcom: q6v5: Avoid handling handover twice
sparc/module: Add R_SPARC_UA64 relocation handling
net: intel: fm10k: Fix parameter idx set but not used
jfs: fix uninitialized waitqueue in transaction manager
jfs: Verify inode mode when loading from disk
ipv6: np->rxpmtu race annotation
usb: xhci: plat: Facilitate using autosuspend for xhci plat devices
usb: mon: Increase BUFF_MAX to 64 MiB to support multi-MB URBs
allow finish_no_open(file, ERR_PTR(-E...))
scsi: lpfc: Define size of debugfs entry for xri rebalancing
scsi: lpfc: Check return status of lpfc_reset_flush_io_context during TGT_RESET
selftests/Makefile: include $(INSTALL_DEP_TARGETS) in clean target to clean net/lib dependency
net/cls_cgroup: Fix task_get_classid() during qdisc run
selftests: Replace sleep with slowwait
selftests: Disable dad for ipv6 in fcnal-test.sh
media: redrat3: use int type to store negative error codes
net: sh_eth: Disable WoL if system can not suspend
phy: cadence: cdns-dphy: Enable lower resolutions in dphy
usb: gadget: f_hid: Fix zero length packet transfer
net: call cond_resched() less often in __release_sock()
ALSA: usb-audio: apply quirk for MOONDROP Quark2
net: nfc: nci: Increase NCI_DATA_TIMEOUT to 3000 ms
dmaengine: dw-edma: Set status for callback_result
dmaengine: mv_xor: match alloc_wc and free_wc
dmaengine: sh: setup_xref error handling
scsi: pm8001: Use int instead of u32 to store error codes
mips: lantiq: xway: sysctrl: rename stp clock
mips: lantiq: danube: add missing device_type in pci node
mips: lantiq: danube: add missing properties to cpu node
media: fix uninitialized symbol warnings
drm/amdkfd: Tie UNMAP_LATENCY to queue_preemption
extcon: adc-jack: Fix wakeup source leaks on device unbind
rds: Fix endianness annotation for RDS_MPATH_HASH
PCI/P2PDMA: Fix incorrect pointer usage in devm_kfree() call
net: Call trace_sock_exceed_buf_limit() for memcg failure with SK_MEM_RECV.
net: When removing nexthops, don't call synchronize_net if it is not necessary
char: misc: Does not request module for miscdevice with dynamic minor
usb: gadget: f_ncm: Fix MAC assignment NCM ethernet
iio: adc: spear_adc: mask SPEAR_ADC_STATUS channel and avg sample before setting register
media: imon: make send_packet() more robust
net: ipv6: fix field-spanning memcpy warning in AH output
bridge: Redirect to backup port when port is administratively down
powerpc/eeh: Use result of error_detected() in uevent
x86/vsyscall: Do not require X86_PF_INSTR to emulate vsyscall
media: pci: ivtv: Don't create fake v4l2_fh
drm/amdkfd: return -ENOTTY for unsupported IOCTLs
selftests/net: Ensure assert() triggers in psock_tpacket.c
selftests/net: Replace non-standard __WORDSIZE with sizeof(long) * 8
PCI: Disable MSI on RDC PCI to PCIe bridges
drm/nouveau: replace snprintf() with scnprintf() in nvkm_snprintbf()
mfd: madera: Work around false-positive -Wininitialized warning
mfd: stmpe-i2c: Add missing MODULE_LICENSE
mfd: stmpe: Remove IRQ domain upon removal
tools/power x86_energy_perf_policy: Prefer driver HWP limits
tools/power x86_energy_perf_policy: Enhance HWP enable
tools/cpupower: Fix incorrect size in cpuidle_state_disable()
hwmon: (dell-smm) Add support for Dell OptiPlex 7040
uprobe: Do not emulate/sstep original instruction when ip is changed
clocksource/drivers/vf-pit: Replace raw_readl/writel to readl/writel
video: backlight: lp855x_bl: Set correct EPROM start for LP8556
tee: allow a driver to allocate a tee_device without a pool
ACPICA: dispatcher: Use acpi_ds_clear_operands() in acpi_ds_call_control_method()
mmc: sdhci-msm: Enable tuning for SDR50 mode for SD card
irqchip/gic-v2m: Handle Multiple MSI base IRQ Alignment
arc: Fix __fls() const-foldability via __builtin_clzl()
cpufreq/longhaul: handle NULL policy in longhaul_exit
selftests/bpf: Fix bpf_prog_detach2 usage in test_lirc_mode2
ACPI: video: force native for Lenovo 82K8
memstick: Add timeout to prevent indefinite waiting
mmc: host: renesas_sdhi: Fix the actual clock
bpf: Don't use %pK through printk
spi: loopback-test: Don't use %pK through printk
soc: qcom: smem: Fix endian-unaware access of num_entries
usb: gadget: f_fs: Fix epfile null pointer access after ep enable.
serial: 8250_dw: handle reset control deassert error
serial: 8250_dw: Use devm_add_action_or_reset()
serial: 8250_dw: Use devm_clk_get_optional() to get the input clock
can: gs_usb: increase max interface to U8_MAX
devcoredump: Fix circular locking dependency with devcd->mutex.
net: ravb: Enforce descriptor type ordering
x86/resctrl: Fix miscount of bandwidth event when reactivating previously unavailable RMID
wifi: brcmfmac: fix crash while sending Action Frames in standalone AP Mode
net: phy: dp83867: Disable EEE support as not implemented
regmap: slimbus: fix bus_context pointer in regmap init calls
drm/etnaviv: fix flush sequence logic
usbnet: Prevents free active kevent
wifi: ath10k: Fix memory leak on unsupported WMI command
ASoC: qdsp6: q6asm: do not sleep while atomic
fbdev: valkyriefb: Fix reference count leak in valkyriefb_init
fbdev: pvr2fb: Fix leftover reference to ONCHIP_NR_DMA_CHANNELS
fbdev: bitblit: bound-check glyph index in bit_putcs*
ACPI: video: Fix use-after-free in acpi_video_switch_brightness()
fbdev: atyfb: Check if pll_ops->init_pll failed
net: usb: asix_devices: Check return value of usbnet_get_endpoints
btrfs: use smp_mb__after_atomic() when forcing COW in create_pending_snapshot()
x86/bugs: Fix reporting of LFENCE retpoline
net/sched: sch_qfq: Fix null-deref in agg_dequeue
Conflicts:
drivers/mmc/host/sdhci-msm.c
drivers/usb/host/xhci-plat.c
Change-Id: I30739fea8840ace3825ea8955598d9efa687bb27
https://source.android.com/docs/security/bulletin/2025-12-01
CVE-2025-48623
CVE-2025-48624
CVE-2025-48637
CVE-2025-48638
CVE-2024-35970
CVE-2025-38236
CVE-2025-38349
CVE-2025-48610
CVE-2025-38500
* tag 'ASB-2025-12-01_11-5.4' of https://android.googlesource.com/kernel/common:
UPSTREAM: crypto: essiv - Check ssize for decryption and in-place encryption
ANDROID: GKI: fix up build break where timer_delete_sync() was used
Revert "net: rtnetlink: remove redundant assignment to variable err"
Revert "net: rtnetlink: add msg kind names"
Revert "net: rtnetlink: add helper to extract msg type's kind"
Revert "net: rtnetlink: use BIT for flag values"
Revert "net: netlink: add NLM_F_BULK delete request modifier"
Revert "net: rtnetlink: add bulk delete support flag"
Revert "net: rtnetlink: fix module reference count leak issue in rtnetlink_rcv_msg"
Revert "net: add ndo_fdb_del_bulk"
Revert "net: rtnetlink: add NLM_F_BULK support to rtnl_fdb_del"
Revert "rtnetlink: Allow deleting FDB entries in user namespace"
Linux 5.4.301
net: rtnetlink: fix module reference count leak issue in rtnetlink_rcv_msg
media: s5p-mfc: remove an unused/uninitialized variable
NFSD: Fix last write offset handling in layoutcommit
NFSD: Minor cleanup in layoutcommit processing
padata: Reset next CPU when reorder sequence wraps around
KEYS: trusted_tpm1: Compare HMAC values in constant time
NFSD: Define a proc_layoutcommit for the FlexFiles layout type
vfs: Don't leak disconnected dentries on umount
jbd2: ensure that all ongoing I/O complete before freeing blocks
ext4: detect invalid INLINE_DATA + EXTENTS flag combination
drm/amdgpu: use atomic functions with memory barriers for vm fault info
ext4: avoid potential buffer over-read in parse_apply_sb_mount_options()
spi: cadence-quadspi: Flush posted register writes before DAC access
spi: cadence-quadspi: Flush posted register writes before INDAC access
memory: samsung: exynos-srom: Fix of_iomap leak in exynos_srom_probe
memory: samsung: exynos-srom: Correct alignment
arm64: errata: Apply workarounds for Neoverse-V3AE
arm64: cputype: Add Neoverse-V3AE definitions
comedi: fix divide-by-zero in comedi_buf_munge()
binder: remove "invalid inc weak" check
xhci: dbc: enable back DbC in resume if it was enabled before suspend
usb/core/quirks: Add Huawei ME906S to wakeup quirk
USB: serial: option: add Telit FN920C04 ECM compositions
USB: serial: option: add Quectel RG255C
USB: serial: option: add UNISOC UIS7720
net: ravb: Ensure memory write completes before ringing TX doorbell
net: usb: rtl8150: Fix frame padding
ocfs2: clear extent cache after moving/defragmenting extents
MIPS: Malta: Fix keyboard resource preventing i8042 driver from registering
Revert "cpuidle: menu: Avoid discarding useful information"
net: bonding: fix possible peer notify event loss or dup issue
sctp: avoid NULL dereference when chunk data buffer is missing
arm64, mm: avoid always making PTE dirty in pte_mkwrite()
net: enetc: correct the value of ENETC_RXB_TRUESIZE
rtnetlink: Allow deleting FDB entries in user namespace
net: rtnetlink: add NLM_F_BULK support to rtnl_fdb_del
net: add ndo_fdb_del_bulk
net: rtnetlink: add bulk delete support flag
net: netlink: add NLM_F_BULK delete request modifier
net: rtnetlink: use BIT for flag values
net: rtnetlink: add helper to extract msg type's kind
net: rtnetlink: add msg kind names
net: rtnetlink: remove redundant assignment to variable err
m68k: bitops: Fix find_*_bit() signatures
hfsplus: return EIO when type of hidden directory mismatch in hfsplus_fill_super()
hfs: fix KMSAN uninit-value issue in hfs_find_set_zero_bits()
dlm: check for defined force value in dlm_lockspace_release
hfsplus: fix KMSAN uninit-value issue in hfsplus_delete_cat()
hfs: validate record offset in hfsplus_bmap_alloc
hfsplus: fix KMSAN uninit-value issue in __hfsplus_ext_cache_extent()
hfs: make proper initalization of struct hfs_find_data
hfs: clear offset and space out of valid records in b-tree node
exec: Fix incorrect type for ret
hfsplus: fix slab-out-of-bounds read in hfsplus_strcasecmp()
ALSA: firewire: amdtp-stream: fix enum kernel-doc warnings
sched/fair: Fix pelt lost idle time detection
sched/balancing: Rename newidle_balance() => sched_balance_newidle()
sched/fair: Trivial correction of the newidle_balance() comment
sched: Make newidle_balance() static again
tls: don't rely on tx_work during send()
tls: always set record_type in tls_process_cmsg
tg3: prevent use of uninitialized remote_adv and local_adv variables
tcp: fix tcp_tso_should_defer() vs large RTT
amd-xgbe: Avoid spurious link down messages during interface toggle
net/ip6_tunnel: Prevent perpetual tunnel growth
net: dlink: handle dma_map_single() failure properly
net: dl2k: switch from 'pci_' to 'dma_' API
media: pci: ivtv: Add missing check after DMA map
media: pci/ivtv: switch from 'pci_' to 'dma_' API
xen/events: Update virq_to_irq on migration
media: lirc: Fix error handling in lirc_register()
media: rc: Directly use ida_free()
drm/exynos: exynos7_drm_decon: remove ctx->suspended
btrfs: avoid potential out-of-bounds in btrfs_encode_fh()
pwm: berlin: Fix wrong register in suspend/resume
media: cx18: Add missing check after DMA map
xen/events: Cleanup find_virq() return codes
cramfs: Verify inode mode when loading from disk
fs: Add 'initramfs_options' to set initramfs mount options
pid: Add a judgment for ns null in pid_nr_ns
minixfs: Verify inode mode when loading from disk
tracing: Fix race condition in kprobe initialization causing NULL pointer dereference
dm: fix NULL pointer dereference in __dm_suspend()
mfd: intel_soc_pmic_chtdc_ti: Set use_single_read regmap_config flag
mfd: intel_soc_pmic_chtdc_ti: Drop unneeded assignment for cache_type
mfd: intel_soc_pmic_chtdc_ti: Fix invalid regmap-config max_register value
Squashfs: reject negative file sizes in squashfs_read_inode()
Squashfs: add additional inode sanity checking
media: mc: Clear minor number before put device
mfd: vexpress-sysreg: Check the return value of devm_gpiochip_add_data()
fs: udf: fix OOB read in lengthAllocDescs handling
KVM: x86: Don't (re)check L1 intercepts when completing userspace I/O
net/9p: fix double req put in p9_fd_cancelled
ext4: guard against EA inode refcount underflow in xattr update
ext4: correctly handle queries for metadata mappings
ext4: increase i_disksize to offset + len in ext4_update_disksize_before_punch()
nfsd: nfserr_jukebox in nlm_fopen should lead to a retry
x86/umip: Fix decoding of register forms of 0F 01 (SGDT and SIDT aliases)
x86/umip: Check that the instruction opcode is at least two bytes
PCI: keystone: Use devm_request_irq() to free "ks-pcie-error-irq" on exit
PCI/AER: Fix missing uevent on recovery when a reset is requested
PCI/IOV: Add PCI rescan-remove locking when enabling/disabling SR-IOV
rseq/selftests: Use weak symbol reference, not definition, to link with glibc
rtc: interface: Fix long-standing race when setting alarm
rtc: interface: Ensure alarm irq is enabled when UIE is enabled
mmc: core: SPI mode remove cmd7
mtd: rawnand: fsmc: Default to autodetect buswidth
sparc: fix error handling in scan_one_device()
sparc64: fix hugetlb for sun4u
sctp: Fix MAC comparison to be constant-time
scsi: hpsa: Fix potential memory leak in hpsa_big_passthru_ioctl()
parisc: don't reference obsolete termio struct for TC* constants
lib/genalloc: fix device leak in of_gen_pool_get()
iio: frequency: adf4350: Fix prescaler usage.
iio: dac: ad5421: use int type to store negative error codes
iio: dac: ad5360: use int type to store negative error codes
crypto: atmel - Fix dma_unmap_sg() direction
cpufreq: intel_pstate: Fix object lifecycle issue in update_qos_request()
drm/nouveau: fix bad ret code in nouveau_bo_move_prep
media: i2c: mt9v111: fix incorrect type for ret
firmware: meson_sm: fix device leak at probe
xen/manage: Fix suspend error path
arm64: dts: qcom: msm8916: Add missing MDSS reset
ACPI: debug: fix signedness issues in read/write helpers
ACPI: TAD: Add missing sysfs_remove_group() for ACPI_TAD_RT
tpm_tis: Fix incorrect arguments in tpm_tis_probe_irq_single
tpm, tpm_tis: Claim locality before writing interrupt registers
crypto: essiv - Check ssize for decryption and in-place encryption
mailbox: zynqmp-ipi: Remove dev.parent check in zynqmp_ipi_free_mboxes
mailbox: zynqmp-ipi: Remove redundant mbox_controller_unregister() call
tools build: Align warning options with perf
net: fsl_pq_mdio: Fix device node reference leak in fsl_pq_mdio_probe
tcp: Don't call reqsk_fastopen_remove() in tcp_conn_request().
net/sctp: fix a null dereference in sctp_disposition sctp_sf_do_5_1D_ce()
drm/vmwgfx: Fix Use-after-free in validation
net/mlx4: prevent potential use after free in mlx4_en_do_uc_filter()
scsi: mvsas: Fix use-after-free bugs in mvs_work_queue
scsi: mvsas: Use sas_task_find_rq() for tagging
scsi: mvsas: Delete mvs_tag_init()
scsi: libsas: Add sas_task_find_rq()
clk: nxp: Fix pll0 rate check condition in LPC18xx CGU driver
clk: nxp: lpc18xx-cgu: convert from round_rate() to determine_rate()
perf session: Fix handling when buffer exceeds 2 GiB
rtc: x1205: Fix Xicor X1205 vendor prefix
perf util: Fix compression checks returning -1 as bool
iio: frequency: adf4350: Fix ADF4350_REG3_12BIT_CLKDIV_MODE
clocksource/drivers/clps711x: Fix resource leaks in error paths
pinctrl: check the return value of pinmux_ops::get_function_name()
Input: uinput - zero-initialize uinput_ff_upload_compat to avoid info leak
mm: hugetlb: avoid soft lockup when mprotect to large memory area
uio_hv_generic: Let userspace take care of interrupt mask
Squashfs: fix uninit-value in squashfs_get_parent
Revert "net/mlx5e: Update and set Xon/Xoff upon MTU set"
net: ena: return 0 in ena_get_rxfh_key_size() when RSS hash key is not configurable
nfp: fix RSS hash key size when RSS is not supported
drivers/base/node: fix double free in register_one_node()
ocfs2: fix double free in user_cluster_connect()
net: usb: Remove disruptive netif_wake_queue in rtl8150_set_multicast
RDMA/siw: Always report immediate post SQ errors
usb: vhci-hcd: Prevent suspending virtually attached devices
scsi: mpt3sas: Fix crash in transport port remove by using ioc_info()
ipvs: Defer ip_vs_ftp unregister during netns cleanup
NFSv4.1: fix backchannel max_resp_sz verification check
remoteproc: qcom: q6v5: Avoid disabling handover IRQ twice
sparc: fix accurate exception reporting in copy_{from,to}_user for M7
sparc: fix accurate exception reporting in copy_to_user for Niagara 4
sparc: fix accurate exception reporting in copy_{from_to}_user for Niagara
sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC III
sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC
IB/sa: Fix sa_local_svc_timeout_ms read race
RDMA/core: Resolve MAC of next-hop device without ARP support
wifi: mt76: fix potential memory leak in mt76_wmac_probe()
drivers/base/node: handle error properly in register_one_node()
watchdog: mpc8xxx_wdt: Reload the watchdog timer when enabling the watchdog
netfilter: ipset: Remove unused htable_bits in macro ahash_region
iio: consumers: Fix offset handling in iio_convert_raw_to_processed()
ASoC: Intel: bytcr_rt5651: Fix invalid quirk input mapping
ASoC: Intel: bytcr_rt5640: Fix invalid quirk input mapping
ASoC: Intel: bytcht_es8316: Fix invalid quirk input mapping
pps: fix warning in pps_register_cdev when register device fail
misc: genwqe: Fix incorrect cmd field being reported in error
usb: gadget: configfs: Correctly set use_os_string at bind
usb: phy: twl6030: Fix incorrect type for ret
tcp: fix __tcp_close() to only send RST when required
PCI: tegra: Fix devm_kcalloc() argument order for port->phys allocation
wifi: mwifiex: send world regulatory domain to driver
ALSA: lx_core: use int type to store negative error codes
media: rj54n1cb0c: Fix memleak in rj54n1_probe()
scsi: myrs: Fix dma_alloc_coherent() error check
scsi: pm80xx: Fix array-index-out-of-of-bounds on rmmod
serial: max310x: Add error checking in probe()
usb: host: max3421-hcd: Fix error pointer dereference in probe cleanup
drm/radeon/r600_cs: clean up of dead code in r600_cs
i2c: designware: Add disabling clocks when probe fails
i2c: mediatek: fix potential incorrect use of I2C_MASTER_WRRD
bpf: Explicitly check accesses to bpf_sock_addr
selftests: watchdog: skip ping loop if WDIOF_KEEPALIVEPING not supported
pwm: tiehrpwm: Fix corner case in clock divisor calculation
block: use int to store blk_stack_limits() return value
blk-mq: check kobject state_in_sysfs before deleting in blk_mq_unregister_hctx
pinctrl: meson-gxl: add missing i2c_d pinmux
soc: qcom: rpmh-rsc: Unconditionally clear _TRIGGER bit for TCS
ACPI: processor: idle: Fix memory leak when register cpuidle device failed
regmap: Remove superfluous check for !config in __regmap_init()
x86/vdso: Fix output operand size of RDPID
perf: arm_spe: Prevent overflow in PERF_IDX2OFF()
driver core/PM: Set power.no_callbacks along with power.no_pm
staging: axis-fifo: flush RX FIFO on read errors
staging: axis-fifo: fix maximum TX packet length check
perf subcmd: avoid crash in exclude_cmds when excludes is empty
dm-integrity: limit MAX_TAG_SIZE to 255
wifi: rtlwifi: rtl8192cu: Don't claim USB ID 07b8:8188
USB: serial: option: add SIMCom 8230C compositions
media: rc: fix races with imon_disconnect()
media: imon: grab lock earlier in imon_ir_change_protocol()
media: imon: reorganize serialization
media: rc: Add support for another iMON 0xffdc device
media: i2c: tc358743: Fix use-after-free bugs caused by orphan timer in probe
media: tuner: xc5000: Fix use-after-free in xc5000_release
media: tunner: xc5000: Refactor firmware load
udp: Fix memory accounting leak.
media: b2c2: Fix use-after-free causing by irq_check_work in flexcop_pci_remove
scsi: target: target_core_configfs: Add length check to avoid buffer overflow
Conflicts:
drivers/soc/qcom/rpmh-rsc.c
kernel/sched/fair.c
Change-Id: I58ab24a3db8be4c698c41fd47daeb1f1fb7884ee