mirror of
https://github.com/BobTheBlinker/android_kernel_motorola_sm6375.git
synced 2026-10-06 20:03:33 -04:00
981,865 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a8558dff46 |
BACKPORT: mm, kfence: insert KFENCE hooks for SLUB
Inserts KFENCE hooks into the SLUB allocator. To pass the originally requested size to KFENCE, add an argument 'orig_size' to slab_alloc*(). The additional argument is required to preserve the requested original size for kmalloc() allocations, which uses size classes (e.g. an allocation of 272 bytes will return an object of size 512). Therefore, kmem_cache::size does not represent the kmalloc-caller's requested size, and we must introduce the argument 'orig_size' to propagate the originally requested size to KFENCE. Without the originally requested size, we would not be able to detect out-of-bounds accesses for objects placed at the end of a KFENCE object page if that object is not equal to the kmalloc-size class it was bucketed into. When KFENCE is disabled, there is no additional overhead, since slab_alloc*() functions are __always_inline. Link: https://lkml.kernel.org/r/20201103175841.3495947-6-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Signed-off-by: Alexander Potapenko <glider@google.com> Reviewed-by: Dmitry Vyukov <dvyukov@google.com> Reviewed-by: Jann Horn <jannh@google.com> Co-developed-by: Marco Elver <elver@google.com> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Andrey Ryabinin <aryabinin@virtuozzo.com> Cc: Andy Lutomirski <luto@kernel.org> Cc: Borislav Petkov <bp@alien8.de> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Christopher Lameter <cl@linux.com> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: David Rientjes <rientjes@google.com> Cc: Eric Dumazet <edumazet@google.com> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Hillf Danton <hdanton@sina.com> Cc: "H. Peter Anvin" <hpa@zytor.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: Joern Engel <joern@purestorage.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Joonsoo Kim <iamjoonsoo.kim@lge.com> Cc: Kees Cook <keescook@chromium.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Paul E. McKenney <paulmck@kernel.org> Cc: Pekka Enberg <penberg@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: SeongJae Park <sjpark@amazon.de> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Vlastimil Babka <vbabka@suse.cz> Cc: Will Deacon <will@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit 5aa4d4f541f500cd63f814ae39bea55024ed5c6b https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) [glider: dropped changes to include/linux/slub_def.h, fixed conflict in mm/slub.c, always pass root cache pointer to kfence_alloc()] Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: Iacfc22088087cb920107a595f71b09e21d5f2f47 |
||
|
|
5ae04c6867 |
BACKPORT: mm, kfence: insert KFENCE hooks for SLAB
Inserts KFENCE hooks into the SLAB allocator. To pass the originally requested size to KFENCE, add an argument 'orig_size' to slab_alloc*(). The additional argument is required to preserve the requested original size for kmalloc() allocations, which uses size classes (e.g. an allocation of 272 bytes will return an object of size 512). Therefore, kmem_cache::size does not represent the kmalloc-caller's requested size, and we must introduce the argument 'orig_size' to propagate the originally requested size to KFENCE. Without the originally requested size, we would not be able to detect out-of-bounds accesses for objects placed at the end of a KFENCE object page if that object is not equal to the kmalloc-size class it was bucketed into. When KFENCE is disabled, there is no additional overhead, since slab_alloc*() functions are __always_inline. Link: https://lkml.kernel.org/r/20201103175841.3495947-5-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Signed-off-by: Alexander Potapenko <glider@google.com> Reviewed-by: Dmitry Vyukov <dvyukov@google.com> Co-developed-by: Marco Elver <elver@google.com> Cc: Christoph Lameter <cl@linux.com> Cc: Pekka Enberg <penberg@kernel.org> Cc: David Rientjes <rientjes@google.com> Cc: Joonsoo Kim <iamjoonsoo.kim@lge.com> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Andrey Ryabinin <aryabinin@virtuozzo.com> Cc: Andy Lutomirski <luto@kernel.org> Cc: Borislav Petkov <bp@alien8.de> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: Eric Dumazet <edumazet@google.com> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Hillf Danton <hdanton@sina.com> Cc: "H. Peter Anvin" <hpa@zytor.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: Jann Horn <jannh@google.com> Cc: Joern Engel <joern@purestorage.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Kees Cook <keescook@chromium.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Paul E. McKenney <paulmck@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: SeongJae Park <sjpark@amazon.de> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Vlastimil Babka <vbabka@suse.cz> Cc: Will Deacon <will@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit 840c0553e89413319d67971a321bcc07114da9b8 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) [glider: resolved minor API change in mm/slab_common.c, dropped changes to include/linux/slab_def.h, resolved conflicts in mm/slab.c, always pass root cache pointer to kfence_alloc] Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: Iab2ba9c7b06b9a234d93ba892be639941861f8ab |
||
|
|
40edd642cb |
FROMGIT: mm/slab: rerform init_on_free earlier
Currently in CONFIG_SLAB init_on_free happens too late, and heap objects go to the heap quarantine not being erased. Lets move init_on_free clearing before calling kasan_slab_free(). In that case heap quarantine will store erased objects, similarly to CONFIG_SLUB=y behavior. Link: https://lkml.kernel.org/r/20201210183729.1261524-1-alex.popov@linux.com Signed-off-by: Alexander Popov <alex.popov@linux.com> Reviewed-by: Alexander Potapenko <glider@google.com> Acked-by: David Rientjes <rientjes@google.com> Acked-by: Joonsoo Kim <iamjoonsoo.kim@lge.com> Cc: Christoph Lameter <cl@linux.com> Cc: Pekka Enberg <penberg@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org> Bug: 177201466 (cherry picked from commit a32d654db543843a5ffb248feaec1a909718addd https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I2bf5e70c1524619526efd792bbdd959b813af1e4 |
||
|
|
b5c9938cb4 |
FROMGIT: kfence: use pt_regs to generate stack trace on faults
Instead of removing the fault handling portion of the stack trace based on the fault handler's name, just use struct pt_regs directly. Change kfence_handle_page_fault() to take a struct pt_regs, and plumb it through to kfence_report_error() for out-of-bounds, use-after-free, or invalid access errors, where pt_regs is used to generate the stack trace. If the kernel is a DEBUG_KERNEL, also show registers for more information. Link: https://lkml.kernel.org/r/20201105092133.2075331-1-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Suggested-by: Mark Rutland <mark.rutland@arm.com> Acked-by: Mark Rutland <mark.rutland@arm.com> Cc: Alexander Potapenko <glider@google.com> Cc: Dmitry Vyukov <dvyukov@google.com> Cc: Alexander Potapenko <glider@google.com> Cc: Jann Horn <jannh@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit 54a5abe9b5d542ee71836439cc662efe178c8211 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I3a60060b24f0efb4faee2e6c953973bc1263e8d1 |
||
|
|
3670dc7688 |
FROMGIT: kfence, arm64: add missing copyright and description header
Add missing copyright and description header to KFENCE source file. Link: https://lkml.kernel.org/r/20210118092159.145934-3-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Reviewed-by: Alexander Potapenko <glider@google.com> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Dmitry Vyukov <dvyukov@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit ce19707fadfe62b7441f2a9455b6866161fbc1f0 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I75c06e57e4bb906626c13318d00cfc7d47b76bde |
||
|
|
b5e331effb |
BACKPORT: arm64, kfence: enable KFENCE for ARM64
Add architecture specific implementation details for KFENCE and enable KFENCE for the arm64 architecture. In particular, this implements the required interface in <asm/kfence.h>. KFENCE requires that attributes for pages from its memory pool can individually be set. Therefore, force the entire linear map to be mapped at page granularity. Doing so may result in extra memory allocated for page tables in case rodata=full is not set; however, currently CONFIG_RODATA_FULL_DEFAULT_ENABLED=y is the default, and the common case is therefore not affected by this change. Link: https://lkml.kernel.org/r/20201103175841.3495947-4-elver@google.com Signed-off-by: Alexander Potapenko <glider@google.com> Signed-off-by: Marco Elver <elver@google.com> Reviewed-by: Dmitry Vyukov <dvyukov@google.com> Co-developed-by: Alexander Potapenko <glider@google.com> Reviewed-by: Jann Horn <jannh@google.com> Reviewed-by: Mark Rutland <mark.rutland@arm.com> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Andrey Ryabinin <aryabinin@virtuozzo.com> Cc: Andy Lutomirski <luto@kernel.org> Cc: Borislav Petkov <bp@alien8.de> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Christopher Lameter <cl@linux.com> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: David Rientjes <rientjes@google.com> Cc: Eric Dumazet <edumazet@google.com> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Hillf Danton <hdanton@sina.com> Cc: "H. Peter Anvin" <hpa@zytor.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: Joern Engel <joern@purestorage.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Joonsoo Kim <iamjoonsoo.kim@lge.com> Cc: Kees Cook <keescook@chromium.org> Cc: Paul E. McKenney <paulmck@kernel.org> Cc: Pekka Enberg <penberg@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: SeongJae Park <sjpark@amazon.de> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Vlastimil Babka <vbabka@suse.cz> Cc: Will Deacon <will@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit d3cb0555da9b469c3f9ebe663d4c0d6265757175 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) [glider: resolved a minor conflict in arch/arm64/Kconfig] Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: Ibe9d26c2c9088862ba8855c18e0c5210a68c7f50 |
||
|
|
097baa9924 |
FROMGIT: kfence, x86: add missing copyright and description header
Add missing copyright and description header to KFENCE source file. Link: https://lkml.kernel.org/r/20210118092159.145934-2-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Reviewed-by: Alexander Potapenko <glider@google.com> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Dmitry Vyukov <dvyukov@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit 8470d53956b7e1b6f3f2dd17f277b9d2bb1472ae https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I90aef6f80e7b5a46f5168d1a24c02a5737785831 |
||
|
|
1a8ea8501d |
BACKPORT: x86, kfence: enable KFENCE for x86
Add architecture specific implementation details for KFENCE and enable KFENCE for the x86 architecture. In particular, this implements the required interface in <asm/kfence.h> for setting up the pool and providing helper functions for protecting and unprotecting pages. For x86, we need to ensure that the pool uses 4K pages, which is done using the set_memory_4k() helper function. Link: https://lkml.kernel.org/r/20201103175841.3495947-3-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Signed-off-by: Alexander Potapenko <glider@google.com> Reviewed-by: Dmitry Vyukov <dvyukov@google.com> Co-developed-by: Marco Elver <elver@google.com> Reviewed-by: Jann Horn <jannh@google.com> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Andrey Ryabinin <aryabinin@virtuozzo.com> Cc: Andy Lutomirski <luto@kernel.org> Cc: Borislav Petkov <bp@alien8.de> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Christopher Lameter <cl@linux.com> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: David Rientjes <rientjes@google.com> Cc: Eric Dumazet <edumazet@google.com> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Hillf Danton <hdanton@sina.com> Cc: "H. Peter Anvin" <hpa@zytor.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: Joern Engel <joern@purestorage.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Joonsoo Kim <iamjoonsoo.kim@lge.com> Cc: Kees Cook <keescook@chromium.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Paul E. McKenney <paulmck@kernel.org> Cc: Pekka Enberg <penberg@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: SeongJae Park <sjpark@amazon.de> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Vlastimil Babka <vbabka@suse.cz> Cc: Will Deacon <will@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit cb99fcd83140d0d58ea36db6c1c2034abc95f983 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) [glider: resolved a minor conflict in arch/x86/Kconfig, s/flush_tlb_one_kernel/__flush_tlb_one_kernel] Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I111caffb0b88c34ed9ff57b95f127b08eacedcb9 |
||
|
|
0c26e79c37 |
FROMGIT: kfence: add missing copyright and description headers
Add missing copyright and description headers to KFENCE source files. Link: https://lkml.kernel.org/r/20210118092159.145934-1-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Reviewed-by: Alexander Potapenko <glider@google.com> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Dmitry Vyukov <dvyukov@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit b86d1ed1155ce1d2420057bfbdcc62b9fd53c1d6 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I95eb6756baaa0d8ca1dbc656708667372cff546d |
||
|
|
9339667c95 |
FROMGIT: kfence: add option to use KFENCE without static keys
For certain usecases, specifically where the sample interval is always set to a very low value such as 1ms, it can make sense to use a dynamic branch instead of static branches due to the overhead of toggling a static branch. Therefore, add a new Kconfig option to remove the static branches and instead check kfence_allocation_gate if a KFENCE allocation should be set up. Link: https://lkml.kernel.org/r/20210111091544.3287013-1-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Suggested-by: Jörn Engel <joern@purestorage.com> Reviewed-by: Jörn Engel <joern@purestorage.com> Cc: Alexander Potapenko <glider@google.com> Cc: Dmitry Vyukov <dvyukov@google.com> Cc: Jann Horn <jannh@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit c01761611b325c1e4ec7d3e236cc9db003cb82fd https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I68a112a8ff68fa24742b198e036f130a9757c27f |
||
|
|
fa80023f3c |
FROMGIT: kfence: fix potential deadlock due to wake_up()
Lockdep reports that we may deadlock when calling wake_up() in
__kfence_alloc(), because we may already hold base->lock. This can happen
if debug objects are enabled:
...
__kfence_alloc+0xa0/0xbc0 mm/kfence/core.c:710
kfence_alloc include/linux/kfence.h:108 [inline]
...
kmem_cache_zalloc include/linux/slab.h:672 [inline]
fill_pool+0x264/0x5c0 lib/debugobjects.c:171
__debug_object_init+0x7a/0xd10 lib/debugobjects.c:560
debug_object_init lib/debugobjects.c:615 [inline]
debug_object_activate+0x32c/0x3e0 lib/debugobjects.c:701
debug_timer_activate kernel/time/timer.c:727 [inline]
__mod_timer+0x77d/0xe30 kernel/time/timer.c:1048
...
Therefore, switch to an open-coded wait loop. The difference to before is
that the waiter wakes up and rechecks the condition after 1 jiffy;
however, given the infrequency of kfence allocations, the difference is
insignificant.
Link: https://lkml.kernel.org/r/000000000000c0645805b7f982e4@google.com
Link: https://lkml.kernel.org/r/20210104130749.1768991-1-elver@google.com
Reported-by: syzbot+8983d6d4f7df556be565@syzkaller.appspotmail.com
Signed-off-by: Marco Elver <elver@google.com>
Suggested-by: Hillf Danton <hdanton@sina.com>
Cc: Alexander Potapenko <glider@google.com>
Cc: Dmitry Vyukov <dvyukov@google.com>
Cc: Jann Horn <jannh@google.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Bug: 177201466
(cherry picked from commit c5fb1ab1a3c6d0ee02d1054a10d51ffcac57aed5
https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10)
Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish
Signed-off-by: Alexander Potapenko <glider@google.com>
Change-Id: Iee40e9f216afbc3fce8e43c0e2a4bc807fdddf39
|
||
|
|
c28a209911 |
FROMGIT: kfence: avoid stalling work queue task without allocations
To toggle the allocation gates, we set up a delayed work that calls toggle_allocation_gate(). Here we use wait_event() to await an allocation and subsequently disable the static branch again. However, if the kernel has stopped doing allocations entirely, we'd wait indefinitely, and stall the worker task. This may also result in the appropriate warnings if CONFIG_DETECT_HUNG_TASK=y. Therefore, introduce a 1 second timeout and use wait_event_timeout(). If the timeout is reached, the static branch is disabled and a new delayed work is scheduled to try setting up an allocation at a later time. Note that, this scenario is very unlikely during normal workloads once the kernel has booted and user space tasks are running. It can, however, happen during early boot after KFENCE has been enabled, when e.g. running tests that do not result in any allocations. Link: https://lkml.kernel.org/r/CADYN=9J0DQhizAGB0-jz4HOBBh+05kMBXb4c0cXMS7Qi5NAJiw@mail.gmail.com Link: https://lkml.kernel.org/r/20201110135320.3309507-1-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Reported-by: Anders Roxell <anders.roxell@linaro.org> Cc: Alexander Potapenko <glider@google.com> Cc: Dmitry Vyukov <dvyukov@google.com> Cc: SeongJae Park <sjpark@amazon.de> Cc: Jann Horn <jannh@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit 80d4693491f6f20de01437319b081fdda2079e67 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I2332ff8144b8bce5c4574b01ea2863e0e71e6124 |
||
|
|
ed71026447 |
FROMGIT: kfence: Fix parameter description for kfence_object_start()
Describe parameter @addr correctly by delimiting with ':'. Link: https://lkml.kernel.org/r/20201106092149.GA2851373@elver.google.com Signed-off-by: Marco Elver <elver@google.com> Reviewed-by: Alexander Potapenko <glider@google.com> Reported-by: Stephen Rothwell <sfr@canb.auug.org.au> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit 65f5b471bfd099a54862e14a895724e982a381c9 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: Ieb71f57f82351eaaabd3c63cd52e97fbfbfca2f1 |
||
|
|
1eee152fce |
BACKPORT: mm: add Kernel Electric-Fence infrastructure
Patch series "KFENCE: A low-overhead sampling-based memory safety error detector", v7. This adds the Kernel Electric-Fence (KFENCE) infrastructure. KFENCE is a low-overhead sampling-based memory safety error detector of heap use-after-free, invalid-free, and out-of-bounds access errors. This series enables KFENCE for the x86 and arm64 architectures, and adds KFENCE hooks to the SLAB and SLUB allocators. KFENCE is designed to be enabled in production kernels, and has near zero performance overhead. Compared to KASAN, KFENCE trades performance for precision. The main motivation behind KFENCE's design, is that with enough total uptime KFENCE will detect bugs in code paths not typically exercised by non-production test workloads. One way to quickly achieve a large enough total uptime is when the tool is deployed across a large fleet of machines. KFENCE objects each reside on a dedicated page, at either the left or right page boundaries. The pages to the left and right of the object page are "guard pages", whose attributes are changed to a protected state, and cause page faults on any attempted access to them. Such page faults are then intercepted by KFENCE, which handles the fault gracefully by reporting a memory access error. Guarded allocations are set up based on a sample interval (can be set via kfence.sample_interval). After expiration of the sample interval, the next allocation through the main allocator (SLAB or SLUB) returns a guarded allocation from the KFENCE object pool. At this point, the timer is reset, and the next allocation is set up after the expiration of the interval. To enable/disable a KFENCE allocation through the main allocator's fast-path without overhead, KFENCE relies on static branches via the static keys infrastructure. The static branch is toggled to redirect the allocation to KFENCE. The KFENCE memory pool is of fixed size, and if the pool is exhausted no further KFENCE allocations occur. The default config is conservative with only 255 objects, resulting in a pool size of 2 MiB (with 4 KiB pages). We have verified by running synthetic benchmarks (sysbench I/O, hackbench) and production server-workload benchmarks that a kernel with KFENCE (using sample intervals 100-500ms) is performance-neutral compared to a non-KFENCE baseline kernel. KFENCE is inspired by GWP-ASan [1], a userspace tool with similar properties. The name "KFENCE" is a homage to the Electric Fence Malloc Debugger [2]. For more details, see Documentation/dev-tools/kfence.rst added in the series -- also viewable here: https://raw.githubusercontent.com/google/kasan/kfence/Documentation/dev-tools/kfence.rst [1] http://llvm.org/docs/GwpAsan.html [2] https://linux.die.net/man/3/efence This patch (of 9): This adds the Kernel Electric-Fence (KFENCE) infrastructure. KFENCE is a low-overhead sampling-based memory safety error detector of heap use-after-free, invalid-free, and out-of-bounds access errors. KFENCE is designed to be enabled in production kernels, and has near zero performance overhead. Compared to KASAN, KFENCE trades performance for precision. The main motivation behind KFENCE's design, is that with enough total uptime KFENCE will detect bugs in code paths not typically exercised by non-production test workloads. One way to quickly achieve a large enough total uptime is when the tool is deployed across a large fleet of machines. KFENCE objects each reside on a dedicated page, at either the left or right page boundaries. The pages to the left and right of the object page are "guard pages", whose attributes are changed to a protected state, and cause page faults on any attempted access to them. Such page faults are then intercepted by KFENCE, which handles the fault gracefully by reporting a memory access error. To detect out-of-bounds writes to memory within the object's page itself, KFENCE also uses pattern-based redzones. The following figure illustrates the page layout: ---+-----------+-----------+-----------+-----------+-----------+--- | xxxxxxxxx | O : | xxxxxxxxx | : O | xxxxxxxxx | | xxxxxxxxx | B : | xxxxxxxxx | : B | xxxxxxxxx | | x GUARD x | J : RED- | x GUARD x | RED- : J | x GUARD x | | xxxxxxxxx | E : ZONE | xxxxxxxxx | ZONE : E | xxxxxxxxx | | xxxxxxxxx | C : | xxxxxxxxx | : C | xxxxxxxxx | | xxxxxxxxx | T : | xxxxxxxxx | : T | xxxxxxxxx | ---+-----------+-----------+-----------+-----------+-----------+--- Guarded allocations are set up based on a sample interval (can be set via kfence.sample_interval). After expiration of the sample interval, a guarded allocation from the KFENCE object pool is returned to the main allocator (SLAB or SLUB). At this point, the timer is reset, and the next allocation is set up after the expiration of the interval. To enable/disable a KFENCE allocation through the main allocator's fast-path without overhead, KFENCE relies on static branches via the static keys infrastructure. The static branch is toggled to redirect the allocation to KFENCE. To date, we have verified by running synthetic benchmarks (sysbench I/O, hackbench) that a kernel compiled with KFENCE is performance-neutral compared to the non-KFENCE baseline. For more details, see Documentation/dev-tools/kfence.rst (added later in the series). Link: https://lkml.kernel.org/r/20201103175841.3495947-2-elver@google.com Signed-off-by: Marco Elver <elver@google.com> Signed-off-by: Alexander Potapenko <glider@google.com> Reviewed-by: Dmitry Vyukov <dvyukov@google.com> Reviewed-by: SeongJae Park <sjpark@amazon.de> Co-developed-by: Marco Elver <elver@google.com> Reviewed-by: Jann Horn <jannh@google.com> Cc: "H. Peter Anvin" <hpa@zytor.com> Cc: Paul E. McKenney <paulmck@kernel.org> Cc: Andrey Konovalov <andreyknvl@google.com> Cc: Andrey Ryabinin <aryabinin@virtuozzo.com> Cc: Andy Lutomirski <luto@kernel.org> Cc: Borislav Petkov <bp@alien8.de> Cc: Catalin Marinas <catalin.marinas@arm.com> Cc: Christopher Lameter <cl@linux.com> Cc: Dave Hansen <dave.hansen@linux.intel.com> Cc: David Rientjes <rientjes@google.com> Cc: Eric Dumazet <edumazet@google.com> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org> Cc: Hillf Danton <hdanton@sina.com> Cc: Ingo Molnar <mingo@redhat.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Joonsoo Kim <iamjoonsoo.kim@lge.com> Cc: Joern Engel <joern@purestorage.com> Cc: Kees Cook <keescook@chromium.org> Cc: Mark Rutland <mark.rutland@arm.com> Cc: Pekka Enberg <penberg@kernel.org> Cc: Peter Zijlstra <peterz@infradead.org> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Vlastimil Babka <vbabka@suse.cz> Cc: Will Deacon <will@kernel.org> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Bug: 177201466 (cherry picked from commit 2a8dede73c3496bbd917644657f3735a4f508cb9 https://github.com/hnaz/linux-mm v5.11-rc4-mmots-2021-01-21-20-10) [glider: dropped KCSAN bits missing in 5.4, resolved minor conflict in init/main.c] Test: CONFIG_KFENCE_KUNIT_TEST=y passes on Cuttlefish Signed-off-by: Alexander Potapenko <glider@google.com> Change-Id: I6b474675cc9732c31118df53fa06c3997f577218 |
||
|
|
695569241b |
devfreq: bimc-bwmon: Don't free/reallocate IRQ during suspend/resume
In __resume_bw_hwmon(), request_threaded_irq() tries to make a big allocation before everything is thawed and oom killer is re-enabled, which triggers an order-2 memory allocation failure from kthreadd. This causes the devfreq device to fail to create the IRQ thread, and therefore, it can't resume. At this point, IRQs like qcom,cpu-cpu-llcc-bwmon, qcom,cpu-llcc-ddr-bwmon, qcom,snoop-l3-bwmon are not allocated. But __suspend_bw_hwmon() tries to free them anyway, causing the following series of call traces: [49231.949317] kthreadd: page allocation failure: order:2, mode:0x400d00(GFP_NOIO|__GFP_ZERO|__GFP_ACCOUNT), nodemask=(null),cpuset=/,mems_allowed=0 [49231.949341] CPU: 2 PID: 2 Comm: kthreadd Not tainted 5.4.278-Scarlet-v1.0-beta8 #1 [49231.949345] Hardware name: redwood based Qualcomm Technologies, Inc. SM7325 (DT) [49231.949349] Call trace: [49231.949366] dump_backtrace+0x0/0x1a0 [49231.949373] show_stack+0x14/0x20 [49231.949384] dump_stack+0x8c/0xcc [49231.949394] warn_alloc+0xec/0x110 [49231.949401] __alloc_pages_slowpath+0xe80/0xf00 [49231.949409] __alloc_pages_nodemask+0x150/0x180 [49231.949421] dup_task_struct+0x90/0x1e0 [49231.949427] copy_process+0x150/0xb30 [49231.949433] _do_fork+0x88/0x350 [49231.949439] kernel_thread+0x48/0x70 [49231.949448] kthreadd+0x1ec/0x280 [49231.949454] ret_from_fork+0x10/0x18 [49231.949470] Mem-Info: [49231.949488] active_anon:348299 inactive_anon:132548 isolated_anon:0 active_file:97104 inactive_file:98252 isolated_file:0 unevictable:61330 dirty:0 writeback:0 unstable:0 slab_reclaimable:23532 slab_unreclaimable:64025 mapped:208056 shmem:15179 pagetables:31685 bounce:0 free:36974 free_pcp:0 free_cma:46 [49231.949502] Node 0 active_anon:1393196kB inactive_anon:530192kB active_file:388416kB inactive_file:393008kB unevictable:245320kB isolated(anon):0kB isolated(file):0kB mapped:832224kB dirty:0kB writeback:0kB shmem:60716kB writeback_tmp:0kB unstable:0kB all_unreclaimable? no [49231.949519] Normal free:147896kB min:9312kB low:44784kB high:50268kB active_anon:1393080kB inactive_anon:530372kB active_file:388236kB inactive_file:392776kB unevictable:245320kB writepending:0kB present:5642556kB managed:5484672kB mlocked:244888kB kernel_stack:82704kB pagetables:126740kB bounce:0kB free_pcp:0kB local_pcp:0kB free_cma:184kB [49231.949522] lowmem_reserve[]: 0 0 [49231.949530] Normal: 32941*4kB (UMECH) 2020*8kB (UMEH) 35*16kB (H) 7*32kB (H) 0*64kB 0*128kB 0*256kB 0*512kB 0*1024kB 0*2048kB 0*4096kB = 148708kB [49231.949568] 257736 total pagecache pages [49231.949573] 3431 pages in swap cache [49231.949578] Swap cache stats: add 7482872, delete 7479455, find 338640/3769326 [49231.949581] Free swap = 2294120kB [49231.949584] Total swap = 4194300kB [49231.949588] 1410639 pages RAM [49231.949591] 0 pages HighMem/MovableOnly [49231.949594] 39471 pages reserved [49231.949597] 72704 pages cma reserved [49231.949606] cma: cma-0 pages: => 579 used of 5120 total pages [49231.949613] cma: cma-1 pages: => 512 used of 3072 total pages [49231.949619] cma: cma-2 pages: => 0 used of 4096 total pages [49231.949627] cma: cma-3 pages: => 826 used of 8192 total pages [49231.949635] cma: cma-4 pages: => 0 used of 7168 total pages [49231.949669] cma: cma-5 pages: => 0 used of 35840 total pages [49231.949676] cma: cma-6 pages: => 0 used of 4096 total pages [49231.949683] cma: cma-7 pages: => 0 used of 5120 total pages [49231.949725] bimc-bwmon 90b9100.qcom,snoop-l3-bwmon: Unable to register interrupt handler! (-12) [49231.949738] devfreq-icc 18590100.qcom,snoop-l3-bw: Unable to start HW monitor! (-12) [49231.949741] devfreq-icc 18590100.qcom,snoop-l3-bw: Unable to resume BW HW mon governor (-12) [49231.949744] devfreq 18590100.qcom,snoop-l3-bw: failed to resume devfreq device [49231.949880] bimc-bwmon 9091000.qcom,cpu-llcc-ddr-bwmon: Unable to register interrupt handler! (-12) [49231.949884] devfreq-icc soc:qcom,cpu-llcc-ddr-bw: Unable to start HW monitor! (-12) [49231.949885] devfreq-icc soc:qcom,cpu-llcc-ddr-bw: Unable to resume BW HW mon governor (-12) [49231.949888] devfreq soc:qcom,cpu-llcc-ddr-bw: failed to resume devfreq device [49231.950002] bimc-bwmon 90b6400.qcom,cpu-cpu-llcc-bwmon: Unable to register interrupt handler! (-12) [49231.950004] devfreq-icc soc:qcom,cpu-cpu-llcc-bw: Unable to start HW monitor! (-12) [49231.950006] devfreq-icc soc:qcom,cpu-cpu-llcc-bw: Unable to resume BW HW mon governor (-12) [49231.950009] devfreq soc:qcom,cpu-cpu-llcc-bw: failed to resume devfreq device [49232.171457] Filesystems sync: 0.005 seconds [49232.398473] Filesystems sync: 0.008 seconds [49232.417668] ------------[ cut here ]------------ [49232.417674] Trying to free already-free IRQ 53 [49232.417702] WARNING: CPU: 6 PID: 3656 at kernel/irq/manage.c:1765 __free_irq+0x2fc/0x400 [49232.417710] CPU: 6 PID: 3656 Comm: binder:502_7 Not tainted 5.4.278-Scarlet-v1.0-beta8 #1 [49232.417714] Hardware name: redwood based Qualcomm Technologies, Inc. SM7325 (DT) [49232.417717] pstate: 60000085 (nZCv daIf -PAN -UAO) [49232.417722] pc : __free_irq+0x2fc/0x400 [49232.417726] lr : __free_irq+0x2f8/0x400 [49232.417728] sp : ffffffaa166669e0 [49232.417731] x29: ffffffaa166669e0 x28: ffffffa9ae80d400 [49232.417735] x27: 0000000000000000 x26: 0000000000000000 [49232.417739] x25: ffffffa91ff3e268 x24: ffffffa9ae80d400 [49232.417742] x23: ffffffaa1948dc40 x22: 0000000000000035 [49232.417746] x21: ffffffa91ff3e2a0 x20: 0000000000000000 [49232.417749] x19: ffffffa91ff3e200 x18: 00000000ffed03ec [49232.417753] x17: 00000000000d03ec x16: ffffffda6c1a54e8 [49232.417756] x15: 0000000000000000 x14: 0000000000000082 [49232.417759] x13: 0000000000000034 x12: 0000000000000004 [49232.417761] x11: ffffffda6c1a54e0 x10: 00000000ffffffff [49232.417764] x9 : 115043787f31d900 x8 : 0000000000000000 [49232.417767] x7 : 0000000000000000 x6 : 7420676e69797254 [49232.417770] x5 : ffffffda6c2d5126 x4 : 0000000000000000 [49232.417773] x3 : 0000000000000000 x2 : ffffffffffffc9ff [49232.417776] x1 : 0000000000000000 x0 : 0000000000000022 [49232.417780] Call trace: [49232.417785] __free_irq+0x2fc/0x400 [49232.417789] free_irq+0x30/0x90 [49232.417802] suspend_bw_hwmon2+0xb8/0x140 [49232.417809] devfreq_bw_hwmon_ev_handler+0x1b4/0x520 [49232.417820] devfreq_suspend_device+0x58/0xe0 [49232.417825] devfreq_suspend+0x54/0x90 [49232.417838] dpm_suspend+0x2c/0x340 [49232.417841] dpm_suspend_start+0x7c/0xa0 [49232.417846] suspend_devices_and_enter+0xe0/0x5c0 [49232.417849] pm_suspend+0x258/0x280 [49232.417858] state_store+0x100/0x150 [49232.417869] kobj_attr_store+0x14/0x30 [49232.417883] sysfs_kf_write+0x38/0x50 [49232.417888] kernfs_fop_write+0x168/0x290 [49232.417896] __vfs_write+0x34/0x190 [49232.417899] vfs_write+0x11c/0x2d0 [49232.417903] __arm64_sys_write+0x78/0x110 [49232.417910] el0_svc_common+0xb0/0x120 [49232.417913] do_el0_svc+0x18/0x20 [49232.417918] el0_sync_handler+0x148/0x190 [49232.417922] el0_sync+0x140/0x180 [49232.417925] ---[ end trace b44a82c335d29d61 ]--- [49232.417952] ------------[ cut here ]------------ [49232.417954] Trying to free already-free IRQ 52 [49232.417962] WARNING: CPU: 6 PID: 3656 at kernel/irq/manage.c:1765 __free_irq+0x2fc/0x400 [49232.417966] CPU: 6 PID: 3656 Comm: binder:502_7 Tainted: G W 5.4.278-Scarlet-v1.0-beta8 #1 [49232.417968] Hardware name: redwood based Qualcomm Technologies, Inc. SM7325 (DT) [49232.417970] pstate: 60000085 (nZCv daIf -PAN -UAO) [49232.417974] pc : __free_irq+0x2fc/0x400 [49232.417977] lr : __free_irq+0x2f8/0x400 [49232.417979] sp : ffffffaa166669e0 [49232.417980] x29: ffffffaa166669e0 x28: ffffffa9ae80d400 [49232.417984] x27: 0000000000000000 x26: 0000000000000000 [49232.417987] x25: ffffffa91ff3e068 x24: ffffffa9ae80d400 [49232.417990] x23: ffffffaa1948d840 x22: 0000000000000034 [49232.417992] x21: ffffffa91ff3e0a0 x20: 0000000000000000 [49232.417995] x19: ffffffa91ff3e000 x18: 00000000ffecfa44 [49232.417998] x17: 00000000000cfa44 x16: ffffffda6c1a54e8 [49232.418001] x15: 0000000000000000 x14: 0000000000000082 [49232.418004] x13: 0000000000000034 x12: 0000000000000004 [49232.418006] x11: ffffffda6c1a54e0 x10: 00000000ffffffff [49232.418009] x9 : 115043787f31d900 x8 : 0000000000000000 [49232.418012] x7 : 0000000000000000 x6 : 7420676e69797254 [49232.418015] x5 : ffffffda6c2d5ace x4 : 0000000000000000 [49232.418018] x3 : 0000000000000000 x2 : ffffffffffffc9d0 [49232.418021] x1 : 0000000000000000 x0 : 0000000000000022 [49232.418023] Call trace: [49232.418027] __free_irq+0x2fc/0x400 [49232.418030] free_irq+0x30/0x90 [49232.418034] suspend_bw_hwmon3+0x94/0x110 [49232.418038] devfreq_bw_hwmon_ev_handler+0x1b4/0x520 [49232.418043] devfreq_suspend_device+0x58/0xe0 [49232.418047] devfreq_suspend+0x54/0x90 [49232.418050] dpm_suspend+0x2c/0x340 [49232.418053] dpm_suspend_start+0x7c/0xa0 [49232.418056] suspend_devices_and_enter+0xe0/0x5c0 [49232.418058] pm_suspend+0x258/0x280 [49232.418062] state_store+0x100/0x150 [49232.418066] kobj_attr_store+0x14/0x30 [49232.418070] sysfs_kf_write+0x38/0x50 [49232.418074] kernfs_fop_write+0x168/0x290 [49232.418077] __vfs_write+0x34/0x190 [49232.418080] vfs_write+0x11c/0x2d0 [49232.418083] __arm64_sys_write+0x78/0x110 [49232.418085] el0_svc_common+0xb0/0x120 [49232.418087] do_el0_svc+0x18/0x20 [49232.418091] el0_sync_handler+0x148/0x190 [49232.418094] el0_sync+0x140/0x180 [49232.418095] ---[ end trace b44a82c335d29d62 ]--- [49232.418112] ------------[ cut here ]------------ [49232.418113] Trying to free already-free IRQ 51 [49232.418120] WARNING: CPU: 6 PID: 3656 at kernel/irq/manage.c:1765 __free_irq+0x2fc/0x400 [49232.418122] CPU: 6 PID: 3656 Comm: binder:502_7 Tainted: G W 5.4.278-Scarlet-v1.0-beta8 #1 [49232.418124] Hardware name: redwood based Qualcomm Technologies, Inc. SM7325 (DT) [49232.418126] pstate: 60000085 (nZCv daIf -PAN -UAO) [49232.418130] pc : __free_irq+0x2fc/0x400 [49232.418133] lr : __free_irq+0x2f8/0x400 [49232.418135] sp : ffffffaa166669e0 [49232.418136] x29: ffffffaa166669e0 x28: ffffffa9ae80d400 [49232.418139] x27: 0000000000000000 x26: 0000000000000000 [49232.418142] x25: ffffffa91ff3de68 x24: ffffffa9ae80d400 [49232.418146] x23: ffffffaa1948d440 x22: 0000000000000033 [49232.418148] x21: ffffffa91ff3dea0 x20: 0000000000000000 [49232.418151] x19: ffffffa91ff3de00 x18: 00000000ffecf08c [49232.418154] x17: 00000000000cf08c x16: ffffffda6c1a54e8 [49232.418156] x15: 0000000000000000 x14: 0000000000000082 [49232.418159] x13: 0000000000000034 x12: 0000000000000004 [49232.418162] x11: ffffffda6c1a54e0 x10: 00000000ffffffff [49232.418165] x9 : 115043787f31d900 x8 : 0000000000000000 [49232.418168] x7 : 0000000000000000 x6 : 7420676e69797254 [49232.418171] x5 : ffffffda6c2d6486 x4 : 0000000000000000 [49232.418174] x3 : 0000000000000000 x2 : ffffffffffffc9a1 [49232.418176] x1 : 0000000000000000 x0 : 0000000000000022 [49232.418179] Call trace: [49232.418182] __free_irq+0x2fc/0x400 [49232.418186] free_irq+0x30/0x90 [49232.418188] suspend_bw_hwmon2+0xb8/0x140 [49232.418192] devfreq_bw_hwmon_ev_handler+0x1b4/0x520 [49232.418197] devfreq_suspend_device+0x58/0xe0 [49232.418201] devfreq_suspend+0x54/0x90 [49232.418204] dpm_suspend+0x2c/0x340 [49232.418207] dpm_suspend_start+0x7c/0xa0 [49232.418209] suspend_devices_and_enter+0xe0/0x5c0 [49232.418212] pm_suspend+0x258/0x280 [49232.418216] state_store+0x100/0x150 [49232.418219] kobj_attr_store+0x14/0x30 [49232.418223] sysfs_kf_write+0x38/0x50 [49232.418227] kernfs_fop_write+0x168/0x290 [49232.418230] __vfs_write+0x34/0x190 [49232.418233] vfs_write+0x11c/0x2d0 [49232.418236] __arm64_sys_write+0x78/0x110 [49232.418238] el0_svc_common+0xb0/0x120 [49232.418240] do_el0_svc+0x18/0x20 [49232.418243] el0_sync_handler+0x148/0x190 [49232.418246] el0_sync+0x140/0x180 [49232.418247] ---[ end trace b44a82c335d29d63 ]--- There is no need to reallocate and free the IRQs during suspend and resume. Replace IRQ thread creation with a simpler [enable|disable]_irq() approach to fix the issue. Co-authored-by: Sultan Alsawaf <sultan@kerneltoast.com> Change-Id: Ia12b8d24feddaef75792988ea55a6e8e29bf687a Signed-off-by: Sultan Alsawaf <sultan@kerneltoast.com> Signed-off-by: Tashfin Shakeer Rhythm <tashfinshakeerrhythm@gmail.com> |
||
|
|
d4d6a087de |
FROMGIT: procfs: prevent unpriveleged processes accessing fdinfo dir
The file permissions on the fdinfo dir from were changed from S_IRUSR|S_IXUSR to S_IRUGO|S_IXUGO, and a PTRACE_MODE_READ check was added for opening the fdinfo files [1]. However, the ptrace permission check was not added to the directory, allowing anyone to get the open FD numbers by reading the fdinfo directory. Add the missing ptrace permission check for opening the fdinfo directory. [1] https://lkml.kernel.org/r/20210308170651.919148-1-kaleshsingh@google.com Link: https://lkml.kernel.org/r/20210713162008.1056986-1-kaleshsingh@google.com Fixes: 7bc3fa0172a4 ("procfs: allow reading fdinfo with PTRACE_MODE_READ") Signed-off-by: Kalesh Singh <kaleshsingh@google.com> Cc: Kees Cook <keescook@chromium.org> Cc: Eric W. Biederman <ebiederm@xmission.com> Cc: Christian Brauner <christian.brauner@ubuntu.com> Cc: Christian König <christian.koenig@amd.com> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Hridya Valsaraju <hridya@google.com> Cc: Jann Horn <jannh@google.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Signed-off-by: Mark Brown <broonie@kernel.org> Bug: 151772539 Change-Id: I274b30aa0a5ce8412eae7161d31c6ee955035da9 (cherry picked from commit fc73829fa54b0c7af32d6da7c972eb3390957da4 git://git.kernel.org/pub/scm/linux/kernel/git/next/linux-next.git master) Signed-off-by: Kalesh Singh <kaleshsingh@google.com> |
||
|
|
7062e8847a |
UPSTREAM: BACKPORT: procfs/dmabuf: add inode number to /proc/*/fdinfo
And 'ino' field to /proc/<pid>/fdinfo/<FD> and /proc/<pid>/task/<tid>/fdinfo/<FD>. The inode numbers can be used to uniquely identify DMA buffers in user space and avoids a dependency on /proc/<pid>/fd/* when accounting per-process DMA buffer sizes. Link: https://lkml.kernel.org/r/20210308170651.919148-2-kaleshsingh@google.com Signed-off-by: Kalesh Singh <kaleshsingh@google.com> Acked-by: Randy Dunlap <rdunlap@infradead.org> Acked-by: Christian König <christian.koenig@amd.com> Cc: Jann Horn <jannh@google.com> Cc: Jeff Vander Stoep <jeffv@google.com> Cc: Kees Cook <keescook@chromium.org> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Minchan Kim <minchan@kernel.org> Cc: Hridya Valsaraju <hridya@google.com> Cc: Matthew Wilcox <willy@infradead.org> Cc: Alexander Viro <viro@zeniv.linux.org.uk> Cc: Kalesh Singh <kaleshsingh@google.com> Cc: Alexey Dobriyan <adobriyan@gmail.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Mauro Carvalho Chehab <mchehab+huawei@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Alexey Gladkov <gladkov.alexey@gmail.com> Cc: Szabolcs Nagy <szabolcs.nagy@arm.com> Cc: Eric W. Biederman <ebiederm@xmission.com> Cc: Christian Brauner <christian.brauner@ubuntu.com> Cc: Michel Lespinasse <walken@google.com> Cc: Bernd Edlinger <bernd.edlinger@hotmail.de> Cc: Andrei Vagin <avagin@gmail.com> Cc: Helge Deller <deller@gmx.de> Cc: James Morris <jamorris@linux.microsoft.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org> (cherry picked from commit 3845f256a8b527127bfbd4ced21e93d9e89aa6d7) [Kalesh Singh - Resolve conflict in Documentation/filesystems/proc.txt] Bug: 159126739 Bug: 167141117 Signed-off-by: Kalesh Singh <kaleshsingh@google.com> Change-Id: Id07beac3edcc95c0b42805e24e5486965acbb46e |
||
|
|
fac0966223 |
UPSTREAM: procfs: allow reading fdinfo with PTRACE_MODE_READ
Android captures per-process system memory state when certain low memory events (e.g a foreground app kill) occur, to identify potential memory hoggers. In order to measure how much memory a process actually consumes, it is necessary to include the DMA buffer sizes for that process in the memory accounting. Since the handle to DMA buffers are raw FDs, it is important to be able to identify which processes have FD references to a DMA buffer. Currently, DMA buffer FDs can be accounted using /proc/<pid>/fd/* and /proc/<pid>/fdinfo -- both are only readable by the process owner, as follows: 1. Do a readlink on each FD. 2. If the target path begins with "/dmabuf", then the FD is a dmabuf FD. 3. stat the file to get the dmabuf inode number. 4. Read/ proc/<pid>/fdinfo/<fd>, to get the DMA buffer size. Accessing other processes' fdinfo requires root privileges. This limits the use of the interface to debugging environments and is not suitable for production builds. Granting root privileges even to a system process increases the attack surface and is highly undesirable. Since fdinfo doesn't permit reading process memory and manipulating process state, allow accessing fdinfo under PTRACE_MODE_READ_FSCRED. Link: https://lkml.kernel.org/r/20210308170651.919148-1-kaleshsingh@google.com Signed-off-by: Kalesh Singh <kaleshsingh@google.com> Suggested-by: Jann Horn <jannh@google.com> Acked-by: Christian König <christian.koenig@amd.com> Cc: Alexander Viro <viro@zeniv.linux.org.uk> Cc: Alexey Dobriyan <adobriyan@gmail.com> Cc: Alexey Gladkov <gladkov.alexey@gmail.com> Cc: Andrei Vagin <avagin@gmail.com> Cc: Bernd Edlinger <bernd.edlinger@hotmail.de> Cc: Christian Brauner <christian.brauner@ubuntu.com> Cc: Eric W. Biederman <ebiederm@xmission.com> Cc: Helge Deller <deller@gmx.de> Cc: Hridya Valsaraju <hridya@google.com> Cc: James Morris <jamorris@linux.microsoft.com> Cc: Jeff Vander Stoep <jeffv@google.com> Cc: Jonathan Corbet <corbet@lwn.net> Cc: Kees Cook <keescook@chromium.org> Cc: Matthew Wilcox <willy@infradead.org> Cc: Mauro Carvalho Chehab <mchehab+huawei@kernel.org> Cc: Michal Hocko <mhocko@suse.com> Cc: Michel Lespinasse <walken@google.com> Cc: Minchan Kim <minchan@kernel.org> Cc: Randy Dunlap <rdunlap@infradead.org> Cc: Suren Baghdasaryan <surenb@google.com> Cc: Szabolcs Nagy <szabolcs.nagy@arm.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org> (cherry picked from commit 7bc3fa0172a423afb34e6df7a3998e5f23b1a94a) Bug: 159126739 Bug: 167141117 Signed-off-by: Kalesh Singh <kaleshsingh@google.com> Change-Id: I842b689670f731138592f45c7124ef446d9aa59a |
||
|
|
eab1bfc6b7 |
BACKPORT: net: core: enable SO_BINDTODEVICE for non-root users
Currently, SO_BINDTODEVICE requires CAP_NET_RAW. This change allows a non-root user to bind a socket to an interface if it is not already bound. This is useful to allow an application to bind itself to a specific VRF for outgoing or incoming connections. Currently, an application wanting to manage connections through several VRF need to be privileged. Previously, IP_UNICAST_IF and IPV6_UNICAST_IF were added for Wine ( |
||
|
|
95dcc82b80 |
ANDROID: fix ABI by undoing atomic64_t -> u64 type conversion
This is pretty much a no-op, but avoids changing struct net ABI. Bug: 274789652 Test: builds, net_test Signed-off-by: Maciej Żenczykowski <maze@google.com> Original-Change-Id: Ia744bdf0a026adccaef8382aaecc771a8d0763a6 Change-Id: Ic22d6bb5e35a7634c1dd4d4c3ed3b142b885ec6f |
||
|
|
ae6919a83c |
UPSTREAM: net: retrieve netns cookie via getsocketopt
It's getting more common to run nested container environments for
testing cloud software. One of such examples is Kind [1] which runs a
Kubernetes cluster in Docker containers on a single host. Each container
acts as a Kubernetes node, and thus can run any Pod (aka container)
inside the former. This approach simplifies testing a lot, as it
eliminates complicated VM setups.
Unfortunately, such a setup breaks some functionality when cgroupv2 BPF
programs are used for load-balancing. The load-balancer BPF program
needs to detect whether a request originates from the host netns or a
container netns in order to allow some access, e.g. to a service via a
loopback IP address. Typically, the programs detect this by comparing
netns cookies with the one of the init ns via a call to
bpf_get_netns_cookie(NULL). However, in nested environments the latter
cannot be used given the Kubernetes node's netns is outside the init ns.
To fix this, we need to pass the Kubernetes node netns cookie to the
program in a different way: by extending getsockopt() with a
SO_NETNS_COOKIE option, the orchestrator which runs in the Kubernetes
node netns can retrieve the cookie and pass it to the program instead.
Thus, this is following up on Eric's commit 3d368ab87cf6 ("net:
initialize net->net_cookie at netns setup") to allow retrieval via
SO_NETNS_COOKIE. This is also in line in how we retrieve socket cookie
via SO_COOKIE.
[1] https://kind.sigs.k8s.io/
Signed-off-by: Lorenz Bauer <lmb@cloudflare.com>
Signed-off-by: Martynas Pumputis <m@lambda.lt>
Cc: Eric Dumazet <edumazet@google.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
(cherry picked from commit e8b9eab99232c4e62ada9d7976c80fd5e8118289)
Bug: 274789652
Tested: builds, net_test passes
Signed-off-by: Maciej Żenczykowski <maze@google.com>
Change-Id: If784a592450af38d70f16da61e36cbbaff80ebca
|
||
|
|
cb8dc8a108 |
Revert "UPSTREAM: seccomp: Remove bogus __user annotations"
This reverts commit 5444477e8a4d31f6e6ff720c2d018d06e405bcc1. Bug: 176068146 Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Change-Id: Ic35b23f2f3ad99093b7df5e82633bba90acbe82a |
||
|
|
2d2e514151 |
UPSTREAM: seccomp: Fix CONFIG tests for Seccomp_filters
Strictly speaking, seccomp filters are only used
when CONFIG_SECCOMP_FILTER.
This patch fixes the condition to enable "Seccomp_filters"
in /proc/$pid/status.
Signed-off-by: Kenta Tada <Kenta.Tada@sony.com>
Fixes: c818c03b661c ("seccomp: Report number of loaded filters in /proc/$pid/status")
Signed-off-by: Kees Cook <keescook@chromium.org>
Link: https://lore.kernel.org/r/OSBPR01MB26772D245E2CF4F26B76A989F5669@OSBPR01MB2677.jpnprd01.prod.outlook.com
(cherry picked from commit 64bdc0244054f7d4bb621c8b4455e292f4e421bc)
Bug: 187129171
Signed-off-by: Connor O'Brien <connoro@google.com>
Change-Id: I44541ba5bac773b10b2593e47be943536b5ce3dd
|
||
|
|
c83d6f169d |
UPSTREAM: seccomp: Fix setting loaded filter count during TSYNC
The desired behavior is to set the caller's filter count to thread's. This value is reported via /proc, so this fixes the inaccurate count exposed to userspace; it is not used for reference counting, etc. Signed-off-by: Hsuan-Chi Kuo <hsuanchikuo@gmail.com> Link: https://lore.kernel.org/r/20210304233708.420597-1-hsuanchikuo@gmail.com Co-developed-by: Wiktor Garbacz <wiktorg@google.com> Signed-off-by: Wiktor Garbacz <wiktorg@google.com> Link: https://lore.kernel.org/lkml/20210810125158.329849-1-wiktorg@google.com Signed-off-by: Kees Cook <keescook@chromium.org> Cc: stable@vger.kernel.org Fixes: c818c03b661c ("seccomp: Report number of loaded filters in /proc/$pid/status") (cherry picked from commit b4d8a58f8dcfcc890f296696cadb76e77be44b5f) Bug: 187129171 Signed-off-by: Connor O'Brien <connoro@google.com> Change-Id: Ia3ee7ec71e9fdbb8d958f9b42b1c6e02c761503f |
||
|
|
b3dbf34600 |
UPSTREAM: xtensa: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for xtensa. Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/79669648ba167d668ea6ffb4884250abcd5ed254.1605101222.git.yifeifz2@illinois.edu (cherry picked from commit 445247b02342a05b7d528bba6d85d2d418875b69) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I7f8ecb0da495552c062e39e1b9bd5ba0aace3b01 |
||
|
|
c632efe06e |
UPSTREAM: sh: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for sh. Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/61ae084cd4783b9b50860d9dedb4a348cf1b7b6f.1605101222.git.yifeifz2@illinois.edu (cherry picked from commit 4c18bc054bffe415bec9e0edaa9ff1a84c1a6973) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I4cdb3b9fda0af5e5d1e4eede11661c828f41aad5 |
||
|
|
d02da3450d |
UPSTREAM: s390: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for s390. Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Acked-by: Heiko Carstens <hca@linux.ibm.com> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/a381b10aa2c5b1e583642f3cd46ced842d9d4ce5.1605101222.git.yifeifz2@illinois.edu (cherry picked from commit c09058eda2654c37fd7ac28c2004c3aae8b988e9) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I643cdd2a001f80bd5f6a64298e4a42412d66651e |
||
|
|
ced31049ee |
UPSTREAM: powerpc: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for powerpc. __LITTLE_ENDIAN__ is used here instead of CONFIG_CPU_LITTLE_ENDIAN to keep it consistent with asm/syscall.h. Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/0b64925362671cdaa26d01bfe50b3ba5e164adfd.1605101222.git.yifeifz2@illinois.edu (cherry picked from commit e7bcb4622ddf4473da6c03fa8423919a568c57dc) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I917b30a0c8cc6697513cc7f12bc84691e3166745 |
||
|
|
f33c0d2fed |
UPSTREAM: parisc: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for parisc. Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Acked-by: Helge Deller <deller@gmx.de> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/9bb86c546eda753adf5270425e7353202dbce87c.1605101222.git.yifeifz2@illinois.edu (cherry picked from commit 6aa7923c8737d1f8fd2a06154155d68dec646464) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I48ef86a4127b9cc87dff5235ded3733c098db7e1 |
||
|
|
48350fd86e |
UPSTREAM: csky: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for csky. Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/f9219026d4803b22f3e57e3768b4e42e004ef236.1605101222.git.yifeifz2@illinois.edu (cherry picked from commit 6e9ae6f98809e0d123ff4d769ba2e6f652119138) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I1fea89150f06be98ea4ee4357ad441e60aa6589f |
||
|
|
a6ecb44045 |
UPSTREAM: arm: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for arm. Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit 424c9102fa7b2a5c15afe47fd14278c849f4eefb) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Change-Id: I90eaafdc43f618ce7dcf76e1cbb8ae1ff542ead9 Bug: 176068146 |
||
|
|
17d2800a6b |
UPSTREAM: arm64: Enable seccomp architecture tracking
To enable seccomp constant action bitmaps, we need to have a static mapping to the audit architecture and system call table size. Add these for arm64. Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit ffde703470b03b1000017ed35c4f90a90caa22cf) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Change-Id: Ib21059de0928a61bd76202a67732432e88c5a5f0 Bug: 176068146 |
||
|
|
a8ea366f56 |
UPSTREAM: selftests/seccomp: Compare bitmap vs filter overhead
As part of the seccomp benchmarking, include the expectations with regard to the timing behavior of the constant action bitmaps, and report inconsistencies better. Example output with constant action bitmaps on x86: $ sudo ./seccomp_benchmark 100000000 Current BPF sysctl settings: net.core.bpf_jit_enable = 1 net.core.bpf_jit_harden = 0 Benchmarking 200000000 syscalls... 129.359381409 - 0.008724424 = 129350656985 (129.4s) getpid native: 646 ns 264.385890006 - 129.360453229 = 135025436777 (135.0s) getpid RET_ALLOW 1 filter (bitmap): 675 ns 399.400511893 - 264.387045901 = 135013465992 (135.0s) getpid RET_ALLOW 2 filters (bitmap): 675 ns 545.872866260 - 399.401718327 = 146471147933 (146.5s) getpid RET_ALLOW 3 filters (full): 732 ns 696.337101319 - 545.874097681 = 150463003638 (150.5s) getpid RET_ALLOW 4 filters (full): 752 ns Estimated total seccomp overhead for 1 bitmapped filter: 29 ns Estimated total seccomp overhead for 2 bitmapped filters: 29 ns Estimated total seccomp overhead for 3 full filters: 86 ns Estimated total seccomp overhead for 4 full filters: 106 ns Estimated seccomp entry overhead: 29 ns Estimated seccomp per-filter overhead (last 2 diff): 20 ns Estimated seccomp per-filter overhead (filters / 4): 19 ns Expectations: native ≤ 1 bitmap (646 ≤ 675): ✔️ native ≤ 1 filter (646 ≤ 732): ✔️ per-filter (last 2 diff) ≈ per-filter (filters / 4) (20 ≈ 19): ✔️ 1 bitmapped ≈ 2 bitmapped (29 ≈ 29): ✔️ entry ≈ 1 bitmapped (29 ≈ 29): ✔️ entry ≈ 2 bitmapped (29 ≈ 29): ✔️ native + entry + (per filter * 4) ≈ 4 filters total (755 ≈ 752): ✔️ [YiFei: Changed commit message to show stats for this patch series] Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/1b61df3db85c5f7f1b9202722c45e7b39df73ef2.1602431034.git.yifeifz2@illinois.edu (cherry picked from commit 192cf32243ce39af65bd095625aec374b38c03df) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Change-Id: Idd30139b4fbb2c06f4b043756bbb09bbacf3b123 Bug: 176068146 |
||
|
|
df0d2c8824 |
UPSTREAM: selftests/seccomp: use 90s as timeout
As seccomp_benchmark tries to calibrate how many samples will take more than 5 seconds to execute, it may end up picking up a number of samples that take 10 (but up to 12) seconds. As the calibration will take double that time, it takes around 20 seconds. Then, it executes the whole thing again, and then once more, with some added overhead. So, the thing might take more than 40 seconds, which is too close to the 45s timeout. That is very dependent on the system where it's executed, so may not be observed always, but it has been observed on x86 VMs. Using a 90s timeout seems safe enough. Signed-off-by: Thadeu Lima de Souza Cascardo <cascardo@canonical.com> Link: https://lore.kernel.org/r/20200601123202.1183526-1-cascardo@canonical.com Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit bc32c9c86581abf7baacf71342df3b0affe367db) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I52379d62bddc9163c6cbbb934483ceabb10cc909 |
||
|
|
43e7581e2e |
UPSTREAM: selftests/seccomp: Improve calibration loop
The seccomp benchmark calibration loop did not need to take so long. Instead, use a simple 1 second timeout and multiply up to target. It does not need to be accurate. Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit 81a0c8bc82be7c15dbf3e54832334552e6b76e2b) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I5583b28bb635c01413548dbe05b2bed8d50def00 |
||
|
|
f8da9865ce |
UPSTREAM: selftests/seccomp: Expand benchmark to per-filter measurements
It's useful to see how much (at a minimum) each filter adds to the syscall overhead. Add additional calculations. Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit d3a37ea9f6e548388b83fe895c7a037bc2ec3f7f) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: Ibc36b007b580cc129b2880b12b151a659590abb9 |
||
|
|
6faba4c168 |
UPSTREAM: x86: Enable seccomp architecture tracking
Provide seccomp internals with the details to calculate which syscall table the running kernel is expecting to deal with. This allows for efficient architecture pinning and paves the way for constant-action bitmaps. Co-developed-by: YiFei Zhu <yifeifz2@illinois.edu> Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/da58c3733d95c4f2115dd94225dfbe2573ba4d87.1602431034.git.yifeifz2@illinois.edu (cherry picked from commit 25db91209a910a0ccf8b093743088d0f4bf5659f) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Change-Id: I48a434063e401b27834e4ba37b88a852da51300b Bug: 176068146 |
||
|
|
c639e51eef |
UPSTREAM: seccomp/cache: Add "emulator" to check if filter is constant allow
SECCOMP_CACHE will only operate on syscalls that do not access any syscall arguments or instruction pointer. To facilitate this we need a static analyser to know whether a filter will return allow regardless of syscall arguments for a given architecture number / syscall number pair. This is implemented here with a pseudo-emulator, and stored in a per-filter bitmap. In order to build this bitmap at filter attach time, each filter is emulated for every syscall (under each possible architecture), and checked for any accesses of struct seccomp_data that are not the "arch" nor "nr" (syscall) members. If only "arch" and "nr" are examined, and the program returns allow, then we can be sure that the filter must return allow independent from syscall arguments. Nearly all seccomp filters are built from these cBPF instructions: BPF_LD | BPF_W | BPF_ABS BPF_JMP | BPF_JEQ | BPF_K BPF_JMP | BPF_JGE | BPF_K BPF_JMP | BPF_JGT | BPF_K BPF_JMP | BPF_JSET | BPF_K BPF_JMP | BPF_JA BPF_RET | BPF_K BPF_ALU | BPF_AND | BPF_K Each of these instructions are emulated. Any weirdness or loading from a syscall argument will cause the emulator to bail. The emulation is also halted if it reaches a return. In that case, if it returns an SECCOMP_RET_ALLOW, the syscall is marked as good. Emulator structure and comments are from Kees [1] and Jann [2]. Emulation is done at attach time. If a filter depends on more filters, and if the dependee does not guarantee to allow the syscall, then we skip the emulation of this syscall. [1] https://lore.kernel.org/lkml/20200923232923.3142503-5-keescook@chromium.org/ [2] https://lore.kernel.org/lkml/CAG48ez1p=dR_2ikKq=xVxkoGg0fYpTBpkhJSv1w-6BG=76PAvw@mail.gmail.com/ Suggested-by: Jann Horn <jannh@google.com> Signed-off-by: YiFei Zhu <yifeifz2@illinois.edu> Reviewed-by: Jann Horn <jannh@google.com> Co-developed-by: Kees Cook <keescook@chromium.org> Signed-off-by: Kees Cook <keescook@chromium.org> Link: https://lore.kernel.org/r/71c7be2db5ee08905f41c3be5c1ad6e2601ce88f.1602431034.git.yifeifz2@illinois.edu (cherry picked from commit 8e01b51a31a1e08e2c3e8fcc0ef6790441be2f61) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Change-Id: I5047f7f0d6502e5de6c047743f1053fda3025a6e Bug: 176068146 |
||
|
|
993d91ceea |
UPSTREAM: seccomp/cache: Lookup syscall allowlist bitmap for fast path
The overhead of running Seccomp filters has been part of some past discussions [1][2][3]. Oftentimes, the filters have a large number of instructions that check syscall numbers one by one and jump based on that. Some users chain BPF filters which further enlarge the overhead. A recent work [6] comprehensively measures the Seccomp overhead and shows that the overhead is non-negligible and has a non-trivial impact on application performance. We observed some common filters, such as docker's [4] or systemd's [5], will make most decisions based only on the syscall numbers, and as past discussions considered, a bitmap where each bit represents a syscall makes most sense for these filters. The fast (common) path for seccomp should be that the filter permits the syscall to pass through, and failing seccomp is expected to be an exceptional case; it is not expected for userspace to call a denylisted syscall over and over. When it can be concluded that an allow must occur for the given architecture and syscall pair (this determination is introduced in the next commit), seccomp will immediately allow the syscall, bypassing further BPF execution. Each architecture number has its own bitmap. The architecture number in seccomp_data is checked against the defined architecture number constant before proceeding to test the bit against the bitmap with the syscall number as the index of the bit in the bitmap, and if the bit is set, seccomp returns allow. The bitmaps are all clear in this patch and will be initialized in the next commit. When only one architecture exists, the check against architecture number is skipped, suggested by Kees Cook [7]. [1] https://lore.kernel.org/linux-security-module/c22a6c3cefc2412cad00ae14c1371711@huawei.com/T/ [2] https://lore.kernel.org/lkml/202005181120.971232B7B@keescook/T/ [3] https://github.com/seccomp/libseccomp/issues/116 [4] |
||
|
|
e24e8cf802 |
UPSTREAM: seccomp: Use current_pt_regs() instead of task_pt_regs(current)
As described in commit
|
||
|
|
c6279c1cb7 |
UPSTREAM: seccomp: kill process instead of thread for unknown actions
Asynchronous termination of a thread outside of the userspace thread library's knowledge is an unsafe operation that leaves the process in an inconsistent, corrupt, and possibly unrecoverable state. In order to make new actions that may be added in the future safe on kernels not aware of them, change the default action from SECCOMP_RET_KILL_THREAD to SECCOMP_RET_KILL_PROCESS. Signed-off-by: Rich Felker <dalias@libc.org> Link: https://lore.kernel.org/r/20200829015609.GA32566@brightrain.aerifal.cx [kees: Fixed up coredump selection logic to match] Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit 4d671d922d51907bc41f1f7f2dc737c928ae78fd) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I23140e1efbb4346de8566421c35d2810de26b209 |
||
|
|
43c11dbfa8 |
UPSTREAM: seccomp: don't leave dangling ->notif if file allocation fails
Christian and Kees both pointed out that this is a bit sloppy to open-code both places, and Christian points out that we leave a dangling pointer to ->notif if file allocation fails. Since we check ->notif for null in order to determine if it's ok to install a filter, this means people won't be able to install a filter if the file allocation fails for some reason, even if they subsequently should be able to. To fix this, let's hoist this free+null into its own little helper and use it. Reported-by: Kees Cook <keescook@chromium.org> Reported-by: Christian Brauner <christian.brauner@ubuntu.com> Signed-off-by: Tycho Andersen <tycho@tycho.pizza> Acked-by: Christian Brauner <christian.brauner@ubuntu.com> Link: https://lore.kernel.org/r/20200902140953.1201956-1-tycho@tycho.pizza Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit e839317900e9f13c83d8711d684de88c625b307a) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: Ie20e79fa15b6891895b7ada1f6ef7d08fdf81e01 |
||
|
|
0506d711a3 |
UPSTREAM: seccomp: don't leak memory when filter install races
In seccomp_set_mode_filter() with TSYNC | NEW_LISTENER, we first initialize
the listener fd, then check to see if we can actually use it later in
seccomp_may_assign_mode(), which can fail if anyone else in our thread
group has installed a filter and caused some divergence. If we can't, we
partially clean up the newly allocated file: we put the fd, put the file,
but don't actually clean up the *memory* that was allocated at
filter->notif. Let's clean that up too.
To accomplish this, let's hoist the actual "detach a notifier from a
filter" code to its own helper out of seccomp_notify_release(), so that in
case anyone adds stuff to init_listener(), they only have to add the
cleanup code in one spot. This does a bit of extra locking and such on the
failure path when the filter is not attached, but it's a slow failure path
anyway.
Fixes: 51891498f2da ("seccomp: allow TSYNC and USER_NOTIF together")
Reported-by: syzbot+3ad9614a12f80994c32e@syzkaller.appspotmail.com
Signed-off-by: Tycho Andersen <tycho@tycho.pizza>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Link: https://lore.kernel.org/r/20200902014017.934315-1-tycho@tycho.pizza
Signed-off-by: Kees Cook <keescook@chromium.org>
(cherry picked from commit a566a9012acd7c9a4be7e30dc7acb7a811ec2260)
Signed-off-by: Jeff Vander Stoep <jeffv@google.com>
Bug: 176068146
Change-Id: I8e6ce0f1646ff997623458f25b733ba79c0a47a2
|
||
|
|
bcce8defc3 |
UPSTREAM: seccomp: add SECCOMP_USER_NOTIF_FLAG_CONTINUE
This allows the seccomp notifier to continue a syscall. A positive
discussion about this feature was triggered by a post to the
ksummit-discuss mailing list (cf. [3]) and took place during KSummit
(cf. [1]) and again at the containers/checkpoint-restore
micro-conference at Linux Plumbers.
Recently we landed seccomp support for SECCOMP_RET_USER_NOTIF (cf. [4])
which enables a process (watchee) to retrieve an fd for its seccomp
filter. This fd can then be handed to another (usually more privileged)
process (watcher). The watcher will then be able to receive seccomp
messages about the syscalls having been performed by the watchee.
This feature is heavily used in some userspace workloads. For example,
it is currently used to intercept mknod() syscalls in user namespaces
aka in containers.
The mknod() syscall can be easily filtered based on dev_t. This allows
us to only intercept a very specific subset of mknod() syscalls.
Furthermore, mknod() is not possible in user namespaces toto coelo and
so intercepting and denying syscalls that are not in the whitelist on
accident is not a big deal. The watchee won't notice a difference.
In contrast to mknod(), a lot of other syscall we intercept (e.g.
setxattr()) cannot be easily filtered like mknod() because they have
pointer arguments. Additionally, some of them might actually succeed in
user namespaces (e.g. setxattr() for all "user.*" xattrs). Since we
currently cannot tell seccomp to continue from a user notifier we are
stuck with performing all of the syscalls in lieu of the container. This
is a huge security liability since it is extremely difficult to
correctly assume all of the necessary privileges of the calling task
such that the syscall can be successfully emulated without escaping
other additional security restrictions (think missing CAP_MKNOD for
mknod(), or MS_NODEV on a filesystem etc.). This can be solved by
telling seccomp to resume the syscall.
One thing that came up in the discussion was the problem that another
thread could change the memory after userspace has decided to let the
syscall continue which is a well known TOCTOU with seccomp which is
present in other ways already.
The discussion showed that this feature is already very useful for any
syscall without pointer arguments. For any accidentally intercepted
non-pointer syscall it is safe to continue.
For syscalls with pointer arguments there is a race but for any cautious
userspace and the main usec cases the race doesn't matter. The notifier
is intended to be used in a scenario where a more privileged watcher
supervises the syscalls of lesser privileged watchee to allow it to get
around kernel-enforced limitations by performing the syscall for it
whenever deemed save by the watcher. Hence, if a user tricks the watcher
into allowing a syscall they will either get a deny based on
kernel-enforced restrictions later or they will have changed the
arguments in such a way that they manage to perform a syscall with
arguments that they would've been allowed to do anyway.
In general, it is good to point out again, that the notifier fd was not
intended to allow userspace to implement a security policy but rather to
work around kernel security mechanisms in cases where the watcher knows
that a given action is safe to perform.
/* References */
[1]: https://linuxplumbersconf.org/event/4/contributions/560
[2]: https://linuxplumbersconf.org/event/4/contributions/477
[3]: https://lore.kernel.org/r/20190719093538.dhyopljyr5ns33qx@brauner.io
[4]: commit
|
||
|
|
2e686a7319 |
UPSTREAM: seccomp: Use -1 marker for end of mode 1 syscall list
The terminator for the mode 1 syscalls list was a 0, but that could be a valid syscall number (e.g. x86_64 __NR_read). By luck, __NR_read was listed first and the loop construct would not test it, so there was no bug. However, this is fragile. Replace the terminator with -1 instead, and make the variable name for mode 1 syscall lists more descriptive. Cc: Andy Lutomirski <luto@amacapital.net> Cc: Will Drewry <wad@chromium.org> Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit fe4bfff86ec54773df3db79e8112e3b0f820c799) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I3d91bf57236f2c20d71f22fa9c0ef7b0b8869bcf |
||
|
|
d3720689a9 |
UPSTREAM: seccomp: Use pr_fmt
Avoid open-coding "seccomp: " prefixes for pr_*() calls. Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit e68f9d49dda1744d548426c8b4335a8d693a36d0) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I2fa08d91c11b59dbb962cc767aa4cd25a42f761d |
||
|
|
246a019c84 |
UPSTREAM: seccomp: notify about unused filter
We've been making heavy use of the seccomp notifier to intercept and handle certain syscalls for containers. This patch allows a syscall supervisor listening on a given notifier to be notified when a seccomp filter has become unused. A container is often managed by a singleton supervisor process the so-called "monitor". This monitor process has an event loop which has various event handlers registered. If the user specified a seccomp profile that included a notifier for various syscalls then we also register a seccomp notify even handler. For any container using a separate pid namespace the lifecycle of the seccomp notifier is bound to the init process of the pid namespace, i.e. when the init process exits the filter must be unused. If a new process attaches to a container we force it to assume a seccomp profile. This can either be the same seccomp profile as the container was started with or a modified one. If the attaching process makes use of the seccomp notifier we will register a new seccomp notifier handler in the monitor's event loop. However, when the attaching process exits we can't simply delete the handler since other child processes could've been created (daemons spawned etc.) that have inherited the seccomp filter and so we need to keep the seccomp notifier fd alive in the event loop. But this is problematic since we don't get a notification when the seccomp filter has become unused and so we currently never remove the seccomp notifier fd from the event loop and just keep accumulating fds in the event loop. We've had this issue for a while but it has recently become more pressing as more and larger users make use of this. To fix this, we introduce a new "users" reference counter that tracks any tasks and dependent filters making use of a filter. When a notifier is registered waiting tasks will be notified that the filter is now empty by receiving a (E)POLLHUP event. The concept in this patch introduces is the same as for signal_struct, i.e. reference counting for life-cycle management is decoupled from reference counting taks using the object. There's probably some trickery possible but the second counter is just the correct way of doing this IMHO and has precedence. Cc: Tycho Andersen <tycho@tycho.ws> Cc: Kees Cook <keescook@chromium.org> Cc: Matt Denton <mpdenton@google.com> Cc: Sargun Dhillon <sargun@sargun.me> Cc: Jann Horn <jannh@google.com> Cc: Chris Palmer <palmer@google.com> Cc: Aleksa Sarai <cyphar@cyphar.com> Cc: Robert Sesek <rsesek@google.com> Cc: Jeffrey Vander Stoep <jeffv@google.com> Cc: Linux Containers <containers@lists.linux-foundation.org> Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com> Link: https://lore.kernel.org/r/20200531115031.391515-3-christian.brauner@ubuntu.com Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit 99cdb8b9a57393b5978e7a6310a2cba511dd179b) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: I1501f5134bf738fa641c8b35ed65a4bed01e6fef |
||
|
|
902f009956 |
UPSTREAM: seccomp: Lift wait_queue into struct seccomp_filter
Lift the wait_queue from struct notification into struct seccomp_filter. This is cleaner overall and lets us avoid having to take the notifier mutex in the future for EPOLLHUP notifications since we need to neither read nor modify the notifier specific aspects of the seccomp filter. In the exit path I'd very much like to avoid having to take the notifier mutex for each filter in the task's filter hierarchy. Cc: Tycho Andersen <tycho@tycho.ws> Cc: Kees Cook <keescook@chromium.org> Cc: Matt Denton <mpdenton@google.com> Cc: Sargun Dhillon <sargun@sargun.me> Cc: Jann Horn <jannh@google.com> Cc: Chris Palmer <palmer@google.com> Cc: Aleksa Sarai <cyphar@cyphar.com> Cc: Robert Sesek <rsesek@google.com> Cc: Jeffrey Vander Stoep <jeffv@google.com> Cc: Linux Containers <containers@lists.linux-foundation.org> Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com> Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit 76194c4e830d570d9e369d637bb907591d2b3111) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: Id87b359f1155d26068cd62d5295dc39849d46495 |
||
|
|
31e7b9c9d8 |
BACKPORT: seccomp: release filter after task is fully dead
The seccomp filter used to be released in free_task() which is called asynchronously via call_rcu() and assorted mechanisms. Since we need to inform tasks waiting on the seccomp notifier when a filter goes empty we will notify them as soon as a task has been marked fully dead in release_task(). To not split seccomp cleanup into two parts, move filter release out of free_task() and into release_task() after we've unhashed struct task from struct pid, exited signals, and unlinked it from the threadgroups' thread list. We'll put the empty filter notification infrastructure into it in a follow up patch. This also renames put_seccomp_filter() to seccomp_filter_release() which is a more descriptive name of what we're doing here especially once we've added the empty filter notification mechanism in there. We're also NULL-ing the task's filter tree entrypoint which seems cleaner than leaving a dangling pointer in there. Note that this shouldn't need any memory barriers since we're calling this when the task is in release_task() which means it's EXIT_DEAD. So it can't modify its seccomp filters anymore. You can also see this from the point where we're calling seccomp_filter_release(). It's after __exit_signal() and at this point, tsk->sighand will already have been NULLed which is required for thread-sync and filter installation alike. Cc: Tycho Andersen <tycho@tycho.ws> Cc: Kees Cook <keescook@chromium.org> Cc: Matt Denton <mpdenton@google.com> Cc: Sargun Dhillon <sargun@sargun.me> Cc: Jann Horn <jannh@google.com> Cc: Chris Palmer <palmer@google.com> Cc: Aleksa Sarai <cyphar@cyphar.com> Cc: Robert Sesek <rsesek@google.com> Cc: Jeffrey Vander Stoep <jeffv@google.com> Cc: Linux Containers <containers@lists.linux-foundation.org> Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com> Link: https://lore.kernel.org/r/20200531115031.391515-2-christian.brauner@ubuntu.com Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit 3a15fb6ed92cb32b0a83f406aa4a96f28c9adbc3) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Resolved minor merge conflict in kernel/exit.c where 5.4 does not have commits 7bc3e6e55acf0 and 6ade99ec6175a. Change-Id: I4a0113f3f64a86937ba5c9ac6e2537926e2827be |
||
|
|
07bd1e5332 |
UPSTREAM: seccomp: rename "usage" to "refs" and document
Naming the lifetime counter of a seccomp filter "usage" suggests a little too strongly that its about tasks that are using this filter while it also tracks other references such as the user notifier or ptrace. This also updates the documentation to note this fact. We'll be introducing an actual usage counter in a follow-up patch. Cc: Tycho Andersen <tycho@tycho.ws> Cc: Kees Cook <keescook@chromium.org> Cc: Matt Denton <mpdenton@google.com> Cc: Sargun Dhillon <sargun@sargun.me> Cc: Jann Horn <jannh@google.com> Cc: Chris Palmer <palmer@google.com> Cc: Aleksa Sarai <cyphar@cyphar.com> Cc: Robert Sesek <rsesek@google.com> Cc: Jeffrey Vander Stoep <jeffv@google.com> Cc: Linux Containers <containers@lists.linux-foundation.org> Signed-off-by: Christian Brauner <christian.brauner@ubuntu.com> Link: https://lore.kernel.org/r/20200531115031.391515-1-christian.brauner@ubuntu.com Signed-off-by: Kees Cook <keescook@chromium.org> (cherry picked from commit b707ddee11d1dc4518ab7f1aa5e7af9ceaa23317) Signed-off-by: Jeff Vander Stoep <jeffv@google.com> Bug: 176068146 Change-Id: If4c78f85b9873aea2dffa1c8afb250f534d0fedc |