https://source.android.com/docs/security/bulletin/2025-12-01
CVE-2025-48623
CVE-2025-48624
CVE-2025-48637
CVE-2025-48638
CVE-2024-35970
CVE-2025-38236
CVE-2025-38349
CVE-2025-48610
CVE-2025-38500
* tag 'ASB-2025-12-01_11-5.4' of https://android.googlesource.com/kernel/common:
UPSTREAM: crypto: essiv - Check ssize for decryption and in-place encryption
ANDROID: GKI: fix up build break where timer_delete_sync() was used
Revert "net: rtnetlink: remove redundant assignment to variable err"
Revert "net: rtnetlink: add msg kind names"
Revert "net: rtnetlink: add helper to extract msg type's kind"
Revert "net: rtnetlink: use BIT for flag values"
Revert "net: netlink: add NLM_F_BULK delete request modifier"
Revert "net: rtnetlink: add bulk delete support flag"
Revert "net: rtnetlink: fix module reference count leak issue in rtnetlink_rcv_msg"
Revert "net: add ndo_fdb_del_bulk"
Revert "net: rtnetlink: add NLM_F_BULK support to rtnl_fdb_del"
Revert "rtnetlink: Allow deleting FDB entries in user namespace"
Linux 5.4.301
net: rtnetlink: fix module reference count leak issue in rtnetlink_rcv_msg
media: s5p-mfc: remove an unused/uninitialized variable
NFSD: Fix last write offset handling in layoutcommit
NFSD: Minor cleanup in layoutcommit processing
padata: Reset next CPU when reorder sequence wraps around
KEYS: trusted_tpm1: Compare HMAC values in constant time
NFSD: Define a proc_layoutcommit for the FlexFiles layout type
vfs: Don't leak disconnected dentries on umount
jbd2: ensure that all ongoing I/O complete before freeing blocks
ext4: detect invalid INLINE_DATA + EXTENTS flag combination
drm/amdgpu: use atomic functions with memory barriers for vm fault info
ext4: avoid potential buffer over-read in parse_apply_sb_mount_options()
spi: cadence-quadspi: Flush posted register writes before DAC access
spi: cadence-quadspi: Flush posted register writes before INDAC access
memory: samsung: exynos-srom: Fix of_iomap leak in exynos_srom_probe
memory: samsung: exynos-srom: Correct alignment
arm64: errata: Apply workarounds for Neoverse-V3AE
arm64: cputype: Add Neoverse-V3AE definitions
comedi: fix divide-by-zero in comedi_buf_munge()
binder: remove "invalid inc weak" check
xhci: dbc: enable back DbC in resume if it was enabled before suspend
usb/core/quirks: Add Huawei ME906S to wakeup quirk
USB: serial: option: add Telit FN920C04 ECM compositions
USB: serial: option: add Quectel RG255C
USB: serial: option: add UNISOC UIS7720
net: ravb: Ensure memory write completes before ringing TX doorbell
net: usb: rtl8150: Fix frame padding
ocfs2: clear extent cache after moving/defragmenting extents
MIPS: Malta: Fix keyboard resource preventing i8042 driver from registering
Revert "cpuidle: menu: Avoid discarding useful information"
net: bonding: fix possible peer notify event loss or dup issue
sctp: avoid NULL dereference when chunk data buffer is missing
arm64, mm: avoid always making PTE dirty in pte_mkwrite()
net: enetc: correct the value of ENETC_RXB_TRUESIZE
rtnetlink: Allow deleting FDB entries in user namespace
net: rtnetlink: add NLM_F_BULK support to rtnl_fdb_del
net: add ndo_fdb_del_bulk
net: rtnetlink: add bulk delete support flag
net: netlink: add NLM_F_BULK delete request modifier
net: rtnetlink: use BIT for flag values
net: rtnetlink: add helper to extract msg type's kind
net: rtnetlink: add msg kind names
net: rtnetlink: remove redundant assignment to variable err
m68k: bitops: Fix find_*_bit() signatures
hfsplus: return EIO when type of hidden directory mismatch in hfsplus_fill_super()
hfs: fix KMSAN uninit-value issue in hfs_find_set_zero_bits()
dlm: check for defined force value in dlm_lockspace_release
hfsplus: fix KMSAN uninit-value issue in hfsplus_delete_cat()
hfs: validate record offset in hfsplus_bmap_alloc
hfsplus: fix KMSAN uninit-value issue in __hfsplus_ext_cache_extent()
hfs: make proper initalization of struct hfs_find_data
hfs: clear offset and space out of valid records in b-tree node
exec: Fix incorrect type for ret
hfsplus: fix slab-out-of-bounds read in hfsplus_strcasecmp()
ALSA: firewire: amdtp-stream: fix enum kernel-doc warnings
sched/fair: Fix pelt lost idle time detection
sched/balancing: Rename newidle_balance() => sched_balance_newidle()
sched/fair: Trivial correction of the newidle_balance() comment
sched: Make newidle_balance() static again
tls: don't rely on tx_work during send()
tls: always set record_type in tls_process_cmsg
tg3: prevent use of uninitialized remote_adv and local_adv variables
tcp: fix tcp_tso_should_defer() vs large RTT
amd-xgbe: Avoid spurious link down messages during interface toggle
net/ip6_tunnel: Prevent perpetual tunnel growth
net: dlink: handle dma_map_single() failure properly
net: dl2k: switch from 'pci_' to 'dma_' API
media: pci: ivtv: Add missing check after DMA map
media: pci/ivtv: switch from 'pci_' to 'dma_' API
xen/events: Update virq_to_irq on migration
media: lirc: Fix error handling in lirc_register()
media: rc: Directly use ida_free()
drm/exynos: exynos7_drm_decon: remove ctx->suspended
btrfs: avoid potential out-of-bounds in btrfs_encode_fh()
pwm: berlin: Fix wrong register in suspend/resume
media: cx18: Add missing check after DMA map
xen/events: Cleanup find_virq() return codes
cramfs: Verify inode mode when loading from disk
fs: Add 'initramfs_options' to set initramfs mount options
pid: Add a judgment for ns null in pid_nr_ns
minixfs: Verify inode mode when loading from disk
tracing: Fix race condition in kprobe initialization causing NULL pointer dereference
dm: fix NULL pointer dereference in __dm_suspend()
mfd: intel_soc_pmic_chtdc_ti: Set use_single_read regmap_config flag
mfd: intel_soc_pmic_chtdc_ti: Drop unneeded assignment for cache_type
mfd: intel_soc_pmic_chtdc_ti: Fix invalid regmap-config max_register value
Squashfs: reject negative file sizes in squashfs_read_inode()
Squashfs: add additional inode sanity checking
media: mc: Clear minor number before put device
mfd: vexpress-sysreg: Check the return value of devm_gpiochip_add_data()
fs: udf: fix OOB read in lengthAllocDescs handling
KVM: x86: Don't (re)check L1 intercepts when completing userspace I/O
net/9p: fix double req put in p9_fd_cancelled
ext4: guard against EA inode refcount underflow in xattr update
ext4: correctly handle queries for metadata mappings
ext4: increase i_disksize to offset + len in ext4_update_disksize_before_punch()
nfsd: nfserr_jukebox in nlm_fopen should lead to a retry
x86/umip: Fix decoding of register forms of 0F 01 (SGDT and SIDT aliases)
x86/umip: Check that the instruction opcode is at least two bytes
PCI: keystone: Use devm_request_irq() to free "ks-pcie-error-irq" on exit
PCI/AER: Fix missing uevent on recovery when a reset is requested
PCI/IOV: Add PCI rescan-remove locking when enabling/disabling SR-IOV
rseq/selftests: Use weak symbol reference, not definition, to link with glibc
rtc: interface: Fix long-standing race when setting alarm
rtc: interface: Ensure alarm irq is enabled when UIE is enabled
mmc: core: SPI mode remove cmd7
mtd: rawnand: fsmc: Default to autodetect buswidth
sparc: fix error handling in scan_one_device()
sparc64: fix hugetlb for sun4u
sctp: Fix MAC comparison to be constant-time
scsi: hpsa: Fix potential memory leak in hpsa_big_passthru_ioctl()
parisc: don't reference obsolete termio struct for TC* constants
lib/genalloc: fix device leak in of_gen_pool_get()
iio: frequency: adf4350: Fix prescaler usage.
iio: dac: ad5421: use int type to store negative error codes
iio: dac: ad5360: use int type to store negative error codes
crypto: atmel - Fix dma_unmap_sg() direction
cpufreq: intel_pstate: Fix object lifecycle issue in update_qos_request()
drm/nouveau: fix bad ret code in nouveau_bo_move_prep
media: i2c: mt9v111: fix incorrect type for ret
firmware: meson_sm: fix device leak at probe
xen/manage: Fix suspend error path
arm64: dts: qcom: msm8916: Add missing MDSS reset
ACPI: debug: fix signedness issues in read/write helpers
ACPI: TAD: Add missing sysfs_remove_group() for ACPI_TAD_RT
tpm_tis: Fix incorrect arguments in tpm_tis_probe_irq_single
tpm, tpm_tis: Claim locality before writing interrupt registers
crypto: essiv - Check ssize for decryption and in-place encryption
mailbox: zynqmp-ipi: Remove dev.parent check in zynqmp_ipi_free_mboxes
mailbox: zynqmp-ipi: Remove redundant mbox_controller_unregister() call
tools build: Align warning options with perf
net: fsl_pq_mdio: Fix device node reference leak in fsl_pq_mdio_probe
tcp: Don't call reqsk_fastopen_remove() in tcp_conn_request().
net/sctp: fix a null dereference in sctp_disposition sctp_sf_do_5_1D_ce()
drm/vmwgfx: Fix Use-after-free in validation
net/mlx4: prevent potential use after free in mlx4_en_do_uc_filter()
scsi: mvsas: Fix use-after-free bugs in mvs_work_queue
scsi: mvsas: Use sas_task_find_rq() for tagging
scsi: mvsas: Delete mvs_tag_init()
scsi: libsas: Add sas_task_find_rq()
clk: nxp: Fix pll0 rate check condition in LPC18xx CGU driver
clk: nxp: lpc18xx-cgu: convert from round_rate() to determine_rate()
perf session: Fix handling when buffer exceeds 2 GiB
rtc: x1205: Fix Xicor X1205 vendor prefix
perf util: Fix compression checks returning -1 as bool
iio: frequency: adf4350: Fix ADF4350_REG3_12BIT_CLKDIV_MODE
clocksource/drivers/clps711x: Fix resource leaks in error paths
pinctrl: check the return value of pinmux_ops::get_function_name()
Input: uinput - zero-initialize uinput_ff_upload_compat to avoid info leak
mm: hugetlb: avoid soft lockup when mprotect to large memory area
uio_hv_generic: Let userspace take care of interrupt mask
Squashfs: fix uninit-value in squashfs_get_parent
Revert "net/mlx5e: Update and set Xon/Xoff upon MTU set"
net: ena: return 0 in ena_get_rxfh_key_size() when RSS hash key is not configurable
nfp: fix RSS hash key size when RSS is not supported
drivers/base/node: fix double free in register_one_node()
ocfs2: fix double free in user_cluster_connect()
net: usb: Remove disruptive netif_wake_queue in rtl8150_set_multicast
RDMA/siw: Always report immediate post SQ errors
usb: vhci-hcd: Prevent suspending virtually attached devices
scsi: mpt3sas: Fix crash in transport port remove by using ioc_info()
ipvs: Defer ip_vs_ftp unregister during netns cleanup
NFSv4.1: fix backchannel max_resp_sz verification check
remoteproc: qcom: q6v5: Avoid disabling handover IRQ twice
sparc: fix accurate exception reporting in copy_{from,to}_user for M7
sparc: fix accurate exception reporting in copy_to_user for Niagara 4
sparc: fix accurate exception reporting in copy_{from_to}_user for Niagara
sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC III
sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC
IB/sa: Fix sa_local_svc_timeout_ms read race
RDMA/core: Resolve MAC of next-hop device without ARP support
wifi: mt76: fix potential memory leak in mt76_wmac_probe()
drivers/base/node: handle error properly in register_one_node()
watchdog: mpc8xxx_wdt: Reload the watchdog timer when enabling the watchdog
netfilter: ipset: Remove unused htable_bits in macro ahash_region
iio: consumers: Fix offset handling in iio_convert_raw_to_processed()
ASoC: Intel: bytcr_rt5651: Fix invalid quirk input mapping
ASoC: Intel: bytcr_rt5640: Fix invalid quirk input mapping
ASoC: Intel: bytcht_es8316: Fix invalid quirk input mapping
pps: fix warning in pps_register_cdev when register device fail
misc: genwqe: Fix incorrect cmd field being reported in error
usb: gadget: configfs: Correctly set use_os_string at bind
usb: phy: twl6030: Fix incorrect type for ret
tcp: fix __tcp_close() to only send RST when required
PCI: tegra: Fix devm_kcalloc() argument order for port->phys allocation
wifi: mwifiex: send world regulatory domain to driver
ALSA: lx_core: use int type to store negative error codes
media: rj54n1cb0c: Fix memleak in rj54n1_probe()
scsi: myrs: Fix dma_alloc_coherent() error check
scsi: pm80xx: Fix array-index-out-of-of-bounds on rmmod
serial: max310x: Add error checking in probe()
usb: host: max3421-hcd: Fix error pointer dereference in probe cleanup
drm/radeon/r600_cs: clean up of dead code in r600_cs
i2c: designware: Add disabling clocks when probe fails
i2c: mediatek: fix potential incorrect use of I2C_MASTER_WRRD
bpf: Explicitly check accesses to bpf_sock_addr
selftests: watchdog: skip ping loop if WDIOF_KEEPALIVEPING not supported
pwm: tiehrpwm: Fix corner case in clock divisor calculation
block: use int to store blk_stack_limits() return value
blk-mq: check kobject state_in_sysfs before deleting in blk_mq_unregister_hctx
pinctrl: meson-gxl: add missing i2c_d pinmux
soc: qcom: rpmh-rsc: Unconditionally clear _TRIGGER bit for TCS
ACPI: processor: idle: Fix memory leak when register cpuidle device failed
regmap: Remove superfluous check for !config in __regmap_init()
x86/vdso: Fix output operand size of RDPID
perf: arm_spe: Prevent overflow in PERF_IDX2OFF()
driver core/PM: Set power.no_callbacks along with power.no_pm
staging: axis-fifo: flush RX FIFO on read errors
staging: axis-fifo: fix maximum TX packet length check
perf subcmd: avoid crash in exclude_cmds when excludes is empty
dm-integrity: limit MAX_TAG_SIZE to 255
wifi: rtlwifi: rtl8192cu: Don't claim USB ID 07b8:8188
USB: serial: option: add SIMCom 8230C compositions
media: rc: fix races with imon_disconnect()
media: imon: grab lock earlier in imon_ir_change_protocol()
media: imon: reorganize serialization
media: rc: Add support for another iMON 0xffdc device
media: i2c: tc358743: Fix use-after-free bugs caused by orphan timer in probe
media: tuner: xc5000: Fix use-after-free in xc5000_release
media: tunner: xc5000: Refactor firmware load
udp: Fix memory accounting leak.
media: b2c2: Fix use-after-free causing by irq_check_work in flexcop_pci_remove
scsi: target: target_core_configfs: Add length check to avoid buffer overflow
Conflicts:
drivers/soc/qcom/rpmh-rsc.c
kernel/sched/fair.c
Change-Id: I58ab24a3db8be4c698c41fd47daeb1f1fb7884ee
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEZH8oZUiU471FcZm+ONu9yGCSaT4FAmkCD9kACgkQONu9yGCS
aT756BAAwEcL3lKO/0QKjMUrvgv7FRoRm9TH9Usq7Nswmiax6YKI3oPgZKfZZL6u
yLGSnYqV8f+Jo8kALMJAybu+Oc3Z8WbBZJG/dvlAkvR7iN19ZCuKa2NwWmh8BTGd
5ZP4/nyXlNlg0/IIZ8xJa3067U895y9IhgMDmSvOZoUFGFecqUmJNBoj8bHgJaPL
/MENYb7D0CcKWbgCwzsqN0EEm9UKb7xgtaaC52sfZD+GHUPF/RLq9QNyUGtyL0m8
y20HG6zXmBiVP+zdtYpTaaEP6OSNOCmiOweSZOFafnatd2XO1hRdnAoEk7+t2Lj+
Xa5GowVUuglmBd8bPbeY8cwVihMxr81DxMns+OpQJTiDD8Q6Whcxbaqwv+tVe7z4
cVDh1xlKG4U2rsDTeiGhbZVNb6gqqy3LnVFqImQE+dcZqlwC0pbqwO2lLdHMa8NF
f+l1U3ab0k9qYx8HQegYDTufW5Eb0Rki/3YyHk8o3alOYd2Oc3r16lw9rpSw2Nq3
Rku+VMSy6z8uvmIP7Ywobla6j73SGFrDc61F0nsU1x+Gd/dMIgkn36dfxciBCgPa
/AmkK1Psx04s7dIPMowx/nXy5Jk8J4k3JlVMeBnMCuFVUNCNp3qzq5WXdeL1G5f5
GFPJPblaZrfKQNSSlpUAmstl8v5t8aJcjxobmHn28fVliFAJdDI=
=XPvH
-----END PGP SIGNATURE-----
Merge 5.4.301 into android11-5.4-lts
Changes in 5.4.301
scsi: target: target_core_configfs: Add length check to avoid buffer overflow
media: b2c2: Fix use-after-free causing by irq_check_work in flexcop_pci_remove
udp: Fix memory accounting leak.
media: tunner: xc5000: Refactor firmware load
media: tuner: xc5000: Fix use-after-free in xc5000_release
media: i2c: tc358743: Fix use-after-free bugs caused by orphan timer in probe
media: rc: Add support for another iMON 0xffdc device
media: imon: reorganize serialization
media: imon: grab lock earlier in imon_ir_change_protocol()
media: rc: fix races with imon_disconnect()
USB: serial: option: add SIMCom 8230C compositions
wifi: rtlwifi: rtl8192cu: Don't claim USB ID 07b8:8188
dm-integrity: limit MAX_TAG_SIZE to 255
perf subcmd: avoid crash in exclude_cmds when excludes is empty
staging: axis-fifo: fix maximum TX packet length check
staging: axis-fifo: flush RX FIFO on read errors
driver core/PM: Set power.no_callbacks along with power.no_pm
perf: arm_spe: Prevent overflow in PERF_IDX2OFF()
x86/vdso: Fix output operand size of RDPID
regmap: Remove superfluous check for !config in __regmap_init()
ACPI: processor: idle: Fix memory leak when register cpuidle device failed
soc: qcom: rpmh-rsc: Unconditionally clear _TRIGGER bit for TCS
pinctrl: meson-gxl: add missing i2c_d pinmux
blk-mq: check kobject state_in_sysfs before deleting in blk_mq_unregister_hctx
block: use int to store blk_stack_limits() return value
pwm: tiehrpwm: Fix corner case in clock divisor calculation
selftests: watchdog: skip ping loop if WDIOF_KEEPALIVEPING not supported
bpf: Explicitly check accesses to bpf_sock_addr
i2c: mediatek: fix potential incorrect use of I2C_MASTER_WRRD
i2c: designware: Add disabling clocks when probe fails
drm/radeon/r600_cs: clean up of dead code in r600_cs
usb: host: max3421-hcd: Fix error pointer dereference in probe cleanup
serial: max310x: Add error checking in probe()
scsi: pm80xx: Fix array-index-out-of-of-bounds on rmmod
scsi: myrs: Fix dma_alloc_coherent() error check
media: rj54n1cb0c: Fix memleak in rj54n1_probe()
ALSA: lx_core: use int type to store negative error codes
wifi: mwifiex: send world regulatory domain to driver
PCI: tegra: Fix devm_kcalloc() argument order for port->phys allocation
tcp: fix __tcp_close() to only send RST when required
usb: phy: twl6030: Fix incorrect type for ret
usb: gadget: configfs: Correctly set use_os_string at bind
misc: genwqe: Fix incorrect cmd field being reported in error
pps: fix warning in pps_register_cdev when register device fail
ASoC: Intel: bytcht_es8316: Fix invalid quirk input mapping
ASoC: Intel: bytcr_rt5640: Fix invalid quirk input mapping
ASoC: Intel: bytcr_rt5651: Fix invalid quirk input mapping
iio: consumers: Fix offset handling in iio_convert_raw_to_processed()
netfilter: ipset: Remove unused htable_bits in macro ahash_region
watchdog: mpc8xxx_wdt: Reload the watchdog timer when enabling the watchdog
drivers/base/node: handle error properly in register_one_node()
wifi: mt76: fix potential memory leak in mt76_wmac_probe()
RDMA/core: Resolve MAC of next-hop device without ARP support
IB/sa: Fix sa_local_svc_timeout_ms read race
sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC
sparc: fix accurate exception reporting in copy_{from_to}_user for UltraSPARC III
sparc: fix accurate exception reporting in copy_{from_to}_user for Niagara
sparc: fix accurate exception reporting in copy_to_user for Niagara 4
sparc: fix accurate exception reporting in copy_{from,to}_user for M7
remoteproc: qcom: q6v5: Avoid disabling handover IRQ twice
NFSv4.1: fix backchannel max_resp_sz verification check
ipvs: Defer ip_vs_ftp unregister during netns cleanup
scsi: mpt3sas: Fix crash in transport port remove by using ioc_info()
usb: vhci-hcd: Prevent suspending virtually attached devices
RDMA/siw: Always report immediate post SQ errors
net: usb: Remove disruptive netif_wake_queue in rtl8150_set_multicast
ocfs2: fix double free in user_cluster_connect()
drivers/base/node: fix double free in register_one_node()
nfp: fix RSS hash key size when RSS is not supported
net: ena: return 0 in ena_get_rxfh_key_size() when RSS hash key is not configurable
Revert "net/mlx5e: Update and set Xon/Xoff upon MTU set"
Squashfs: fix uninit-value in squashfs_get_parent
uio_hv_generic: Let userspace take care of interrupt mask
mm: hugetlb: avoid soft lockup when mprotect to large memory area
Input: uinput - zero-initialize uinput_ff_upload_compat to avoid info leak
pinctrl: check the return value of pinmux_ops::get_function_name()
clocksource/drivers/clps711x: Fix resource leaks in error paths
iio: frequency: adf4350: Fix ADF4350_REG3_12BIT_CLKDIV_MODE
perf util: Fix compression checks returning -1 as bool
rtc: x1205: Fix Xicor X1205 vendor prefix
perf session: Fix handling when buffer exceeds 2 GiB
clk: nxp: lpc18xx-cgu: convert from round_rate() to determine_rate()
clk: nxp: Fix pll0 rate check condition in LPC18xx CGU driver
scsi: libsas: Add sas_task_find_rq()
scsi: mvsas: Delete mvs_tag_init()
scsi: mvsas: Use sas_task_find_rq() for tagging
scsi: mvsas: Fix use-after-free bugs in mvs_work_queue
net/mlx4: prevent potential use after free in mlx4_en_do_uc_filter()
drm/vmwgfx: Fix Use-after-free in validation
net/sctp: fix a null dereference in sctp_disposition sctp_sf_do_5_1D_ce()
tcp: Don't call reqsk_fastopen_remove() in tcp_conn_request().
net: fsl_pq_mdio: Fix device node reference leak in fsl_pq_mdio_probe
tools build: Align warning options with perf
mailbox: zynqmp-ipi: Remove redundant mbox_controller_unregister() call
mailbox: zynqmp-ipi: Remove dev.parent check in zynqmp_ipi_free_mboxes
crypto: essiv - Check ssize for decryption and in-place encryption
tpm, tpm_tis: Claim locality before writing interrupt registers
tpm_tis: Fix incorrect arguments in tpm_tis_probe_irq_single
ACPI: TAD: Add missing sysfs_remove_group() for ACPI_TAD_RT
ACPI: debug: fix signedness issues in read/write helpers
arm64: dts: qcom: msm8916: Add missing MDSS reset
xen/manage: Fix suspend error path
firmware: meson_sm: fix device leak at probe
media: i2c: mt9v111: fix incorrect type for ret
drm/nouveau: fix bad ret code in nouveau_bo_move_prep
cpufreq: intel_pstate: Fix object lifecycle issue in update_qos_request()
crypto: atmel - Fix dma_unmap_sg() direction
iio: dac: ad5360: use int type to store negative error codes
iio: dac: ad5421: use int type to store negative error codes
iio: frequency: adf4350: Fix prescaler usage.
lib/genalloc: fix device leak in of_gen_pool_get()
parisc: don't reference obsolete termio struct for TC* constants
scsi: hpsa: Fix potential memory leak in hpsa_big_passthru_ioctl()
sctp: Fix MAC comparison to be constant-time
sparc64: fix hugetlb for sun4u
sparc: fix error handling in scan_one_device()
mtd: rawnand: fsmc: Default to autodetect buswidth
mmc: core: SPI mode remove cmd7
rtc: interface: Ensure alarm irq is enabled when UIE is enabled
rtc: interface: Fix long-standing race when setting alarm
rseq/selftests: Use weak symbol reference, not definition, to link with glibc
PCI/IOV: Add PCI rescan-remove locking when enabling/disabling SR-IOV
PCI/AER: Fix missing uevent on recovery when a reset is requested
PCI: keystone: Use devm_request_irq() to free "ks-pcie-error-irq" on exit
x86/umip: Check that the instruction opcode is at least two bytes
x86/umip: Fix decoding of register forms of 0F 01 (SGDT and SIDT aliases)
nfsd: nfserr_jukebox in nlm_fopen should lead to a retry
ext4: increase i_disksize to offset + len in ext4_update_disksize_before_punch()
ext4: correctly handle queries for metadata mappings
ext4: guard against EA inode refcount underflow in xattr update
net/9p: fix double req put in p9_fd_cancelled
KVM: x86: Don't (re)check L1 intercepts when completing userspace I/O
fs: udf: fix OOB read in lengthAllocDescs handling
mfd: vexpress-sysreg: Check the return value of devm_gpiochip_add_data()
media: mc: Clear minor number before put device
Squashfs: add additional inode sanity checking
Squashfs: reject negative file sizes in squashfs_read_inode()
mfd: intel_soc_pmic_chtdc_ti: Fix invalid regmap-config max_register value
mfd: intel_soc_pmic_chtdc_ti: Drop unneeded assignment for cache_type
mfd: intel_soc_pmic_chtdc_ti: Set use_single_read regmap_config flag
dm: fix NULL pointer dereference in __dm_suspend()
tracing: Fix race condition in kprobe initialization causing NULL pointer dereference
minixfs: Verify inode mode when loading from disk
pid: Add a judgment for ns null in pid_nr_ns
fs: Add 'initramfs_options' to set initramfs mount options
cramfs: Verify inode mode when loading from disk
xen/events: Cleanup find_virq() return codes
media: cx18: Add missing check after DMA map
pwm: berlin: Fix wrong register in suspend/resume
btrfs: avoid potential out-of-bounds in btrfs_encode_fh()
drm/exynos: exynos7_drm_decon: remove ctx->suspended
media: rc: Directly use ida_free()
media: lirc: Fix error handling in lirc_register()
xen/events: Update virq_to_irq on migration
media: pci/ivtv: switch from 'pci_' to 'dma_' API
media: pci: ivtv: Add missing check after DMA map
net: dl2k: switch from 'pci_' to 'dma_' API
net: dlink: handle dma_map_single() failure properly
net/ip6_tunnel: Prevent perpetual tunnel growth
amd-xgbe: Avoid spurious link down messages during interface toggle
tcp: fix tcp_tso_should_defer() vs large RTT
tg3: prevent use of uninitialized remote_adv and local_adv variables
tls: always set record_type in tls_process_cmsg
tls: don't rely on tx_work during send()
sched: Make newidle_balance() static again
sched/fair: Trivial correction of the newidle_balance() comment
sched/balancing: Rename newidle_balance() => sched_balance_newidle()
sched/fair: Fix pelt lost idle time detection
ALSA: firewire: amdtp-stream: fix enum kernel-doc warnings
hfsplus: fix slab-out-of-bounds read in hfsplus_strcasecmp()
exec: Fix incorrect type for ret
hfs: clear offset and space out of valid records in b-tree node
hfs: make proper initalization of struct hfs_find_data
hfsplus: fix KMSAN uninit-value issue in __hfsplus_ext_cache_extent()
hfs: validate record offset in hfsplus_bmap_alloc
hfsplus: fix KMSAN uninit-value issue in hfsplus_delete_cat()
dlm: check for defined force value in dlm_lockspace_release
hfs: fix KMSAN uninit-value issue in hfs_find_set_zero_bits()
hfsplus: return EIO when type of hidden directory mismatch in hfsplus_fill_super()
m68k: bitops: Fix find_*_bit() signatures
net: rtnetlink: remove redundant assignment to variable err
net: rtnetlink: add msg kind names
net: rtnetlink: add helper to extract msg type's kind
net: rtnetlink: use BIT for flag values
net: netlink: add NLM_F_BULK delete request modifier
net: rtnetlink: add bulk delete support flag
net: add ndo_fdb_del_bulk
net: rtnetlink: add NLM_F_BULK support to rtnl_fdb_del
rtnetlink: Allow deleting FDB entries in user namespace
net: enetc: correct the value of ENETC_RXB_TRUESIZE
arm64, mm: avoid always making PTE dirty in pte_mkwrite()
sctp: avoid NULL dereference when chunk data buffer is missing
net: bonding: fix possible peer notify event loss or dup issue
Revert "cpuidle: menu: Avoid discarding useful information"
MIPS: Malta: Fix keyboard resource preventing i8042 driver from registering
ocfs2: clear extent cache after moving/defragmenting extents
net: usb: rtl8150: Fix frame padding
net: ravb: Ensure memory write completes before ringing TX doorbell
USB: serial: option: add UNISOC UIS7720
USB: serial: option: add Quectel RG255C
USB: serial: option: add Telit FN920C04 ECM compositions
usb/core/quirks: Add Huawei ME906S to wakeup quirk
xhci: dbc: enable back DbC in resume if it was enabled before suspend
binder: remove "invalid inc weak" check
comedi: fix divide-by-zero in comedi_buf_munge()
arm64: cputype: Add Neoverse-V3AE definitions
arm64: errata: Apply workarounds for Neoverse-V3AE
memory: samsung: exynos-srom: Correct alignment
memory: samsung: exynos-srom: Fix of_iomap leak in exynos_srom_probe
spi: cadence-quadspi: Flush posted register writes before INDAC access
spi: cadence-quadspi: Flush posted register writes before DAC access
ext4: avoid potential buffer over-read in parse_apply_sb_mount_options()
drm/amdgpu: use atomic functions with memory barriers for vm fault info
ext4: detect invalid INLINE_DATA + EXTENTS flag combination
jbd2: ensure that all ongoing I/O complete before freeing blocks
vfs: Don't leak disconnected dentries on umount
NFSD: Define a proc_layoutcommit for the FlexFiles layout type
KEYS: trusted_tpm1: Compare HMAC values in constant time
padata: Reset next CPU when reorder sequence wraps around
NFSD: Minor cleanup in layoutcommit processing
NFSD: Fix last write offset handling in layoutcommit
media: s5p-mfc: remove an unused/uninitialized variable
net: rtnetlink: fix module reference count leak issue in rtnetlink_rcv_msg
Linux 5.4.301
Change-Id: Ie2685625a0f630cb01ac8d8c9ddcd68c64ea6ed7
Signed-off-by: Greg Kroah-Hartman <gregkh@google.com>
- Add EXPORT_SYMBOL_GPL for find_task_by_vpid() so that drivers
can be loadable as a module.
- This API is required by loadable driver module from samsung to
read process related information based on pid and thread id.
To get information on when a certain process or thread was started,
duration of run, Average load contributed by it.
Signed-off-by: Abhilasha Rao <abhilasha.hv@samsung.corp-partner.google.com>
Bug: 158067689
Change-Id: I0db9cc50c93eedff0f3e9dea0ac09a5d17d118f0
(cherry picked from commit bee18dd57e89f4e7aa79f6054b9238b99e45c191)
struct pid's count is an atomic_t field used as a refcount. Use
refcount_t for it which is basically atomic_t but does additional
checking to prevent use-after-free bugs.
For memory ordering, the only change is with the following:
- if ((atomic_read(&pid->count) == 1) ||
- atomic_dec_and_test(&pid->count)) {
+ if (refcount_dec_and_test(&pid->count)) {
kmem_cache_free(ns->pid_cachep, pid);
Here the change is from: Fully ordered --> RELEASE + ACQUIRE (as per
refcount-vs-atomic.rst) This ACQUIRE should take care of making sure the
free happens after the refcount_dec_and_test().
The above hunk also removes atomic_read() since it is not needed for the
code to work and it is unclear how beneficial it is. The removal lets
refcount_dec_and_test() check for cases where get_pid() happened before
the object was freed.
Link: http://lkml.kernel.org/r/20190701183826.191936-1-joel@joelfernandes.org
Signed-off-by: Joel Fernandes (Google) <joel@joelfernandes.org>
Reviewed-by: Andrea Parri <andrea.parri@amarulasolutions.com>
Reviewed-by: Kees Cook <keescook@chromium.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Will Deacon <will.deacon@arm.com>
Cc: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Cc: Elena Reshetova <elena.reshetova@intel.com>
Cc: Jann Horn <jannh@google.com>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Cc: KJ Tsanaktsidis <ktsanaktsidis@zendesk.com>
Cc: Michal Hocko <mhocko@suse.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQRAhzRXHqcMeLMyaSiRxhvAZXjcogUCXSMhUgAKCRCRxhvAZXjc
okkiAQC3Hlg/O2JoIb4PqgEvBkpHSdVxyuWagn0ksjACW9ANKQEAl5OadMhvOq16
UHGhKlpE/M8HflknIffoEGlIAWHrdwU=
=7kP5
-----END PGP SIGNATURE-----
Merge tag 'pidfd-updates-v5.3' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux
Pull pidfd updates from Christian Brauner:
"This adds two main features.
- First, it adds polling support for pidfds. This allows process
managers to know when a (non-parent) process dies in a race-free
way.
The notification mechanism used follows the same logic that is
currently used when the parent of a task is notified of a child's
death. With this patchset it is possible to put pidfds in an
{e}poll loop and get reliable notifications for process (i.e.
thread-group) exit.
- The second feature compliments the first one by making it possible
to retrieve pollable pidfds for processes that were not created
using CLONE_PIDFD.
A lot of processes get created with traditional PID-based calls
such as fork() or clone() (without CLONE_PIDFD). For these
processes a caller can currently not create a pollable pidfd. This
is a problem for Android's low memory killer (LMK) and service
managers such as systemd.
Both patchsets are accompanied by selftests.
It's perhaps worth noting that the work done so far and the work done
in this branch for pidfd_open() and polling support do already see
some adoption:
- Android is in the process of backporting this work to all their LTS
kernels [1]
- Service managers make use of pidfd_send_signal but will need to
wait until we enable waiting on pidfds for full adoption.
- And projects I maintain make use of both pidfd_send_signal and
CLONE_PIDFD [2] and will use polling support and pidfd_open() too"
[1] https://android-review.googlesource.com/q/topic:%22pidfd+polling+support+4.9+backport%22https://android-review.googlesource.com/q/topic:%22pidfd+polling+support+4.14+backport%22https://android-review.googlesource.com/q/topic:%22pidfd+polling+support+4.19+backport%22
[2] aab6e3eb73/src/lxc/start.c (L1753)
* tag 'pidfd-updates-v5.3' of git://git.kernel.org/pub/scm/linux/kernel/git/brauner/linux:
tests: add pidfd_open() tests
arch: wire-up pidfd_open()
pid: add pidfd_open()
pidfd: add polling selftests
pidfd: add polling support
This adds the pidfd_open() syscall. It allows a caller to retrieve pollable
pidfds for a process which did not get created via CLONE_PIDFD, i.e. for a
process that is created via traditional fork()/clone() calls that is only
referenced by a PID:
int pidfd = pidfd_open(1234, 0);
ret = pidfd_send_signal(pidfd, SIGSTOP, NULL, 0);
With the introduction of pidfds through CLONE_PIDFD it is possible to
created pidfds at process creation time.
However, a lot of processes get created with traditional PID-based calls
such as fork() or clone() (without CLONE_PIDFD). For these processes a
caller can currently not create a pollable pidfd. This is a problem for
Android's low memory killer (LMK) and service managers such as systemd.
Both are examples of tools that want to make use of pidfds to get reliable
notification of process exit for non-parents (pidfd polling) and race-free
signal sending (pidfd_send_signal()). They intend to switch to this API for
process supervision/management as soon as possible. Having no way to get
pollable pidfds from PID-only processes is one of the biggest blockers for
them in adopting this api. With pidfd_open() making it possible to retrieve
pidfds for PID-based processes we enable them to adopt this api.
In line with Arnd's recent changes to consolidate syscall numbers across
architectures, I have added the pidfd_open() syscall to all architectures
at the same time.
Signed-off-by: Christian Brauner <christian@brauner.io>
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Acked-by: Arnd Bergmann <arnd@arndb.de>
Cc: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: Kees Cook <keescook@chromium.org>
Cc: Joel Fernandes (Google) <joel@joelfernandes.org>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Jann Horn <jannh@google.com>
Cc: Andy Lutomirsky <luto@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Aleksa Sarai <cyphar@cyphar.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: linux-api@vger.kernel.org
This patch adds polling support to pidfd.
Android low memory killer (LMK) needs to know when a process dies once
it is sent the kill signal. It does so by checking for the existence of
/proc/pid which is both racy and slow. For example, if a PID is reused
between when LMK sends a kill signal and checks for existence of the
PID, since the wrong PID is now possibly checked for existence.
Using the polling support, LMK will be able to get notified when a process
exists in race-free and fast way, and allows the LMK to do other things
(such as by polling on other fds) while awaiting the process being killed
to die.
For notification to polling processes, we follow the same existing
mechanism in the kernel used when the parent of the task group is to be
notified of a child's death (do_notify_parent). This is precisely when the
tasks waiting on a poll of pidfd are also awakened in this patch.
We have decided to include the waitqueue in struct pid for the following
reasons:
1. The wait queue has to survive for the lifetime of the poll. Including
it in task_struct would not be option in this case because the task can
be reaped and destroyed before the poll returns.
2. By including the struct pid for the waitqueue means that during
de_thread(), the new thread group leader automatically gets the new
waitqueue/pid even though its task_struct is different.
Appropriate test cases are added in the second patch to provide coverage of
all the cases the patch is handling.
Cc: Andy Lutomirski <luto@amacapital.net>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Daniel Colascione <dancol@google.com>
Cc: Jann Horn <jannh@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: Jonathan Kowalski <bl0pbl33p@gmail.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Kees Cook <keescook@chromium.org>
Cc: David Howells <dhowells@redhat.com>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: kernel-team@android.com
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Co-developed-by: Daniel Colascione <dancol@google.com>
Signed-off-by: Daniel Colascione <dancol@google.com>
Signed-off-by: Joel Fernandes (Google) <joel@joelfernandes.org>
Signed-off-by: Christian Brauner <christian@brauner.io>
Add SPDX license identifiers to all files which:
- Have no license information of any form
- Have EXPORT_.*_SYMBOL_GPL inside which was used in the
initial scan/conversion to ignore the file
These files fall under the project license, GPL v2 only. The resulting SPDX
license identifier is:
GPL-2.0-only
Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Hash functions are not needed since idr is used now. Let's remove hash
header file for cleanup.
Link: http://lkml.kernel.org/r/20190430053319.95913-1-scuttimmy@gmail.com
Signed-off-by: Timmy Li <scuttimmy@gmail.com>
Cc: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Mike Rapoport <rppt@linux.vnet.ibm.com>
Cc: KJ Tsanaktsidis <ktsanaktsidis@zendesk.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
The failure path removes the allocated PIDs from the wrong namespace.
This could lead to us inadvertently reusing PIDs in the leaf namespace
and leaking PIDs in parent namespaces.
Fixes: 95846ecf9d ("pid: replace pid bitmap implementation with IDR API")
Cc: <stable@vger.kernel.org>
Signed-off-by: Matthew Wilcox <willy@infradead.org>
Acked-by: "Eric W. Biederman" <ebiederm@xmission.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Move remaining definitions and declarations from include/linux/bootmem.h
into include/linux/memblock.h and remove the redundant header.
The includes were replaced with the semantic patch below and then
semi-automated removal of duplicated '#include <linux/memblock.h>
@@
@@
- #include <linux/bootmem.h>
+ #include <linux/memblock.h>
[sfr@canb.auug.org.au: dma-direct: fix up for the removal of linux/bootmem.h]
Link: http://lkml.kernel.org/r/20181002185342.133d1680@canb.auug.org.au
[sfr@canb.auug.org.au: powerpc: fix up for removal of linux/bootmem.h]
Link: http://lkml.kernel.org/r/20181005161406.73ef8727@canb.auug.org.au
[sfr@canb.auug.org.au: x86/kaslr, ACPI/NUMA: fix for linux/bootmem.h removal]
Link: http://lkml.kernel.org/r/20181008190341.5e396491@canb.auug.org.au
Link: http://lkml.kernel.org/r/1536927045-23536-30-git-send-email-rppt@linux.vnet.ibm.com
Signed-off-by: Mike Rapoport <rppt@linux.vnet.ibm.com>
Signed-off-by: Stephen Rothwell <sfr@canb.auug.org.au>
Acked-by: Michal Hocko <mhocko@suse.com>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Chris Zankel <chris@zankel.net>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: Geert Uytterhoeven <geert@linux-m68k.org>
Cc: Greentime Hu <green.hu@gmail.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Guan Xuetao <gxt@pku.edu.cn>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: "James E.J. Bottomley" <jejb@parisc-linux.org>
Cc: Jonas Bonn <jonas@southpole.se>
Cc: Jonathan Corbet <corbet@lwn.net>
Cc: Ley Foon Tan <lftan@altera.com>
Cc: Mark Salter <msalter@redhat.com>
Cc: Martin Schwidefsky <schwidefsky@de.ibm.com>
Cc: Matt Turner <mattst88@gmail.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Michal Simek <monstr@monstr.eu>
Cc: Palmer Dabbelt <palmer@sifive.com>
Cc: Paul Burton <paul.burton@mips.com>
Cc: Richard Kuo <rkuo@codeaurora.org>
Cc: Richard Weinberger <richard@nod.at>
Cc: Rich Felker <dalias@libc.org>
Cc: Russell King <linux@armlinux.org.uk>
Cc: Serge Semin <fancer.lancer@gmail.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Tony Luck <tony.luck@intel.com>
Cc: Vineet Gupta <vgupta@synopsys.com>
Cc: Yoshinori Sato <ysato@users.sourceforge.jp>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Make the clone and fork syscalls return EAGAIN when the limit on the
number of pids /proc/sys/kernel/pid_max is exceeded.
Currently, when the pid_max limit is exceeded, the kernel will return
ENOSPC from the fork and clone syscalls. This is contrary to the
documented behaviour, which explicitly calls out the pid_max case as one
where EAGAIN should be returned. It also leads to really confusing error
messages in userspace programs which will complain about a lack of disk
space when they fail to create processes/threads for this reason.
This error is being returned because alloc_pid() uses the idr api to find
a new pid; when there are none available, idr_alloc_cyclic() returns
-ENOSPC, and this is being propagated back to userspace.
This behaviour has been broken before, and was explicitly fixed in
commit 35f71bc0a0 ("fork: report pid reservation failure properly"),
so I think -EAGAIN is definitely the right thing to return in this case.
The current behaviour change dates from commit 95846ecf9d ("pid:
replace pid bitmap implementation with IDR AIP") and was I believe
unintentional.
This patch has no impact on the case where allocating a pid fails because
the child reaper for the namespace is dead; that case will still return
-ENOMEM.
Link: http://lkml.kernel.org/r/20180903111016.46461-1-ktsanaktsidis@zendesk.com
Fixes: 95846ecf9d ("pid: replace pid bitmap implementation with IDR AIP")
Signed-off-by: KJ Tsanaktsidis <ktsanaktsidis@zendesk.com>
Reviewed-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Michal Hocko <mhocko@suse.com>
Cc: Gargi Sharma <gs051095@gmail.com>
Cc: Rik van Riel <riel@redhat.com>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Everywhere except in the pid array we distinguish between a tasks pid and
a tasks tgid (thread group id). Even in the enumeration we want that
distinction sometimes so we have added __PIDTYPE_TGID. With leader_pid
we almost have an implementation of PIDTYPE_TGID in struct signal_struct.
Add PIDTYPE_TGID as a first class member of the pid_type enumeration and
into the pids array. Then remove the __PIDTYPE_TGID special case and the
leader_pid in signal_struct.
The net size increase is just an extra pointer added to struct pid and
an extra pair of pointers of an hlist_node added to task_struct.
The effect on code maintenance is the removal of a number of special
cases today and the potential to remove many more special cases as
PIDTYPE_TGID gets used to it's fullest. The long term potential
is allowing zombie thread group leaders to exit, which will remove
a lot more special cases in the code.
Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
To access these fields the code always has to go to group leader so
going to signal struct is no loss and is actually a fundamental simplification.
This saves a little bit of memory by only allocating the pid pointer array
once instead of once for every thread, and even better this removes a
few potential races caused by the fact that group_leader can be changed
by de_thread, while signal_struct can not.
Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
The cost is the the same and this removes the need
to worry about complications that come from de_thread
and group_leader changing.
__task_pid_nr_ns has been updated to take advantage of this change.
Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
This results in no change in structure size on 64-bit machines as it
fits in the padding between the gfp_t and the void *. 32-bit machines
will grow the structure from 8 to 12 bytes. Almost all radix trees are
protected with (at least) a spinlock, so as they are converted from
radix trees to xarrays, the data structures will shrink again.
Initialising the spinlock requires a name for the benefit of lockdep, so
RADIX_TREE_INIT() now needs to know the name of the radix tree it's
initialising, and so do IDR_INIT() and IDA_INIT().
Also add the xa_lock() and xa_unlock() family of wrappers to make it
easier to use the lock. If we could rely on -fplan9-extensions in the
compiler, we could avoid all of this syntactic sugar, but that wasn't
added until gcc 4.6.
Link: http://lkml.kernel.org/r/20180313132639.17387-8-willy@infradead.org
Signed-off-by: Matthew Wilcox <mawilcox@microsoft.com>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Cc: Darrick J. Wong <darrick.wong@oracle.com>
Cc: Dave Chinner <david@fromorbit.com>
Cc: Ryusuke Konishi <konishi.ryusuke@lab.ntt.co.jp>
Cc: Will Deacon <will.deacon@arm.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
There are several functions that do find_task_by_vpid() followed by
get_task_struct(). We can use a helper function instead.
Link: http://lkml.kernel.org/r/1509602027-11337-1-git-send-email-rppt@linux.vnet.ibm.com
Signed-off-by: Mike Rapoport <rppt@linux.vnet.ibm.com>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
-----BEGIN PGP SIGNATURE-----
iQIVAwUAWl80tvSw1s6N8H32AQJq8A//ViRN5fExrd678Eh2Bz1ytrJYMUfYY3Hv
QTH5TH9zFyLFyWLB1Iwe13sdLVTTM88O0qcDb54Lx9fWUqeMZyYvBhLtWPc00lTU
0m3EyYR87MFWaEV+VxaVWgWaWkMDkd39KubDitcS+YIBDszTuMpYodhPUsgLt7lr
pePX7eurXKdQPTh4NUOjGA2NaZot3tga76J6D8NKruGYUstQCGxpP1ryiFfACnwf
NLWNO8ZBMtlDwX1mHYOOMFMaBzFzXorPm7jY4HJDf3mUM84xI3ach6CuH9RTSzfq
A+qB1U3QILPVFo2HtqOHui4bFjRwqOf6uIrI/KcnioJ37w1O+KFcMJeDnX2I211q
f2lXehJLQA7kPmxQw8T3//HDRaLXc0Qxt7IPZRFinrlkcN4oh3DD5euMfCFBSoZG
PTbjxlgMfzJPoZtqAcy0rV5L54a/F4h915OQPJCKLwujIsXD2nT993vNmGDyq4zh
BzNMxSXJC8p+jYvQpNhWyyxwDBBT/YsVQo/ACwg4eJnD3blVTAioRT9ZZcAcsY0F
0z1eWW5RiknzIaXQWvjfK0gYKpO+aMSu9+gipHfMbU3yXG+sPj/H6zAHYzqX3uQZ
jb5Iujjnu49W/YD+RiMenuu59lNXUnLSeRnlV7dw0qxGK1FzGo24+ZzKFhJhKvzG
tdfUsev1Mc8=
=jhWg
-----END PGP SIGNATURE-----
Merge tag 'init_task-20180117' of git://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs
Pull init_task initializer cleanups from David Howells:
"It doesn't seem useful to have the init_task in a header file rather
than in a normal source file. We could consolidate init_task handling
instead and expand out various macros.
Here's a series of patches that consolidate init_task handling:
(1) Make THREAD_SIZE available to vmlinux.lds for cris, hexagon and
openrisc.
(2) Alter the INIT_TASK_DATA linker script macro to set
init_thread_union and init_stack rather than defining these in C.
Insert init_task and init_thread_into into the init_stack area in
the linker script as appropriate to the configuration, with
different section markers so that they end up correctly ordered.
We can then get merge ia64's init_task.c into the main one.
We then have a bunch of single-use INIT_*() macros that seem only
to be macros because they used to be used per-arch. We can then
expand these in place of the user and get rid of a few lines and
a lot of backslashes.
(3) Expand INIT_TASK() in place.
(4) Expand in place various small INIT_*() macros that are defined
conditionally. Expand them and surround them by #if[n]def/#endif
in the .c file as it takes fewer lines.
(5) Expand INIT_SIGNALS() and INIT_SIGHAND() in place.
(6) Expand INIT_STRUCT_PID in place.
These macros can then be discarded"
* tag 'init_task-20180117' of git://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs:
Expand INIT_STRUCT_PID and remove
Expand the INIT_SIGNALS and INIT_SIGHAND macros and remove
Expand various INIT_* macros and remove
Expand INIT_TASK() in init/init_task.c and remove
Construct init thread stack in the linker script rather than by union
openrisc: Make THREAD_SIZE available to vmlinux.lds
hexagon: Make THREAD_SIZE available to vmlinux.lds
cris: Make THREAD_SIZE available to vmlinux.lds
Expand INIT_STRUCT_PID in the single place that uses it and then remove it.
There doesn't seem any point in the macro.
Signed-off-by: David Howells <dhowells@redhat.com>
Tested-by: Tony Luck <tony.luck@intel.com>
Tested-by: Will Deacon <will.deacon@arm.com> (arm64)
Tested-by: Palmer Dabbelt <palmer@sifive.com>
Acked-by: Thomas Gleixner <tglx@linutronix.de>
pidhash is no longer required as all the information can be looked up
from idr tree. nr_hashed represented the number of pids that had been
hashed. Since, nr_hashed and PIDNS_HASH_ADDING are no longer relevant,
it has been renamed to pid_allocated and PIDNS_ADDING respectively.
[gs051095@gmail.com: v6]
Link: http://lkml.kernel.org/r/1507760379-21662-3-git-send-email-gs051095@gmail.com
Link: http://lkml.kernel.org/r/1507583624-22146-3-git-send-email-gs051095@gmail.com
Signed-off-by: Gargi Sharma <gs051095@gmail.com>
Reviewed-by: Rik van Riel <riel@redhat.com>
Tested-by: Tony Luck <tony.luck@intel.com> [ia64]
Cc: Julia Lawall <julia.lawall@lip6.fr>
Cc: Ingo Molnar <mingo@kernel.org>
Cc: Pavel Tatashin <pasha.tatashin@oracle.com>
Cc: Kirill Tkhai <ktkhai@virtuozzo.com>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Cc: Christoph Hellwig <hch@infradead.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Patch series "Replacing PID bitmap implementation with IDR API", v4.
This series replaces kernel bitmap implementation of PID allocation with
IDR API. These patches are written to simplify the kernel by replacing
custom code with calls to generic code.
The following are the stats for pid and pid_namespace object files
before and after the replacement. There is a noteworthy change between
the IDR and bitmap implementation.
Before
text data bss dec hex filename
8447 3894 64 12405 3075 kernel/pid.o
After
text data bss dec hex filename
3397 304 0 3701 e75 kernel/pid.o
Before
text data bss dec hex filename
5692 1842 192 7726 1e2e kernel/pid_namespace.o
After
text data bss dec hex filename
2854 216 16 3086 c0e kernel/pid_namespace.o
The following are the stats for ps, pstree and calling readdir on /proc
for 10,000 processes.
ps:
With IDR API With bitmap
real 0m1.479s 0m2.319s
user 0m0.070s 0m0.060s
sys 0m0.289s 0m0.516s
pstree:
With IDR API With bitmap
real 0m1.024s 0m1.794s
user 0m0.348s 0m0.612s
sys 0m0.184s 0m0.264s
proc:
With IDR API With bitmap
real 0m0.059s 0m0.074s
user 0m0.000s 0m0.004s
sys 0m0.016s 0m0.016s
This patch (of 2):
Replace the current bitmap implementation for Process ID allocation.
Functions that are no longer required, for example, free_pidmap(),
alloc_pidmap(), etc. are removed. The rest of the functions are
modified to use the IDR API. The change was made to make the PID
allocation less complex by replacing custom code with calls to generic
API.
[gs051095@gmail.com: v6]
Link: http://lkml.kernel.org/r/1507760379-21662-2-git-send-email-gs051095@gmail.com
[avagin@openvz.org: restore the old behaviour of the ns_last_pid sysctl]
Link: http://lkml.kernel.org/r/20171106183144.16368-1-avagin@openvz.org
Link: http://lkml.kernel.org/r/1507583624-22146-2-git-send-email-gs051095@gmail.com
Signed-off-by: Gargi Sharma <gs051095@gmail.com>
Reviewed-by: Rik van Riel <riel@redhat.com>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Cc: Julia Lawall <julia.lawall@lip6.fr>
Cc: Ingo Molnar <mingo@kernel.org>
Cc: Pavel Tatashin <pasha.tatashin@oracle.com>
Cc: Kirill Tkhai <ktkhai@virtuozzo.com>
Cc: Eric W. Biederman <ebiederm@xmission.com>
Cc: Christoph Hellwig <hch@infradead.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
This was reported many times, and this was even mentioned in commit
52ee2dfdd4 ("pids: refactor vnr/nr_ns helpers to make them safe") but
somehow nobody bothered to fix the obvious problem: task_tgid_nr_ns() is
not safe because task->group_leader points to nowhere after the exiting
task passes exit_notify(), rcu_read_lock() can not help.
We really need to change __unhash_process() to nullify group_leader,
parent, and real_parent, but this needs some cleanups. Until then we
can turn task_tgid_nr_ns() into another user of __task_pid_nr_ns() and
fix the problem.
Reported-by: Troy Kensinger <tkensinger@google.com>
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
After commit 3d375d7859 ("mm: update callers to use HASH_ZERO flag"),
drop unused pidhash_size in pidhash_init().
Link: http://lkml.kernel.org/r/1500389267-49222-1-git-send-email-wangkefeng.wang@huawei.com
Signed-off-by: Kefeng Wang <wangkefeng.wang@huawei.com>
Reviewed-by: Pavel Tatashin <Pasha.Tatashin@Oracle.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
alloc_pidmap() advances pid_namespace::last_pid. When first pid
allocation fails, then next created process will have pid 2 and
pid_ns_prepare_proc() won't be called. So, pid_namespace::proc_mnt will
never be initialized (not to mention that there won't be a child
reaper).
I saw crash stack of such case on kernel 3.10:
BUG: unable to handle kernel NULL pointer dereference at (null)
IP: proc_flush_task+0x8f/0x1b0
Call Trace:
release_task+0x3f/0x490
wait_consider_task.part.10+0x7ff/0xb00
do_wait+0x11f/0x280
SyS_wait4+0x7d/0x110
We may fix this by restore of last_pid in 0 or by prohibiting of futher
allocations. Since there was a similar issue in Oleg Nesterov's commit
314a8ad0f1 ("pidns: fix free_pid() to handle the first fork failure").
and it was fixed via prohibiting allocation, let's follow this way, and
do the same.
Link: http://lkml.kernel.org/r/149201021004.4863.6762095011554287922.stgit@localhost.localdomain
Signed-off-by: Kirill Tkhai <ktkhai@virtuozzo.com>
Acked-by: Cyrill Gorcunov <gorcunov@openvz.org>
Cc: Andrei Vagin <avagin@virtuozzo.com>
Cc: Andreas Gruenbacher <agruenba@redhat.com>
Cc: Kees Cook <keescook@chromium.org>
Cc: Michael Kerrisk <mtk.manpages@googlemail.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Paul Moore <paul@paul-moore.com>
Cc: Eric Biederman <ebiederm@xmission.com>
Cc: Andy Lutomirski <luto@amacapital.net>
Cc: Ingo Molnar <mingo@kernel.org>
Cc: Serge Hallyn <serge@hallyn.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
We are going to split <linux/sched/task.h> out of <linux/sched.h>, which
will have to be picked up from other headers and a couple of .c files.
Create a trivial placeholder <linux/sched/task.h> file that just
maps to <linux/sched.h> to make this patch obviously correct and
bisectable.
Include the new header in the files that are going to need it.
Acked-by: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Mike Galbraith <efault@gmx.de>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: linux-kernel@vger.kernel.org
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Since we need to change the implementation, stop exposing internals.
Provide KREF_INIT() to allow static initialization of struct kref.
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: linux-kernel@vger.kernel.org
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Most users of IS_ERR_VALUE() in the kernel are wrong, as they
pass an 'int' into a function that takes an 'unsigned long'
argument. This happens to work because the type is sign-extended
on 64-bit architectures before it gets converted into an
unsigned type.
However, anything that passes an 'unsigned short' or 'unsigned int'
argument into IS_ERR_VALUE() is guaranteed to be broken, as are
8-bit integers and types that are wider than 'unsigned long'.
Andrzej Hajda has already fixed a lot of the worst abusers that
were causing actual bugs, but it would be nice to prevent any
users that are not passing 'unsigned long' arguments.
This patch changes all users of IS_ERR_VALUE() that I could find
on 32-bit ARM randconfig builds and x86 allmodconfig. For the
moment, this doesn't change the definition of IS_ERR_VALUE()
because there are probably still architecture specific users
elsewhere.
Almost all the warnings I got are for files that are better off
using 'if (err)' or 'if (err < 0)'.
The only legitimate user I could find that we get a warning for
is the (32-bit only) freescale fman driver, so I did not remove
the IS_ERR_VALUE() there but changed the type to 'unsigned long'.
For 9pfs, I just worked around one user whose calling conventions
are so obscure that I did not dare change the behavior.
I was using this definition for testing:
#define IS_ERR_VALUE(x) ((unsigned long*)NULL == (typeof (x)*)NULL && \
unlikely((unsigned long long)(x) >= (unsigned long long)(typeof(x))-MAX_ERRNO))
which ends up making all 16-bit or wider types work correctly with
the most plausible interpretation of what IS_ERR_VALUE() was supposed
to return according to its users, but also causes a compile-time
warning for any users that do not pass an 'unsigned long' argument.
I suggested this approach earlier this year, but back then we ended
up deciding to just fix the users that are obviously broken. After
the initial warning that caused me to get involved in the discussion
(fs/gfs2/dir.c) showed up again in the mainline kernel, Linus
asked me to send the whole thing again.
[ Updated the 9p parts as per Al Viro - Linus ]
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
Cc: Andrzej Hajda <a.hajda@samsung.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Link: https://lkml.org/lkml/2016/1/7/363
Link: https://lkml.org/lkml/2016/5/27/486
Acked-by: Srinivas Kandagatla <srinivas.kandagatla@linaro.org> # For nvmem part
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Pull scheduler fixes from Thomas Gleixner:
"Three small fixes in the scheduler/core:
- use after free in the numa code
- crash in the numa init code
- a simple spelling fix"
* 'sched-urgent-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
pid: Fix spelling in comments
sched/numa: Fix use-after-free bug in the task_numa_compare
sched: Fix crash in sched_init_numa()
Mark those kmem allocations that are known to be easily triggered from
userspace as __GFP_ACCOUNT/SLAB_ACCOUNT, which makes them accounted to
memcg. For the list, see below:
- threadinfo
- task_struct
- task_delay_info
- pid
- cred
- mm_struct
- vm_area_struct and vm_region (nommu)
- anon_vma and anon_vma_chain
- signal_struct
- sighand_struct
- fs_struct
- files_struct
- fdtable and fdtable->full_fds_bits
- dentry and external_name
- inode for all filesystems. This is the most tedious part, because
most filesystems overwrite the alloc_inode method.
The list is far from complete, so feel free to add more objects.
Nevertheless, it should be close to "account everything" approach and
keep most workloads within bounds. Malevolent users will be able to
breach the limit, but this was possible even with the former "account
everything" approach (simply because it did not account everything in
fact).
[akpm@linux-foundation.org: coding-style fixes]
Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>
Acked-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: Michal Hocko <mhocko@suse.com>
Cc: Tejun Heo <tj@kernel.org>
Cc: Greg Thelen <gthelen@google.com>
Cc: Christoph Lameter <cl@linux.com>
Cc: Pekka Enberg <penberg@kernel.org>
Cc: David Rientjes <rientjes@google.com>
Cc: Joonsoo Kim <iamjoonsoo.kim@lge.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
I got a crash during a "perf top" session that was caused by a race in
__task_pid_nr_ns() :
pid_nr_ns() was inlined, but apparently compiler chose to read
task->pids[type].pid twice, and the pid->level dereference crashed
because we got a NULL pointer at the second read :
if (pid && ns->level <= pid->level) { // CRASH
Just use RCU API properly to solve this race, and not worry about "perf
top" crashing hosts :(
get_task_pid() can benefit from same fix.
Signed-off-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
This commit renames rcu_lockdep_assert() to RCU_LOCKDEP_WARN() for
consistency with the WARN() series of macros. This also requires
inverting the sense of the conditional, which this commit also does.
Reported-by: Ingo Molnar <mingo@kernel.org>
Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Reviewed-by: Ingo Molnar <mingo@kernel.org>
copy_process will report any failure in alloc_pid as ENOMEM currently
which is misleading because the pid allocation might fail not only when
the memory is short but also when the pid space is consumed already.
The current man page even mentions this case:
: EAGAIN
:
: A system-imposed limit on the number of threads was encountered.
: There are a number of limits that may trigger this error: the
: RLIMIT_NPROC soft resource limit (set via setrlimit(2)), which
: limits the number of processes and threads for a real user ID, was
: reached; the kernel's system-wide limit on the number of processes
: and threads, /proc/sys/kernel/threads-max, was reached (see
: proc(5)); or the maximum number of PIDs, /proc/sys/kernel/pid_max,
: was reached (see proc(5)).
so the current behavior is also incorrect wrt. documentation. POSIX man
page also suggest returing EAGAIN when the process count limit is reached.
This patch simply propagates error code from alloc_pid and makes sure we
return -EAGAIN due to reservation failure. This will make behavior of
fork closer to both our documentation and POSIX.
alloc_pid might alsoo fail when the reaper in the pid namespace is dead
(the namespace basically disallows all new processes) and there is no
good error code which would match documented ones. We have traditionally
returned ENOMEM for this case which is misleading as well but as per
Eric W. Biederman this behavior is documented in man pid_namespaces(7)
: If the "init" process of a PID namespace terminates, the kernel
: terminates all of the processes in the namespace via a SIGKILL signal.
: This behavior reflects the fact that the "init" process is essential for
: the correct operation of a PID namespace. In this case, a subsequent
: fork(2) into this PID namespace will fail with the error ENOMEM; it is
: not possible to create a new processes in a PID namespace whose "init"
: process has terminated.
and introducing a new error code would be too risky so let's stick to
ENOMEM for this case.
Signed-off-by: Michal Hocko <mhocko@suse.cz>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: Michael Kerrisk <mtk.manpages@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Pull vfs pile #2 from Al Viro:
"Next pile (and there'll be one or two more).
The large piece in this one is getting rid of /proc/*/ns/* weirdness;
among other things, it allows to (finally) make nameidata completely
opaque outside of fs/namei.c, making for easier further cleanups in
there"
* 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs:
coda_venus_readdir(): use file_inode()
fs/namei.c: fold link_path_walk() call into path_init()
path_init(): don't bother with LOOKUP_PARENT in argument
fs/namei.c: new helper (path_cleanup())
path_init(): store the "base" pointer to file in nameidata itself
make default ->i_fop have ->open() fail with ENXIO
make nameidata completely opaque outside of fs/namei.c
kill proc_ns completely
take the targets of /proc/*/ns/* symlinks to separate fs
bury struct proc_ns in fs/proc
copy address of proc_ns_ops into ns_common
new helpers: ns_alloc_inum/ns_free_inum
make proc_ns_operations work with struct ns_common * instead of void *
switch the rest of proc_ns_operations to working with &...->ns
netns: switch ->get()/->put()/->install()/->inum() to working with &net->ns
make mntns ->get()/->put()/->install()/->inum() work with &mnt_ns->ns
common object embedded into various struct ....ns
alloc_pid() does get_pid_ns() beforehand but forgets to put_pid_ns() if it
fails because disable_pid_allocation() was called by the exiting
child_reaper.
We could simply move get_pid_ns() down to successful return, but this fix
tries to be as trivial as possible.
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: Aaron Tomlin <atomlin@redhat.com>
Cc: Pavel Emelyanov <xemul@parallels.com>
Cc: Serge Hallyn <serge.hallyn@ubuntu.com>
Cc: Sterling Alexander <stalexan@redhat.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
for now - just move corresponding ->proc_inum instances over there
Acked-by: "Eric W. Biederman" <ebiederm@xmission.com>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
"case 0" in free_pid() assumes that disable_pid_allocation() should
clear PIDNS_HASH_ADDING before the last pid goes away.
However this doesn't happen if the first fork() fails to create the
child reaper which should call disable_pid_allocation().
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: "Serge E. Hallyn" <serge@hallyn.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Serge Hallyn <serge.hallyn@ubuntu.com> writes:
> Since commit af4b8a83ad it's been
> possible to get into a situation where a pidns reaper is
> <defunct>, reparented to host pid 1, but never reaped. How to
> reproduce this is documented at
>
> https://bugs.launchpad.net/ubuntu/+source/lxc/+bug/1168526
> (and see
> https://bugs.launchpad.net/ubuntu/+source/lxc/+bug/1168526/comments/13)
> In short, run repeated starts of a container whose init is
>
> Process.exit(0);
>
> sysrq-t when such a task is playing zombie shows:
>
> [ 131.132978] init x ffff88011fc14580 0 2084 2039 0x00000000
> [ 131.132978] ffff880116e89ea8 0000000000000002 ffff880116e89fd8 0000000000014580
> [ 131.132978] ffff880116e89fd8 0000000000014580 ffff8801172a0000 ffff8801172a0000
> [ 131.132978] ffff8801172a0630 ffff88011729fff0 ffff880116e14650 ffff88011729fff0
> [ 131.132978] Call Trace:
> [ 131.132978] [<ffffffff816f6159>] schedule+0x29/0x70
> [ 131.132978] [<ffffffff81064591>] do_exit+0x6e1/0xa40
> [ 131.132978] [<ffffffff81071eae>] ? signal_wake_up_state+0x1e/0x30
> [ 131.132978] [<ffffffff8106496f>] do_group_exit+0x3f/0xa0
> [ 131.132978] [<ffffffff810649e4>] SyS_exit_group+0x14/0x20
> [ 131.132978] [<ffffffff8170102f>] tracesys+0xe1/0xe6
>
> Further debugging showed that every time this happened, zap_pid_ns_processes()
> started with nr_hashed being 3, while we were expecting it to drop to 2.
> Any time it didn't happen, nr_hashed was 1 or 2. So the reaper was
> waiting for nr_hashed to become 2, but free_pid() only wakes the reaper
> if nr_hashed hits 1.
The issue is that when the task group leader of an init process exits
before other tasks of the init process when the init process finally
exits it will be a secondary task sleeping in zap_pid_ns_processes and
waiting to wake up when the number of hashed pids drops to two. This
case waits forever as free_pid only sends a wake up when the number of
hashed pids drops to 1.
To correct this the simple strategy of sending a possibly unncessary
wake up when the number of hashed pids drops to 2 is adopted.
Sending one extraneous wake up is relatively harmless, at worst we
waste a little cpu time in the rare case when a pid namespace
appropaches exiting.
We can detect the case when the pid namespace drops to just two pids
hashed race free in free_pid.
Dereferencing pid_ns->child_reaper with the pidmap_lock held is safe
without out the tasklist_lock because it is guaranteed that the
detach_pid will be called on the child_reaper before it is freed and
detach_pid calls __change_pid which calls free_pid which takes the
pidmap_lock. __change_pid only calls free_pid if this is the
last use of the pid. For a thread that is not the thread group leader
the threads pid will only ever have one user because a threads pid
is not allowed to be the pid of a process, of a process group or
a session. For a thread that is a thread group leader all of
the other threads of that process will be reaped before it is allowed
for the thread group leader to be reaped ensuring there will only
be one user of the threads pid as a process pid. Furthermore
because the thread is the init process of a pid namespace all of the
other processes in the pid namespace will have also been already freed
leading to the fact that the pid will not be used as a session pid or
a process group pid for any other running process.
CC: stable@vger.kernel.org
Acked-by: Serge Hallyn <serge.hallyn@canonical.com>
Tested-by: Serge Hallyn <serge.hallyn@canonical.com>
Reported-by: Serge Hallyn <serge.hallyn@ubuntu.com>
Signed-off-by: "Eric W. Biederman" <ebiederm@xmission.com>
copy_process() adds the new child to thread_group/init_task.tasks list and
then does attach_pid(child, PIDTYPE_PID). This means that the lockless
next_thread() or next_task() can see this thread with the wrong pid. Say,
"ls /proc/pid/task" can list the same inode twice.
We could move attach_pid(child, PIDTYPE_PID) up, but in this case
find_task_by_vpid() can find the new thread before it was fully
initialized.
And this is already true for PIDTYPE_PGID/PIDTYPE_SID, With this patch
copy_process() initializes child->pids[*].pid first, then calls
attach_pid() to insert the task into the pid->tasks list.
attach_pid() no longer need the "struct pid*" argument, it is always
called after pid_link->pid was already set.
Signed-off-by: Oleg Nesterov <oleg@redhat.com>
Cc: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: Michal Hocko <mhocko@suse.cz>
Cc: Pavel Emelyanov <xemul@parallels.com>
Cc: Sergey Dyasly <dserrg@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Pull VFS updates from Al Viro,
Misc cleanups all over the place, mainly wrt /proc interfaces (switch
create_proc_entry to proc_create(), get rid of the deprecated
create_proc_read_entry() in favor of using proc_create_data() and
seq_file etc).
7kloc removed.
* 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs: (204 commits)
don't bother with deferred freeing of fdtables
proc: Move non-public stuff from linux/proc_fs.h to fs/proc/internal.h
proc: Make the PROC_I() and PDE() macros internal to procfs
proc: Supply a function to remove a proc entry by PDE
take cgroup_open() and cpuset_open() to fs/proc/base.c
ppc: Clean up scanlog
ppc: Clean up rtas_flash driver somewhat
hostap: proc: Use remove_proc_subtree()
drm: proc: Use remove_proc_subtree()
drm: proc: Use minor->index to label things, not PDE->name
drm: Constify drm_proc_list[]
zoran: Don't print proc_dir_entry data in debug
reiserfs: Don't access the proc_dir_entry in r_open(), r_start() r_show()
proc: Supply an accessor for getting the data from a PDE's parent
airo: Use remove_proc_subtree()
rtl8192u: Don't need to save device proc dir PDE
rtl8187se: Use a dir under /proc/net/r8180/
proc: Add proc_mkdir_data()
proc: Move some bits from linux/proc_fs.h to linux/{of.h,signal.h,tty.h}
proc: Move PDE_NET() to fs/proc/proc_net.c
...
Split the proc namespace stuff out into linux/proc_ns.h.
Signed-off-by: David Howells <dhowells@redhat.com>
cc: netdev@vger.kernel.org
cc: Serge E. Hallyn <serge.hallyn@ubuntu.com>
cc: Eric W. Biederman <ebiederm@xmission.com>
Signed-off-by: Al Viro <viro@zeniv.linux.org.uk>
Move BITS_PER_PAGE from pid_namespace.c to pid_namespace.h, since we can
simplify the define PID_MAP_ENTRIES by using the BITS_PER_PAGE.
[akpm@linux-foundation.org: kernel/pid.c:54:1: warning: "BITS_PER_PAGE" redefined]
Signed-off-by: Raphael S.Carvalho <raphael.scarv@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
find_next_offset() searches for an available "cleaned bit" in the
respective pid bitmap (page), so returns the offset if found, otherwise
it returns a value equals to BITS_PER_PAGE.
For example, suppose find_next_offset didn't find any available bit, so
there's no purpose to call mk_pid (Wasteful Cpu Cycles).
Therefore, I found it could be better to call mk_pid after the checking
(offset < BITS_PER_PAGE) returned sucessfully! Another point: If (offset
< BITS_PER_PAGE) results in a "failure", then mk_pid would be called
again afterwards.
[akpm@linux-foundation.org: simplify code]
Signed-off-by: Raphael S. Carvalho <raphael.scarv@gmail.com>
Cc: "Eric W. Biederman" <ebiederm@xmission.com>
Cc: Serge Hallyn <serge.hallyn@canonical.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>