Commit graph

884,342 commits

Author SHA1 Message Date
Guru Das Srinagesh
1b68d5edc0 soc: qcom: qti_battery_debug: Add NULL check
Ensure that valid memory is allocated for the array of all votables
before an attempt is made to populate it.

Change-Id: I9a0c3e35e345a88560e39d47484348b6c476628d
Signed-off-by: Guru Das Srinagesh <gurus@codeaurora.org>
2020-06-05 15:19:43 -07:00
Patrick Daly
58e71e6912 iommu: arm-smmu: Add support for new attributes
Increase client control of iommu fault handling by allowing
separate configuration of the HUPCF and CFCFG bits.

Change-Id: Ifc6071775f171ecfe60a9f0824491536b3295ec4
Signed-off-by: Patrick Daly <pdaly@codeaurora.org>
Signed-off-by: Isaac J. Manjarres <isaacm@codeaurora.org>
2020-06-05 15:19:42 -07:00
Isaac J. Manjarres
6476e65d66 soc: qcom: mem-buf: Fix error handling when releasing memory
When releasing memory, notify the owner of the memory that it
can be reclaimed only if the recipient VM has successfully
released the memory. There is also no reason to continue
to hold onto the mem-buf descriptor memory if a failure
occurs when relinquishing the memory, so free those structures
as well, independent of the result of the relinquish operation.
While we're here, touch up some of the error logging in the
error paths as well.

Change-Id: I29bb1774507fd75f206c0537b6007182b4d6cb48
Signed-off-by: Isaac J. Manjarres <isaacm@codeaurora.org>
2020-06-05 14:54:24 -07:00
Raghavendra Rao Ananta
beb519f70a haven: dbl: Fix use-after-free in tx/rx unregister
The hh_dbl_tx_unregister() and hh_dbl_rx_unregister() functions
tries to dereference client_desc in pr_debug(), which would
have already been freed if the tx/rx's counterpart was unregistered.
Hence, move the pr_debug() statements right before the section
where kfree() is called on client_desc to avoid use-after-free.

Change-Id: I183b0a3df0665ab90a85e0907b39641cd19f4923
Signed-off-by: Raghavendra Rao Ananta <rananta@codeaurora.org>
2020-06-05 14:40:44 -07:00
Isaac J. Manjarres
dbe59ab3f8 soc: qcom: mem-buf: Do not free memory if hyp_assign() fails
The state of memory--with respect to access control--is unknown
when a hyp_assign() call fails because the underlying SCM call
fails. In these cases, it is not safe to access the memory, so
we should not free it back to the system.

Change-Id: I992225cdd5f26fbba60d2966d051d026326cd593
Signed-off-by: Isaac J. Manjarres <isaacm@codeaurora.org>
2020-06-05 14:33:21 -07:00
Isaac J. Manjarres
8340c9f96e soc: qcom: mem-buf: Align allocation sizes to MHP subsection size
The allocation code hot-adds the memory from another VM through
the memory hotplug code to create the S1 CPU MMU mappings. The memory
hotplug code can only add memory at a subsection granule, meaning
that all memory requests must be subsection size aligned. Thus,
enforce all allocation requests to be subsection size aligned.

Change-Id: Ia4d3294e36265f0bd956fa42670f2ac8d6974908
Signed-off-by: Isaac J. Manjarres <isaacm@codeaurora.org>
2020-06-05 14:32:52 -07:00
Subbaraman Narayanamurthy
4c9179cd12 defconfig: lahaina: Enable AMOLED ECM driver
AMOLED ECM (Embedded Current Measurement) driver helps measure
the display current for OLED panels. Enable it.

Change-Id: I94d729dbfa8b46792afd424333a0bf9c106ffef1
Signed-off-by: Subbaraman Narayanamurthy <subbaram@codeaurora.org>
2020-06-05 13:55:16 -07:00
Elliot Berman
f7c7a20027 haven: irq: Support lending from other domains
Client drivers may lend IRQs with knowledge of only their GPIO interrupt
number, which would not directly have an underlying GIC hwirq. Thus,
tweak RM's understanding of IRQs to be aware of IRQ domains. Now, the
irq backed by GPIO will be translated to GIC domain.

Change-Id: I191d9e072b14f4501f5bbeb5d5263f50e0eef8c1
Signed-off-by: Elliot Berman <eberman@codeaurora.org>
2020-06-05 13:48:40 -07:00
Jeevan Shriram
cab5940348 include: linux: remove unused APIs when CORESIGHT is disabled
Remove unused apis when the config CORESIGHT is disabled to avoid
compilation error.

Change-Id: I1a05d26d8f07fc70c3f061c5d5debe1f4fb7edaa
Signed-off-by: Jeevan Shriram <jshriram@codeaurora.org>
2020-06-05 12:15:31 -07:00
Srinivas Rao L
2c212faa23 cpuidle: lpm_levels: Wakeup biased cpu
If a biased cpu entered shallowest LPM state and
there are no wakeups for it, can stay in the shallowest
state for long.

Program wakeup for the biased CPU to wakeup after the
expected bias window is completed, so that the cpu can
enter a deeper state.

Change-Id: Ic92c779f0f8b1fa85aa8b3afa68d075f8d5d7dd6
Signed-off-by: Srinivas Rao L <lsrao@codeaurora.org>
2020-06-05 23:57:27 +05:30
qctecmdr
28cf0d0908 Merge "kernel: sound: update codec options with block size" 2020-06-05 08:56:45 -07:00
qctecmdr
ca65ecc5c8 Merge "config: Enable TOS and DSCP target support" 2020-06-05 08:56:44 -07:00
qctecmdr
51baf5ad2c Merge "clk: qcom: clk-rcg2: Add support to print rcg's CMD_DFSR register" 2020-06-05 05:06:33 -07:00
qctecmdr
86bb23c4e5 Merge "dwc3: gadget: Don't block doorbell before halting USB controller" 2020-06-05 05:06:32 -07:00
qctecmdr
a455d70f08 Merge "dwc3-msm: Move override usb speed functionality outside edev check" 2020-06-05 05:06:32 -07:00
qctecmdr
65fbcfac6f Merge "usb: dwc3: gsi: Disable GSI wrapper on clearing run_stop" 2020-06-05 05:06:28 -07:00
qctecmdr
5313030298 Merge "sched: Compile cpu_isolated_mask in SCHED_WALT only" 2020-06-05 05:06:27 -07:00
qctecmdr
9c7599702e Merge "defconfig: arm64: Enable Global clock controller for HOLI" 2020-06-05 05:06:27 -07:00
qctecmdr
25a07ed45c Merge "usb: phy: Reset and initialize HSPHY in host mode when EUD is enable" 2020-06-05 05:06:26 -07:00
qctecmdr
4b99be94b6 Merge "hwmon: Add QTI AMOLED ECM driver" 2020-06-05 05:06:26 -07:00
Sharath Chandra Vurukala
bec8ff53c4 config: Enable TOS and DSCP target support
Enable TOS and DSCP target support for Holi target

Change-Id: Ib7fbe8442007160f8158f41ba29613fdf1b7fcd1
Signed-off-by: Sharath Chandra Vurukala <sharathv@codeaurora.org>
2020-06-05 16:16:37 +05:30
Sumukh Hallymysore Ravindra
2b0acb8d64 msm: synx: default user callback fix
Avoid race with enqueue and dequeue of user callback
data. This is necessary since the callback added in the
event queue could be dequeued and variable cleaned up
before wake up function is completed.

Change-Id: I70aa09d8fc88bcb6e7e4b7aebd0f854a9de075d0
Signed-off-by: Sumukh Hallymysore Ravindra <shallymy@codeaurora.org>
2020-06-05 15:39:04 +05:30
Vinayak Menon
262e96de91 taskstats: handle NULL nla case in taskstats2
Avoid NULL dereference in taskstats2_foreach.

Fixes: 75aad5da3593 ("taskstats: add a option to send all tasks data to user")
Change-Id: I0c1860c003b73bcecee7a7c7db2939668517ac82
Signed-off-by: Vinayak Menon <vinmenon@codeaurora.org>
Signed-off-by: Prakash Gupta <guptap@codeaurora.org>
2020-06-05 14:38:07 +05:30
Vinayak Menon
5a5b186dfb taskstats: add support for system stats
Add a new command to get system wide information from
userspace. This patch adds system command to export
memory statistics.

Change-Id: I20c1c24f49024f49b21e3d8e88799559fb5b2058
Signed-off-by: Vinayak Menon <vinmenon@codeaurora.org>
[guptap@codeaurora.org: rename nr_indirectly_reclaimable_bytes]
Signed-off-by: Prakash Gupta <guptap@codeaurora.org>
Signed-off-by: Mukesh Ojha <mojha@codeaurora.org>
2020-06-05 14:38:07 +05:30
Vinayak Menon
acffd6f45a taskstats: add a option to send all tasks data to user
Adds a new taskstats command to share task memory statistics
to userspace. The command works in two modes. It can share
information per pid, and also send the statistics for all
tasks within a given oom_score_adj range.

Change-Id: Iae742ea4f96754022bc634d285c5ce140b32749d
Signed-off-by: Vinayak Menon <vinmenon@codeaurora.org>
[guptap@codeaurora.org: enforce policy later using pre_doit]
Signed-off-by: Prakash Gupta <guptap@codeaurora.org>
2020-06-05 14:38:07 +05:30
Vinayak Menon
efa5ce2c9a mm: skip rss check on MM_UNRECLAIMABLE
MM_UNRECLAIMABLE rss counter can be updated by drivers
on exit_files. But since exit_mm is called early, there
is a chance of false bad rss messages. Skip the check
for MM_UNRECLAIMABLE.

Change-Id: Id9a79db20f1ae711ec801a646d7c28d92e94f70b
Signed-off-by: Vinayak Menon <vinmenon@codeaurora.org>
[guptap@codeaurora.org: Add Kconfig entry to prevent ABI breakages in GKI]
Signed-off-by: Prakash Gupta <guptap@codeaurora.org>
2020-06-05 14:38:07 +05:30
Vijayanand Jitta
ff60d1d2db ion: add ion pages to NR_UNRECLAIMABLE_PAGES
add the ion pages allocated through system and system secure
heaps to NR_UNRECLAIMABLE_PAGES memory counter.

Change-Id: I4bf846fcd0454a5ba0e1cad08e9b7e882f92ceac
Signed-off-by: Vijayanand Jitta <vjitta@codeaurora.org>
[guptap@codeaurora.org: Add Kconfig entry to prevent ABI breakages in GKI]
Signed-off-by: Prakash Gupta <guptap@codeaurora.org>
2020-06-05 14:38:07 +05:30
Vijayanand Jitta
bfe3637f89 mm: introduce NR_UNRECLAIMABLE_PAGES
Introduce NR_UNRECLAIMABLE_PAGES memory counter which accounts
the pages that cannot be reclaimed under memory pressure.

Change-Id: I9afe50537b0d3c2e7ffc07916b23cce4329e3679
Signed-off-by: Vijayanand Jitta <vjitta@codeaurora.org>
[guptap@codeaurora.org: split resident_page_types change]
Signed-off-by: Prakash Gupta <guptap@codeaurora.org>
2020-06-05 14:38:07 +05:30
Vinayak Menon
c99ba85df8 mm: add rss counter for unreclaimable pages
Add a per mm rss counter to hold the unreclaimable
pages. This can include the pages allocated by a
task and shared with hardware for DMA etc.

Change-Id: Iec77d69eca0a4f8f6e23f866c80c0143620fcaf2
Signed-off-by: Vinayak Menon <vinmenon@codeaurora.org>
[guptap@codeaurora.org: update resident_page_types dependency, add kconfig]
Signed-off-by: Prakash Gupta <guptap@codeaurora.org>
2020-06-05 14:38:06 +05:30
qctecmdr
7b7d13e1a5 Merge "drivers: soc: qcom: handle system sleep activities" 2020-06-05 00:39:02 -07:00
qctecmdr
1fa725da96 Merge "clk: qcom: lahaina: Add pll test ctl regs" 2020-06-05 00:39:02 -07:00
qctecmdr
732adf0e7c Merge "clk: qcom: gcc-lahaina: Add USB force_mem_core_on clocks" 2020-06-05 00:39:01 -07:00
qctecmdr
31cb055051 Merge "arm64: configs: Disable DCC console for Lahaina" 2020-06-05 00:39:01 -07:00
Minchan Kim
f51a23e5bf mm/madvise: pass task and mm to do_madvise
Patch series "introduce memory hinting API for external process", v7.

Now, we have MADV_PAGEOUT and MADV_COLD as madvise hinting API.  With
that, application could give hints to kernel what memory range are
preferred to be reclaimed.  However, in some platform(e.g., Android), the
information required to make the hinting decision is not known to the app.
Instead, it is known to a centralized userspace daemon(e.g.,
ActivityManagerService), and that daemon must be able to initiate reclaim
on its own without any app involvement.

To solve the concern, this patch introduces new syscall -
process_madvise(2).  Bascially, it's same with madvise(2) syscall but it
has some differences.

1. It needs pidfd of target process to provide the hint

2.  It supports only MADV_{COLD|PAGEOUT|MERGEABLE|UNMEREABLE} at this
   moment.  Other hints in madvise will be opened when there are explicit
   requests from community to prevent unexpected bugs we couldn't support.

3.  Only privileged processes can do something for other process's
   address space.

For more detail of the new API, please see "mm: introduce external memory
hinting API" description in this patchset.

This patch (of 7):

In upcoming patches, do_madvise will be called from external process
context so we shouldn't asssume "current" is always hinted process's
task_struct.

Furthermore, we must not access mm_struct via task->mm, but obtain it
via access_mm() once (in the following patch) and only use that pointer
[1], so pass it to do_madvise() as well.  Note the vma->vm_mm pointers
are safe, so we can use them further down the call stack.

And let's pass *current* and current->mm as arguments of do_madvise so
it shouldn't change existing behavior but prepare next patch to make
review easy.

Note: io_madvise passes NULL as target_task argument of do_madvise because
it couldn't know who is target.

[1] http://lore.kernel.org/r/CAG48ez27=pwm5m_N_988xT1huO7g7h6arTQL44zev6TD-h-7Tg@mail.gmail.com

[vbabka@suse.cz: changelog tweak]
[minchan@kernel.org: use current->mm for io_uring]
  Link: http://lkml.kernel.org/r/20200423145215.72666-1-minchan@kernel.org
[akpm@linux-foundation.org: fix it for upstream changes]
[akpm@linux-foundation.org: whoops]
[rdunlap@infradead.org: add missing includes]
Link: http://lkml.kernel.org/r/20200302193630.68771-2-minchan@kernel.org
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Jann Horn <jannh@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: Daniel Colascione <dancol@google.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: John Dias <joaodias@google.com>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Kirill Tkhai <ktkhai@virtuozzo.com>
Cc: Oleksandr Natalenko <oleksandr@redhat.com>
Cc: SeongJae Park <sjpark@amazon.de>
Cc: Christian Brauner <christian.brauner@ubuntu.com>
Cc: <linux-man@vger.kernel.org>
From: Minchan Kim <minchan@kernel.org>
Subject: mm/madvise: introduce process_madvise() syscall: an external memory hinting API

There is usecase that System Management Software(SMS) want to give a
memory hint like MADV_[COLD|PAGEEOUT] to other processes and in the
case of Android, it is the ActivityManagerService.

The information required to make the reclaim decision is not known to
the app.  Instead, it is known to the centralized userspace
daemon(ActivityManagerService), and that daemon must be able to
initiate reclaim on its own without any app involvement.

To solve the issue, this patch introduces a new syscall
process_madvise(2).  It uses pidfd of an external process to give the
hint.

 int process_madvise(int pidfd, void *addr, size_t length, int advice,
			unsigned long flags);

Since it could affect other process's address range, only privileged
process(CAP_SYS_PTRACE) or something else(e.g., being the same UID)
gives it the right to ptrace the process could use it successfully.
The flag argument is reserved for future use if we need to extend the
API.

I think supporting all hints madvise has/will supported/support to
process_madvise is rather risky.  Because we are not sure all hints
make sense from external process and implementation for the hint may
rely on the caller being in the current context so it could be
error-prone.  Thus, I just limited hints as MADV_[COLD|PAGEOUT] in this
patch.

If someone want to add other hints, we could hear hear the usecase and
review it for each hint.  It's safer for maintenance rather than
introducing a buggy syscall but hard to fix it later.

Q.1 - Why does any external entity have better knowledge?

Quote from Sandeep

"For Android, every application (including the special SystemServer)
are forked from Zygote.  The reason of course is to share as many
libraries and classes between the two as possible to benefit from the
preloading during boot.

After applications start, (almost) all of the APIs end up calling into
this SystemServer process over IPC (binder) and back to the
application.

In a fully running system, the SystemServer monitors every single
process periodically to calculate their PSS / RSS and also decides
which process is "important" to the user for interactivity.

So, because of how these processes start _and_ the fact that the
SystemServer is looping to monitor each process, it does tend to *know*
which address range of the application is not used / useful.

Besides, we can never rely on applications to clean things up
themselves.  We've had the "hey app1, the system is low on memory,
please trim your memory usage down" notifications for a long time[1].
They rely on applications honoring the broadcasts and very few do.

So, if we want to avoid the inevitable killing of the application and
restarting it, some way to be able to tell the OS about unimportant
memory in these applications will be useful.

- ssp

Q.2 - How to guarantee the race(i.e., object validation) between when
giving a hint from an external process and get the hint from the target
process?

process_madvise operates on the target process's address space as it
exists at the instant that process_madvise is called.  If the space
target process can run between the time the process_madvise process
inspects the target process address space and the time that
process_madvise is actually called, process_madvise may operate on
memory regions that the calling process does not expect.  It's the
responsibility of the process calling process_madvise to close this
race condition.  For example, the calling process can suspend the
target process with ptrace, SIGSTOP, or the freezer cgroup so that it
doesn't have an opportunity to change its own address space before
process_madvise is called.  Another option is to operate on memory
regions that the caller knows a priori will be unchanged in the target
process.  Yet another option is to accept the race for certain
process_madvise calls after reasoning that mistargeting will do no
harm.  The suggested API itself does not provide synchronization.  It
also apply other APIs like move_pages, process_vm_write.

The race isn't really a problem though.  Why is it so wrong to require
that callers do their own synchronization in some manner?  Nobody
objects to write(2) merely because it's possible for two processes to
open the same file and clobber each other's writes --- instead, we tell
people to use flock or something.  Think about mmap.  It never
guarantees newly allocated address space is still valid when the user
tries to access it because other threads could unmap the memory right
before.  That's where we need synchronization by using other API or
design from userside.  It shouldn't be part of API itself.  If someone
needs more fine-grained synchronization rather than process level,
there were two ideas suggested - cookie[2] and anon-fd[3].  Both are
applicable via using last reserved argument of the API but I don't
think it's necessary right now since we have already ways to prevent
the race so don't want to add additional complexity with more
fine-grained optimization model.

To make the API extend, it reserved an unsigned long as last argument
so we could support it in future if someone really needs it.

Q.3 - Why doesn't ptrace work?

Injecting an madvise in the target process using ptrace would not work
for us because such injected madvise would have to be executed by the
target process, which means that process would have to be runnable and
that creates the risk of the abovementioned race and hinting a wrong
VMA.  Furthermore, we want to act the hint in caller's context, not the
callee's, because the callee is usually limited in cpuset/cgroups or
even freezed state so they can't act by themselves quick enough, which
causes more thrashing/kill.  It doesn't work if the target process are
ptraced(e.g., strace, debugger, minidump) because a process can have at
most one ptracer.

[1] https://developer.android.com/topic/performance/memory"

[2] process_getinfo for getting the cookie which is updated whenever
    vma of process address layout are changed - Daniel Colascione -
    https://lore.kernel.org/lkml/20190520035254.57579-1-minchan@kernel.org/T/#m7694416fd179b2066a2c62b5b139b14e3894e224

[3] anonymous fd which is used for the object(i.e., address range)
    validation - Michal Hocko -
    https://lore.kernel.org/lkml/20200120112722.GY18451@dhcp22.suse.cz/

Link: http://lkml.kernel.org/r/20200302193630.68771-3-minchan@kernel.org
Link: http://lkml.kernel.org/r/20200508183320.GA125527@google.com
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
Cc: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Daniel Colascione <dancol@google.com>
Cc: Jann Horn <jannh@google.com>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: John Dias <joaodias@google.com>
Cc: Kirill Tkhai <ktkhai@virtuozzo.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Oleksandr Natalenko <oleksandr@redhat.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: SeongJae Park <sjpark@amazon.de>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: <linux-man@vger.kernel.org>
From: Minchan Kim <minchan@kernel.org>
Subject: fix process_madvise build break for arm64

0-day reported build break from process_madvise on ARM64.

   aarch64-linux-ld: arch/arm64/kernel/head.o: relocation R_AARCH64_ABS32 against `_kernel_offset_le_lo32' can not be used when making a shared object
   aarch64-linux-ld: arch/arm64/kernel/efi-entry.stub.o: relocation R_AARCH64_ABS32 against `__efistub_stext_offset' can not be used when making a shared object
   arch/arm64/kernel/head.o: In function `kimage_vaddr':
   (.idmap.text+0x0): dangerous relocation: unsupported relocation
   arch/arm64/kernel/head.o: In function `__primary_switch':
   (.idmap.text+0x378): dangerous relocation: unsupported relocation
   (.idmap.text+0x380): dangerous relocation: unsupported relocation
>> arch/arm64/kernel/sys32.o:(.rodata+0xdb8): undefined reference to `__arm64_process_madvise'

This patch should fix it.

Link: http://lkml.kernel.org/r/20200303145756.GA219683@google.com
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reported-by: kbuild test robot <lkp@intel.com>
From: Minchan Kim <minchan@kernel.org>
Subject: mm: fix build error for mips of process_madvise

kbuild test rebot reported build break of process_madvise for mips[1].
This patch should fix it.

[1] https://lore.kernel.org/linux-mm/202005080716.cUcbCQ3i%25lkp@intel.com/
Link: http://lkml.kernel.org/r/20200508052517.GA197378@google.com
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reported-by: kbuild test robot <lkp@intel.com>
Cc: Stephen Rothwell <sfr@canb.auug.org.au>
From: Andrew Morton <akpm@linux-foundation.org>
Subject: mm-introduce-external-memory-hinting-api-fix-2-fix

the compat bit comes later

Cc: Minchan Kim <minchan@kernel.org>

From: Minchan Kim <minchan@kernel.org>
Subject: mm/madvise: check fatal signal pending of target process

Bail out to prevent unnecessary CPU overhead if target process has pending
fatal signal during (MADV_COLD|MADV_PAGEOUT) operation.

Link: http://lkml.kernel.org/r/20200302193630.68771-4-minchan@kernel.org
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
Cc: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Daniel Colascione <dancol@google.com>
Cc: Jann Horn <jannh@google.com>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: John Dias <joaodias@google.com>
Cc: Kirill Tkhai <ktkhai@virtuozzo.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Oleksandr Natalenko <oleksandr@redhat.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: SeongJae Park <sjpark@amazon.de>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: <linux-man@vger.kernel.org>
From: Minchan Kim <minchan@kernel.org>
Subject: pid: move pidfd_get_pid() to pid.c

process_madvise syscall needs pidfd_get_pid function to translate pidfd to
pid so this patch move the function to kernel/pid.c.

Link: http://lkml.kernel.org/r/20200302193630.68771-5-minchan@kernel.org
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Suggested-by: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Reviewed-by: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Acked-by: Christian Brauner <christian.brauner@ubuntu.com>
Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Jann Horn <jannh@google.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Daniel Colascione <dancol@google.com>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: John Dias <joaodias@google.com>
Cc: Kirill Tkhai <ktkhai@virtuozzo.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Oleksandr Natalenko <oleksandr@redhat.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: SeongJae Park <sjpark@amazon.de>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: <linux-man@vger.kernel.org>
From: Minchan Kim <minchan@kernel.org>
Subject: mm/madvise: support both pid and pidfd for process_madvise

There is a demand[1] to support pid as well pidfd for process_madvise
to reduce unnecessary syscall to get pidfd if the user has control of
the target process (ie, they could guarantee the process is not gone or
pid is not reused).

This patch aims for supporting both options like waitid(2).  So, the
syscall is currently,

        int process_madvise(idtype_t idtype, id_t id, void *addr,
                size_t length, int advice, unsigned long flags);

@which is actually idtype_t for userspace library and currently, it
supports P_PID and P_PIDFD.

[1]  https://lore.kernel.org/linux-mm/9d849087-3359-c4ab-fbec-859e8186c509@virtuozzo.com/

Link: http://lkml.kernel.org/r/20200302193630.68771-6-minchan@kernel.org
Signed-off-by: Minchan Kim <minchan@kernel.org>
Suggested-by: Kirill Tkhai <ktkhai@virtuozzo.com>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Vlastimil Babka <vbabka@suse.cz>
Cc: Christian Brauner <christian@brauner.io>
Cc: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Daniel Colascione <dancol@google.com>
Cc: Jann Horn <jannh@google.com>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: John Dias <joaodias@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Oleksandr Natalenko <oleksandr@redhat.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: SeongJae Park <sjpark@amazon.de>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: <linux-man@vger.kernel.org>
From: Oleksandr Natalenko <oleksandr@redhat.com>
Subject: mm/madvise: allow KSM hints for remote API

It all began with the fact that KSM works only on memory that is marked by
madvise().  And the only way to get around that is to either:

  * use LD_PRELOAD; or
  * patch the kernel with something like UKSM or PKSM.

(i skip ptrace can of worms here intentionally)

To overcome this restriction, lets employ a new remote madvise API.  This
can be used by some small userspace helper daemon that will do auto-KSM
job for us.

I think of two major consumers of remote KSM hints:

  * hosts, that run containers, especially similar ones and especially in
    a trusted environment, sharing the same runtime like Node.js;

  * heavy applications, that can be run in multiple instances, not
    limited to opensource ones like Firefox, but also those that cannot be
    modified since they are binary-only and, maybe, statically linked.

Speaking of statistics, more numbers can be found in the very first
submission, that is related to this one [1].  For my current setup with
two Firefox instances I get 100 to 200 MiB saved for the second instance
depending on the amount of tabs.

1 FF instance with 15 tabs:

   $ echo "$(cat /sys/kernel/mm/ksm/pages_sharing) * 4 / 1024" | bc
   410

2 FF instances, second one has 12 tabs (all the tabs are different):

   $ echo "$(cat /sys/kernel/mm/ksm/pages_sharing) * 4 / 1024" | bc
   592

At the very moment I do not have specific numbers for containerised
workload, but those should be comparable in case the containers share
similar/same runtime.

[1] https://lore.kernel.org/patchwork/patch/1012142/

Link: http://lkml.kernel.org/r/20200302193630.68771-8-minchan@kernel.org
Signed-off-by: Oleksandr Natalenko <oleksandr@redhat.com>
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reviewed-by: SeongJae Park <sjpark@amazon.de>
Cc: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Daniel Colascione <dancol@google.com>
Cc: Jann Horn <jannh@google.com>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: John Dias <joaodias@google.com>
Cc: Kirill Tkhai <ktkhai@virtuozzo.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: Vlastimil Babka <vbabka@suse.cz>
Cc: <linux-man@vger.kernel.org>
From: Minchan Kim <minchan@kernel.org>
Subject: mm: support vector address ranges for process_madvise

This patch extends a) process_madvise(2) support vector address ranges in
a system call and then b) support the vector address ranges to local
process as well as external process.

Android app has thousands of vmas due to zygote so it's totally waste of
CPU and power if we should call the syscall one by one for each vma.
(With testing 2000-vma syscall vs 1-vector syscall, it showed 15%
performance improvement.  I think it would be bigger in real practice
because the testing ran very cache friendly environment).

Another potential use case for the vector range is to amortize the cost of
TLB shootdowns for multiple ranges when using MADV_DONTNEED; this could
benefit users like TCP receive zerocopy and malloc implementations.  In
future, we could find more usecases for other advises so let's make it
happens as API since we introduce a new syscall at this moment.  With
that, existing madvise(2) user could replace it with process_madvise(2)
with their own pid if they want to have batch address ranges support
feature.

So finally, the API is as follows,

  ssize_t process_madvise(idtype_t idtype, id_t id,
		const struct iovec *iovec, unsigned long vlen,
                int advice, unsigned long flags);

DESCRIPTION
  The process_madvise() system call is used to give advice or directions
  to the kernel about the address ranges from external process as well as
  local process. It provides the advice to address ranges of process
  described by iovec and vlen. The goal of such advice is to improve system
  or application performance.

  The idtype and id arguments select the target process to be advised as
  follows:

    idtype == P_PID
      select the process whose process ID matches id

    idtype == P_PIDFD
      select the process referred to by the PID file descriptor
      specified in id. (See pidofd_open(2) for further information)

  The pointer iovec points to an array of iovec structures, defined in
  <sys/uio.h> as:

    struct iovec {
    	void *iov_base;		/* starting address */
	size_t iov_len;		/* number of bytes to be advised */
    };

  The iovec describes address ranges beginning at address(iov_base)
  and with size length of bytes(iov_len).

  The vlen represents the number of elements in iovec.

  The advice is indicated in the advice argument, which is one of the
  following at this moment if the target process specified by idtype and
  id is external.

    MADV_COLD
    MADV_PAGEOUT
    MADV_MERGEABLE
    MADV_UNMERGEABLE

  Permission to provide a hint to external process is governed by a
  ptrace access mode PTRACE_MODE_ATTACH_FSCREDS check; see ptrace(2).

  The process_madvise supports every advice madvise(2) has if target
  process is in same thread group with calling process so user could
  use process_madvise(2) to extend existing madvise(2) to support
  vector address ranges.

RETURN VALUE
  On success, process_madvise() returns the number of bytes advised.
  This return value may be less than the total number of requested
  bytes, if an error occurred. The caller should check return value
  to determine whether a partial advice occurred.

Link: http://lkml.kernel.org/r/20200423145215.72666-2-minchan@kernel.org
Signed-off-by: Minchan Kim <minchan@kernel.org>
Cc: David Rientjes <rientjes@google.com>
Cc: Arjun Roy <arjunroy@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: Daniel Colascione <dancol@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: John Dias <joaodias@google.com>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: Oleksandr Natalenko <oleksandr@redhat.com>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Vlastimil Babka <vbabka@suse.cz>
From: Minchan Kim <minchan@kernel.org>
Subject: mm: support compat_sys_process_madvise

This patch supports compat syscall for process_madvise

Link: http://lkml.kernel.org/r/20200423195835.GA46847@google.com
Signed-off-by: Minchan Kim <minchan@kernel.org>
Cc: Arjun Roy <arjunroy@google.com>
Cc: Brian Geffon <bgeffon@google.com>
Cc: Daniel Colascione <dancol@google.com>
Cc: David Rientjes <rientjes@google.com>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: John Dias <joaodias@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Oleksandr Natalenko <oleksandr@redhat.com>
Cc: Sandeep Patil <sspatil@google.com>
Cc: SeongJae Park <sj38.park@gmail.com>
Cc: Shakeel Butt <shakeelb@google.com>
Cc: Sonny Rao <sonnyrao@google.com>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Tim Murray <timmurray@google.com>
Cc: Vlastimil Babka <vbabka@suse.cz>
From: Randy Dunlap <rdunlap@infradead.org>
Subject: mm-support-vector-address-ranges-for-process_madvise-fix-fix

fix process_madvise prototype

Cc: Minchan Kim <minchan@kernel.org>
From: Zheng Bin <zhengbin13@huawei.com>
Subject: mm/madvise: make function 'do_process_madvise' static

Fix sparse warnings:

mm/madvise.c:1233:9: warning: symbol 'do_process_madvise' was not declared. Should it be static?

Link: http://lkml.kernel.org/r/20200429014030.41147-1-zhengbin13@huawei.com
Signed-off-by: Zheng Bin <zhengbin13@huawei.com>
Reported-by: Hulk Robot <hulkci@huawei.com>
Cc: Minchan Kim <minchan@kernel.org>
From: Minchan Kim <minchan@kernel.org>
Subject: mm: fix s390 compat build error

Nathan reported build error with sys_compat_process_madvise.
This patch should fix it.

Link: http://lkml.kernel.org/r/20200429012421.GA132200@google.com
Signed-off-by: Minchan Kim <minchan@kernel.org>
Reported-by: Nathan Chancellor <natechancellor@gmail.com>
Tested-by: Nathan Chancellor <natechancellor@gmail.com>	[build]
From: Andrew Morton <akpm@linux-foundation.org>
Subject: mm-support-vector-address-ranges-for-process_madvise-fix-fix-fix-fix-fix

add compat_sys_process_madvise to mips syscall table

Conflicts:
	fs/io_uring.c
	mm/madvise.c

Change-Id: I89b92904043c6d7fbf9747746d20b823dbc20410
Cc: Minchan Kim <minchan@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Stephen Rothwell <sfr@canb.auug.org.au>
Git-commit: 82f576dd0298d675df9b19ac0638d79b5ca79e59
Git-Repo: git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
[charante@codeaurora.org: Fixed merge conflicts]
Signed-off-by: Charan Teja Reddy <charante@codeaurora.org>
2020-06-05 11:09:36 +05:30
Linus Torvalds
4b3fb647e2 mm: check that mm is still valid in madvise()
IORING_OP_MADVISE can end up basically doing mprotect() on the VM of
another process, which means that it can race with our crazy core dump
handling which accesses the VM state without holding the mmap_sem
(because it incorrectly thinks that it is the final user).

This is clearly a core dumping problem, but we've never fixed it the
right way, and instead have the notion of "check that the mm is still
ok" using mmget_still_valid() after getting the mmap_sem for writing in
any situation where we're not the original VM thread.

See commit 04f5866e41 ("coredump: fix race condition between
mmget_not_zero()/get_task_mm() and core dumping") for more background on
this whole mmget_still_valid() thing.  You might want to have a barf bag
handy when you do.

We're discussing just fixing this properly in the only remaining core
dumping routines.  But even if we do that, let's make do_madvise() do
the right thing, and then when we fix core dumping, we can remove all
these mmget_still_valid() checks.

Change-Id: I503f628ca80bd759e101d03a96f3600294a0726a
Reported-and-tested-by: Jann Horn <jannh@google.com>
Fixes: c1ca757bd6f4 ("io_uring: add IORING_OP_MADVISE")
Acked-by: Jens Axboe <axboe@kernel.dk>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Git-commit: d208cbe90b0f39abbd019a383514306997ba47e3
Git-repo: git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
Signed-off-by: Charan Teja Reddy <charante@codeaurora.org>
2020-06-05 11:09:35 +05:30
Jens Axboe
4f8f769a62 mm: make do_madvise() available internally
This is in preparation for enabling this functionality through io_uring.
Add a helper that is just exporting what sys_madvise() does, and have the
system call use it.

No functional changes in this patch.

Change-Id: I77394d51a06aeb36e0ba733787dbd71a8e06a582
Reviewed-by: Pavel Begunkov <asml.silence@gmail.com>
Signed-off-by: Jens Axboe <axboe@kernel.dk>
Git-commit: db08ca25253d56f1f76eb4b3fe32a7ac1fbab741
Git-repo: git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
Signed-off-by: Charan Teja Reddy <charante@codeaurora.org>
2020-06-05 11:09:34 +05:30
qctecmdr
485cd69967 Merge "cpuidle: lpm-levels: Track and predict next rescheduling ipi" 2020-06-04 20:01:38 -07:00
qctecmdr
f071bb5252 Merge "defconfig: lahaina: Enable linux bridge" 2020-06-04 20:01:38 -07:00
qctecmdr
92daa813ad Merge "radio: RTC6226: implement file read for rtc6226 driver" 2020-06-04 20:01:38 -07:00
qctecmdr
6275165e1e Merge "qseecom: Enable APIs only when module is enabled" 2020-06-04 20:01:37 -07:00
qctecmdr
e49657f5cf Merge "radio: RTC6226: remove the V4L2_CAP_DEVICE_CAPS cap as device_caps" 2020-06-04 20:01:37 -07:00
qctecmdr
14e6052999 Merge "proc: update perms of node "reclaim"" 2020-06-04 20:01:37 -07:00
Vivek Aknurwar
0b066d6321 clk: qcom: clk-rcg2: Add support to print rcg's CMD_DFSR register
Add support to print rcg's CMD_DFSR register. DFS_EN bit 0 of CMD_DFSR
register indicates clk DFS mode.

Change-Id: I5d98550f89bdd14496f94edc455d67f61a1566a4
Signed-off-by: Vivek Aknurwar <viveka@codeaurora.org>
2020-06-04 15:03:49 -07:00
Vivek Aknurwar
da558fdefd clk: qcom: clk-alpha-pll: Add support to print PLL SSC registers
Add support to print PLL SSC registers.

Change-Id: Id2f6c19f4a768b3f30370a1b6a573d2b41d1daa9
Signed-off-by: Vivek Aknurwar <viveka@codeaurora.org>
2020-06-04 15:03:37 -07:00
qctecmdr
c2ac24a1a9 Merge "mhi: core: Read transfer length from an event properly" 2020-06-04 14:48:28 -07:00
qctecmdr
0efc14c3da Merge "mhi: core: Fix out of bound channel id handling" 2020-06-04 14:48:28 -07:00
qctecmdr
5a4580bd6b Merge "abi: Update qcom whitelist for cnss and netif" 2020-06-04 14:48:27 -07:00
qctecmdr
ab1b160dd8 Merge "coresight: etm4x: Fix use-after-free of per-cpu etm drvdata" 2020-06-04 14:48:27 -07:00
Mayank Rana
f1984dceeb dwc3-msm: Add support to vote USB FORCE_MEM_CORE_ON
USB FORCE_MEM_CORE_ON is need to set 1 with GCC_USB30_MASTER_CLK to
retain USB controller CSR when system is into CXPC (i.e. using MX
instead of CX). Without this bit set, USB controller CSR is not
retain once coming out of CXPC. Add support to vote this as clock
and tie with USB GDSC functionality.

Change-Id: I25e65a2d06848ef84ec5fd2041db8c6a8d9a7b4e
Signed-off-by: Mayank Rana <mrana@codeaurora.org>
2020-06-04 13:14:26 -07:00
Bhaumik Bhatt
036d20a74d mhi: core: Trigger host resume if client requests device vote
When an MHI client on the host requests a device resume/vote
using the mhi_device_get() API, host should also resume from
its own suspended state to allow efficient voting as MHI host
could be in DRV suspended state. Host must ensure that link
control is brought back to the application processor once a
vote is initiated.

Change-Id: I4dcf8f0d8079ea3fa4c354d5c8ee98c6ca0a4394
Signed-off-by: Bhaumik Bhatt <bbhatt@codeaurora.org>
2020-06-04 12:49:20 -07:00